NewYour coding agent can read the release notes before it upgrades.Set up the MCP server →
pub.dev · #1505 most downloaded on pub.dev
Moved to flutter_edge_ai. This is the final release under the name flutter_gemma; new versions ship as flutter_edge_ai, part of Flutter Edge AI.
Last release 3 days ago
05 Oct 2026
Ships on a steady schedule
a new release about every 1 weeks
Nearly every release is documented
notes for 59 of the last 60 stable releases
Nothing withdrawn
no release was ever pulled
3 years old
142 releases · first in 2024
The codelabs carry the litert_embeddings.js that flutter_gemma_litertlm 1.8.4 ships by @DenisovAV in #562
Full Changelog: v1.11.0...v1.11.4
One column per month.
flutter_edge_ai; this is the final release under the name flutter_gemma.New flutter-gemma-diagnostics agent skill for flutter_gemma_diagnostics.
flutter-gemma-diagnostics agent skill for flutter_gemma_diagnostics.Web setup pins @mediapipe/tasks-genai 0.10.29, @litert-lm/core 0.17.1, ORT-web 1.30.0, Transformers.js 4.3.0.
@mediapipe/tasks-genai 0.10.29, @litert-lm/core 0.17.1, ORT-web 1.30.0, Transformers.js 4.3.0.The built-in AI skill describes flutter_gemma_builtin_ai 0.3.0: Windows, Linux and setup.
flutter_gemma_builtin_ai 0.3.0: Windows, Linux and setup.LiteRT-LM v0.17.0 (native-v0.17.0): litertlm 1.7.0, speech 0.5.1 by @DenisovAV in #522
Full Changelog: v1.8.3...v1.11.0
EmbeddingModel.activeBackend: getActiveEmbedder(preferredBackend:) no longer goes unremarked.implements EmbeddingModel: add activeBackend and isClosed; extends inherits defaults.Add activationDataType to getActiveModel; float32 fixes wrong digits on some GPUs.
activationDataType to getActiveModel; float32 fixes wrong digits on some GPUs.New initialize(embeddingTokenizers:) — register one or embeddings throw; engines no longer bundle a tokenizer.
initialize(embeddingTokenizers:) — register one or embeddings throw; engines no longer bundle a tokenizer.web/rag/ build from the published archive.FunctionGemma on .litertlm answers a tool result instead of repeating the call.
.litertlm answers a tool result instead of repeating the call.fix(litertlm): don't clamp maxTokens on NPU, and say which models the NPU runs by @DenisovAV in #508
Full Changelog: v1.8.2...v1.8.3
.litertlm on iOS wrapping every prompt in turn markers twice; deprecate StopTokenFilter.docs(website): full content audit + Built-in AI page + Engines nav group by @DenisovAV in #455
Full Changelog: v1.6.5...v1.8.2
dart run skills@ get --all teaches your AI assistant this package.Add VectorStoreRepository.flush(); custom implementations must declare it (#492).
VectorStoreRepository.flush(); custom implementations must declare it (#492).Whisper output language on getActiveStt and transcribe (#500).
getActiveStt and transcribe (#500).SpeechRecognizer implementations: transcribe gained language: and the type gained a language field.Example: LFM2.5-230M and the fromHuggingFace install path.
fromHuggingFace install path.Fix createChat dropping tools — function calling on every non-FFI engine.
createChat dropping tools — function calling on every non-FFI engine.@litert-lm/core setup snippet to 0.17.0.Fix web sessions failing right after install because the active model was only half-saved.
macOS setup snippet is now a thin shim; the staging logic ships with flutter_gemma_litertlm (#457).
flutter_gemma_litertlm (#457).kotlinx.coroutines keep from the Android consumer rules so R8 can shrink them (#486).Hugging Face install (#454): resolveHuggingFace / one-call installModel(…).fromHuggingFace(repo), with engines auto-registering their resolver.
resolveHuggingFace / one-call installModel(…).fromHuggingFace(repo), with engines auto-registering their resolver.feat(builtin-ai): web arm via Chrome Prompt API (LanguageModel) — 0.2.0 by @DenisovAV in #446
Full Changelog: v1.6.4...v1.6.5
InferenceChat flag, not a hardcoded gemma4 check (no behavior change).fix(core): stop claiming background_downloader's updates stream from the host by @DenisovAV in #450
Full Changelog: v1.6.2...v1.6.4
flutter_gemma_mediapipe still needs 16.0 (#441).Download updates are scoped to our own task group; FileDownloader().updates stays the host's (#445).
FileDownloader().updates stays the host's (#445).docs(website): vision/audio per-encoder backend (1.5.9 / litertlm 1.4.2) by @DenisovAV in #430
Full Changelog: v1.5.9...v1.6.2
ModelFileType.onnx install is fileless — Transformers.js owns the repo, install just sets it active.getActiveModel compares every runtime param before reusing the cached model.
getActiveModel compares every runtime param before reusing the cached model.getActiveModel calls no longer share a model built for other params.FilterField.validateSchema — rejects duplicate field names, checked in release.FieldEquals / FieldRange semantics both stores must honour.Add ModelFileType.onnx + ONNX engine chat-format routing — enables the new flutter_gemma_onnx engine.
ModelFileType.onnx + ONNX engine chat-format routing — enables the new flutter_gemma_onnx engine.fix(litertlm): detect stream-callback ABI at runtime (LiteRT-LM v0.15.0) by @DenisovAV in #410
Full Changelog: v1.5.7...v1.5.9
docs: Windows discrete GPU is fixed in litertlm 1.4.0 — the README said it was still broken.
chore(website): docs → flutter_gemma 1.5.6; document VoiceSession streamAudio by @DenisovAV in #420
🙏 Credit: the #391 fix was also submitted independently by @o-mid in #421, opened before #422 landed — thank you for the fix and the report.
Full Changelog: v1.5.6...v1.5.7
fix(agent): require flutter_gemma ^1.5.5 (onMaxToolTurns) by @DenisovAV in #418
Full Changelog: v1.5.5...v1.5.6
consolidate: AgentLoop delegates its tool loop to core (core 1.5.5 + agent 0.2.2) by @DenisovAV in #417
Full Changelog: v1.5.4...v1.5.5
feat(agent): stream final-answer tokens in AgentLoop (VoiceSession parity) by @DenisovAV in #415
Full Changelog: v1.5.3...v1.5.4
Add InferenceChat.generateChatResponseWithTools — the reusable function-calling driver loop.
- Add TtsModelType.qwen3.
fix: namespace companion install files (tokenizers, TTS bundle aux) per model, fixing STT/embedding tokenizer collisions.
Add uninstallEmbedder/uninstallStt/uninstallTts (delete all files); route uninstallModel off the deprecated path.
package:flutter_gemma/genai.dart — use ChatMessage with InferenceChat (#181).sendMessage/generateContent (+ streams) covering text, vision, audio, thinking, and tool calls (#181).activeModelSpec/activeEmbedderSpec/activeSttSpec/activeTtsSpec + getModelPath.getStorageInfo/getOrphanedFiles/cleanupStorage/performCleanup — surface real errors, not success-shaped defaults.FlutterGemma.rag namespace (add/search/removeDocument/stats/clear); plumb removeDocument (#390 audit).uninstallEmbedder/uninstallStt/uninstallTts (delete all files); route uninstallModel off the deprecated path.ModelSource + FileSystemService from the barrel (required public parameter types).FlutterGemmaPlugin as the low-level SPI tier; FlutterGemma facade is the canonical entry.docs: document the on-device voice loop (VoiceSession, STT → LLM → TTS) shipped in flutter_gemma_speech 0.3.0.
feat: on-device text-to-speech — selectable TTS model (Matcha today) via installTts()/getActiveTts()/synthesize(), raw PCM output.
installTts()/getActiveTts()/synthesize(), raw PCM output.RuntimeConfig.artifactPaths.feat: on-device speech-to-text — new opt-in flutter_gemma_speech runs a selectable ASR model (moonshine today) on all native platforms via installStt(
flutter_gemma_speech runs a selectable ASR model (moonshine today) on all native platforms via installStt()/getActiveStt()/transcribe().$PWD-relative dir, not %LOCALAPPDATA% — model installed but not found at load.Docs: add community models to the README model list (SmolLM3-3B, Phi-4-mini-reasoning, Qwen2-VL, SmolVLM2, LLaVA-OneVision).
@litert-lm/core setup snippet to 0.14.0.Fix multi-GB background_downloader temp-file leak on Android (#383).
Add ModelFileType.builtIn for OS system models (Gemini Nano / Apple FM)
Cancel native .litertlm decoding before session teardown to avoid an ANR (#364, #373).
.litertlm decoding before session teardown to avoid an ANR (#364, #373).} (#366).chat_template (#367)._TypeError (#367).parameters:{type:OBJECT} for no-argument FunctionGemma tools (#367).required:[] inside FunctionGemma items, as the template does (#367).<end_function_call> leaking into FunctionGemma chat text and history (#366).properties/items, else require it (#367).None the way the template's Python does (#367).None argument as null instead of the string "None" (#366).ToolChoice.required on FunctionGemma (#367).result:{json} (#367).Fix AGP 9 build on android.builtInKotlin=false — apply KGP unless built-in Kotlin is on, so the plugin's Kotlin compiles (#360).
android.builtInKotlin=false — apply KGP unless built-in Kotlin is on, so the plugin's Kotlin compiles (#360).Fix SmartDownloader resume-loop hang — cap resume attempts + arm a watchdog timer (#355).
POST_NOTIFICATIONS at runtime (#356).FOREGROUND_SERVICE_DATA_SYNC + service merge (#356).resume()'s bool, arm reattach watchdog (#355 #356).Add skillExecutors: to FlutterGemma.initialize + the SkillExecutorProvider/SkillExecutorRegistry seam for the new opt-in flutter_gemma_agent package.
skillExecutors: to FlutterGemma.initialize + the SkillExecutorProvider/SkillExecutorRegistry seam for the new opt-in flutter_gemma_agent package.libOpenCL.so (+ -car/-pixel ICDs) in the plugin manifest — completes the #324 Mali GPU fix (#349; 1.1.1 only opened the SP-HAL door).<|tool_call> tokens the browser runtime leaves unconverted, and guard a crash when the SDK returns content as a String.Migrate Android module to Built-in Kotlin — apply KGP only on AGP < 9, ready for AGP 9+ (#323).
downloadUpdatesStream/fileSystemService injection to FlutterGemma.initialize (#342, thanks @hughesyadaddy).Declare libvndksupport.so for the Android GPU backend — fixes Mali GPU hard-freeze on Android 12+ (#324).
libvndksupport.so for the Android GPU backend — fixes Mali GPU hard-freeze on Android 12+ (#324).Deprecate enableHnsw (no-op; vector search now runs inside the store engine).
Filter support: FilterSchema/FilterField/FilterFieldType + configure(FilterSchema) on VectorStoreRepository.filterSchema: to FlutterGemma.initialize (threaded to the vector store at registration).enableHnsw (no-op; vector search now runs inside the store engine).Add clearActiveInferenceIdentity/clearActiveEmbeddingIdentity (non-breaking defaults on ModelFileManager).
clearActiveInferenceIdentity/clearActiveEmbeddingIdentity (non-breaking defaults on ModelFileManager).DownloadError/DownloadException from flutter_gemma.dart.Add maxOutputTokens on createSession/openSession/createChat/openChat to cap generation length (#318).
maxOutputTokens on createSession/openSession/createChat/openChat to cap generation length (#318).maxTokens dartdoc: it's the context window (input + output), not reply length.Fix empty assistant turn polluting chat history after stopGeneration() — cancelled/empty responses are no longer recorded (#325).
stopGeneration() — cancelled/empty responses are no longer recorded (#325).homepage to fluttergemma.dev. No code change.Stable 1.0 of the modular package split (see 1.0.0-rc.1 below for full notes).
dart:io is off the web/wasm import graph.core/domain/platform_types.dart.large_file_handler ^0.5.0.Modular package split: core flutter_gemma + opt-in flutter_gemma_litertlm / flutter_gemma_mediapipe / flutter_gemma_embeddings / flutter_gemma_rag_qdr
flutter_gemma + opt-in flutter_gemma_litertlm / flutter_gemma_mediapipe / flutter_gemma_embeddings / flutter_gemma_rag_qdrant / flutter_gemma_rag_sqlite.FlutterGemma.initialize(inferenceEngines:, embeddingBackends:, vectorStore:) — register the opt-in packages you added; core registers none by default.Nothing published for this version
Fix embedding freezing the UI thread (#299): forward pass runs on a background isolate.
/utf-8 to the MSVC plugin target.Android Qualcomm NPU (PreferredBackend.npu): auto-extracts QNN dispatch libs from APK at runtime (#293).
PreferredBackend.npu): auto-extracts QNN dispatch libs from APK at runtime (#293).wal_options now native in upstream.Concurrent sessions (#226): openSession()/openChat() run independent dialogues on one loaded model.
openSession()/openChat() run independent dialogues on one loaded model..litertlm inference via @litert-lm/core early preview (WebGPU/WASM, text-only).getActiveModel() after app restart (#227): mobile + web auto-restore from prefs.InferenceModel.activeBackend getter + NPU→GPU→CPU fallback on the FFI path with BackendInitException carrying per-attempt details.large_file_handler ^0.3.1 → ^0.4.0.LiteRT-LM v0.12.0 native bump (commit ffed38a): NPU dispatch now available on Linux/macOS as well as Windows.
ffed38a): NPU dispatch now available on Linux/macOS as well as Windows.libLiteRtLm.dylib with -Wl,-x so xcrun strip -x -S succeeds during release builds.%LOCALAPPDATA% env var is relative — falls back to USERPROFILE\AppData\Local then getApplicationSupportDirectory().Native LiteRT-LM prebuilts for flutter_gemma, built from LiteRT-LM 924e79c9 (v0.16.0) with LiteRT 0ff28117 .
Native LiteRT-LM prebuilts for flutter_gemma, built from LiteRT-LM 924e79c9 (v0.16.0) with LiteRT 0ff28117.
Consumed automatically by flutter_gemma_litertlm/hook/build.dart (Native Assets) at pub get time; SHA256-verified against the map baked into that hook. Previous release: native-v0.14.0 — there was no native-v0.15.0.
libStreamProxy resolves the shape at runtime, so both old and new hosts work.--define=litert_link_capi_so=true, a name upstream had deleted. Bazel accepts unknown defines silently, so the LiteRt runtime was being linked statically, which conflicts with the separately shipped WebGPU accelerator once Dawn became its own library. Corrected to litert_runtime_link_mode=dynamic + resolve_symbols_in_exec=false. That issue has been retracted.LiteRtDispatch.dll plus a version-matched OpenVino runtime (2026.3.0.dev20260622). The carried-forward pair shipped OpenVino 2026.2.0 against a runtime pinned to 2026.3.0, which is what broke backend=npu.libLiteRtDispatch_Qualcomm.so rebuilt from the derived LiteRT ref, and the ten QNN runtime libraries refreshed from the same QAIRT 2.44.0.260225. The stale pair failed with Qnn System library version 1.8.0 is mismatched. The minimum supported version is 1.11.0.libStreamProxy.dylib had been inheriting the build host's OS since native-v0.14.0 and shipped with minos 26.0; it is now built with -mmacosx-version-min=11.0.Engine initialized successfully, NPU engine_create in 498 ms.install_name_tool rewrite clean on every one (Native Assets re-runs it on each pub get); iOS minos 13.0, macOS 11.0; gpu_registry @executable_path patch present, basename dlopen absent.| Archive | Files |
|---|---|
litertlm-android_arm64.tar.gz |
19 — core + Qualcomm QNN NPU stack |
litertlm-ios_arm64.tar.gz |
4 |
litertlm-ios_sim_arm64.tar.gz |
4 |
litertlm-macos_arm64.tar.gz |
4 |
litertlm-linux_x86_64.tar.gz |
7 |
litertlm-linux_arm64.tar.gz |
7 |
litertlm-windows_x86_64.tar.gz |
28 — core + DXC runtime + Intel NPU stack |
SHA256 sums for every archive are in checksums_litertlm.txt.
@Deprecated, removal in 1.0.searchSimilar(... filter: Filter(must: [...], should: [...], mustNot: [...])). Honored on native; silently ignored on Web.FileSystemService; legacy Documents/ reads kept as fallback.example: add TranslateGemma 4B translation demo via task-first home navigation (#177).
Unified embedding on LiteRT C API + Dart FFI on all native platforms (#264).
Fix Android GPU sampler dlopen failure (#270, thanks @prithidevghosh): patchelf --add-needed libLiteRtLm.so on libLiteRtTopK{OpenCl,WebGpu}Sampler.so.
patchelf --add-needed libLiteRtLm.so on libLiteRtTopK{OpenCl,WebGpu}Sampler.so.Message.withImages([...]) for multi-image input, chat/Conversation session metrics via FFI, persistent prefix messages replayed on session rebuild after history truncation. Backward-compatible (Message.withImage(...) still works).temperature/topK/topP/seed are silently ignored when PreferredBackend.npu is selected — LiteRT-LM NPU executor only supports internal greedy sampling.LiteRtDispatch.dll), OpenVino runtime, TBB (~30 MB); FFI client passes dispatch_lib_dir + use_hw_masking_for_npu=false at engine_create so PreferredBackend.npu works on Intel LunarLake/PantherLake silicon without manual DLL placement.LiteRtLmFfiClient stub on web was missing the enableSpeculativeDecoding parameter — dart2js failed compilation when web target was actually built.build-litertlm-native-windows.yml workflow for Windows-only rebuilds.Your coding agent can read these notes before it upgrades. Set up the MCP server →