NewYour coding agent can read the release notes before it upgrades.Set up the MCP server →
pub.dev · #1505 most downloaded on pub.dev
Run Gemma and other LLMs on-device in Flutter (Android, iOS, Web, Desktop). Multimodal vision/audio, function calling, thinking mode, GPU, embeddings, RAG.
Last release 6 days ago
29 Sep 2026
Ships on a steady schedule
a new release about every 1 weeks
Nearly every release is documented
notes for 59 of the last 60 stable releases
Nothing withdrawn
no release was ever pulled
3 years old
141 releases · first in 2024
Fix desktop embedding on pub.dev installs (#250 follow-up): tensorflowlite_c.{dll,so,dylib} now bundled via Native Assets — regression from 0.14.0 set
tensorflowlite_c.{dll,so,dylib} now bundled via Native Assets — regression from 0.14.0 setup-script removal.Fix macOS dylib loading on pub.dev installs (#255).
fromAsset install on desktop (#250 mode 2).One column per month.
Fix App Store ITMS-90208 rejection on iOS (#245): downgraded patched libGemmaModelConstraintProvider.dylib minos 26.2 → 14.0 to match the other compan
libGemmaModelConstraintProvider.dylib minos 26.2 → 14.0 to match the other companion dylibs.libLiteRtLm.so with -Wl,-z,max-page-size=16384.npm run build, bundles cache_api.js + LiteRT WASM into dist/, ships package.json + vite.config.js in pub tarball.Fix macOS `install_name_tool` failure (#247): dart run / build_runner / flutter test on a pure-Dart library aborted with larger updated load commands
install_name_tool failure (#247): dart run / build_runner / flutter test on a pure-Dart library aborted with larger updated load commands do not fit because upstream Apple companion dylibs lack -headerpad_max_install_names. Skip them from Native Assets on macOS and bundle via Podfile post_install instead.<|"|> tokens from tool_calls.arguments were written to history raw, making the model echo them on later turns. Strip recursively before persist via new SdkResponseParser.cleanRawForHistory.modelFromAsset install on desktop (#250): the large_file_handler channel call threw MissingPluginException on macOS / Windows / Linux because the package only ships Android + iOS plugins. Now catches it and falls back to the in-memory loadAsset → writeFile path. Embedding (localagents-rag) and .litertlm FFI paths additionally fail fast with a typed message on non-arm64 Android (x86_64 emulator / armeabi-v7a) instead of a generic JVM crash — see new "Platform & Architecture Support" section in README.macos/Podfile post_install to the new snippet (see README → macOS Setup). iOS / Linux / Windows / Android unaffected.[*/perf] lines break out cost of dylib load, engine_create, prefill, decode — visible via debugPrint in both debug and release builds.Web build fix (#244): 0.14.0 broke web compilation by statically importing core/ffi/litert_lm_client.dart (which imports dart:ffi) through the flutter
core/ffi/litert_lm_client.dart (which imports dart:ffi) through the flutter_gemma_interface.dart → mobile/flutter_gemma_mobile.dart chain. JS/Wasm targets cannot compile dart:ffi. Added conditional imports in mobile/flutter_gemma_mobile.dart that swap litert_lm_client.dart and ffi_inference_model.dart for *_stub.dart shims on dart.library.js_interop. The web plugin (FlutterGemmaWeb) registers itself as FlutterGemmaPlugin.instance before any FFI code path runs, so the stubs' constructors (which throw UnsupportedError) are never actually invoked on web.lib*.dylib symlinks alongside the bundled .framework/ accelerators in Runner.app/Frameworks/ so LiteRT-LM's gpu_registry could dlopen them by basename — App Store Connect rejected those builds with "Unexpected file found in Frameworks". 0.14.1 patches the upstream LiteRT-LM source (runtime/components/sampler_factory.cc and litert/runtime/accelerators/gpu_registry.cc) to load Apple platform accelerators via @executable_path/../Frameworks/<X>.framework/<X> (macOS) / @executable_path/Frameworks/<X>.framework/<X> (iOS) instead of libX.dylib. Native Assets bundles the framework bundles correctly out of the box; no host-side Podfile symlinks needed. Patch applied during local Bazel rebuild — see native/litert_lm/patch_c_api.sh section 10.ModelType.gemma4 routes tool definitions to LiteRT-LM SDK via litert_lm_conversation_config_set_tools (OpenAI Chat Completions JSON). SDK applies chat_template.jinja through minja, renders native <|tool>declaration:...<tool|> tokens, and parses the model's <|tool_call>...<tool_call|> response back into structured tool_calls JSON. flutter_gemma reads the result via SdkResponseParser.extractToolCalls (handles parallel calls and the multimodal content[] path) and returns FunctionCallResponse to the app — no Dart-side prompt engineering needed.<|"|> Gemma 4 escape tokens from string arguments (recursively, including nested maps/lists).example/lib/models/model.dart: Gemma 4 E2B / E4B entries switched to modelType: ModelType.gemma4.Desktop FFI rewrite: macOS, Linux, Windows now run LiteRT-LM directly via dart:ffi against the C API. Removed Kotlin/JVM gRPC server, Azul Zulu JRE 24
dart:ffi against the C API. Removed Kotlin/JVM gRPC server, Azul Zulu JRE 24 download, and litertlm-server.jar bundling. Engine creation ~2 s (was ~10–15 s incl. JVM cold-start).litertlm models on iPhone (Gemma 3 1B, Gemma 3n E2B, Gemma 4 E2B). Multimodal vision + audio work on devicedxil.dll + dxcompiler.dll v1.9.2602) bundled in the Windows native archive — no manual install required.litertlm models on Android now go through the same Dart FFI path as desktop/iOS (was com.google.ai.edge.litertlm:litertlm-android AAR before). MediaPipe stays for .task/.bin modelsLiteRtLmFfiClient (lib/core/ffi/)stream_proxy_redirect_stderr exposes glog/abseil output on iOS/Android via temp file; helps diagnose engine init failureshook/build.dart from GitHub release native-v0.10.2; SHA256-verified, bundled via Native Assetspost_install block in your Podfile to symlink lib*.dylib next to the bundled .frameworks — gpu_registry calls dlopen by basename. See the macOS Setup section in README for the exact snippet (iOS works via the same pattern in example/ios/Podfile)ModelType.qwen3: New model type for Qwen3 models with thinking support
/no_think appended automatically when isThinking: false — faster TTFTcreateChat(maxFunctionBufferLength: 2048) for long function call argsFileSource now handles backslash paths correctlyVectorStoreRepository.removeDocument(id:) to delete documents from vector storeFix Qwen3 thinking mode (#224): Qwen3 tags now stripped automatically
<think> tags now stripped automaticallyFix iOS compile error (#222): XNNPack delegate type mismatch in EmbeddingModel.swift
EmbeddingModel.swiftTensorFlowLiteSelectTfOps — simulator builds work on Apple SiliconFix macOS SIGSEGV (#219): Per-conversation mutex in gRPC server prevents conversation.close() racing with sendMessageAsync on a native thread → use-af
conversation.close() racing with sendMessageAsync on a native thread → use-after-free in C++ fixedsetup_desktop.sh now downloads libLiteRtMetalAccelerator.dylib from GitHub Release so GPU inference uses the Metal delegate instead of falling back to static C APITensorFlowLiteSwift (source pod — cloned entire TensorFlow repo) with direct TensorFlowLiteC C API in EmbeddingModel.swiftFileSource absolute paths: Accept both Unix (/path) and Windows (C:\path) absolute paths in FileSource validation
/path) and Windows (C:\path) absolute paths in FileSource validationLiteRT-LM 0.10.0: Updated Android and JVM SDK from 0.9.0 to 0.10.0
isThinking: true now works with Gemma 4 E2B/E4B models (Android, iOS, Desktop; not Web)large_file_handler platform support: Conditional imports for pub.dev platform analysis compatibilityGemma 4 E2B/E4B: Added support for next-gen multimodal models (text + image + audio)
createChat() and createSession() for setting system-level context.litertlm models across platforms.litertlm models now work on iOS.litertlm modelspreferredBackend (Metal delegate now activated)addAudio + enableAudioModality)dart:io imports with conditional imports for WASM compilation supportexample/integration_test/benchmark_comparison_test.dart for comparing model performance on deviceToolChoice enum: auto / required / none parameter in createChat() to control tool calling behavior
auto / required / none parameter in createChat() to control tool calling behaviorParallelFunctionCallResponse for multiple function calls in one responseFunctionCallFormat implementations (Gemma, Qwen, DeepSeek, Llama, Phi, FunctionGemma)<tool_call> Format: Qwen/Mistral-style function call parsing<|tool_calls|> format supportnativeLibraryDir to LiteRT-LM Backend.NPU()Dual-Prefix Embeddings (TaskType): Improved RAG retrieval quality with query/document prefixes
TaskType.retrievalQuery (default) — for search queriesTaskType.retrievalDocument — for document indexingEmbedData.TaskType)addDocument() automatically uses document prefix.tflite embedding models (EmbeddingGemma, Gecko) on macOS, Windows, Linux
dart:ffi — no gRPC, no JVM overheaddart_sentencepiece_tokenizer (BPE + Unigram, auto-detect format)sqlite3 dart:ffi replacing platform-specific codepatrol dependency, migrated all integration tests to standard integration_testLiteRT-LM 0.9.0-beta: Updated from 0.9.0-alpha02 on Android and Desktop (JVM)
Conversation.cancelProcess()CancelGeneration RPCLlmInference.cancelProcessing() (MediaPipe 0.10.26)Removed deprecated package attribute from AndroidManifest.xml
model.type field in tokenizer.jsoniosPath parameter for platform-aware tokenizer downloads
litertlm-server.jar URL to v0.12.5 (#189)package attribute from AndroidManifest.xmlAndroid ProGuard Fix: Added ProGuard rules for LiteRT-LM classes
Android LiteRT-LM Engine: Added LiteRT-LM inference engine for Android
.litertlm → LiteRT-LM, .task/.bin → MediaPipe)supportAudio parameter in session configurationModel Deletion Fix: Fixed model deletion not removing metadata
Web Large Model Support: WebStorageMode for models >2GB via OPFS streaming
WebStorageMode for models >2GB via OPFS streaming (#162)🖥️ Desktop Support: Full support for macOS, Windows, and Linux platforms
.litertlm model format only (MediaPipe .task/.bin not supported on desktop)🐛 iOS Embeddings Fix: Fix crash on repeated embedding inference
🤖 FunctionGemma Single-Turn Mode: FunctionGemma now operates in single-turn mode by design (clears history after each response)
🤖 FunctionGemma Support: Added ModelType.functionGemma for Google's specialized function calling model
ModelType.functionGemma for Google's specialized function calling model
✅ iOS Embeddings Fix: XNNPACK + SentencePiece integration for better results on iOS
@0.11.13/web/*.js)🌐 Web VectorStore: Full RAG support on web with SQLite WASM
tasks-vision-image-generator to 0.10.26.1 for Android 15+ compatibility🐛 Mobile Build Fix: Fixed compilation errors on iOS/Android platforms
⚠️ BREAKING CHANGE: Explicit initialization now required
await FlutterGemma.initialize() in main() before using the plugin🌐 Web Embedding Support: Added support for embedding generation on web platform
🤖 CI/CD Automation: Added GitHub Actions workflows for automated testing and release builds
🚀 VectorStore Optimization ⚠️ BREAKING (RAG only):
🐛 iOS Simulator Fix: Fixed "Filename cannot contain path separators" crash on iOS Simulator
canResume(), resume(), cancel() methods from DownloadService interface⚠️ Deprecated: Marked legacy asset/file management methods as deprecated with migration hints
New: Fluent builder API with FlutterGemma.installModel().fromNetwork/fromAsset/fromBundled/fromFile()
FlutterGemma.installModel().fromNetwork/fromAsset/fromBundled/fromFile()modelManager.downloadModelWithProgress()) still works as facade🌐 Web Multimodal Support: Added full multimodal image processing support for web platform
.litertlm model files optimized for web platform🛡️ Fixed: Updated ProGuard rules for Android release build compatibility
🐛 Fixed: Export missing ModelType and other public API types to resolve import issues
🚀 Embedding Models Support: Added full support for text embedding models
InferenceModelSpec and EmbeddingModelSpec for better model organizationreplace and keep🔧 Model Replace Policy: Added configurable model replacement system with keep/replace options and ensureModelReady() method
ensureModelReady() methodHuggingFaceDownloader wrapper to handle CDN server inconsistencies and resume failures.task files (MediaPipe-handled) and .bin/.tflite files (manual formatting)🛑 Stop Generation: Added Android support for stopping text generation mid-process with session.cancelRequestGenerationAsync() (#89, #19, #34)
session.cancelRequestGenerationAsync() (#89, #19, #34)📚 Documentation: Updated README with comprehensive model information and usage examples
📥 Background Downloads: Added background download support for model files
🚀 New Models: Added support for 4 new compact models:
🧠 Thinking Mode: Added thinking mode support for DeepSeek models with persistent thinking bubbles
✨ Function Calling: Added support for function calling, allowing models to interact with external tools.
🖼️ MULTIMODAL SUPPORT: Added full support for text + image input with Gemma 3 Nano vision models
Message class with support for text, image, and multimodal content
Message.text() - for text-only messagesMessage.withImage() - for text + image messagesMessage.imageOnly() - for image-only messagesmessage.hasImage - to check if message contains image🚀 GEMMA 3 NANO SUPPORT: Added full support for Gemma 3 Nano models
input_pos != nullptr errorsUpgraded Mediapipe to 0.10.24 for iOS and Android
- iOS LoRA support added - iOS topP support added
- Readme updated
- Readme updated
Add web platform support in pubspec.yaml
Upgraded Mediapipe to 0.10.22 for Android and Web
Added Chat functionality for instruction tuned model
Added opportunity to manage inference session
Fixed crash on generation for Android
IMPORTANT: Breaking changes in the API
- Added close method
- Small fixes for Android
Your coding agent can read these notes before it upgrades. Set up the MCP server →