NewYour coding agent can read the release notes before it upgrades.Set up the MCP server →
pub.dev · #1686 most downloaded on pub.dev
Moved to flutter_edge_ai_litertlm. This is the final release under the name flutter_gemma_litertlm; new versions ship as flutter_edge_ai_litertlm, part of Flutter Edge AI.
Last release 3 days ago
05 Oct 2026
Ships fairly regularly
a new release about every 1 weeks
Nearly every release is documented
notes for 30 of 30 stable releases
Nothing withdrawn
no release was ever pulled
4 months old
31 releases · first in 2026
One column per month.
Moved to flutter_edge_ai_litertlm; this is the final release under the name flutter_gemma_litertlm.
flutter_edge_ai_litertlm; this is the final release under the name flutter_gemma_litertlm.Web setup pins @litert-lm/core 0.17.1.
@litert-lm/core 0.17.1.Web reports the accelerator embeddings ran on, and whether it was fully accelerated.
activeBackend no longer claims NPU on macOS, Linux, iOS or non-Qualcomm Android.fix(litertlm): don't clamp maxTokens on NPU, and say which models the NPU runs by @DenisovAV in #508
Full Changelog: v1.8.2...v1.8.3
activationDataType to the engine; float32 fixes wrong digits on some GPUs.docs(website): full content audit + Built-in AI page + Engines nav group by @DenisovAV in #455
Full Changelog: v1.6.5...v1.8.2
A chat stopped mid-reply answered every later message with nothing (#325).
Google Play no longer rejects apps over 16 KB page sizes (#529).
@litertjs/core 2.5.3.flutter_gemma_embeddings; asks core for a tokenizer, so register embeddingTokenizers:.Native runtime native-v0.17.0-a: tool calls no longer crash the app.
native-v0.17.0-a: tool calls no longer crash the app..litertlm takes the runtime's tool path — needs core 1.8.4.Native runtime: LiteRT-LM v0.17.0 (native-v0.17.0).
native-v0.17.0).CreateTensorBufferFromHostMemory status 3.fix(core): stop claiming background_downloader's updates stream from the host by @DenisovAV in #450
Full Changelog: v1.6.2...v1.6.4
Web: move @litert-lm/core 0.14.0 -> 0.17.0.
@litert-lm/core 0.14.0 -> 0.17.0.maxOutputTokens now caps generation instead of being ignored.docs(website): vision/audio per-encoder backend (1.5.9 / litertlm 1.4.2) by @DenisovAV in #430
Full Changelog: v1.5.9...v1.6.2
Docs: the README and library docs show the engine-carried resolver and the one-call fromHuggingFace(repo) install.
fromHuggingFace(repo) install.Add LitertlmManifestResolver: manifest-driven Hugging Face install for repos shipping litertlm_manifest.json, auto-registered from LiteRtLmEngine (#45
LitertlmManifestResolver: manifest-driven Hugging Face install for repos shipping litertlm_manifest.json, auto-registered from LiteRtLmEngine (#454).Fix a rebuilt native library not reaching the build.
Android: embeddings no longer poison the loader, fixing zero-chunk streams and SIGABRT (#447).
Share the host-side native library lookup with core instead of a private copy.
Add LiteRT embedding forward-pass + LiteRtEmbeddingBackend (moved from flutter_gemma_embeddings).
LiteRtEmbeddingBackend (moved from flutter_gemma_embeddings).tokenizerFactory descriptor field; byte-identical vectors, no API change.fix: vision + audio encoders default to CPU (fixes GPU vision hard-fail on Metal/WebGPU); overridable.
Fix Windows NPU: the OpenVino compiler DLLs shipped but were never bundled into the app.
Migrate to LiteRT-LM v0.16.0 (native-v0.16.0) — fixes the Android OpenCL per-turn leak (#348, #402).
libStreamProxy.dylib to 11.0 instead of the build host's.@litert-lm/core 0.14.0 (the 0.16.0 npm publish ships no dist/).Clearer engine-create error for GPU-only .litertlm models run on CPU (#390).
.litertlm models run on CPU (#390).Also expose the LiteRt interpreter (LiteRtBindings) for embeddings/speech; own web litert.js.
LiteRtBindings) for embeddings/speech; own web litert.js.Migrate FFI to LiteRT-LM v0.14.0 — native per-session sampler (opaque session-config); native-v0.14.0.
@litert-lm/core 0.12.1 → 0.14.0 (text path; API-compatible).Smooth UI during Android GPU prefill — flush the OpenCL queue every 2 ops (#364).
Guard native cancel against a freed conversation — fixes a use-after-free SIGSEGV on close-mid-stream (#379).
Create the native conversation off the main isolate to avoid ANRs on multimodal models (#365).
Clamp maxTokens up to 1024 (min context for .litertlm) to fix the DYNAMIC_UPDATE_SLICE crash (#318).
maxTokens up to 1024 (min context for .litertlm) to fix the DYNAMIC_UPDATE_SLICE crash (#318).maxOutputTokens (session + chat) via native set_max_output_tokens; skipped on NPU.Fix PreferredBackend.npu on Android (Qualcomm) + Windows (Intel): native-v0.13.1-a restores the NPU dispatch libs omitted from 1.0.0 (#155).
PreferredBackend.npu on Android (Qualcomm) + Windows (Intel): native-v0.13.1-a restores the NPU dispatch libs omitted from 1.0.0 (#155).homepage to fluttergemma.dev. No code change.Stable 1.0.0; spec imports redirected off the dart:io mobile lib for a wasm-clean web graph.
dart:io mobile lib for a wasm-clean web graph.Initial release: LiteRT-LM (.litertlm) on-device inference engine for flutter_gemma via dart:ffi.
.litertlm) on-device inference engine for flutter_gemma via dart:ffi.LiteRtLmEngine (InferenceEngineProvider). Owns the shared LiteRT-LM native library.@litert-lm/core, early preview).Your coding agent can read these notes before it upgrades. Set up the MCP server →