NewYour coding agent can read the release notes before it upgrades.Set up the MCP server →
pub.dev · #3059 most downloaded on pub.dev
LiteRT (formerly TensorFlow Lite) Flutter plugin. Drop-in on-device ML inference with bundled native libraries for supported native platforms and web runtimes.
Last release 10 days ago
28 Sep 2026
Release timing varies
gaps range from 8 days to 2 months
Nearly every release is documented
notes for 60 of the last 60 stable releases
Nothing withdrawn
no release was ever pulled
7 months old
88 releases · first in 2026
One column per month.
Documentation and example fixes only. No change to the package's code or behavior.
Documentation and example fixes only. No change to the package's code or behavior.
CompiledModel with {Accelerator.npu, Accelerator.gpu, Accelerator.cpu},
which throws on Apple platforms. It now suggests {npu, cpu} and notes that
Apple cannot combine NPU and GPU.verifyCompiledModel compares, its 1% default tolerance, and its limits. The
migration guide already linked to it, but the section did not exist.flutter_litert_flex; it
depends on it for its model-matrix tests.CompiledModel ignores precision on the web and the
setting had no effect. The API parameter is unchanged.prepareCameraFrameFromImage now decides BGRA vs. RGBA from the frame's format.raw instead of the platform, so desktop streams decode correctly on came
prepareCameraFrameFromImage now decides BGRA vs. RGBA from the frame's
format.raw instead of the platform, so desktop streams decode correctly on
camera_desktop 2.0.0, which delivers BGRA on every desktop platform, as well
as on 1.x, which delivered RGBA on Linux and Windows. Before this, Linux and
Windows frames from camera_desktop 2.0.0 were decoded as RGBA, swapping red
and blue. 'BGRA' selects BGRA and 'RGBA' selects RGBA. Any other value,
or no format at all, keeps the previous default (BGRA on macOS, RGBA
elsewhere), and an explicit isBgra still takes precedence, so callers that
pass isBgra: Platform.isMacOS should drop it to get the new behavior.camera_desktop: ^2.0.0.Android's CompiledModel runtime is updated from LiteRT Next 2.1.6 to 2.2.0. The 2.2.0 AAR ships the same libLiteRt.so and libLiteRtClGlAccelerator.so
libLiteRt.so and
libLiteRtClGlAccelerator.so set per ABI, every C API symbol the package
binds is still exported, and the headers it relies on change only
additively. The Qualcomm JIT runtime keeps the same library names and QAIRT
2.47.0.260601 pin, so a flutterLitert.qualcommNpuRuntimeDir bundle must
now come from the LiteRT 2.2.0 JIT runtime release. The Interpreter stays on
LiteRT 1.4.2, and iOS, macOS, Windows, and Linux stay on LiteRT Next 2.1.5.LiteRtInterpreter and CompiledModel) is updated from
LiteRT.js 2.4.0 to 2.5.3 (@litertjs/core@2.5.3, auto-loaded from
jsDelivr). Since 2.5.0, a WebGPU compile that cannot place every op no
longer throws. Chromium browsers with JSPI return a partially delegated
WebGPU model, which CompiledModel reports as
{Accelerator.gpu, Accelerator.cpu} with isFullyAccelerated false.
Browsers without JSPI recompile the model on WASM inside LiteRT.js. Requests
that allow a CPU fallback (auto, gpuWithCpuFallback, {gpu, cpu})
accept that model and report {Accelerator.cpu} with didFallback true.
A strict CompiledModelConfig.gpu() or {Accelerator.gpu} request now
disposes it and throws StateError, preserving the 2.4.0 contract that a
strict GPU request never silently runs on the CPU.Correct the native-only inference utility documentation (issue #17). InterpreterFactory and InterpreterPool examples now import package:flutter_litert
InterpreterFactory and InterpreterPool examples now import
package:flutter_litert/native.dart so their native types resolve consistently
in IDEs and static analysis. Document their supported platforms and the
portable CompiledModelConfig.auto() alternative.dartcv4 2.3.1 removes its hardcoded iOS 12 target and adds a deployment-target
hook option. Document the upgrade, iOS 15 configuration, and Flutter SDK
meta pin compatibility requirement. Upgrade the example to the fixed
dependency with Flutter 3.47.5 and an iOS 15 Runner and hook target.
A clean Xcode 27.0 / iOS 27 simulator build now passes the DartCV CMake
probe and final Runner link. The iOS package also drops the legacy -ObjC
linker option that Xcode 27 rejects; explicit C API symbol anchors retain the
packaged FFI entry points. The bundled SwiftPM TensorFlow Lite simulator
slices are arm64-only, so the example excludes x86_64 when building for the
simulator. The package's minimum iOS version remains 13.0; the example's iOS
15 DartCV target and the macOS Core ML build are independent constraints.CompiledModel construction. CompiledModelConfig and
CompiledModelPolicy add auto, cpu, strict gpu, mixed
gpuWithCpuFallback, strict npu, and mixed npuWithCpuFallback choices.
Use fromFileWithConfig, fromBufferWithConfig, or the portable
fromBufferWithConfigAsync. This provides the PerformanceConfig.auto
equivalent requested in
issue #18
without changing any existing constructor.{gpu, cpu} graph, select NPU, benchmark available
accelerators, or treat successful compilation as proof of numerical
correctness. Production models still need per-model validation.gpu/npu perform no package-level CPU retry;
gpuWithCpuFallback/npuWithCpuFallback first request a mixed graph that can
place operations on CPU, then perform a complete CPU-only retry if that
construction throws. The existing fromBufferWithGpuFallback methods retain
their mixed-graph-first behavior and map to the new explicit policy.requestedConfig, requestedAccelerators, and didFallback getters sit
alongside the existing effective accelerators and isFullyAccelerated
diagnostics. They distinguish what the caller requested, what LiteRT
constructed, whether a complete retry or runtime narrowing occurred, and
whether the whole graph was accelerated. Operation-level placement inside a
successful mixed graph does not set didFallback.fromBufferWithConfigAsync.
Auto uses WebGPU with the existing hung-compile watchdog and retries on WASM;
strict NPU remains unsupported, while npuWithCpuFallback reports the NPU
failure and continues on WASM. Unsupported-platform stubs carry matching
signatures for conditional-import parity.`CompiledModel` now defaults to `Precision.fp32` instead of `fp16`. This changes numeric output and costs about 30% median GPU latency across the five
CompiledModel now defaults to Precision.fp32 instead of fp16. This
changes numeric output and costs about 30% median GPU latency across the five
GPUs measured (four architectures; Apple Metal appears as both M4 and
A17 Pro), so it is a deliberate correctness-over-speed default. The cost
is real and worth stating plainly: in 84 paired same-model comparisons fp32
was slower in 67 of them, with a median of +29.9% and a worst case of
+21.6 ms. Apple M4 is the lone exception, where fp32 is marginally faster
(median -6.5%); every other architecture pays +37% to +43%. Restricting the
comparison to the models where fp16 actually passed parity, which are the only
ones anyone could legitimately keep on fp16, the median cost is +24.3%.
Accuracy is what justifies the default anyway. Across the 29 published detection
models, strict-GPU fp32 matched a plain-CPU reference for every model that
compiled on all four GPU architectures measured, while fp16 matched only 4 of
18 on Adreno 740, 5 of 18 on Xclipse, 4 of 18 on Apple Metal, and 1 of 12 on
Mali-G715. Google's own LiteRT Python API reproduces the Apple figures
through the same underlying switch, so this is upstream numerical behaviour
rather than a binding artefact: fp16 carries about three decimal digits of
mantissa and these graphs emit pixel-space coordinates and landmark
positions. fromBufferWithGpuFallback already defaulted to fp32; the plain
constructors now agree with it. Pass Precision.fp16 explicitly to keep the
old behaviour, ideally per model and validated on the target GPU. Full
results: GPU vendor matrix.test/benchmark gains an
Apple (macOS + iOS) matrix harness, physical Android matrices for Adreno,
Mali, and Xclipse, and a cross-check against the official LiteRT Python API
that agrees with the Dart implementation to 6.3e-14 on CPU reference outputs.Package.swift
pinned TensorFlowLiteCCoreML to a May release predating both the NPU entry
points and the global-MEAN padding patch, so accelerator registration
returned kLiteRtStatusErrorUnsupported for every model and the Interpreter
Core ML delegate rejected four models macOS accepted. The patched framework
had been built but only ever uploaded as a CI artifact, never published, so
the pin was never moved. coreml-ios-v1.1.0 publishes it for SwiftPM.
CocoaPods needed a separate fix: its frameworks are downloaded by the podspec
rather than shipped in the package, and that bundle (libs-v0.1.8) predated
the NPU work entirely, so fixing SwiftPM alone left every CocoaPods consumer,
which is Flutter's default iOS path, with the identical runtime failure.
libs-v0.1.9 carries the patched framework, and the podspec now gates the
download on the NPU symbol actually being present rather than on the file
merely existing, so a stale cache re-downloads instead of silently degrading.
Both channels are now checked in CI against the artifacts consumers really
fetch. Measured on a physical iPhone 15
Pro afterwards, iOS matches macOS exactly: Core ML 24 of 29 models with 13
accurate, strict {npu} 1 of 29, {npu, cpu} 24 of 29 with 12 accurate.{npu, cpu} request no longer fails when Core ML NPU cannot
register. The graceful-degrade path added for Android was gated on
Platform.isAndroid, and macOS never exposed the gap because registration
succeeds there. A caller asking for CPU fallback got a hard failure instead,
which is the one outcome an explicit fallback exists to prevent. Strict
{npu} still throws; a mixed request drops the NPU and reports the effective
set through the accelerators getter.CoreMlDelegateOptions exposes
no precision control. Every model Core ML computed incorrectly across the 29
published models was correct on an fp32 path, and most are the same models
that fail on GPU fp16. Unlike GPU, where Precision.fp32 fixes it, NPU has
no equivalent: validate per model or do not use it.LiteRtLockTensorBuffer fails on managed buffers. Controlled runs place
it on compilation rather than inference, so an application that compiles its
models once and runs them for hours is unaffected, while one that repeatedly
constructs and disposes CompiledModel instances is not. A Galaxy A35 with
Mali-G68 and a Galaxy A56 with Xclipse showed no such failures, so it is not a
Mali-family property.CompiledModel. Android API 31+
arm64 apps can now use an app-provided LiteRT JIT NPU runtime. NPU requests
get a dedicated environment with the official compiler-plugin and dispatch
directories plus a reusable JIT cache; initializing CPU/GPU first no
longer prevents later NPU setup.flutterLitert.qualcommNpuRuntimeDir Gradle property can fuse exactly one
prepared Qualcomm runtime into an arm64 local/Test Lab APK and fails the
build when its nine-library JIT set is incomplete or ambiguous. The plugin
manifest also exposes the device-provided Qualcomm libcdsprpc.so library
through Android's optional native-library allowlist.Android physical NPU (Firebase Test Lab) workflow targets the
Galaxy S23/SM8550 (HTP v73), S24 Ultra/SM8650 (v75), or S25 Ultra/SM8750
(v79). Its default strict smoke gate consumes one Test Lab run only after
runtime preparation, AAB build, and package validation pass; an opt-in full
sweep adds face, segmentation, and pose correctness comparisons. Physical
validation passed strict NPU inference and the full representative sweep on
all three generations. MobileFaceNet passed the default CPU-reference
tolerance, while selfie segmentation and heavy pose were consistently
identified as model-specific accuracy risks.Un-deprecated the GPU, Metal, and CoreML Interpreter delegates. 3.0.0 deprecated them in favour of CompiledModel and announced removal in 4.0.0; that…
Fixes a 3x macOS CPU slowdown, adds a way to detect an upstream LiteRT defect that returns wrong answers silently, and fixes two resource bugs. Additive: no existing symbol changes signature.
Heads-up for bit-exact tests. The bundled macOS arm64 libtensorflowlite_c
is now a bazel build rather than a CMake one, which changes float32 output in the
last few ULPs because ruy multithreading is finally active and reductions
accumulate in a different order. Measured on the bundled face-detection model:
87% of elements differ, by at most 3.8e-05 against an output range of 181.7, i.e.
0.000021%. Tolerance-based comparisons are unaffected; a byte-level golden pinned
on macOS arm64 will need regenerating.
TRANSPOSE_CONV has no kMultithreadOptimized variant and parallelises only
through ruy's gemm, so deconv-heavy models paid the full cost: a 384px landmark
model went from 83.7ms to 26.8ms, now matching iOS exactly. fully_connected
and batch_matmul gain similarly. Intel Macs keep their existing CMake slice
and are unchanged. See doc/macos_transpose_conv_gap.md.CompiledModel.
Accelerator.npu now lazily registers a dedicated Core ML
CPUAndNeuralEngine accelerator on macOS 13+. Strict {npu} compilation
rejects any non-delegated TFLite operation; {npu, cpu} applies Core ML
before XNNPACK and rejects zero-node Core ML delegation rather than silently
returning CPU-only inference. The build carries the required global-MEAN
padding fix and is covered by fixed-input output comparisons across
representative models. NPU+GPU combinations remain unsupported. See
doc/macos_compiled_model_npu.md.CompiledModel path now has its own Core ML accelerator-registration bridge,
uses the same strict {npu} and Core-ML-first {npu, cpu} semantics, and
rejects zero-node delegation. The arm64+x86_64 simulator suite passes strict
inference, a five-model mixed-mode correctness sweep, fallback diagnostics,
and NPU+GPU rejection. This does not yet constitute Neural Engine validation:
simulators have no ANE, physical-iPhone testing remains pending, and SwiftPM
still needs a release artifact containing the patched Core ML entry points.
See doc/ios_compiled_model_npu.md.verifyCompiledModel(bytes, compiled) checks a CompiledModel against
a bare-CPU Interpreter and reports the deviation, returning
BackendVerification. LiteRT Next can return kLiteRtStatusOk while producing
output that is wrong, or never written at all, and neither is visible from a
status code or from timing. Run it once at init before trusting a
CompiledModel. It reports rather than throwing or swapping backends, so the
policy stays with the caller; default tolerance is 1% of the output range,
against measured separation of 0.068% (healthy) versus 42%+ (corrupt). Cost is
one Interpreter build plus one inference, 4-56ms depending on the model.CompiledModel.isFullyAccelerated reports whether the whole graph ran
on a selected accelerator. Note that false is ambiguous: partially delegated
graphs report false even when the accelerator genuinely ran, so this is not a
way to detect a silent CPU fallback. Use verifyCompiledModel for that.InterpreterPool.initialize is now all-or-nothing. A failure part
way through left the interpreters it had already built alive, and because the
dispose-first branch is keyed on isInitialized, which a failed call never
sets, retrying accumulated them: a pool of 3 could end up holding 4, the extra
one live with an XNNPACK threadpool but never used.CoreMlDelegate leaked its options struct when constructed without
explicit options. Caller-supplied options are still left to the caller.LiteRtStatus values in error messages now carry their name, so
LiteRtStatus=3 reads LiteRtStatus=3 (kLiteRtStatusErrorRuntimeFailure).CompiledModel and announced removal in 4.0.0;
that is reversed, and no removal is scheduled. Two reasons. PerformanceConfig.gpu()
and .coreml() are built on these classes and were never deprecated, so the
removal would have broken supported API with no notice (interpreter_factory.dart
was suppressing its own deprecation warning to keep compiling). And
CompiledModel cannot replace them yet: it reports success while leaving the
output buffer unwritten for models whose output tensor ends up dynamic, which
covers heatmap models with a deconvolution head. A deprecation that cannot be
acted on, pointing at a backend that returns wrong numbers, is worse than none.
Prefer PerformanceConfig over constructing delegates directly, and gate any
CompiledModel adoption behind verifyCompiledModel.Adds shared utilities that detector packages were each re-deriving locally. All additive; no existing symbol changes behaviour.
Adds shared utilities that detector packages were each re-deriving locally. All additive; no existing symbol changes behaviour.
aggregateActiveAccelerator(Iterable<String?>) (web) collapses the
per-runner backends of a multi-stage detector into the single accelerator it
should report. It returns 'webgpu' when any runner is still on WebGPU, so
the runtime GPU-error fallback and slow-WebGPU warmup (both gated on the
reported accelerator) stay armed under mixed compile outcomes where some
models fell back to WASM and others did not.compiledModelFromBufferAuto(...) and isDefaultGpuCpuAccelerators(...)
centralize the "is this the permissive {gpu, cpu} default?" branch that
decides between CompiledModel.fromBufferWithGpuFallback and
CompiledModel.fromBuffer. An explicit accelerator set is still honoured
as-is; only the two-way default degrades.iouLTRB(...) is the exact intersection-over-union of two axis-aligned
boxes, for frame-to-frame track matching. It deliberately has no epsilon,
unlike the NMS ratio in nms_utils.dart which adds 1e-7; mixing the two
shifts matches at threshold boundaries.CompiledModel.fromBufferWithGpuFallback now forwards precision to
its CPU paths. Previously only the GPU attempt received it, so the forceCpu
shortcut and the CPU retry after a failed GPU compile silently fell back to
fromBuffer's fp16 default. A single call with no arguments therefore ran
fp32 on GPU and fp16 on CPU, defeating the fp32 default that exists because
pixel-space landmark and box coordinates lose accuracy in fp16. Callers that
passed fp16, including every detector package built on this plugin, are
unaffected; callers that asked for fp32 now get it on the fallback path.
fromBufferWithGpuFallbackAsync delegates and is fixed with it. The web
implementation documents precision as accepted-but-ignored and is unchanged.collectOutputShapes(Interpreter) (native) returns every output tensor's
shape keyed by index, walking indices until getOutputTensor throws. It reads
shapes only and never touches Tensor.data, so no buffer views are
materialized and quantized outputs are safe to enumerate. Use
TensorFloat32Views when the buffers themselves are needed.Adds explicit support for detection models whose confidence tensors are already activated probabilities.
Adds explicit support for detection models whose confidence tensors are already activated probabilities.
postProcessDetections and postProcessDetectionsFlat now accept
scoresAreProbabilities: true, which skips sigmoid for class and
objectness values and compares probability thresholds directly.false, preserving the existing
logits contract and output for all current callers.Adds Android OpenCL/GL acceleration to the LiteRT Next CompiledModel path.
Adds Android OpenCL/GL acceleration to the LiteRT Next CompiledModel path.
libLiteRtClGlAccelerator.so from the pinned
LiteRT 2.1.5 AAR by default for arm64-v8a and x86_64.
armeabi-v7a remains CPU-only.libOpenCL.so, libOpenCL-car.so, libOpenCL-pixel.so, and
libvndksupport.so, all required="false"), so apps targeting
Android 12+ can load them without adding their own
uses-native-library entries.{gpu, cpu}
compilation can fail after the accelerator registers. The
fromBufferWithGpuFallback factories catch that error and retry CPU-only.flutterLitert.bundleGpuAccelerator=false to omit about 3 MB per ABI. The
classic Interpreter runtime and GPU delegate are unchanged.Web CompiledModel robustness fix. No API changes.
Web CompiledModel robustness fix. No API changes.
fromBufferWithGpuFallbackAsync and {gpu, cpu}
accelerator sets now bound the WebGPU attempt with a 60-second watchdog
and fall back to WASM when it trips, honoring their always-yield-a-model
contract. If the abandoned compile settles later, its model is disposed.
Strict {gpu} requests are never timed out and keep surfacing whatever
the runtime does.integration_response_data.json, the custom driver writes that file
on failure too, and CI prints it when a drive fails, so a recurrence
pinpoints the stalled stage instead of reporting an empty failure detail.Brings CompiledModel to the web via Google's LiteRT.js (the same auto-loaded @litertjs/core that powers LiteRtInterpreter), fixes App Store uploads fo
Brings CompiledModel to the web via Google's LiteRT.js (the same
auto-loaded @litertjs/core that powers LiteRtInterpreter), fixes App
Store uploads for SwiftPM installs (#15), and fixes a nondeterministic ARM64
detection decode. Additive and backward compatible.
Web CompiledModel:
CompiledModel.fromBufferAsync and
fromBufferWithGpuFallbackAsync; pair them with the existing runAsync
for portable code. LiteRT.js compilation is Promise-based, so on the web
they are the only way to build a model: the synchronous fromFile,
fromBuffer, fromBufferWithGpuFallback, and run throw
UnsupportedError there.cpu compiles on WASM, gpu on WebGPU, and
{gpu, cpu} tries WebGPU with a WASM fallback; model.accelerators
reports what LiteRT.js actually resolved (including {gpu, cpu} for
partially accelerated WebGPU models). npu throws ArgumentError on the
web, precision is accepted but ignored, and the zero-copy
TensorBufferMode.hostMemory path stays native-only.LiteRtRuntimeError, so callers can dispose the model and rebuild it with
{Accelerator.cpu}.Accelerator/Precision/TensorBufferMode enums moved to a shared
source file (no API change), and the example app now builds its
CompiledModel with fromBufferAsync.Web backend selection and Safari compatibility:
wasm/ directory
instead of a pinned file, so LiteRT.js's feature probe serves Safari the
compat build (relaxed SIMD is default-off there) while Chrome and Firefox
keep the fast relaxed-SIMD build. URLs pinned via
configureLiteRtWebLoader are unaffected.resolveWebAccelerator('auto' | 'webgpu' | 'wasm') in
web_detector_utils.dart: 'auto' picks WebGPU only on Chromium with a
hardware (non-software) adapter, probed once per page load; explicit
values pass through. Firefox's WebGPU works but runs ~22x slower than its
WASM SIMD, so API presence alone must not select it.WebGpuFallback.maybeSwapIfWebGpuSlow: times a few warmup inferences
after an 'auto' init that landed on WebGPU and swaps to WASM past a
budget (default 50ms median), catching slow-but-functional GPU stacks the
error-driven fallback cannot see.WebGpuFallback.withFallback now swaps only on LiteRtRuntimeError, so
logic bugs surface instead of masquerading as GPU fallbacks, and marks
fellBackToWasm only after a successful swap. All compile-time, runtime,
and warmup fallbacks now log their cause via debugPrint.iOS fix (#15): App Store validation rejects the loose libLiteRt.dylib /
libLiteRtMetalAccelerator.dylib files that SwiftPM's bare-dylib
xcframeworks embedded in the app's Frameworks/ directory, surfacing as
ITMS-90426 ("Invalid Swift Support"). SwiftPM now ships the same
framework-wrapped xcframeworks as CocoaPods (identical binaries, release
litert-ios-v1.0.1) and registers the Metal accelerator through the shared
LiteRtRegisterGpuAccelerator shim, so GPU CompiledModel keeps working.
No API change; run flutter clean and rebuild. Note: Flutter's SwiftPM
support independently embeds a framework built at minos iOS 12.0 that can
also trigger ITMS-90426; if uploads still fail, disable SwiftPM
(flutter: config: enable-swift-package-manager: false in pubspec.yaml)
until flutter_tools is fixed.
ARM64 fix: on Apple Silicon, the SIMD decode in postProcessDetectionsFlat
could return a different detection count (or phantom boxes) for
byte-identical model output, because the Dart ARM64 JIT miscompiles the
greaterThan().select() lane-carried argmax it used. The winning class is
now recovered with a scalar argmax over the few anchors that clear the
threshold, so the decode is deterministic and matches the scalar reference.
Affects every downstream detector that decodes channel-major YOLO output; no
API change.
Fixes the two hero demo images stacking vertically on the pub.dev package page. pub.dev's README stylesheet forces img{height:auto}, so they are now s
Fixes the two hero demo images stacking vertically on the pub.dev package page.
pub.dev's README stylesheet forces img{height:auto}, so they are now sized with
percentage width (honored by both pub.dev and GitHub) and stay side by side.
Documentation only; no code, API, or runtime change.
Adds camera-agnostic helpers for building live detection previews, and documents the end-to-end live-camera pipeline in the README. No native or web r
Adds camera-agnostic helpers for building live detection previews, and documents the end-to-end live-camera pipeline in the README. No native or web runtime code changed. Additive and backward compatible.
FrameThrottle: a single-slot gate that drops camera frames arriving while a
previous frame is still being processed, replacing the hand-rolled
bool _isProcessing plus try/finally pattern in downstream apps.CoverFitTransform: maps detector coordinates onto a cover-fitted camera
preview (uniform scale, centered overflow, optional front-camera mirroring),
wrapping the existing coverFitScaleOffset. Use map for points and
scaleLength for radii and stroke widths.Fixes an Android build failure on Android Gradle Plugin (AGP) 9.x (issue #14). AGP 9 changed the default of android.sourceset.disallowProvider to true
Fixes an Android build failure on Android Gradle Plugin (AGP) 9.x (issue #14).
AGP 9 changed the default of android.sourceset.disallowProvider to true,
which rejects passing a Provider to the legacy jniLibs source-set API. The
plugin handed layout.buildDirectory.dir("litert-jni") (a Provider<Directory>)
to jniLibs.srcDir(...), so configuration failed at android/build.gradle.kts
with "You cannot add Provider instances to the Android SourceSet API." AGP 8.x is
unaffected, which is why it only surfaced for consumers on AGP 9.
libLiteRt.so is now contributed as a generated jniLibs source through the AGP
Variant API (androidComponents.onVariants { ... jniLibs.addGeneratedSourceDirectory(...) }) instead of the legacy
sourceSets { ... srcDir(<Provider>) } block. AGP owns the task dependency, so
the manual preBuild hook is removed, and a litertNextVersion bump now
re-downloads because the version is a tracked task input. Verified building the
plugin AAR on both AGP 8.11.1 and AGP 9.2.1.Also includes a performance pass over the Dart inference wrappers, verified with interleaved AOT A/B benchmarks on macOS and a physical iPhone:
Interpreter.run() with typed-data I/O is ~2x faster (771 -> 364 ns wrapper
overhead); CompiledModel.run() in managed mode is ~27% faster; the shared
YOLO-style decode utility is up to 72% faster (SIMD argmax, logit-space
pruning); packYuv420 accepts an optional reuse buffer so camera loops skip
a per-frame ~1.4 MB allocation.CompiledModel.runAsync/dispatchAsync now run the blocking native call on
a lazily spawned per-model helper isolate instead of blocking the calling
isolate, keeping the UI thread responsive during inference. Calls against
the same model serialize in FIFO order, and sync buffer-touching APIs
(run, dispatch, writeInput, readOutput, close) now throw
StateError while an async dispatch is in flight. runAsync with
thread-affine mobile GPU stacks (some Android OpenGL/OpenCL drivers) is
unvalidated; prefer run there.Restores the WASM-ready score on pub.dev (back to 160/160), which dropped to 150 when pub.dev upgraded its analyzer (pana 0.23.13). pana 0.23.13 mis-r
Restores the WASM-ready score on pub.dev (back to 160/160), which dropped to
150 when pub.dev upgraded its analyzer (pana 0.23.13). pana 0.23.13 mis-resolves
conditional export/import directives: it derives the condition name with
name.tokens.map((t) => t.value()).join(), and because library is a Dart
keyword Token.value() returns it upper-cased, so if (dart.library.X) becomes
dart.LIBRARY.X and never matches. Every conditional then resolves to its
default (first) URI. The main flutter_litert.dart barrel defaulted to the
native (dart:ffi / dart:isolate) surface, so the WASM/platform analysis saw
those libraries as reachable.
flutter_litert.dart barrel now defaults to the WASM-safe web
surface and gates the native surface on dart.library.io, so the package is
WASM-compatible again. Runtime behavior is unchanged: real native and web
builds resolve exactly as before.Isolate, SendPort, File, ...) and therefore cannot be WASM-safe is now
published from a new package:flutter_litert/native.dart library instead of
the main barrel: IsolateWorkerBase, IsolateRpcClient,
setupIsolateHandshake, InterpreterPool, ModelCheckpoint. Native code
using these now also needs import 'package:flutter_litert/native.dart';.
TensorFloat32Views and the rest of the API stay on the main barrel.InterpreterOptions on web gains hasDelegate, threads, and
copyWithoutDelegates() to match the native API.Preserve thread tuning and custom-op registrations when delegate application fails and interpreter creation retries on CPU.
Interpreter creation now falls back to CPU when a configured delegate cannot be applied to a model/runtime, instead of failing. This fixes classic Int
Interpreter creation now falls back to CPU when a configured delegate cannot
be applied to a model/runtime, instead of failing. This fixes classic
Interpreter creation for models that cannot use the default iOS Metal
delegate, including on the iOS simulator: it now warns and retries on CPU. The
fallback covers every creation path (fromAsset, fromBuffer, fromBytes, and
the isolate interpreter), and the iOS integration job now also exercises the
classic Interpreter path so this is caught in CI.
Makes the package web- and WASM-compatible. dart:isolate was reachable from the public API (via decode_failure.dart and isolate_rpc_server.dart) but i
Makes the package web- and WASM-compatible. dart:isolate was reachable from
the public API (via decode_failure.dart and isolate_rpc_server.dart) but is
unavailable on web/WASM; the isolate-dependent code now sits behind conditional
imports so none of it is reachable on the web build. No API changes.
The prebuilt LiteRt/LiteRtMetalAccelerator xcframeworks downloaded by the podspec shipped an arm64-only ios-arm64-simulator slice. CocoaPods selects i
The prebuilt LiteRt/LiteRtMetalAccelerator xcframeworks downloaded by the
podspec shipped an arm64-only ios-arm64-simulator slice. CocoaPods selects
ios-arm64_x86_64-simulator on the simulator, so the slice was skipped and the
build failed copying a non-existent slice (rsync: No such file or directory).
SwiftPM builds were unaffected. The litert-ios-v1.0.0 release asset was
re-uploaded with universal ios-arm64_x86_64-simulator slices.
Fixes the iOS CocoaPods build for the LiteRT Next runtime. The prebuilt
LiteRt.xcframework / LiteRtMetalAccelerator.xcframework download shipped an
arm64-only ios-arm64-simulator slice, whose identifier does not match the
ios-arm64_x86_64-simulator slice CocoaPods selects on the simulator. The build
then failed copying a non-existent slice (rsync ... No such file or directory).
ios-arm64_x86_64-simulator slice (arm64 device binary + x86_64 stub), so the
simulator build resolves and links. (SwiftPM builds were unaffected.)…place instead of each carrying its own copy. No breaking changes; the Interpreter and CompiledModel APIs are unchanged.
Additive release: shared isolate, CompiledModel-pooling, and image-RPC
utilities, extracted so the packages built on flutter_litert can maintain them
in one place instead of each carrying its own copy. No breaking changes; the
Interpreter and CompiledModel APIs are unchanged.
serveIsolateRpc: the isolate-side counterpart to IsolateRpcClient.
Drives the {id, op} -> {id, result | error} protocol from a handler map,
replacing the hand-written listen/switch/try-catch envelope each worker
isolate used to carry. IsolateRpcExactError lets a handler send a verbatim
wire-error string when the main side relies on the exact text (e.g. a
startsWith error contract).IsolateWorkerBase.disposeGracefully and
IsolateRpcClient.disposeGracefully: send the dispose op and await the
isolate's acknowledgement before killing it, so the isolate can free native
interpreters / CompiledModels. Isolate.kill(priority: immediate) otherwise
races past the queued dispose message and leaks the native handles.CompiledModelPool: a round-robin pool of CompiledModel slots, each
with its own reusable input buffer and AsyncLock, so concurrent inferences
(e.g. one per detected object) land on distinct models with leak-free init
teardown. A pool of size 1 degrades to a safe single-model-plus-lock.compiled_io_utils: compiledFloatCount, squareSideFromFloats,
compiledSquareInputSide, compiledOutputFloatCounts, and
indexWhereFloatCount for deriving tensor geometry from a CompiledModel,
whose tensor sizes are exposed only in bytes.cameraFrameRpcFields and cameraFrameFromRpcMessage: pack a
CameraFrame into an isolate-request field map and rebuild it on the isolate
side (any image decode stays in the consumer, keeping this dependency-free).decodeFailurePrefix, throwDecodeFailure, and
rethrowOrFormatException: signal an undecodable-image failure from inside
an isolate and surface it as a FormatException on the main side instead of a
cryptic downstream error.Deprecated: manual hardware-acceleration delegates for the Interpreter API, namely GpuDelegateV2 (Android GL/CL), the Metal GpuDelegate, and CoreMlDel…
CompiledModel API: CompiledModel.fromFile,
CompiledModel.fromBuffer, and CompiledModel.fromBufferWithGpuFallback,
with automatic hardware-accelerator selection via
Accelerator.{cpu, gpu, npu}, Precision, and TensorBufferMode. This is
the recommended path for GPU/NPU acceleration going forward, following
Google's LiteRT Next guidance
(https://developers.google.com/edge/litert/next/get_started). Supported on
Android, iOS, macOS, Windows, and Linux.GpuDelegateV2 (Android GL/CL), the Metal GpuDelegate, and
CoreMlDelegate
(with their *Options). They remain fully functional but are superseded by
CompiledModel's automatic accelerator selection and are planned for removal
in 4.0.0. The Interpreter API itself, the CPU XNNPackDelegate, and
FlexDelegate are NOT deprecated and remain fully supported.XNNPackDelegate with XNNPackDelegateOptions no longer
crashes on the arm64 Android emulator. The options struct was initialized by
calling the native TfLiteXNNPackDelegateOptionsDefault(), which returns the
struct by value; that by-value FFI return crashes the Dart VM on the arm64
Android emulator (it works on real devices, macOS, and iOS). The struct is now
built in Dart, matching upstream tflite_flutter, while preserving the QS8/QU8
quantization defaults; the resulting native options are unchanged, so there is
no behavior difference on real devices.TensorFloat32Views input views are now genuinely writable. They were
previously built from the unmodifiable Tensor.data view, so indexed writes
(views.inputs[0][i] = x) threw UnsupportedError, and bulk
setAll/setRange only worked through a Dart VM enforcement gap that a
future SDK could close. Views are now captured via the new
Tensor.asFloat32View(), a mutable Float32List aliasing the tensor's
native buffer (valid until the next resize/allocateTensors).SignatureRunner.run() per-call overhead roughly halved (16-17µs → 7µs per
call on the bundled test/benchmark/signature_runner_benchmark_test.dart):
tensor handles are cached by name between allocations, and the valid-names
error text is built only when a lookup actually fails instead of on every
getInputTensor/getOutputTensor call.IsolateInterpreter.run/runForMultipleInputs no longer
silently drop calls. A call issued while a previous run is in flight is now
queued and completes with real results (previously it returned normally
without writing the output buffers); frame-skipping callers can check
state == IsolateInterpreterState.loading before calling. Running after
close() now throws StateError instead of returning silently.TensorType.fromValue is O(1) instead of scanning all enum values (it runs
on every Tensor.type access), and inference timing uses a reused monotonic
Stopwatch instead of two DateTime.now() calls per run.test/benchmark/engine_overhead_benchmark_test.dart (MediaPipe
face_detection_short_range, macOS host):
run()/runForMultipleInputs() with nested-list input and output drops
from 8.9ms to 1.9ms per inference (native floor 1.0ms) by converting
tensors through a single pre-sized buffer instead of one small allocation
per element, and by reading outputs through typed views instead of a
per-element ByteData.view.Tensor.setTo/copyTo now copy directly between Dart memory and
TfLiteTensorData instead of round-tripping through a native scratch
buffer (two extra copies per tensor per inference).Float32List (or other typed data) as an input no
longer resizes the input tensor to rank 1, which broke models with
rank-sensitive ops (CONV_2D failed to prepare). Flat typed data whose
element count matches the tensor is now staged as-is, and is the fastest
run() input type.Float32List, Int32List,
Int64List, Int16List, Int8List). Bytes are bulk-copied directly
into the buffer; previously this threw a shape-mismatch ArgumentError.
run() with Float32List in/out now measures within ~7% of the
raw tensor-views floor.copyTo(Uint8List)/copyTo(ByteBuffer) now fill and
return the destination instead of returning a separate copy.run, runAsync, lock/unlock paths).Android: support both AGP 8 and AGP 9 by moving the plugin Gradle files to Kotlin DSL and updating the Android tooling plugin declarations (6c332e3b).
Fix GPU and CoreML delegates silently falling back to CPU on macOS and iOS (#11). macOS now bundles the GPU/CoreML dylibs that were previously omitted
Complete the AGP 9 / built-in Kotlin fix from 2.8.0 (#10). 2.8.0 resolved the "Inconsistent JVM-target ... (17) and (21)" error on the AGP 8.11 transi
android.builtInKotlin=true on AGP 9 still
failed with "The 'org.jetbrains.kotlin.android' plugin is no longer required
since AGP 9.0": the Flutter Gradle plugin auto-applies the legacy Kotlin
plugin to this module, and AGP 9 rejects it. The plugin now applies
kotlin-android only on AGP < 9 (which also stops Flutter from auto-applying
it), and keeps the JVM-target pin guarded so the AGP-9-without-built-in-Kotlin
case is skipped. Verified building against AGP 8.11.1 and 9.0.1 with built-in
Kotlin both enabled and disabled.Fix the "Inconsistent JVM-target compatibility detected ... (17) and (21)" Android build failure under AGP 9 / Flutter 3.44+ (#10). The fix pins the K
…(all platform variants). Old name kept as a @Deprecated alias.
InterpreterOptions.addCustomOp(...): high-level method for registering
custom TFLite ops; handles native string allocation and lifetime internally,
replacing the previous raw tfliteBinding call pattern.Interpreter.fromBytes(Uint8List): async cross-platform constructor,
matching the web API. Native platforms complete immediately; unsupported stub
throws UnsupportedError.lastNativeInferenceDurationMicroSeconds →
lastInferenceDurationMicroseconds on Interpreter, SignatureRunner, and
LiteRtInterpreter (all platform variants). Old name kept as a @Deprecated
alias.configureLiteRtLoader → configureLiteRtWebLoader. Old name kept as
a @Deprecated alias and re-exported from all_web.dart.camera_frame.dart: widen .planes cast from List<dynamic> to
Iterable<dynamic> for broader compatibility.Fix iOS Swift Package Manager builds: repackage the bundled TensorFlowLite xcframeworks (correct simulator slice identifiers and framework structure)
flutter_litert and opencv_dart.Raise minimum deployment targets to iOS 13.0 / macOS 10.15 to satisfy Swift Package Manager's FlutterFramework requirement (fixes SPM build failures o
FlutterFramework requirement (fixes SPM build failures on macOS/iOS).flutter_litert_flex: ^1.0.0.Update example to use flutter_litert_flex: ^0.0.8.
flutter_litert_flex: ^0.0.8.Fix SPM: add missing FlutterFramework dependency to iOS and macOS Package.swift.
FlutterFramework dependency to iOS and macOS Package.swift.Add SPM support for iOS: TensorFlowLiteC, TensorFlowLiteCMetal and TensorFlowLiteCCoreML are now declared as binary targets in Package.swift so the pl
Fix WASM compatibility: replace dart:io import in camera_frame.dart with flutter/foundation.dart to allow package to compile under the WASM runtime.
prepareCameraFrameFromImage and prepareCameraFrame now auto-detect isBgra based on platform. macOS uses BGRA, Windows and Linux use RGBA. The isBgra p
* Update documentation
Add decodeBitmap(Uint8List bytes) free function: decodes encoded image bytes (JPEG, PNG, etc.) to a web.ImageBitmap via createImageBitmap, off the mai
decodeBitmap(Uint8List bytes) free function: decodes encoded image bytes (JPEG, PNG, etc.) to a web.ImageBitmap via createImageBitmap, off the main thread.WebGpuFallback mixin: transparent WebGPU-to-WASM runtime fallback for web detector classes. Provides withFallback<T>() which catches GPU errors, swaps all runners to WASM via swapToWasm(), and retries once. Apply with with WebGpuFallback; implement activeAccelerator and swapToWasm().package:flutter_litert/flutter_litert.dart on web.Add LiteRtInterpreter, an alternative web inference path backed by Google's official LiteRT.js runtime (@litertjs/core). Selectable at construction ti
LiteRtInterpreter, an alternative web inference path backed by Google's official LiteRT.js runtime (@litertjs/core). Selectable at construction time via LiteRtInterpreter.fromBytes(bytes, accelerator: 'webgpu' | 'wasm'), with automatic fallback from webgpu to wasm when ops aren't supported by the GPU delegate.
Interpreter hot path used by detector packages: fromBytes, getInputTensor / getOutputTensors, runForMultipleInputs(inputs, outputs). runForMultipleInputs is async (LiteRT.js run returns a Promise).Float32List, ByteBuffer, or the legacy nested List<List<List<double>>> shape used by tflite-js callers; the float-typed buffer paths take a single bulk copy.JSFloat32Array.toDart directly, skipping the dataSync().dartify() round-trip.Interpreter._tensorFromJSTensor: replaces dataSync().dartify() as List<double> + Float32List.fromList(...) with a single bulk copy via JSTensorExtensions.dataSyncFloat32. ~25 ms / call savings on a 705k-element YOLOv8n output.LiteRtInterpreter.fromBytes(...) call programmatically appends a <script type="module"> to <head> that imports @litertjs/core from jsDelivr and calls loadLiteRt(...); consumers don't have to touch their web/index.html. Override URLs (for self-hosting / strict CSP) or disable auto-loading via configureLiteRtLoader(moduleUrl: ..., wasmUrl: ..., autoLoad: ...). Existing host-page loaders that assign window.LiteRt and dispatch a litert-ready event still work.Interpreter remains the default web runtime.Make camera_overlay.dart WASM-compatible on Flutter Web
camera_overlay.dart WASM-compatible on Flutter WebAdd painter primitives drawLandmarkMarker, drawSkeletonConnections, and drawBoundingBoxOutline for reuse by detector example apps and overlay widgets.
drawLandmarkMarker, drawSkeletonConnections, and drawBoundingBoxOutline for reuse by detector example apps and overlay widgets. Pure Dart + dart:ui, no new dependencies.Add camera-overlay helpers used across detector example apps: rotationForFrame, detectionSize, coverFitScaleOffset, barQuarterTurns, and FpsCounter. A
rotationForFrame, detectionSize, coverFitScaleOffset, barQuarterTurns, and FpsCounter. All pure Dart + Flutter SDK, no new dependencies. Lets example apps drop ~200 lines of duplicated orientation / sizing / FPS boilerplate.Add prepareCameraFrameFromImage, a duck-typed wrapper around prepareCameraFrame that accepts a CameraImage-shaped object directly (any object exposing
prepareCameraFrameFromImage, a duck-typed wrapper around prepareCameraFrame that accepts a CameraImage-shaped object directly (any object exposing width, height, planes with bytes/bytesPerRow/bytesPerPixel). Lets detector packages expose one-line camera-stream APIs without adding package:camera as a dependency here. Pure Dart, no new dependencies.Add prepareCameraFrame helper plus CameraFrame, CameraFrameConversion, and CameraFrameRotation types. Describes a camera frame (YUV420 or packed BGRA/
prepareCameraFrame helper plus CameraFrame, CameraFrameConversion, and CameraFrameRotation types. Describes a camera frame (YUV420 or packed BGRA/RGBA) in a pure-Dart descriptor that detector packages can hand to their existing detection isolate, moving the cvtColor / rotate work off the UI thread without adding opencv_dart as a dependency here.CameraPlane typedef (structurally identical to YuvPlane; use whichever name reads better at the call site).TensorFloat32Views (native only): captures Float32List views of an Interpreter's input/output tensors once after allocateTensors, letting detector packages reuse the same view wrappers on every inference instead of recreating them per-call. Pure Dart, no new dependencies.Add packYuv420 helper for packing NV12 / NV21 / I420 camera frames into a contiguous buffer
packYuv420 helper for packing NV12 / NV21 / I420 camera frames into a contiguous bufferMinor performance/accuracy optimizations:
fillNHWC4DFloat32List fast paths for common tensor flattening shapesFix Android JVM target mismatch: bump Java compile target to 17 to match Kotlin target set by Flutter toolchain
Fix Android Flutter beta builds by aligning Kotlin and Java JVM targets to 11
Fix edge case in output buffer allocation
* Update documentation
Enable XNNPACK delegate on Android (ARM NEON SIMD acceleration in auto mode)
PerformanceConfig.xnnpack() on iOSAdd Windows XNNPack delegate support (2-5x CPU inference speedup via SIMD)
Fix Android custom ops library alignment for 16 KB page-size devices
Add useIsolateInterpreter parameter to skip nested isolate creation
Fix native crash during repeated inference by removing unsafe output tensor writeback
Fix macOS native crashes by disabling auto IsolateInterpreter for no-delegate interpreters.
Fix WASM compatibility: move dart:isolate imports behind conditional exports so web compilation path is WASM-safe
dart:isolate imports behind conditional exports so web compilation path is WASM-safeFix: use-after-free when interpreter reads model weights from freed buffer, transfer buffer ownership from Model to Interpreter
Model to InterpreterAdd IsolateWorkerBase for shared isolate lifecycle management
IsolateWorkerBase for shared isolate lifecycle managementRoundRobinPool generic round-robin pool utilityTensorType enum, LandmarkMixin, listUtils shared helpersnms()DelegateLibraryLoaderall_unsupported.dart, version.dart, flutter_litert_method_channel.dart, flutter_litert_platform_interface.dartModel buffer leak, delegate options leak, stale tensor cacheBreaking: Point.x and Point.y changed from int to double.
Breaking: Point.x and Point.y changed from int to double.
Point to double-precision with optional z depth, ==/hashCode, toMap()/fromMap(), is3DBoundingBox class (4-corner Point-based, supports rotated boxes)
BoundingBox.ltrb() factory for axis-aligned boxesleft/top/right/bottom convenience getterswidth, height, center, corners computed propertiestoMap()/fromMap() serializationFix tensor cache bug, add shared Point class, dedup internals
Add NaN handling to clamp01(), returns 0.0 for NaN inputs
clamp01(), returns 0.0 for NaN inputsAdd IsolateRpcClient and setupIsolateHandshake for reusable isolate request/response communication
IsolateRpcClient and setupIsolateHandshake for reusable isolate request/response communicationYour coding agent can read these notes before it upgrades. Set up the MCP server →