NewYour coding agent can read the release notes before it upgrades.Set up the MCP server →
pub.dev · #3112 most downloaded on pub.dev
Cross-platform, on-device LLM inference for Dart and Flutter. Run GGUF (llama.cpp) and LiteRT-LM models locally on Android, iOS, macOS, Windows, Linux and web.
Last release today
07 Oct 2026
Ships on a steady schedule
a new release about every 2 weeks
Nearly every release is documented
notes for 57 of 58 stable releases
Nothing withdrawn
no release was ever pulled
8 months old
58 releases · first in 2026
One column per month.
Disposal now cancels model and LoRA downloads started by deprecated loadModelSource , without waiting for resolution that ignores cancellation ( #895…
Fix a Flutter macOS app aborting in ggml-metal when it quits while a
llama.cpp model loads or after a hot restart during a load, and a Dart
program aborting when it ends or kills an isolate during a load: the
llama.cpp runtime now frees the models, contexts, projectors and decision
heads still allocated at exit
(#813,
llamadart-native#96).
A native host's C exit() while a llama.cpp call is running is not covered.
Update the default llama.cpp runtime to leehack/llamadart-native@v0.5.0-2,
a rebuild of llama.cpp v0.5.0 whose Apple XCFramework carries a privacy
manifest (File Timestamp, reason C617.1). A runtime older than v0.5.0-1
keeps the previous exit behavior and logs a warning at the first model load.
Fix a process aborting when a Dart program dies of an error, or a Flutter
macOS app quits or hot restarts, while an image model loads or generates,
and a macOS Metal process aborting in ggml-metal when a native host exits
with an image model loaded: the stable_diffusion runtime now records
progress instead of calling back into Dart, and frees the image models
still allocated at exit
(stable-diffusion-native#10).
Image progress events now arrive up to about 50 ms after the runtime
reports them. A quit during an image generation waits for the generation
to finish; dispose the engine first to quit at once.
Update the stable_diffusion runtime to
leehack/stable-diffusion-native@v0.2.0-1, a rebuild of the same
stable-diffusion.cpp commit whose Apple XCFramework carries a privacy
manifest (File Timestamp, reasons C617.1 and 3B52.1). Image generation
needs this runtime or a later one.
Documented that LiteRT-LM on iOS has run on a device only on iOS 18.3.2;
iOS 16.4 remains the declared, untested deployment floor
(#831).
Documented Apple privacy manifests: one reaches an app inside a companion
package's XCFramework, never through the hook path, and an app whose
llama.cpp or stable_diffusion runtime has none declares File Timestamp
reason C617.1 (plus 3B52.1 for stable_diffusion) itself.
Flutter iOS and macOS apps pair core 0.11.0 with
llamadart_llama_cpp_flutter 0.0.21, llamadart_litert_lm_flutter
0.0.13 and llamadart_stable_diffusion_flutter 0.0.2, which link the
runtimes this release pins; an older llama.cpp or stable_diffusion companion
fails the Apple build.
Adopt LiteRT-LM v0.17.0-7 with provider-free iOS artifacts; Gemma FST
constrained decoding remains unavailable on iOS.
Update the default LiteRT-LM runtime to leehack/litert-lm-native@v0.17.0-8.
Its iOS SwiftPM frameworks carry Apple privacy manifests (File Timestamp
C617.1 and 3B52.1, System Boot Time 35F9.1, User Defaults CA92.1);
an app that runs LiteRT-LM on the hook path declares them itself.
Fix Android LiteRT-LM GPU engines keeping their graphics memory after
deletion, which got an app killed after one or two model reloads on a
Galaxy S24
(litert-lm-native#59).
Documented that LiteRT-LM's default GPU selection on Android generates wrong
text for Qwen3 0.6B on Adreno 750; load it with ComputeDevice.cpu
(#553).
Documented that llama.cpp Vulkan on Android is experimental and
device-dependent, with known failures on Pixel 9 Pro, Galaxy A53 and Galaxy
S24; auto stays on the CPU
(#948).
LiteRT-LM release sync can remove an obsolete iOS provider target from modern
Swift packages while preserving required macOS runtime libraries.
Reentrant engine disposal now shares one teardown and immediately reports
disposed state, including calls from logging or backend cancellation hooks.
Native requests, generation, and speech synthesis fail promptly when their
worker exits unexpectedly, allowing disposal to finish.
Native recurrent and hybrid models reject nonzero speculative rollback
capacity before context creation, avoiding runtime graph-budget aborts.
Native embeddings reject inputs that exceed a required one-pass micro-batch,
avoiding non-causal attention aborts and incorrect MEAN/CLS pooled vectors.
Disposal now cancels model and LoRA downloads started by deprecated
loadModelSource, without waiting for resolution that ignores cancellation
(#895).
Prevented execution of incomplete parallel tool-call replies when generation
reaches a reported limit; the tool loop rolls back the whole turn.
Automatic tool loops now reject pinned LiteRT-LM native and Web runtimes
before generation because they cannot report token-limit truncation reliably;
manually managed completion remains available (#919).
Web GGUF completions preserve runtime limits as length, so tool loops
report truncated and roll back cut-off turns; typed Web speech recognition
raises a runtime truncation error with the partial transcript.
Redacted known URL secrets repeated in model filenames from download
cancellation messages.
Model unload, replacement, and disposal now report interrupted tool loops as
cancelled, preserving partial answers and rolling back unfinished turns.
Projector loads now reject model changes during loading instead of reporting
success after the model has been unloaded.
Fixed: URL-valued ModelSource.path diagnostics redact credentials and
signed queries while preserving the loading path and cache identity
(#846).
Fixed LiteRT-LM speech recognition startup failures hanging indefinitely;
failed workers now release their resources and allow another attempt.
Kept cancelled replies out of reset chats and rolled back turns cancelled
before any output; structured JSON cancellation now reports a state error.
A model change during draft-model resolution also rolls back the turn;
resets during context preparation preserve the replacement conversation.
Redacted URL credentials and signed query strings from redirected download
failures, download snapshots and invalid model-source errors, including
slashless URLs; LiteRT-LM Web model names reject decoded URL delimiters.
LlamaEngine.load(LlamaModel(source, projector:), params:, download:, onProgress:, store:) creates an engine and loads a model with its
projector in one atomic call, and setModel loads or replaces the model of
an engine: the loaded model keeps serving until every new file has
downloaded, except on the Web, where the runtime unloads it before it
fetches the new one. A class that implements LlamaEngine must add
setModel, and an override of loadMultimodalProjectorSource its new
download parameter
(#846).
Deprecated: LlamaEngine.loadModel, loadModelSource,
loadModelFromUrl and loadMultimodalProjector; use LlamaEngine.load or
setModel. loadMultimodalProjectorSource(options:) is now download:.
They still work until 1.0
(#846).
Fixed: a load checks that it can proceed before it downloads:
LlamaEngine.load and setModel reject a projector for a LiteRT-LM model
and ComputeDevice.npu for a GGUF first, and the deprecated
loadModelSource throws for an already loaded engine first
(#846).
Fixed: LlamaEngine.dispose() stops the downloads of a running
setModel at once, and unloadModel() and dispose() stop a
loadMultimodalProjectorSource download instead of waiting for it; the
load throws LlamaStateException
(#895,
#896).
Behavior change: on the Web, a ModelSource.path loads as a URL
relative to the document, or as a blob: URL, for models, projectors, LoRA
adapters and draft models; it used to throw LlamaUnsupportedException
(#846).
Fixed: the llama.cpp WebGPU bridge resolves a relative model,
projector, LoRA or draft model URL against the document, as LiteRT-LM Web
does; its worker resolved one against webgpu_bridge/, so the fetch
failed (#846).
A local ModelSource.path whose file name holds %2F or %5C, or whose
directory is named %2e or %2e%2e, loads through LlamaEngine.load,
setModel, loadMultimodalProjectorSource, setLoraSource,
ModelParams.loras, draft models and the image, decision and speech
engines' load (#846).
Breaking (Preview): ImageGenerationEngine.generate returns
Future<ImageGenerationTask>; await it before reading events or calling
cancel (#850).
Breaking: ImageGenerationTask, SpeechToTextTask and
TextToSpeechTask no longer report a failure as an error on events; read
it from done, or from the new result, which returns the result or
throws the failure, or LlamaStateException when the task is cancelled
(#850).
Fixed: SpeechToTextTask.cancel() stops only that recognition; it no
longer cancels chat and other requests on the same LlamaEngine
(#850).
ModelParams.device (ComputeDevice) selects the device for every
runtime: auto keeps each runtime's default, and an explicit cpu, gpu
or npu runs there or throws LlamaUnsupportedException instead of
falling back to another device
(#849).
Deprecated: ModelParams.liteRtLmBackend, LiteRtLmBackendPreference
and LiteRtLmBackend(preferredBackend:); use ModelParams.device. They
still work until 1.0
(#849).
Behavior change: DecisionModelParams(device: ComputeDevice.gpu) runs
the encoder on a GPU, Vulkan on Android, or throws
LlamaUnsupportedException from the encoder load; it ran on the CPU on
Android and threw before loading where GPU modules load with the model
(#849).
Behavior change: LlamaEngine loads run ModelParams.validate() before
any download or native call, so an invalid combination throws
LlamaArgumentException instead of LlamaModelException
(#849).
Breaking: ModelParams.validate() throws LlamaArgumentException,
ModelDownloadController throws LlamaArgumentException or
LlamaStateException, and Web backend calls before a model load throw
LlamaStateException, instead of ArgumentError or StateError; a failed
load's details is now the cause's message instead of a {type, message}
map, and the not-ready error names LlamaEngine.load and setModel
(#843).
Deprecated, behavior change: sourceLangCode and targetLangCode on
LlamaEngine.create, createStructuredJson, chatTemplate and
BackendNativeChatGeneration.generateChat; pass
chatTemplateKwargs: {'source_lang_code': 'en', 'target_lang_code': 'ko'},
which TranslateGemma templates now read, as llama.cpp does. A custom
BackendNativeChatGeneration now gets the codes from LlamaEngine only in
chatTemplateKwargs
(#853).
Deprecated: LlamaLogging.configure(level:, nativeLevel:, handler:)
replaces LlamaEngine.configureLogging and the engine's setLogLevel,
setDartLogLevel and setNativeLogLevel; levels are now library-wide, so
the last call wins and reaches every running engine's worker, including the
default native backend's (#845).
Deprecated: speech engines follow the shared engine pattern:
SpeechToTextEngine.load(SpeechToTextModel(...)) and
TextToSpeechEngine.load(TextToSpeechModel(...)) download every
ModelSource and own what they load, attach(engine, adapter:) borrows a
loaded LlamaEngine, dispose() cancels the running task, and
transcribeOnce and synthesizeOnce return the final result. Adapters
(Qwen3AsrAdapter, LiteRtLmAsrAdapter, Qwen3TtsAdapter, or your own
SpeechToTextPromptAdapter) replace SpeechToTextModelProfile,
TextToSpeechModelProfile, the modelProfile constructors and
SpeechToTextEngine.liteRtLm, which keep working until 1.0
(#848).
Breaking: SpeechToTextEngine gains dispose(), isDisposed,
adapter and transcribeOnce, and TextToSpeechEngine gains dispose(),
isDisposed, adapter and synthesizeOnce, so a class that implements
either must add them
(#848).
Breaking: DecisionEngine.load(DecisionModel(encoder:, head:, config:), params:, download:, onProgress:) loads a decision model from
ModelSources into an engine it owns, atomically, and dispose() frees it
all; DecisionEngine.attach(engine, head:, config:) adds a head to a loaded
LlamaEngine; decision engines report an instance capabilities. The
String-path load(engine, headPath:, configPath:) is removed; use
attach (#847).
Breaking (Preview): image generation follows the shared engine
pattern: ImageGenerationEngine.load(ImageGenerationModel(source, components: [...]), params:, download:, onProgress:) downloads every
ModelSource into the model cache, with combined progress, cancellation
and cache reuse, and detects each file's role from its header. Generation
settings move to ImageGenerationRequest. download's bearer token and
headers never reach more than one origin: such remote files throw
LlamaArgumentException. The model presets, ImageGenerationModelFamily,
ImageGenerationModelFiles, ImageGenerationDefaults,
ImageGenerationOptions (now ImageModelParams, and the engine's options
getter params), ImageGenerationDevice (now ComputeDevice) and String
paths are removed; MIGRATION.md maps each former preset to its files and
request settings
(#883).
Breaking: LlamaEngine.dispose() is idempotent and terminal, like
every other engine's: each call returns the same future, and afterwards
loads, requests and DecisionEngine.attach throw LlamaStateException
while capabilities reports the engine as disposed. getBackendName,
getAvailableBackends, isGpuSupported, getVramInfo, listGpuDevices
and getResolvedGpuLayers used to answer after dispose() and now throw
LlamaStateException too. A load running when it is called throws
LlamaStateException, and its model is unloaded. LlamaEngine gains
isDisposed, so a class that implements it must add it
(#851).
Breaking (Preview): ImageGenerationEngine.capabilities is async, and
ImageGenerationEngine.runtimeCapabilities() is removed; use
checkRuntime(). ImageGenerationCapabilities and DecisionCapabilities
implement EngineCapabilities
(#851).
Behavior change: LlamaEngine.supportsVision and supportsAudio report
what capabilities reports: true for a LiteRT-LM bundle that takes media
directly, and false instead of throwing when the runtime cannot probe the
projector (#851).
Behavior change: responseFormat maps with an unknown type or
key, such as json_shema or a misspelled schma, now throw
LlamaUnsupportedException before generation instead of generating
unconstrained output; a null-valued key counts as absent
(#836,
#864).
ChatSession.create takes responseFormat, and the new
ChatSession.createStructuredJson decodes the reply. A turn that fails or
is cancelled before its first chunk, such as a strict format on LiteRT-LM,
removes its user message from the history, and one cancelled or failing
mid-stream keeps the partial reply, so alternating-role templates keep
working
(#836,
#864).
LlamaCompletionChunk.model, observer model names and load logs report
llama_model, and web LiteRT-LM omits general.name, when a URL's last
path segment repeats its userinfo credential; web LiteRT-LM
litert_lm.model_url shows a relative URL as given and a blob: or data:
URL as its scheme (#822).
loadModelSource throws LlamaUnsupportedException instead of
ArgumentError for a local path whose file name holds %2F or %5C or
whose directory is named %2e or %2e%2e; loadModel still loads it
(#822).
Behavior change: on Android and iOS, the default model cache is now
llamadart/models in the app's cache directory instead of the temporary
directory, which Android empties on every app update and iOS purges; add
DefaultModelDownloadManager.globalCacheDirectory to move every default
download (#838).
Native LlamaBackend() picks llama.cpp or LiteRT-LM from the model file's
header, not its extension, so extensionless downloads load in the right
runtime and mislabelled files throw LlamaModelFormatException; name a Web
URL's format with ModelSource.url(..., format: ModelFormat.liteRtLm)
(#837).
Add LlamaEngine.runtime and LlamaEngine.capabilities, one snapshot of
what the loaded model's runtime supports: image and audio input,
embeddings, multi-turn chat, tools, structured output, grammars, every
sampling control including penalty, stream batching and speculative
strategies. Native LiteRT-LM reports the image, audio and speculative
decoding support the bundle declares, best-effort (litert-lm-native#60),
and a request the runtime then fails for lack of one throws
LlamaUnsupportedException naming it instead of an opaque error.
backendGenerationCapabilities is deprecated; on native LiteRT-LM it now
reports streamBatching and, for bundles without a declared drafter, no
speculative strategy (#841).
Read completions without choices.first.delta: chunk.text,
chunk.thinking, chunk.toolCalls and a typed chunk.finishReason
(LlamaFinishReason); stream.text(), stream.textDeltas() and
stream.collect() (a LlamaCompletion with assembled tool calls and an
assistant message); and the one-shot engine.complete(messages) and
session.send('...')
(#840).
session.sendWithTools(text, tools: ...) and completeWithTools(parts, ...) run the model's tool calls with each ToolDefinition.handler,
concurrently for parallel calls, until it answers, and return a
LlamaToolLoopResult whose stopReason also reports maxRounds,
unhandled calls, context overflow, a reply cut off at maxTokens
(truncated, rolled back) and cancellation
(#842).
Breaking: ChatSession gains createStructuredJson, and LlamaEngine
gains runtime, capabilities, setLoraSource and removeLoraSource, so
a class that implements either must add them.
Breaking: ToolDefinition.handler is nullable, so tools the app runs
itself can leave it out; code that calls tool.handler(params) must check
it for null first (#842).
GenerationGrammarTrigger.typed(type: GrammarTriggerType.word, ...)
replaces the raw-int constructor, now deprecated; an unknown raw trigger
type throws LlamaUnsupportedException on llama.cpp instead of being
ignored (#844).
LoRA adapters (setLoraSource, removeLoraSource,
LoraAdapterConfig.source), speculative draft models
(SpeculativeDecodingConfig.draftModel, withDraftModel,
withDraftModelDownload) take a ModelSource, so they download and cache
like models, and LiteRtLmAsrRuntimeConfig.source takes local
ModelSource files
(#852).
Deprecated: the String path forms of LoRA adapters, speculative draft
models and LiteRT-LM ASR files
(#852).
Breaking: package:llamadart/llamadart.dart is the app API. The raw
ffigen bindings move to package:llamadart/llama_cpp_bindings.dart (native
only, outside semantic versioning), and the custom-backend SPI moves to the
new package:llamadart/backend.dart: every Backend* type except
BackendPerfContextData and BackendTextToSpeechModel, LiteRtLmBackend,
LiteRtLmRuntimeClient, LiteRtLmRuntimeMetrics, LiteRtLmRuntimeResult,
LiteRtLmAsrRuntimeSession, LiteRtLmAsrPushResult,
LiteRtLmAsrProcessResult and LiteRtLmAsrProcessState. The app API now
exports TemplateToolCallSerialization
(#355).
Breaking: the LlamaEngine text-to-speech and decision hooks,
modelHandle and contextHandle move to the LlamaEngineBackendHooks
extension in package:llamadart/backend.dart, so neither a subclass nor an
implements LlamaEngine fake can override them; fake a backend that
implements BackendTextToSpeech or BackendDecision instead
(#355).
Breaking: the deprecated LiteRtLmBenchmarkClient,
LiteRtLmBenchmarkMetrics, LiteRtLmBenchmarkResult,
LiteRtLmRuntimeClient.conversationTokenCount and
LiteRtLmRuntimeClient.replaceConversationWithClone are removed
(#355).
v0.1.54, unchanged from 0.10.0:v0.5.0, are qualified against native v0.5.0, and keepv0.5.0@7fe450e19305b828c199d602c23a8337aaa1f03b parity@litert-lm/core@0.15.0. Immutable Web asset manifest:8a9f83c15035eeb034a6563e6f753382d7d7f9be81503ef76902138da7841176.Add the llamadart_stable_diffusion_flutter 0.0.1 companion package: Flutter iOS and macOS apps that add it link the image generation runtime through S
llamadart_stable_diffusion_flutter 0.0.1 companion package:xcodebuild, not by plain flutter build or flutter run output).leehack/stable-diffusion-native@v0.2.0.ModelParams.chatTemplate, which engine.create andengine.chatTemplate ignored in favor of the GGUF templateLlamaBackend.applyChatTemplate now throwsLlamaUnsupportedException for a template override it cannot renderqwen.gguf, in LlamaCompletionChunk.model instead ofLlamaOperation.model, it now reads local paths as pathsdata: and blob: URLslitert_lm.model_url%, such asC:\models\qwen 100%.gguf, or whose URL file name decodes to one, insteadArgumentError; loadModelSource and ModelCacheEntry keep% in local and cache paths literal instead of percent-decoding them intomain returns or throwsImageGenerationEngineAppLifecycleListener.onExitRequested,dart run skills@ get.ModelParams.loras at model load on nativev0.1.54, fail the loadLlamaUnsupportedException instead ofLlamaModelException when native or web LiteRT-LM rejects a ModelParamsstable_diffusion native runtimellamadart_native_runtimes for imageall, andllamadart_stable_diffusion_backends picks its CPU or Vulkan build on LinuxImageGenerationEngine, onstable_diffusion runtime: SDXS and SD-Turbo presets,ImageGenerationDefaults.width and height); not available on the webImageGenerationEngine.warmUp, which compiles the GPU pipelines ofImageGenerationEngine.checkRuntime, which probes the image runtimeload now probes the same way, andLlamaUnsupportedException for image or audio partsLlamaInferenceException. Native llama.cpp and LiteRT-LM also throw itLlamaImageContent.url.LlamaEngine.supportsEmbeddings. Native llama.cpp embeddings of anLlamaUnsupportedException, and inputLlamaInferenceException, instead of aException.llmMemAvailable andMemAvailable leaves out memory the system frees onwarmUp leaves the size unset, and the basicv0.1.54, unchanged from 0.9.0:v0.5.0, are qualified against native v0.5.0, and keepv0.5.0@7fe450e19305b828c199d602c23a8337aaa1f03b parity@litert-lm/core@0.15.0. Immutable Web asset manifest:8a9f83c15035eeb034a6563e6f753382d7d7f9be81503ef76902138da7841176.Document generic JSON tool calling as an intentional fallback, including its prompt and model-reliability limits; runtime behavior is unchanged ( #755
contextSize using smaller processing batches, while preservingLlamaModelException with its real cause, not as a COOP/COEPpreferredBackend GPU module is notcuda with thellamadart_server example exiting at startup on Windows; it stopspresencePenalty, minP and thinkingBudget, and runtime LoRAsetLora, removeLora, clearLoras), on WebGPU with bridgebackendGenerationCapabilities.speculativeDecodingStrategies; other assetsGenerationParams.minP on WebGPU when the bridge lacksLlamaUnsupportedException instead of ignoring it, and ignore apreservedTokens entry there, as native llama.cppLlamaEngine.backendGenerationCapabilities, which reports whether thepresencePenalty, minP and thinkingBudget; the\n and \r in Qwen3-Coder XML reasoning, as llama.cpp doesChatSession history</think>. Only whitespace atChatSession history. Content and reasoning are<tool_call> envelope intoChatSession history; only a possible envelope openingChatSession thinking, is trimmed per thought as" \n Hello there. \n\n" streams as "Hello there.". Before, it streamedLlamaModelException when a WebGPU model load fails with a bridgeLlamaInferenceException orLlamaStateException for such Web embedding, next-token scoring and stateLlamaEngine model and projector load errors and logs and the model field//user:pass@host/... URLs and relative paths with a query, and out ofLlamaException now throws LlamaModelException. The details of a{type, message} map instead ofString in detailsDecisionEngine.load throw LlamaStateException when another model isphys_footprint on macOS and iOS, RssAnon plus RssShmemVmSwap on Linux and Android, read after malloc_trim(0) where the CPrivateUsage plusSharedCommitUsage on Windows (PrivateUsage alone on builds without it).peak_memory_bound without memoryleak_slope_bound when the least-squares footprint<|tool_list_start|>, as llama.cpp does, so LFM2.5-1.2B-Instruct andv0.1.54, on the final create chunk and to observersnull in chat templates, as llama.cpp does: QwQ-32Bnull or <function text>dinja 1.2.0, so more chat prompts match llama.cpp: tojsongetPerformanceContext() evalTokens and sampleCount without speculativecreate chunk asLlamaCompletionChunk.usage on native llama.cppLlamaEngine(observers: ...), which reports chat and text completions,leehack/llamadart-native@v0.5.0 (llama.cpp v0.5.0) with Apple companion0.0.20, regenerated matching Dart FFI bindings, refreshed the llamadart_llama_cpp_flutter Apple SwiftPMAdd LlamaEngine.scoreNextToken(...) for next-token log-probabilities on
native llama.cpp and WebGPU bridge assets v0.1.52+, matching llama-server
n_probs; check
supportsNextTokenScoring first
(#694).
Add example/laya_command_bar, a Flutter text field that reshapes into a
reminder, message, calculation or other command as you type, read by a
Laya decision model, by EmbeddingGemma and labelled examples, or by small
LLMs' next-token scores, including the decision model decider-2b.
Throw LlamaModelException when native llama.cpp cannot find or load a
multimodal projector, and LlamaUnsupportedException when the runtime lacks
the mtmd functions; LlamaEngine.supportsAudio also throws the latter.
Speech-to-text capabilities now say when no projector is loaded, using the
new LlamaEngine.hasMultimodalProjector
(#325).
Honour LlamaEngine.cancelGeneration() issued right after listening to a
create, generate or ChatSession.create stream, before it reaches the
backend, instead of running the whole generation
(#602).
Cancel an active text-to-speech synthesis on LlamaEngine.unloadModel() and
dispose() instead of waiting for it to finish
(#628).
Cancel an active Qwen3-ASR transcription on LlamaEngine.unloadModel() and
dispose() instead of completing it with the transcript cut at the unload
(#670).
Send LiteRT-LM tool calls and tool results in the runtime's own message
format, so Gemma 4 reads tool output and Qwen3 tool histories no longer
fail
(#681).
Start a native llama.cpp generation requested right after a cancel once the
cancelled run stops, instead of failing with generation is already in progress. An overlap with a running generation that was not cancelled now
throws LlamaStateException
(#655).
Render Qwen3 prompts as llama.cpp does: an earlier assistant tool-call
turn without reasoning no longer gets an empty <think> block
(#691).
Require dinja 1.1.0. Its Jinja string comparison makes three more chat
templates render as llama.cpp does: MiniMax-M1 adds no empty
system block for an empty or whitespace-only system message; NVIDIA
Nemotron Nano v2 drops the blank line before a tool call, the blank lines
before its tool instructions when tools come with an empty or
whitespace-only system message, and an empty final assistant turn without
a generation prompt; and Functionary v3.2 tool declarations drop stray
// Format=<|NONE|> lines and spell out nested object parameters
(#351).
Cancel a generation's backend run as soon as its stream subscription is
cancelled, instead of at its next token, which during prompt evaluation
meant after the whole prompt. A native llama.cpp generation requested right
after such a cancel now waits for it instead of throwing
LlamaStateException, and native llama.cpp sees a cancel between text
prompt micro-batches (ModelParams.microBatchSize, 512 tokens by default)
or, with speculative decoding, between batches (ModelParams.batchSize)
(#663,
#660).
Render every result of a tool message holding several
LlamaToolResultContent parts, as one tool message per result like
llama.cpp, instead of only the first; LlamaChatMessage.toJson lists them
all (#683).
Render Gemma 4 tool calls and tool results as llama.cpp does, so Gemma 4
GGUF models can read tool output
(#669).
Stop a Qwen3-TTS audio decode at its next chunk boundary when native
text-to-speech is cancelled, instead of finishing the native step in
progress first. This needs llamadart-native v0.4.1-1 or later; older
runtimes keep the previous behaviour
(llamadart-native#86,
#322).
Add an experimental DecisionEngine for Laya-style decision models (a
ModernBERT encoder GGUF plus a safetensors head) on native llama.cpp, with
typed ChoiceKey, ScoreKey and NoulKey questions
(#604).
Add example/basic_app/bin/llamadart_decision_example.dart, a console demo
that triages a support ticket with DecisionEngine
(#604).
Add example/laya_tetris, a Flutter app in which a Laya decision model
plays real-time Tetris through DecisionEngine
(#604).
Run example/laya_tetris on Web through the WebGPU bridge, with a live demo
at https://leehack-flutter-laya-tetris.static.hf.space.
Add a notebook in example/laya_tetris/training/ that fine-tunes a Laya
decision head for the Tetris example and exports it for DecisionEngine
(#604).
Run DecisionEngine on WebGPU through the bridge decision API
(apiVersion 1), which bridge assets v0.1.47+ include
(#604).
Stop native image and audio requests from seeding the repeat penalty with
leftover memory, which made output depend on the previous request
(#603).
After a failed native prompt decode, the next reusePromptPrefix request no
longer runs on the wrong KV cache or keeps failing
(#601).
Reject embed() and embedBatch() on rank-pooled reranker GGUFs with
LlamaUnsupportedException on native, instead of returning memory read
past llama.cpp's classifier-score buffer
(#583).
Throw LlamaInferenceException from native embed() and embedBatch()
when input to an encoder-only model or a model without a KV cache (such as
BERT-family and ModernBERT GGUFs) does not fit one microBatchSize pass,
instead of aborting the process or embedding only the last chunk
(#607).
v0.1.54 for the decision API,supportsCompletionUsage flag, which llamadart does not use yetv0.5.0, arev0.5.0, and keep Web/native llama.cppv0.5.0@7fe450e19305b828c199d602c23a8337aaa1f03b parity and Web@litert-lm/core@0.15.0. Immutable Web asset manifest:8a9f83c15035eeb034a6563e6f753382d7d7f9be81503ef76902138da7841176.LlamaInferenceException whose details contain (invalid grammar), and theTypeError. Later calls fail withNo model loaded. Call loadModelFromUrl first.; call unloadModel(), thenloadModel() and any projector again to recoverGrammar rejected every candidate token onlyjfk.wav. Every GGUFfunctional_pass;stt, tts and litert-asr speech validation pack run from 8ToolChoice.auto on a prompt thatdecision-gguf-{cpu,metal,vulkan,cuda,webgpu} validation profiles thatDecisionEngine token ids, raw logits and answers against the Layachat-gguf-webgpu validation profile, hash Web validationtts unloads and disposes the enginestt must fail withLlamaSpeechTranscriptTruncatedException at maxOutputTokens and at thestt runs now execute 28 checks and tts runs 25topK: 1 for zero-temperature LiteRT-LM Web generation, matchingModel … loaded from …; native engine creation is deferred until the first generation or tokenizer call instead of loaded successfully whenloadModel, since it creates thedxcompiler.dll and dxil.dll in the Windows x64 LiteRT-LM runtimelibgomp.so.1; libgomp1 on Ubuntu/Debian, libgomp on Fedora and Arch)ggml/wrapper symbol is missing fromLlamaEngine.configureLogging handler. A worker takes the Dart logger levelLlamaEngine.setDartLogLevel/setLogLevel update anone sends nothing and debug records arewarn records gated by that level instead of the native logLlamaLogger.level and the BackendDartLogLevel capabilityBearer <token> and the value in token=, key=,secret=, password=, api_key= and apikey=<value> outside HTTP URLs in<redacted-secret>; URL andhook/build.dart intolib/src/hook/native_release_pins.dart, the only file the pin sync nowModelParams.liteRtLmCacheDir to choose the native LiteRT-LM runtimeModelParams.liteRtLmMaxProgramCacheBytes, which*_mldrift_program_cache.bin files above the cap before each enginetopK: 1 for zero-temperatureLiteRtLmRuntimeClient.createConversation calls, which returned incoherentstartupDiagnostics=[...] suffix overflows: teardown entries, now prefixedteardown: , are dropped first, duplicates are recorded once, entries are...Failed to preload Windows backend module startup diagnostic perpromptTemplate on the non-nativeLiteRtLmRuntimeClient.createConversation placeholder, so callers passingTemplateCaps.detect results in a per-isolate LRU keyed by exactsupportsTools and supportsToolCalls for chat templates thatsupportsParallelToolCalls when it throws[TOOL_CALLS] block unlesspattern keyword when generating GBNF, for anchored(...) groups*/+/?/{m,n}minLength andmaxLength still apply there. An applied pattern replacesminLength/maxLength as it does in llama.cpp, so a schema carrying bothpattern and maxLength is no longer length-bounded. A schema carryingpattern but no explicit type now yields a string rule instead ofUnrecognized schema. Mistral Nemo and Magistral tool-call ids areSpeechToTextEngine now fails a native Qwen3-ASR transcript that reachesmaxOutputTokens withLlamaSpeechTranscriptTruncatedException instead of completing withcreate() reports finishReason: 'length' when nativeToolChoice.auto on WebGPU from forcing a tool call: it now skipsGenerationParams.grammarLazy and a non-root grammarRoot withLlamaUnsupportedException; backends report this through the newBackendLazyGrammarSupportlibcudart.so.12, libcublas.so.12)cuda backend needs on the default loader path; llamadartLD_LIBRARY_PATHpackages/llamadart_validation beforeFAILED with a null error on aGpuBackend.metal or GpuBackend.hipMetal and HIP, butMTL and ROCm, so loading fell back to automatic deviceHIPCPUerror 138 as the documentedUnsupportedError, asthread constructor failed already was, instead of rethrowing the rawLlamaModelException when WebGPU cannot fetch or load a multimodal?query and #fragment, where they directly&-separated part that=, as whole tokens; and bare query values and the fragment of 101 of ?v=1, stay. Other URLs in the details losecontent when Qwen2.5 wraps a Hermes<tool_call>{{"name": ...}}</tool_call>, withAlign native leehack/llamadart-native@v0.4.1 on upstream b29c606e28a01b1bc8c1351026a0fa6e616bf6c4 , with matching Dart bindings and Apple companion 0.
leehack/llamadart-native@v0.4.1 on upstreamb29c606e28a01b1bc8c1351026a0fa6e616bf6c4, with matching Dart bindings0.0.19. This resolves the 0.8.23 grammar limitation:{2000} repetitions are accepted againv0.1.44 for matchingv0.4.1@b29c606e28a01b1bc8c1351026a0fa6e616bf6c4 parity.8d61f453753ac7a7d839ac12318b70986a814748d86029993118c19454293aa9.v0.17.0-6 with Apple companion 0.0.11,dxil.dll and dxcompiler.dll, which D3D12 GPU engine@litert-lm/core@0.15.0.Fail Apple builds before native symbol lookup when the resolved llama.cpp companion does not match the core native runtime, with an actionable upgrade
Fail Apple builds before native symbol lookup when the resolved llama.cpp
companion does not match the core native runtime, with an actionable upgrade
diagnostic instead of allowing ABI-incompatible frameworks.
Updated the default llama.cpp native runtime pin to
leehack/llamadart-native@v0.4.0, regenerated matching Dart FFI bindings, refreshed
the llamadart_llama_cpp_flutter Apple SwiftPM checksum, and
aligned current README/website native override docs. Updated multimodal calls
for the matching ABI. Saved native sessions from older runtimes must be
regenerated. Apple companion 0.0.18 supplies the matching native runtime.
Aligned the default WebGPU bridge assets to v0.1.43 for Web/native
llama.cpp v0.4.0@5266f24da75dc449bd56cbed7addb9c8e4a6a73e parity.
Web @litert-lm/core@0.15.0 and native LiteRT pins are unchanged. Immutable
manifest: 111eefc3588842cebfe665b363378edca34924764610263e1eda5280dfcfaa27.
Known upstream limitation: llama.cpp v0.4.0 can reject large grammar
repetitions, such as root ::= "a"{2000}. The post-v0.4.0 upstream
correction is tracked in llamadart-native#76 and is not part of this release.
Updated llamadart_llama_cpp_flutter to 0.0.17 with the Apple SwiftPM v0.3.0 runtime pin.
Updated llamadart_llama_cpp_flutter to 0.0.17 with the
Apple SwiftPM v0.3.0 runtime pin.
Updated the default llama.cpp native runtime pin to
leehack/llamadart-native@v0.3.0, regenerated matching Dart FFI bindings, refreshed
the llamadart_llama_cpp_flutter Apple SwiftPM checksum, and
aligned current README/website native override docs.
Aligned the default WebGPU bridge assets to v0.1.41 for corrected
TypeScript declarations and TTS recovery guidance, retaining Web/native
llama.cpp v0.3.0@c1d0e7a004015f23bc0233470b747b596f29b264 parity and Web
@litert-lm/core@0.15.0. Immutable manifest:
fe97604daabaad6aefa223a8637d5fd9dcac09dd4a61b2ef19cd6aabb39392b9.
Consolidated native release tag grammar across Dart, Python, Bash, workflows,
and documentation via a machine-readable fixture contract (#404).
Patched the website's vulnerable nanoid and uuid dependency paths. Until Docusaurus replaces its unpatched image parser, automatic local Markdown imag…
Aligned the default WebGPU bridge assets to v0.1.39 (immutable manifest
b355d01040604f6ae2c5c5fe5bb42b858101a96f03f67e4b27b32fe41ce3b2bf),
restoring Web/native llama.cpp upstream v0.2.0@bb4caa7540188872173c44d161602d9271386413
parity with native anchor llamadart-native@v0.2.0-1 while preserving
approved Web @litert-lm/core@0.15.0 packaging.
Fixed native Qwen3-ASR transcription by applying the model chat template to
audio turns, while preserving the validated raw-prompt Web bridge contract.
Empty ASR output now fails explicitly instead of reporting an empty result.
Fixed Qwen 2.5/3 LiteRT-LM ToolChoice.required requests silently running
without their required-call grammar and finishing with no call. They now fail
with an actionable LlamaUnsupportedException before generation when the
backend cannot enforce the declared tool schema; Gemma 4 compatibility and
auto/none tool routing are unchanged.
Fixed Gemma 4 thinking-budget output so split channel controls and tool-call
envelopes stay out of visible assistant content while preserving ordinary
whitespace. The chat example now also validates custom tool declarations,
executes declared host handlers exactly once, appends tool results, and
performs a bounded continuation for both llama.cpp and LiteRT-LM backends.
Patched the website's vulnerable nanoid and uuid dependency paths. Until
Docusaurus replaces its unpatched image parser, automatic local Markdown
images are rejected; website contributors should use static pathname URLs.
Fixed llamadart_native_runtimes values none, off, and the string
false selecting every runtime family instead of none;
the build hook fails with its No native runtimes selected error again, as
it did before 0.8.0. A YAML boolean false clears the selection too. Unset,
empty, and all-unrecognised config still select every family.
The published package no longer ships the doc/ directory; that contributor
and maintainer documentation is maintained on GitHub, and the packaged files
that link to it now use absolute URLs.
Fixed unanchored docs/ and website/ publish-exclusions that matched those
directory names at any depth and dropped tool/docs/ plus the
llamadart_server example's OpenAPI spec and Swagger UI sources from the
package, leaving the published example unable to analyze. Both patterns are
now root-anchored.
Narrowed the dinja dependency constraint to >=1.0.0 <1.1.0 so chat
template capability detection cannot silently resolve against an unverified
Jinja parser minor. A 1.0.x patch can still reorganise the private sources
the analyzer imports; a new coupling test turns that into a named failure.
Fixed Command R7B, Hermes, and Hunyuan V3 tool grammars so distinct tool or
parameter names cannot collide after conversion to internal GBNF rule names.
Fixed DeepSeek V3.2 DSML tool calls using their upstream
<|DSML|function_calls> envelope while preserving DeepSeek V4's distinct
<|DSML|tool_calls> grammar and parser behavior.
Fixed partial GLM 4.5, Poolside Laguna, and Muse Glimmer tool envelopes
leaking into streamed assistant content, while preserving completed calls,
malformed final output, and ordinary text surrounding Muse recipient
channels.
Fixed schema-constrained tool calls for Kimi K3, MiniMax M1/M3, DeepSeek
V3.2/V4, and Muse Glimmer, including exact escaped names, required fields,
declared value types, matching MiniMax M3 element tags, zero-argument calls,
and strings containing delimiter characters. Required-tool mode now accepts
each format's reasoning/content prefix while still requiring a call.
MiniMax M3, DeepSeek DSML, Muse Glimmer, Poolside Laguna, and GLM 4.5 now
reconstruct argument values from the declared tool schema instead of
guessing from text. Added ToolParam.nullType for null-only JSON Schema
properties.
llama.cpp backend initialization failures now complete the worker startup
handshake with a typed LlamaBackendInitializationException and collected
native-loader diagnostics. A failed or incompatible worker is torn down
instead of being reported ready and leaving later requests waiting forever.
Added template-aware parsing for Kimi K3, MiniMax M1/M3, DeepSeek V3.2/V4,
Muse Glimmer, and Poolside Laguna, preventing their native tool calls from
silently falling back to plain content.
Fixed XML-style tool-call parsing to honor raw-versus-JSON argument values
and final-value delimiters. Apriel 1.5 and Xiaomi MiMo now parse nested JSON
values without splitting on inner commas, while malformed payloads remain
ordinary assistant content.
Made native video-input capability truthful without claiming end-to-end
support. LlamaVideoContent requests now fail with an actionable
LlamaUnsupportedException, LlamaEngine.supportsVideo reports public
consumability as false, and the llama.cpp worker uses the behavioral
mtmd_helper_support_video result instead of exported helper symbols. The
published v0.2.0-1 archive has not been qualified for end-to-end video;
full path/byte input remains blocked on cross-platform FFmpeg/ffprobe
packaging and Dart frame lifecycle wiring.
Native release synchronization and build-hook overrides now accept stable
vMAJOR.MINOR.PATCH artifacts and ordered vMAJOR.MINOR.PATCH-N wrapper
rebuilds plus nightly bNNNN-N rebuilds, while preserving historical
bNNNN and bNNNN-llamadart.N artifacts. Sync rejects invalid tags,
leading-zero nightly versions, rollback, wrapper/nightly latest results,
incompatible manifest contracts, missing bundles, and release/manifest
checksum or version skew without changing the default pin.
Fixed Web/native backend API parity. WebAutoBackend now forwards grammar
constraint support from its active runtime, so strict structured output fails
early with an actionable error on unsupported Web backends, and the Web-safe
LiteRtLmRuntimeClient stub now exposes the native client's thinking-tag
configuration method.
A failed llama.cpp model load now reports the startup diagnostics collected
during native library discovery, so a missing or unloadable runtime library
explains itself instead of surfacing as a bare load failure. Platforms that
record no diagnostics keep their previous message unchanged.
llama.cpp worker errors now keep their type. Every backend method routes an
ErrorResponse through the file's own error mapper instead of rebuilding a
bare Exception, tokenize and detokenize no longer discard the worker's
error entirely, and a core UnsupportedError raised for an unavailable native
capability is classified rather than flattened. State-file failures now throw
LlamaStateException.
Fixed multimodal media placeholders being normalized inconsistently. <img>,
<|img|>, <start_of_image> and indexed markers such as <|image_1|> are now
rewritten to the mtmd marker on every path, rather than depending on which
layer rendered the prompt. MiniMax-M2 and MiniCPM-5 also now detect a
forced-open thinking block using the same rule as every other handler.
aLoRA adapters are now rejected with LlamaUnsupportedException instead of
being applied like ordinary LoRA adapters. An aLoRA adapter must activate only
after its invocation tokens appear in the prompt, so applying it from the
start of generation silently changed output. Missing metadata-inspection
symbols in custom native runtimes also fail closed with the same typed error,
and rejected adapters are released when the cleanup ABI is available. LoRA
errors from the worker keep their typed exception instead of arriving as a
bare Exception, and a failed adapter load now throws
LlamaModelException.
Deprecated LiteRtLmRuntimeClient.conversationTokenCount() and
replaceConversationWithClone(). Both are unused and are scheduled for
removal in the next major release; open an issue if you depend on either.
NativeLlamaBackend.modelLoadFromUrl now throws LlamaUnsupportedException
instead of UnimplementedError, bringing it into the LlamaException
hierarchy. It is the same exception type LlamaEngine.loadModelFromUrl
already throws for this condition; each keeps its own message.
Updated the default llama.cpp native runtime to the immutable
leehack/llamadart-native@v0.2.0-1 release (llama.cpp v0.2.0), adding LFM2
DSpark support plus current upstream correctness and backend performance
fixes. Matching Dart FFI bindings, including the new multimodal
projector-device field, and the Apple SwiftPM artifact checksum were
refreshed. Linux libmtmd.so.0 now loads without the old
libmtmd.so.SOVERSION compatibility alias.
Removed the abandoned Dart-side MTP/n-gram speculative-decoding scaffolding
from the llama.cpp backend. Speculative decoding is unchanged: it continues to
run through the native wrapper, and SpeculativeDecodingStrategy keeps every
existing option.
Bumped llamadart_llama_cpp_flutter to 0.0.16 so the v0.2.0-1 Apple
SwiftPM pin actually publishes; 0.0.14 was already on pub.dev, so release
automation skipped it and Apple builds would have kept the b10514 runtime.
Corrected the WebGPU bridge docs, which claimed the pinned v0.1.37 bridge
assets match the default native llama.cpp runtime. They embed b10514 and now
trail the native v0.2.0-1 pin.
Chat-template capability detection now logs a debug message naming the
probe (string-content, typed-content, system-role, tools) when its
render throws, so a template that fails to render is distinguishable from
one that genuinely lacks the capability.
Added code_assets 2.x compatibility while retaining support for 1.x native asset toolchains.
Added code_assets 2.x compatibility while retaining support for 1.x native
asset toolchains.
Updated the default llama.cpp native runtime pin to
leehack/llamadart-native@b10514, adding BailingMoE3,
GraniteSWA/GraniteMoeSWA, speculators-format DSpark checkpoints, and current
upstream multimodal/backend improvements. Matching Dart FFI bindings and the
llamadart_llama_cpp_flutter Apple SwiftPM artifact were refreshed.
Updated the default WebGPU bridge assets to v0.1.37, embedding llama.cpp
b10514 to restore native/Web parity. The bridge provisions an explicit 1 MiB
Wasm stack for wasm32 and memory64, preventing Qwen3-ASR memory64 context
construction from overflowing the default stack.
Improved Web microphone transcription by warming up browser capture before
showing the recording-ready state, trimming the warmup silence, and
rejecting too-short, silent, or unsupported PCM WAV captures before
inference.
Added logical batch-size (n_batch) and micro-batch-size (n_ubatch)
controls for llama.cpp/WebGPU models in the Flutter chat example.
Improved Android Auto backend selection by probing the packaged Vulkan device
before memory planning, avoiding unnecessary CPU fallback on capable models.
Updated the native LiteRT-LM runtime to v0.16.0-native.2. The Apple companion
packages the iOS Gemma constraint provider and Metal accelerator/sampler
plugins required by the published runtime.
Added an experimental SpeechToTextEngine.liteRtLm path with capability
discovery, worker-isolated CPU inference, bounded mono 16 kHz float PCM,
partial/final transcript events, finalization, and cancellation.
Added experimental live dictation to the native Flutter chat example for
chat models using selectable Moonshine Tiny and Parakeet TDT 0.6B sidecars.
Live dictation is CPU-only, English-only, capped at five minutes, and
unavailable on Linux and Web.
Improved Flutter chat example model downloads with bounded retries for
transient network failures, safe resume after truncated responses, and a
distinct integrity-verification state after transfer reaches 100%.
Redesigned the Flutter chat example onboarding and Lab surfaces, preserved a
completed model card's viewport position when downloads reorder the catalog,
and stopped streaming responses from pulling users away from chat history.
Added an experimental typed TextToSpeechEngine for native llama.cpp and
WebGPU with Qwen3-TTS models, including capability discovery, speaker
references, cancellable progress, complete 24 kHz PCM output, and WAV
encoding. The Flutter chat example adds synthesis, playback, replay, and WAV
save controls. Apple apps discover the TTS ABI in the embedded llama
framework; current LiteRT-LM artifacts fail explicitly as unsupported.
Added an experimental typed SpeechToTextEngine with an explicit Qwen3-ASR
adapter profile for whole-file llama.cpp transcription. The Flutter chat
example includes a checksum-pinned Qwen3-ASR 0.6B preset plus file and
microphone transcription on native and WebGPU. Web accepts WAV bytes only;
native LiteRT-LM live dictation remains a separate implementation.
Added Ask with voice to the native Flutter chat example for Gemma 4 E2B
through LiteRT-LM direct media and audio-capable GGUF + projector paths. It
sends a microphone recording through multimodal chat and remains separate
from typed speech-to-text.
Added experimental, opt-in native llama.cpp DSpark speculative decoding
through SpeculativeDecodingConfig.draftDspark(...) with a compatible
external draft model.
Updated the default llama.cpp native runtime pin to leehack/llamadart-native@b10333 , regenerated matching Dart FFI bindings, refreshed the llamadart_
Updated the default llama.cpp native runtime pin to
leehack/llamadart-native@b10333, regenerated matching Dart FFI bindings,
refreshed the llamadart_llama_cpp_flutter Apple SwiftPM checksum, and aligned
current README/website native override docs.
Updated WebGPU bridge assets to v0.1.27 (llama.cpp b10333), keeping the
native and Web GGUF runtimes on the same upstream revision.
Fixed corrupt Qwen3.5 output on Android Vulkan by preserving the KQV
offload required for correct hybrid model inference while retaining the
remaining conservative Android context settings.
Updated the default llama.cpp native runtime pin to leehack/llamadart-native@b10276 , picking up Qwen3-TTS model-loading primitives, explicit bundled-
Updated the default llama.cpp native runtime pin to
leehack/llamadart-native@b10276, picking up Qwen3-TTS model-loading
primitives, explicit bundled-MTP loading, automatic model-specific token
suppression, and recent model, multimodal, speculative-decoding, backend,
and runtime fixes. Regenerated matching Dart FFI bindings, migrated model
loading to llama.cpp's load-mode ABI and the penalty sampler to the new
vocabulary-sized ABI, and refreshed the llamadart_llama_cpp_flutter Apple
SwiftPM checksum. Speech generation is not yet exposed through the public
Dart API.
Updated the default LiteRT-LM runtimes to
leehack/litert-lm-native@v0.15.0-native.3 and
@litert-lm/core@0.15.0. The native artifact includes a corrected v0.15
streaming callback bridge and an Android Dawn rollback for Mali-G715 GPU
device loss; incompatible callback runtimes now fail safely before
generation, and concrete macOS app, framework, and cache libraries take
precedence over process-linked assets.
Disabled automatic WebGPU fetch-backed model loading by default. Streamed
loading remains the safe default; controlled range-capable deployments can
opt in explicitly.
Updated WebGPU bridge assets to v0.1.26 (llama.cpp b10276), refreshing
both WebAssembly runtimes while preserving the existing bridge API.
Updated the default llama.cpp native runtime to leehack/llamadart-native@b10075 with matching bindings and Apple artifacts.
Updated the default llama.cpp native runtime to
leehack/llamadart-native@b10075 with matching bindings and Apple artifacts.
Added Tencent Hunyuan V3 chat-template, reasoning, and tool-call support.
Fixed Gemma 4 LiteRT-LM text generation in the Web chat app after model
loading completed successfully.
Restored GGUF loading in deployed Web chat apps by packaging the pinned
WebGPU runtime assets with Flutter Web builds.
Updated the default llama.cpp native runtime to leehack/llamadart-native@b9982, including regenerated bindings and safer multimodal UTF-8 prompt handl
Updated the default llama.cpp native runtime to
leehack/llamadart-native@b9982, including regenerated bindings and safer
multimodal UTF-8 prompt handling.
Improved llama.cpp batching defaults and ChatSession context management for
more predictable generation under constrained contexts.
Added llama.cpp presencePenalty sampling and thinking-budget controls,
including the server's thinking_budget_tokens extension. Unsupported
WebGPU and LiteRT-LM paths now fail explicitly.
Reworked the runnable TUI coding agent with a focused Pi-style workflow, Unsloth Qwen3.6 defaults, shared model-source loading, and clearer reasoning and final-answer presentation.
Improved the OpenAI-compatible server with standard client-managed tool-call
transcripts, named tool_choice, configurable thinking behavior, and shared
model-source loading.
Added clipboard media attachments to the runnable chat app. Desktop and web users can paste screenshots or copied image/audio files with Cmd/Ctrl+V, m
Added clipboard media attachments to the runnable chat app. Desktop and web
users can paste screenshots or copied image/audio files with Cmd/Ctrl+V,
mobile users can choose Paste attachment, and ordinary text paste remains
unchanged.
Restored the runnable macOS chat app build phase that embeds and signs
LiteRT-LM runtime libraries inside the sandboxed app bundle, and enabled
LiteRT-LM Metal selection on iOS with the consolidated upstream runtime.
Updated the default LiteRT-LM runtime to
leehack/litert-lm-native@v0.14.0-native.2, which fixes Android GPU plugin
symbol resolution and uses the checksum-pinned official Apple XCFrameworks.
Fixed native LiteRT-LM generation incorrectly treating the requested maximum response length as a forced benchmark decode count. Short responses no longer wait for every allowed token before streaming, and LiteRT-LM chat flushes its first token immediately.
Added an app-owned FIFO model-download queue to the runnable chat app. Only one model transfers at a time, queued cards show their position and can leave the queue independently, and a responsive shell progress pill keeps the active download visible after the settings panel closes.
Replaced the runnable chat app's broad built-in model catalog with a focused Unsloth-first set. Added cross-platform Gemma 4 E4B plus native-desktop Gemma 4 12B/26B-A4B/31B and Qwen3.6 35B-A3B presets. The model library now promotes downloaded models, supports name/capability search and Mobile & Web/Desktop filters, explains incompatible choices, and uses quieter cards with compact compatibility, capability, and recommended-setting summaries. Custom entries retain independent remove-from-library and downloaded-file actions.
Enabled Gemma 4 audio attachments in the runnable chat app for the current
native GGUF projector and LiteRT-LM bundle, while keeping LiteRT-LM Web
correctly text-only. Model capability settings now distinguish direct media
input from external mmproj input and persist that distinction across app
launches.
Redesigned the runnable chat app with a quieter responsive shell, compact runtime details, full-screen mobile settings, streamlined message/composer surfaces, accessible controls, protected conversation deletion, and copy/regenerate actions for assistant responses. Model downloads now remain active when settings closes or switches between drawer and pinned layouts.
Hardened the runnable chat app for narrow windows, 200% text scaling, and macOS assistive technology; removed stale assistant placeholders after empty or failed generations; kept context accounting stable when regenerating a response; and redacted signed model URL parameters from labels and generation errors.
Clarified the chat app's backend and GPU-layer controls, exposed the active loaded backend beside its preference, and restored GPU offload when switching from a zero-layer CPU configuration back to Auto or a GPU backend. Native Max now maps to full llama.cpp offload, while Auto uses model size, available device memory, safe system headroom, and requested context to choose full or partial offload and only reduces context when the model does not fit. Auto intent now persists separately from resolved layer/context values so device headroom is recalculated on every model load and after app restarts.
Updated the runnable chat app's web runtimes to pinned WebGPU bridge assets v0.1.18 (llama.cpp b9915) and @litert-lm/core@0.14.0, keeping hosted and l
Updated the runnable chat app's web runtimes to pinned WebGPU bridge assets
v0.1.18 (llama.cpp b9915) and @litert-lm/core@0.14.0, keeping hosted
and local inference on reproducible runtime versions.
Improved the runnable chat app's model-cache UX by reusing cached GGUF and
LiteRT-LM web bundles for benign catalog URLs such as ?download=true,
preserving browser model caches during app cache cleanup, adding a text-only
projector skip path, and polishing download/load progress states.
Updated the default llama.cpp native runtime pin to
leehack/llamadart-native@b9935, regenerated matching Dart FFI bindings, refreshed
the llamadart_llama_cpp_flutter Apple SwiftPM checksum, and
aligned current README/website native override docs.
Fixed load lifecycle guards so repeated model loads preserve the active model state, URL load unsupported-runtime diagnostics stay typed, and unload c
Fixed load lifecycle guards so repeated model loads preserve the active model state, URL load unsupported-runtime diagnostics stay typed, and unload cancels active generation before freeing llama.cpp handles.
Tightened LiteRT-LM runtime validation and local smoke coverage by requiring complete macOS arm64 runtime caches, requiring an explicit model path for the LiteRT-LM chat feature smoke scenario, normalizing the chat app's LiteRT-LM auto context size, and adding the missing iOS-compatible SwiftPM Gemma provider target. Flutter macOS LiteRT-LM companion-package builds now fall back to hook-managed native assets, and the companion package no longer links the incomplete macOS LiteRT-LM SwiftPM artifact set.
Reworked the README into a shorter entry point, fixed stale docs/examples found during the documentation review, and aligned release, Android smoke, WebGPU mem64, native sync, and capability-support wording with the current workflows and runtime behavior. WebGPU runtime LoRA calls now throw an unsupported-operation error instead of reporting no-op success.
Updated the default llama.cpp native runtime pin to
leehack/llamadart-native@b9873-llamadart.2, keeping the b9873 llama.cpp
ABI/bindings while picking up wrapper fixes for native release
provenance and backend-selected speculative sampler acceptance. Refreshed
the llamadart_llama_cpp_flutter Apple SwiftPM checksum and aligned current
README/website native override docs.
Aligned llama.cpp n-gram speculative rejected-tail rollback with upstream
server behavior by trimming rejected target tokens before falling back to
checkpoint replay, avoiding early token-count/EOG drift in accepted-draft
ngram-map-k runs. The speculative benchmark can now emit full generated
text with --include-output for parity investigations.
Fixed llama.cpp n-gram speculative configuration mapping so
draftTokenMax no longer implicitly overrides upstream ngramSizeM, and
documented upstream comparison commands plus measured n-gram benchmark
results.
Added a durable flag-based llama.cpp speculative benchmark runner and wired
it into the local E2E scenario list, testing matrix, README, and website
performance guide. The runner can now generate a llama.cpp-compatible
static n-gram cache file for ngram-cache E2E validation, and renders
benchmark prompts with configured or loaded GGUF chat templates before the
generic fallback instead of silently falling back to a hard-coded Gemma
prompt.
Documented compatible DFlash GGUF metadata, a known-good public target/draft
model pair, and troubleshooting guidance for incompatible dflash-draft or
missing dflash.target_layers artifacts.
Hardened generic llama.cpp external draft-model speculative decoding so
draft-context processing does not request unused logits, avoiding a
draft-simple native abort while preserving target verification logits.
Extended llama.cpp speculative performance diagnostics with draft-attempt, target-verification-token, and replay-token counters. Enabled speculative runs now preserve explicit zero counters instead of collapsing them to null, and local benchmark JSON includes the new fields.
Hardened LiteRT-LM generation validation so llama.cpp-only speculative decoding knobs fail loudly instead of silently degrading to LiteRT-LM's boolean speculative toggle.
Added llama.cpp upstream speculative decoding parity through
SpeculativeDecodingConfig constructors for draft-simple, EAGLE3, MTP,
DFlash, ngram-simple, ngram-map-k, ngram-map-k4v, ngram-mod, ngram-cache,
and mixed n-gram plus one draft-model strategy, including generic native
wrapper bindings, docs, and local benchmark matrix coverage.
Added LlamaStructuredOutput and LlamaEngine.createStructuredJson(...)
helpers for strict JSON-object / JSON-schema generation with final-output
validation and typed decoding.
Added LlamaEngine.loadMultimodalProjectorSource(...) so GGUF
multimodal projector files can use the same ModelSource resolver,
native download/cache manager, authentication, checksum, and progress
options as loadModelSource(...), while preserving the existing
loadMultimodalProjector(...) path/string API.
Improved the runnable chat app's Manage Models cache UX so model and mmproj asset cache states are shown separately, missing multimodal projectors can be re-cached without re-fetching already cached model assets, and runtime media capability mismatches surface as user-readable warnings. Custom signed or tokenized Hugging Face URLs now require confirmation before they are saved.
Updated the default LiteRT-LM native runtime pin to leehack/litert-lm-native@v0.14.0-native.1, refreshed matching native-assets checksums, aligned the
Updated the default LiteRT-LM native runtime pin to
leehack/litert-lm-native@v0.14.0-native.1, refreshed matching native-assets
checksums, aligned the llamadart_litert_lm_flutter Apple SwiftPM checksums,
and documented the newly exposed native LiteRT-LM load/generation controls:
per-request max output tokens, native sampler params, thread count, one
default-scale initial text LoRA adapter, activation data type, prefill chunk
size, parallel file-section loading, and Android LiteRT dispatch library
directory.
Hardened LiteRT-LM runtime packaging and local runtime preparation for the 0.14 line, including Linux/Windows runtime dependency resolution, Android Dawn companion libraries, Apple SwiftPM checksums, and macOS fallback app runtime copying.
Hardened release automation by adding CODEOWNERS coverage for publication-sensitive files and making pub.dev/GitHub Release propagation waits configurable with longer defaults.
Added post-merge release automation so a merged release-prep PR can publish missing companion package versions, push the core release tag, wait for pub.dev, and confirm the GitHub Release without a separate manual tag step.
Added LlamaEngine.getModelFileType() for llama.cpp/GGUF models, exposing
native model file type / quantization metadata from llama_model_ftype and
llama_ftype_name when available.
Refreshed llama.cpp b9860 template parity for DeepSeek V4 and MiniCPM5,
including MiniCPM5 XML tool-call detection, rendering, and parsing.
Updated the default llama.cpp native runtime pin to
leehack/llamadart-native@b9860, regenerated matching Dart FFI bindings, refreshed
the llamadart_llama_cpp_flutter Apple SwiftPM checksum, and
aligned current README/website native override docs.
Updated the default llama.cpp native runtime pin to leehack/llamadart-native@b9829, refreshed the llamadart_llama_cpp_flutter Apple SwiftPM checksum,
leehack/llamadart-native@b9829, refreshed the
llamadart_llama_cpp_flutter Apple SwiftPM checksum, and aligned current
README/website native override docs for the release.Potentially breaking behavior change: native model cache defaults changed without breaking Dart source compatibility. DefaultModelDownloadManager() no
Potentially breaking behavior change: native model cache defaults changed
without breaking Dart source compatibility. DefaultModelDownloadManager() now
prefers the platform shared cache on desktop/server instead of the process temp
directory, and mobile DefaultModelDownloadManager.auto() without an explicit
app-private directory now uses a best-effort temporary/cache fallback instead
of throwing. Apps or tests that asserted the old temp path or mobile exception
should pass an explicit cache directory or follow MIGRATION.md.
Added optional androidAppPrivateCacheDirectory and
iosAppPrivateCacheDirectory arguments to
DefaultModelDownloadManager.auto(...) so apps can provide platform-specific
mobile cache roots without constructor-level Platform.isAndroid /
Platform.isIOS branching.
Updated the default native DefaultModelDownloadManager() constructor to use
the per-user shared model cache on desktop/server platforms and the mobile
app-private cache fallback, so plain LlamaEngine(...) remote source loads use
a platform-appropriate default while preserving a temporary fallback for hosts
that cannot expose a desktop cache environment.
Broadened the hooks dependency constraint to support both the existing build-hooks package family and the latest stable release, restoring the pub.dev
Broadened the hooks dependency constraint to support both the existing
build-hooks package family and the latest stable release, restoring the
pub.dev dependency freshness score without breaking downstream packages that
still resolve hooks 1.x.
Made web-safe backend stubs the default conditional import/export targets,
preserving native dart:io selection while avoiding false WASM compatibility
deductions in pub.dev analysis.
Added a CI release-doc version consistency check so current README/website install snippets and companion package READMEs stay aligned with package pu
Added a CI release-doc version consistency check so current README/website
install snippets and companion package READMEs stay aligned with package
pubspec.yaml versions, and documented that companion/core package publishing
happens only after release-prep merge with explicit maintainer approval for
each tag.
Updated the default llama.cpp native runtime pin to
leehack/llamadart-native@b9803, regenerated matching Dart FFI bindings, refreshed
the llamadart_llama_cpp_flutter Apple SwiftPM checksum, and
aligned current README/website native override docs.
Added DefaultModelDownloadManager.auto(...) plus explicit model cache root
constructors for shared desktop caches, app-private mobile caches,
user-selected model libraries, and App Group containers. Implicit shared cache
resolution now fails loudly on mobile and web where the OS cannot provide a
hidden cross-developer model folder.
Fixed multimodal chat-template rendering so templates that force-open reasoning (for example Qwen3.5 VLM prompts ending with ) preserve enable_thinkin
<think>) preserve
enable_thinking and stream generated reasoning through delta.thinking
instead of delta.content.Updated the default llama.cpp native runtime pin to leehack/llamadart-native@b9776, regenerated matching Dart FFI bindings, refreshed the llamadart_ll
leehack/llamadart-native@b9776, regenerated matching Dart FFI bindings,
refreshed the llamadart_llama_cpp_flutter Apple SwiftPM checksum, and
aligned current README/website native override docs.Fixed the split-library mtmd fallback ABI for image and byte-buffer multimodal inputs so Windows mtmd.dll and other split mtmd native bundles use the
mtmd.dll and other split mtmd native bundles
use the same bitmap helper signature as the generated native binding path.
This avoids corrupting the first mtmd bitmap-helper call for Gemma 4/MMProj
style multimodal loads and adds native symbol regression coverage for the
fallback ABI.Updated the default llama.cpp native runtime pin to leehack/llamadart-native@b9744, regenerated matching Dart FFI bindings, refreshed the llamadart_ll
Updated the default llama.cpp native runtime pin to
leehack/llamadart-native@b9744, regenerated matching Dart FFI bindings,
refreshed the llamadart_llama_cpp_flutter Apple SwiftPM checksum, and
aligned current README/website native override docs.
Expanded llama.cpp chat-template parity coverage for the latest upstream
template fixtures, including Cohere2 MoE, LFM2.5 tool-call, and Granite 4.1
templates. LFM2.5 templates that use plain List of tools: [...] prompts
with <|tool_call_start|> / <|tool_call_end|> now route through the LFM2
handler like upstream llama.cpp, and ToolChoice.required now uses
grammar-constrained LFM2 tool-call generation.
Fixed streaming tool-call parsing so partial North/Cohere bare action arrays are not emitted as content before the complete tool call is parsed.
Expanded the local GGUF chat feature smoke to cover thinking, tool-call, and optional multimodal turns through the unified local E2E runner.
Fixed Windows CUDA backend discovery when the native asset bundle directory is not on the app PATH. llama.cpp backend modules are now loaded from thei
PATH. llama.cpp backend modules are now loaded from their
resolved bundle path in a way that lets colocated CUDA redistributables such
as cudart64_12.dll, cublas64_12.dll, and cublasLt64_12.dll resolve
correctly.Updated the default llama.cpp native runtime pin to leehack/llamadart-native@b9694, regenerated matching Dart FFI bindings, refreshed the llamadart_ll
Updated the default llama.cpp native runtime pin to
leehack/llamadart-native@b9694, regenerated matching Dart FFI bindings,
refreshed the llamadart_llama_cpp_flutter Apple SwiftPM checksum, and
updated the default WebGPU bridge asset pin to
leehack/llama-web-bridge-assets@v0.1.17 (llama.cpp b9699). The WebGPU
backend now caps unset large-model browser batches so Gemma 4 mem64 loads do
not fall back to context-sized compute buffers.
Added BackendGpuEnumeration.listGpuDevices({probeBackends}) (exposed via
LlamaEngine.listGpuDevices) to enumerate GPU-class devices for offload
selection — backend, per-backend mainGpu index, name, description, device
id, type, and free/total memory per device. With an empty probeBackends
only already-registered backends are inspected, so an unsupported GPU runtime
cannot crash the process during enumeration; pass specific backends to opt
into loading just those modules first. Web/WebGPU return an empty list.
Added Cohere2 MoE / North Code chat-template detection and parsing so
<|START_TEXT|> responses and <|START_ACTION|> tool-call arrays are
handled separately from older Command-R templates.
Fixed docs references that still pointed at llamadart_litert_lm_flutter 0.0.1 and the pre-native.1 LiteRT-LM release after the 0.8.0 native pin sync m
llamadart_litert_lm_flutter 0.0.1 and
the pre-native.1 LiteRT-LM release after the 0.8.0 native pin sync moved
LiteRT-LM Apple/runtime artifacts to v0.13.1-native.1..litertlm image/audio chat parts through LiteRT-LM
Conversation message JSON so bundles with native media processors can accept
LlamaImageContent / LlamaAudioContent path and encoded-byte inputs without
a separate mmproj projector.Added responseFormat routing to LlamaEngine.create(...) for grammar-capable backends, deprecated the legacy chatTemplate(...) jsonSchema shortcut, and…
llamadart_llama_cpp_flutter for GGUF/llama.cpp and
llamadart_litert_lm_flutter for .litertlm/LiteRT-LM. These companion
packages live under packages/ in this repository and publish as separate
pub.dev packages.llamadart so pure Dart/native-assets
consumers can keep using the core package without taking a Flutter SDK
constraint.0.0.1; native pin sync bumps only the
affected companion package patch version. Companion package publishing uses
package-specific tags after the first manual pub.dev publish, and skips
companion versions that already exist on pub.dev.llamadart_native_runtimes to include all available
runtime families. For Flutter iOS/macOS app builds, installed companion
packages decide Apple SPM runtimes; for every other build,
llamadart_native_runtimes remains the selector.leehack/llamadart-native@b9587, regenerated matching Dart FFI bindings,
and refreshed the llamadart_llama_cpp_flutter Apple SwiftPM checksum.SpeculativeDecodingConfig.mtp(draftModelPath: ...) for llama.cpp
external draft-model MTP sessions, with draft model caching and cleanup
tied to the target model lifetime.responseFormat routing to LlamaEngine.create(...) for
grammar-capable backends, deprecated the legacy chatTemplate(...)
jsonSchema shortcut, and made strict response-format requests fail early
on LiteRT-LM instead of silently degrading to unconstrained generation..litertlm text chat through LiteRT-LM Conversation
APIs so structured history, system messages, tool declarations, and
template extra context reach the runtime without a Dart-rendered prompt.
Unsupported cases still fall back to the existing Dart chat-template path..litertlm ModelParams for
liteRtLmActivationDataType, liteRtLmPrefillChunkSize,
liteRtLmParallelFileSectionLoading, and liteRtLmDispatchLibDir,
forwarding the pinned LiteRT-LM v0.13.1-native.1 engine-settings C APIs
while keeping defaults unchanged.Added explicit pub.dev platform metadata for Android, iOS, Linux, macOS, web, and Windows. This keeps the package listing aligned with the actual cros
Compatibility note: no Dart API breaking changes. Flutter Apple apps must target iOS 16.4/macOS 14.0 or newer, and non-Android native apps that ship .…
leehack/llamadart-native and leehack/litert-lm-native
XCFramework artifacts through darwin/llamadart/Package.swift.MinimumOSVersion mismatches in App
Store uploads. Standalone Dart macOS keeps the native-assets dylib fallback.llama_cpp and litert_lm by
default; iOS, macOS, Linux, and Windows now default to llama_cpp only.SpeculativeDecodingConfig as a backend-neutral generation option for
selecting speculative decoding strategies such as MTP while keeping the
existing GenerationParams.speculativeDecoding flag as a compatibility
switch.SpeculativeDecodingConfig.mtp(...), defaulting to a conservative
one-token draft depth unless callers tune draftTokenMax.leehack/llamadart-native@b9547, including the MTP wrapper exports and
llama-common runtime packaging.ModelParams.speculativeRollbackTokenMax so llama.cpp contexts can
reserve recurrent-state rollback snapshots required by Qwen3.5 MTP-style
models.draft-mtp backend-sampling path can abort with vk::DeviceLostError;
CPU and other supported backends remain available, and a dart-define debug
override is available for reproductions..litertlm models should opt in with llamadart_native_runtimes.Compatibility note: no public API breaking changes for existing GGUF / llama.cpp callers. LiteRT-LM support is additive, with deprecated benchmark wra…
.litertlm routing through LlamaBackend() on native and
web targets, with native bundle downloads from leehack/litert-lm-native
and web loading through @litert-lm/core.ModelParams.liteRtLmBackend so callers can select LiteRT-LM CPU,
GPU, or Android NPU execution where supported. auto chooses GPU on
Android/macOS and CPU elsewhere on native targets..litertlm loading through
loadModelSource(...), preserving the selected LiteRT-LM backend after the
cache manager resolves the local file.ChatSession token counting support.hooks.user_defines.llamadart.llamadart_native_tag,
llamadart_native_repository, and llamadart_native_path so apps can test
a different compatible native runtime source without patching llamadart..dart_tool/lib is unavailable.GenerationParams.speculativeDecoding for native LiteRT-LM. The
default remains disabled; llama.cpp, WebGPU, and LiteRT-LM web reject the
option until their speculative paths are implemented..litertlm thinking and tool calling by replacing the stub
template with the canonical Gemma 4 chat template, parsing the runtime
thought channel as reasoning, and suppressing reasoning deltas when callers
set enableThinking: false..litertlm chat-template registry seeded with
Gemma 4/3/3n and Qwen 2.5/3. Pass ModelParams.chatTemplate to override
detection for other models.supportsGrammarConstraints from the active NativeAutoBackend delegate.<|tool_call> markers as assistant content before the final
tool_calls chunk.ModelParams.preferMemory64 and ModelParams.modelBytesHint so
large WebGPU GGUF models such as Gemma 4 E2B can choose the 64-bit bridge
core before hitting the wasm32 address-space limit..litertlm chat-app turns by swallowing unsupported token-count
refreshes, avoiding unsupported minP/penalty parameters for LiteRT-LM
web generation, and replacing the stuck "Loading model 0%" label with an
indeterminate load message..litertlm load time by skipping WebGPU CacheStorage prefetch
for LiteRT-LM models, which are fetched directly by @litert-lm/core.window.__llamadartBridgeReadyPromise, requiring the bridge
prefetch API, and surfacing actionable errors for old bridge assets.?download=true URLs to be prefetched into the
browser cache while still skipping credentialed or signed URLs..litertlm loading by resolving embedded LiteRtLm and
StreamProxy frameworks from the app bundle, matching the macOS runtime
path behavior.ChatSession now forwards empty-choices completion chunks instead of
throwing, strips multiple <think> blocks, and trims history only on
user-message turn boundaries.LlamaEngine.generate wraps unexpected backend errors in
LlamaInferenceException so callers catching LlamaException see the
documented error type.$refs nested inside other
$ref targets and fails loudly on unresolvable or external $refs.minItems/maxItems, model downloads
use connection and idle-read timeouts, and partial-download resume is
restricted to files with stored validators.tool/gguf_chat_features_smoke.dart and the
chat-app-web-gemma4-webgpu-smoke E2E scenario for real-model parser and
WebGPU mem64 validation.doc/litert_lm_templates.md for backend
selection, platform/runtime support, package-size controls, benchmark
results, model templates, and current LiteRT-LM capability limits..litertlm loads instead of being silently ignored.Compatibility note: no public API breaking changes in 0.6.17; existing 0.6.16 callers remain compatible. The release only refreshes the pinned native…
leehack/llamadart-native@b9371, picking up llama.cpp b9371.MTLLibraryErrorDomain Code=3.0.6.17;
existing 0.6.16 callers remain compatible. The release only refreshes
the pinned native runtime and generated low-level bindings.Compatibility note: no public API breaking changes in 0.6.16; existing 0.6.15 callers remain compatible. The changes improve native VRAM diagnostics,…
getVramInfo() so it reports free/total VRAM from
llama.cpp GPU-class backend devices when available, using props-based
memory reporting first and the legacy memory probe as a fallback.CacheStorage failures fall back to direct network loading, and
credentialed/signed model URLs skip persistent browser cache storage.0.6.16;
existing 0.6.15 callers remain compatible. The changes improve native VRAM
diagnostics, WebGPU browser recovery, and chat app download lifecycle
behavior.Compatibility note: no public API breaking changes in 0.6.15; existing 0.6.14 callers remain compatible. The chat-template changes fix multimodal seri…
tool/testing/run_local_e2e.dart as a discovery and orchestration
entry point for heavyweight local-only Dart E2E, Flutter device, and
Web/Playwright smoke scenarios.test-chat server/mtmd build requirements.--list and
--dry-run first.0.6.15;
existing 0.6.14 callers remain compatible. The chat-template changes fix
multimodal serialization behavior for affected templates, and the local E2E
runner is additive.Compatibility note: no public API breaking changes in 0.6.14; the WebGPU bridge asset update and ModelDownloadController are additive, and existing 0.…
leehack/llama-web-bridge-assets@v0.1.16 (llama.cpp b9165),
picking up the published JS bridge build, TypeScript declaration asset,
and refreshed bridge docs.ModelDownloadController, a dependency-free helper that turns
ModelDownloadManager cache/download work into app-facing lifecycle states
for resolving, cache checks, downloads, verification, ready, failed,
cancelled, and retry flows.ModelDownloadManager adapter
so its model-management UI demonstrates the controller while preserving the
example's multi-asset and web-cache service behavior.0.6.14;
the WebGPU bridge asset update and ModelDownloadController are additive, and
existing 0.6.13 callers remain compatible.Compatibility note: no public API breaking changes in 0.6.13; existing loadModel(...) callers are unchanged. Code that probes state persistence suppor…
ModelSource for local paths, HTTP(S) URLs, and Hugging Face
hf://owner/repo/path/to/model.gguf references, including deterministic
cache keys and redacted metadata/log identities for signed URLs.ModelLoadOptions, ModelCachePolicy, resolver targets, and
download/cache metadata/progress value models for package-managed model
download and cache management.DefaultModelDownloadManager support for streaming
HTTP downloads, .part files, atomic promotion, persisted metadata,
authenticated bearer/custom headers, cancellation, retry, Range resume,
cache hit/refresh/cache-only/no-cache policies, SHA-256 verification,
cache listing, removal, clearing, and age/size pruning.hf:// references now accept
?revision=... for branch/ref names containing slashes, and docs clarify
current single-file behavior, private/gated bearer-token usage, separate
mmproj asset handling, sharded-GGUF limitations, and redaction guarantees..part
files or metadata, while distinct cache entries can still download in
parallel and waiting-caller cancellation does not cancel the active download.ModelSource.path(...) option semantics: local paths now reject
remote/download-only options (non-default cache policies, cache directories,
authenticated headers, resume, and retry overrides) while continuing to
support cancellation and optional local SHA-256 verification.LlamaEngine.loadModelSource(...) to route local sources through the
existing native local loader, remote sources through the native download
cache before local loading, and simple remote sources through URL-capable web
backends when available.LlamaEngine.supportsStatePersistence,
LlamaEngine.stateSaveFile(...), and
LlamaEngine.stateLoadFile(...) so callers can persist and restore
llama.cpp KV-cache state for fast raw-prompt resume/fork workflows.BackendStatePersistence, BackendStatePersistenceSupport, and
StateLoadResult for custom backend implementers and diagnostics.ChatSession
message history must be persisted separately.v0.1.15+,
including Dart JS interop, backend forwarding, and browser integration test
coverage.0.6.13;
existing loadModel(...) callers are unchanged. Code that probes state
persistence support should prefer LlamaEngine.supportsStatePersistence over
structural backend type checks so web/router backends can report
bridge-version-dependent support accurately.Compatibility note: no public API breaking changes in 0.6.12.
leehack/llamadart-native@b9016,
picking up the CUDA 12.8 Blackwell-capable native bundles.leehack/llama-web-bridge-assets@v0.1.14 (llama.cpp b9016) so
native and web runtimes track the same upstream revision.ModelParams.useMmap (default true) and
ModelParams.useMlock (default false), wired to
llama_model_params.use_mmap / use_mlock. Lets callers turn off mmap
for platforms where memory-mapped weights hurt throughput, or pin
weights in RAM to avoid first-token paging spikes.ModelParams.flashAttention with the FlashAttention.{auto, enabled, disabled} enum, wired to
llama_context_params.flash_attn_type. Explicit settings win over the
existing automatic Android/Vulkan heuristics; auto preserves prior
behavior.ModelParams.cacheTypeK and ModelParams.cacheTypeV with the
KvCacheType.{f16, q8_0, q4_0} enum, wired to
llama_context_params.type_k / type_v. Enables KV-cache
quantization (Q8_0 ≈ halves KV memory; Q4_0 ≈ quarters it). When the
user requests a non-F16 KV type with flashAttention: auto, the
service auto-promotes flash attention to enabled — llama.cpp requires
it for KV quantization.ModelParams.kvUnified (nullable) for explicit override of
llama_context_params.kv_unified. null keeps the existing
auto-enable-when-multi-sequence behavior.ModelParams.ropeFrequencyBase and
ModelParams.ropeFrequencyScale (both nullable) for
context-extension overrides on llama_context_params.rope_freq_base /
rope_freq_scale. null keeps the model's trained values.ModelParams load tuning knobs through the
WebGPU bridge path, including maxParallelSequences, flash attention,
KV-cache type, KV-unified, RoPE, split-mode, and main-GPU options.batchSize /
microBatchSize cascade to n_batch = n_ctx and n_ubatch = n_batch,
avoiding first-embedding aborts for BERT-class/non-causal encoder models
while preserving explicit caller values and Qwen3.5 web tuning.ModelParams.mainGpu and wired it to llama.cpp
llama_model_params.main_gpu.ModelParams.splitMode and wired it to llama.cpp
llama_model_params.split_mode, enabling explicit single-GPU selection
with ModelSplitMode.none.0.6.12.Compatibility note: no public API breaking changes in 0.6.11.
leehack/llamadart-native@b8955.<|channel>thought ... <channel|> blocks into thinking deltas instead of leaking Gemma 4 thought markers into content output.0.6.11.Compatibility note: no public API breaking changes in 0.6.10.
leehack/llamadart-native@b8638.384px max edge across Android, iOS, macOS, and Web to reduce multimodal context pressure.supportsVision / supportsAudio instead of model-family assumptions.0.6.10.Compatibility note: no public API breaking changes in 0.6.9.
16.4 or newer across the README, docs site, and example docs.example/chat_app iOS Podfile and Runner project settings to use deployment target 16.4.ggml_backend_score during asset-based backend fallback so unsupported Android CPU variant libraries are skipped before initialization.auto backend resolution to prefer CPU by default while keeping Vulkan available for explicit opt-in.hooks.user_defines requires flutter clean && flutter pub get before rebuilding.0.6.9.Compatibility note: no public API breaking changes in 0.6.8.
leehack/llamadart-native@b8480.0.6.8.Compatibility note: no public API breaking changes in 0.6.7.
leehack/llamadart-native@b8373.libllamadart mappings so colocated native dependencies resolve more reliably at runtime.<tool_call> and the JSON payload.0.6.7.Compatibility note: no public API breaking changes in 0.6.6.
leehack/llamadart-native@b8216.leehack/llama-web-bridge-assets@v0.1.10 (llama.cpp b8216).Q4_K_M GGUFs across the example catalog and tooling.p_eval, eval, sample, reuse) backed by llama.cpp context timings with manual timing fallback when built-in counters report zero.0.8B / 2B / 4B models by re-enabling KQV/op-offload/flash-attention where stable.0.8B and 2B, and reduced Android 0.8B context to 2048 for lower first-token latency.0.8B projector work onto CPU on Android.0.6.6.Compatibility note: no public API breaking changes in 0.6.5.
LlamaEngine.embed(...) and LlamaEngine.embedBatch(...) for direct vector generation.BackendEmbeddings for custom backend implementers.BackendBatchEmbeddings and worker-side batch embedding request/response path to reduce isolate round-trip overhead in embedBatch(...).ModelParams.maxParallelSequences (n_seq_max) so contexts can reserve multiple sequence slots for true multi-sequence embedding batches.example/basic_app/bin/llamadart_embedding_example.dart.example/basic_app/bin/llamadart_sqlite_vector_example.dart for local embedding retrieval with SQLite vector search.tool/testing/native_embedding_benchmark.dart to compare sequential embedding calls vs embedBatch(...) throughput (with optional --json-out).tool/testing/native_embedding_sweep.dart to run max-seq sweeps and dump CSV speedup reports for plotting.LlamaEngine.embed(...) / embedBatch(...).leehack/llama-web-bridge-assets@v0.1.8.v0.1.8 bridge bundle through local fetch-script checksum verification.example/chat_app recommended Qwen presets to the Qwen3.5 lineup (0.8B, 2B, 4B, 9B) and removed older Qwen2.5/Qwen3 defaults from the in-app library.mmproj) wiring for Qwen3.5 model cards and tuned safer multimodal defaults (contextSize: 8192, maxTokens: 1024).0.6.5.Compatibility note: no public API breaking changes. android-arm64 now defaults to cpu_profile: full, which may increase package size compared with bas…
Multimodal projector offload alignment:
preferredBackend: cpu or gpuLayers: 0) now also disable mmproj GPU offload.Package metadata cleanup:
pubspec.yaml (environment.flutter, flutter, path_provider, json_rpc_2, integration_test) to keep the core package pure Dart.Backend selection safety and status accuracy:
preferredBackend: cpu no longer initializes optional GPU backends during startup/model load probing.offload_kqv, op_offload, flash-attention auto path) when effective GPU layers resolve to zero, preventing GPU allocation attempts during context creation in CPU mode.ModelParams.batchSize (n_batch) and ModelParams.microBatchSize (n_ubatch) so context batch sizing can be tuned independently from contextSize while preserving legacy defaults.getAvailableBackends) vs active runtime backend (getBackendName).BackendAvailability capability and LlamaEngine.getAvailableBackends() to support safe settings UIs without forcing GPU initialization.BackendRuntimeDiagnostics capability and LlamaEngine.getResolvedGpuLayers() to expose resolved native load-time layer count for runtime diagnostics.example/chat_app to populate backend selector options from safe availability discovery while keeping active-backend status bound to effective runtime backend.Web model cache + large-model UX improvements (chat app):
ArrayBuffer pressure.Web model-load resilience:
WebGpuLlamaBackend to retry web model loads with reduced context sizes (and CPU fallback as last attempt) when bridge errors indicate browser memory pressure.llama_webgpu_core_mem64) with automatic fallback to wasm32 when unsupported.leehack/llama-web-bridge-assets@v0.1.5.custom_headers in generated Space README frontmatter.Android arm64 CPU variant policy and loader hardening:
b8138 to b8157 to consume Android arm64 CPU-variant runtime bundles.cpu_profile (full default, compact) and advanced cpu_variants override.scripts/android_runtime_smoke.sh) and smoke-plan docs for device verification.android-arm64 now defaults to cpu_profile: full, which may increase package size compared with baseline-only CPU packaging.Native runtime sync (llama.cpp b8138):
b8099 to b8138.example/tui_coding_agent, a nocterm-based terminal coding agent with tool-calling loop, workspace-scoped file/command tools, and runtime model switching.unsloth/GLM-4.7-Flash-GGUF:UD-Q4_K_XL) with support for custom local paths/URLs/Hugging Face shorthand.--native-tool-calling for experimentation).Native inference performance improvements:
create(...).streamBatchTokenThreshold,
streamBatchByteThreshold).reusePromptPrefix, enabled by default) with conservative full-replay
fallback to preserve deterministic parity.ChatSession context trimming using bounded turn-offset
search to avoid repeated linear recount loops on long histories.tool/testing/native_inference_benchmark.dart for TTFT,
throughput, and latency measurement with tunable generation settings.tool/testing/native_prompt_reuse_parity.dart and curated prompt
sets for deterministic prompt-reuse parity validation.Moved hook backend-config support code out of hook/src/ into lib/src/hook/ because pub.dev currently only allows hook/build.dart under hook files.
Publishing compatibility fix:
hook/src/ into
lib/src/hook/ because pub.dev currently only allows hook/build.dart
under hook files.llama.cpp parity expansion (Dart-native template/parser pipeline):
peg_parser_builder, peg_chat_parser) and integrated parser-carrying render/parse flow for PEG-native/constructed formats.ChatTemplateMatcher, ChatTemplateRoutingContext,
ChatTemplateEngine.registerHandler(...),
ChatTemplateEngine.unregisterHandler(...),
ChatTemplateEngine.clearCustomHandlers(...),
ChatTemplateEngine.registerTemplateOverride(...),
ChatTemplateEngine.unregisterTemplateOverride(...),
ChatTemplateEngine.clearTemplateOverrides(...), and
per-call customHandlerId / parse handlerId routing.chatTemplateKwargs and templateNow.Parity test coverage and tooling:
run_llama_cpp_chat_tests.sh, run_template_parity_suites.sh).peg_parser_builder, template_internal_metadata) to satisfy structure guards.Test cleanup and maintainability:
Native integration cleanup (llamadart-native migration):
tool/testing/prepare_llama_cpp_source.sh to fetch/refresh ggml-org/llama.cpp into .dart_tool/llama_cpp (or LLAMA_CPP_SOURCE_DIR) pinned to a resolved ref (LLAMA_CPP_REF, default latest release tag).tool/testing/run_llama_cpp_chat_tests.sh to use prepared .dart_tool source instead of third_party/llama_cpp, so local upstream chat-suite runs no longer depend on vendored source.LLAMA_CPP_TEMPLATES_DIR or .dart_tool/llama_cpp/models/templates instead of third_party/llama_cpp.KleidiAI/ZenDNN are CPU-path optimizations, not selectable runtime backend modules.Linux runtime/link validation and backend loader hardening:
linux-arm64 vs linux-x64) when preparing runtime dependencies.cpu/vulkan/blas) and optional cuda/hip module dependency resolution.scripts/check_native_link_deps.sh helper plus dedicated validation images:
docker/validation/Dockerfile.cuda-linkcheck and
docker/validation/Dockerfile.hip-linkcheck.Chat example backend UX cleanup:
Auto backend option from settings; only concrete runtime-detected backends are shown.Auto preference to the best detected backend at runtime.ChatTemplateEngine now preserves handler-provided tokens even when grammar is attached via params, avoiding token-loss regressions in tool/thinking fo
llama.cpp parity hardening:
ChatTemplateEngine now preserves handler-provided tokens even when grammar is attached via params, avoiding token-loss regressions in tool/thinking formats.Chat example loop/lifecycle hardening:
fmt:*, think:*, content:json, fallback:tool-result) and strengthened detach/exit disposal paths.Parity/integration test robustness:
tool_calling_integration_test now accepts both structured tool_calls deltas and XML-style <tool_call> payloads.Documentation updates:
{"response":"..."}) and documented content:json diagnostics.penalty=1.0, top_p=0.95, min_p=0.05) and added a CLI README batch parity-matrix usage example.Chat app backend/status fixes:
gpuLayers while still allowing load-time CPU enforcement.Context size auto mode:
Context Size: Auto by preserving 0 in persisted settings and passing auto behavior through to session context-limit resolution.Tool-call parsing fixes (Hermes):
{{...}} layer second, and only fall back to full _normalizeDoubleBraces when all braces are consistently doubled._normalizeDoubleBraces that bails out on mixed single/double brace payloads to prevent corruption of valid nested JSON.Tool-call parsing fixes (Magistral):
_extractJsonObject to handle \n, \r, and \t between [ARGS] and the JSON body.Example app (basic_app):
toList() buffering with await for streaming for real-time token yield.tools parameter to every follow-up create() call and bounded tool-execution loop with _maxToolRounds = 10.Test coverage:
[ARGS] with newline/nested arguments.Example rename (server):
example/api_server to example/llamadart_server.llamadart_server.example/llamadart_server.GLM 4.5 template parity:
<tool_call> payloads with <arg_key>/<arg_value> pairs.<|user|> stop handling for tool-call flows.Template/native runtime fixes:
Added minP to GenerationParams with a default value of 0.0 and copyWith support.
minP to GenerationParams with a default value of 0.0 and copyWith support.min_p sampler initialization in LlamaCppService when minP > 0.GenerationParams.minP default and copyWith behavior.Chat template parity hardening:
ToolCallGrammarUtils helpers for wrapped object/array tool-call grammar generation and root-rule wrapping.llama_grammar_init_impl parse failures during tool-calling generations.Updated README internal links to absolute GitHub URLs so they resolve reliably on pub.dev.
.pubignore to exclude local build outputs, large model/test artifacts, and checked-out third_party sources from package uploads.Root exports were tightened; previously exposed internals such as ToolRegistry, LlamaTokenizer, and ChatTemplateProcessor are no longer part of the pu
[BREAKING] Public API Changes:
ToolRegistry, LlamaTokenizer, and ChatTemplateProcessor are no longer part of the public package API.ChatSession now centers on create(...) streaming LlamaCompletionChunk; legacy chat(...) / chatText(...) style usage must migrate.LlamaChatMessage constructor names were standardized (.fromText, .withContent) in place of older named constructors.maxTokens in GenerationParams increased from 512 to 4096.LlamaChatMessage.toJson() no longer includes name on tool role messages.ModelParams.logLevel was removed; logging control now lives on LlamaEngine via setDartLogLevel(...) and setNativeLogLevel(...).LlamaBackend interface changed for custom backend implementers (notably getVramInfo and updated applyChatTemplate).loadModel(...) now requires unloading first.MIGRATION.md.Template/Parser Parity Expansion:
<|python_tag|> parsing for Llama 3 flows.Template Extensibility APIs:
ChatTemplateEngine.customTemplate and customHandlerId routing support and threaded handler identity into parse paths.Logging Controls:
LlamaEngine: setDartLogLevel and setNativeLogLevel, while keeping setLogLevel as a convenience method.none log level suppression so llama.cpp/ggml logs are fully muted when requested.Chat App Improvements:
Test Suite Overhaul:
Default LlamaChatMessage constructor (string-based) is now deprecated; use .fromText() or .withContent() instead.
LlamaBackend for strict Web isolation using "Native-First" conditional exports, ensuring native performance and full web safety.LlamaBackend() factory across all examples and scripts.getLoadedContextInfo() and robust GGUF metadata fallback in LlamaEngine.llama.cpp native backend.dart_test.yaml and @TestOn tags to enable seamless execution of all tests across VM and Chrome with a single dart test command.dup2 to /dev/null) for LlamaLogLevel.none on native platforms.dart analyze across the core library and all example applications.ModelService architecture..meta files to track download progress across app restarts.ModelCard with a visual Pause/Resume toggle.mtmd module from llama.cpp for native platforms.
loadMultimodalProjector to LlamaEngine.LlamaChatMessage.withContent and LlamaContentPart (Text, Image, Audio).mtmd module.Question: / Answer: chat template fallback for Moondream models.chat() and chatWithTools() logic from LlamaEngine to ChatSession.LlamaEngine is now a dedicated low-level orchestrator for model loading, tokenization, and raw inference.dispose() in ChatService.ChatSession class to automatically manage conversation history and system prompts.ChatSession now implements an automated sliding window to truncate history when the model's context limit is approached.llama.cpp updates, regression testing, and release artifact generation.LlamaChatMessage.role now returns a LlamaChatRole enum instead of a String. All manual role string comparisons should be updated to use the enum.LlamaChatMessage constructor (string-based) is now deprecated; use .fromText() or .withContent() instead.LlamaChatMessage.roleString is deprecated and will be removed in v1.0.llama.cpp to tag b7898.[BREAKING] Removal of `LlamaService`: The legacy LlamaService facade has been removed. Use LlamaEngine with LlamaBackend() instead for all platforms.
LlamaService: The legacy LlamaService facade has been removed. Use LlamaEngine with LlamaBackend() instead for all platforms.wllama v2 features, including native chat templating and threading info.LlamaLogLevel.none suppresses all output; other levels enable default stderr logging.NativeFinalizer dependency to avoid race conditions. Explicitly call dispose() to release native resources.ModelParams to include initial LoRA configurations and introduced supportsUrlLoading for better platform abstraction.basic_app example to support testing LoRA adapters via the --lora flag.Nothing published for this version
WASM Support: Full support for running the Flutter app and LLM inference in WASM on the web.
wllama integration with better error handling and progress reporting.Supported platforms: iOS, macOS, Android, Linux, Windows, Web.
llama.cpp backend.Your coding agent can read these notes before it upgrades. Set up the MCP server →