NewYour coding agent can read the release notes before it upgrades.Set up the MCP server →
pub.dev · #4633 most downloaded on pub.dev
Dart binding for llama.cpp --- high level wrappers for both Dart and Flutter
Last release 2 months ago
01 Aug 2026
Release timing varies
gaps range from 2 weeks to 9 months
Nearly every release is documented
notes for 14 of 15 stable releases
Nothing withdrawn
no release was ever pulled
3 years old
22 releases · first in 2024
0.9.0-dev.12: Flutter plugin packaging, mobile API, deterministic tea…
0.9.0-dev.12: Flutter plugin packaging, mobile API, deterministic tea…
…rdown
Supersedes 0.9.0-dev.11, which was tagged and built but rejected by
pub.dev at upload: only build.dart and link.dart are currently
allowed under hook/. The AAR-extraction helper moved to
lib/src/hook/android_native_assets.dart and hook/build.dart imports
it by package URI. dev.11 was never published; its GitHub release
assets are identical in content to this one.
Same llama.cpp pin as dev.10 (afeebe10, tag b10182) — no native
rebuild required for behavior, but the artifacts are re-cut so the
SwiftPM manifest and the bundled Android AAR match this tag.
hook/build.dart) that extracts the
verified release AAR's jni/<abi>/*.so and registers them as bundled
code assets. Flutter apps no longer add anything to Gradle. Opt out
with bundle_android: false, or point at a different artifact (for
example the Snapdragon Hexagon AAR) with android_aar:, under
hooks.user_defines.llama_cpp_dart in the app's pubspec.darwin/llama_cpp_dart, whose
binary target is pinned to this release's llama-xcframework.zip.
This is how Flutter resolves the plugin now that SwiftPM is enabled by
default on stable. CocoaPods is intentionally not supported.ContextParams.mobile() — small nCtx / nBatch / nUbatch
defaults for phones and tablets, with the KV cache types overridable.LlamaModel.estimateVramBytes({int nCtx}) — planning estimate for
weights plus an f16 KV cache plus a 15% runtime-buffer allowance.isDisposed on LlamaModel and LlamaEngine.shiftPolicy / shift on EngineChat.generate(), so long-reasoning
chats can slide the context. Throws for multimodal histories, where
media embeddings cannot be reconstructed after a shift.LlamaEngine.spawn(libraryPath:) is now optional, defaulting to
the platform library name. On Android that resolves the bundled
libllama.so. Existing calls that pass a path are unaffected.LlamaEngine.dispose() waits up to 30 seconds (was 2) for native
teardown and now surfaces a teardown failure instead of discarding it.classes.jar. The placeholder entry was created from /tmp/...,
so the cleanup zip -d never matched it and the entry shipped in
every AAR; on a Windows host it would have been the builder's home
directory. Reported in #107.tool/package_apple_xcframework.sh replaces the inline zip step and
uses ditto to preserve versioned-framework symlinks, then verifies
the extracted archive's Info.plist, Versions/Current symlink, and
code signature before it can be published.tool/check_android_aar_alignment.sh gates every Android build on
16 KB ELF LOAD alignment (see #107); CI additionally builds a
throwaway Flutter app and asserts all five .so files reach the APK,
and greps pub publish --dry-run so the AAR cannot drop out of the
published package.One column per month.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
release: 0.9.0-dev.10
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Native rebuild required — src/llama.cpp moved d6d0ce82 → afeebe10
(tag b10182), about seven weeks of upstream work.
e6dd0e29a, which collapsed the use_mmap /
use_direct_io / use_mlock booleans in llama_model_params into a
single llama_load_mode enum. ModelParams keeps its three
booleans — they are mapped at the FFI boundary, so callers are
unaffected. One semantic caveat: the enum has no direct-I/O-plus-mlock
value, so when both are requested direct I/O wins and mlock is dropped.
useDirectIo keeps its documented precedence over useMmap.llama_model_n_layer_nextn,
llama_model_ftype, llama_ftype_name,
llama_vocab_get_suppress_tokens, and the mtmd batch-encoding API
(mtmd_batch_init / _add_chunk / _encode / _get_output_embd).mtmd_encode is deprecated upstream in favor of mtmd_encode_chunk.
This package reaches multimodal via mtmd_helper_eval_chunks and never
called it, so no change was needed.LlamaLibrary.dispose now clears the log callback.
LlamaLog.silence installs a Pointer.fromFunction bound to the
isolate that registered it, but the slot it occupies lives in
process-global llama.cpp/ggml state and outlives that isolate. The
stale pointer stayed installed, so the next isolate to emit a log line
invoked a callback owned by a dead isolate and the VM aborted with
"Cannot invoke native callback from a different isolate". Surfaced by
Dart 3.12's stricter cross-isolate check.
Known remaining issue: parallel dart test still hits the concurrent
variant of this race, where one isolate holds a live callback while
another loads a model. Run the model-backed suite with -j 1 until
silence() stops using a Dart callback altogether.
ffigen 20.1.1 → 21.0.0, lints 5.0.0 → 6.1.0 (dev dependencies).
Note ffigen 21 requires Dart SDK ≥ 3.10 to run the generator; the
package's own sdk: ^3.5.0 constraint for consumers is unchanged.-resource-dir compiler-opt, which pointed at a clang
17 toolchain directory that no longer exists.fix(build): disable MTMD_VIDEO across all native builds
fix(build): disable MTMD_VIDEO across all native builds
The d6d0ce82 bump pulled in mtmd video decoding, which #includes
vendor/sheredom/subprocess.h and calls posix_spawn to launch ffmpeg. That
broke the Android CPU AAR build (posix_spawn isn't in the NDK at API 26),
and the feature is useless in every target we ship (it needs an ffmpeg
binary in PATH at runtime). MTMD_VIDEO defaults ON upstream; force it OFF
in all four build scripts. Only the mtmd_helper_video_* path drops out;
the bitmap-init symbols the binding uses are unaffected.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Native rebuild required — src/llama.cpp moved 6b4e4bd58 → d6d0ce82
(picks up Gemma-4 E2B/E4B MTP support, #24282).
MtpSpeculativeDecoder is gone, along with ContextType /
ContextParams.ctxType. It relied on the NextN hidden-state staging
API (llama_set_embeddings_nextn / llama_get_embeddings_nextn_ith,
formerly *_pre_norm), which lives only in llama.cpp's private C++
header and had to be resolved by hand via C++-mangled symbols. That
approach is ABI-fragile (broke on the upstream pre_norm→nextn
rename; would not work under MSVC) and reaches past the public C API.
The binding now uses only the ffigen-generated public C headers
(llama.h + mtmd). Classic target+draft SpeculativeDecoder (pure C
API) is unaffected.llama_context_params
gains ctx_other + n_outputs_max; llama_set_warmup is deprecated.mtmd_helper_bitmap_init_from_{file,buf} now take a placeholder bool
and return a mtmd_helper_bitmap_wrapper (use .bitmap).-force_load, which fixes
the iOS dlsym/dead-strip failure (#104) and a static-framework clean-build
ordering trap. It bundles common and links Metal/Accelerate internally;
the macOS slice uses the versioned bundle layout and each slice is signed
with the org.ggml.llama identifier so on-device installs validate.Pure-Dart release on the same llama.cpp pin (tag `b9360`, sha 6b4e4bd58) — no native rebuild required.
Pure-Dart release on the same llama.cpp pin (tag b9360, sha
6b4e4bd58) — no native rebuild required.
MtpSpeculativeDecoder — self-speculative decoding driven by a
model's own Multi-Token Prediction / NextN heads (no separate draft
model). Pair a normal target LlamaContext with a draft context
created ctxType: ContextType.mtp off the same model. Mirrors
upstream llama.cpp's MTP loop (PR #22673): the target emits pre-norm
hidden states, the NextN head proposes tokens conditioned on them, and
the target verifies a round in one pass. Output is byte-identical to
plain greedy decoding.PARTIAL_ONLY | ON_DEVICE state checkpoints
rather than partial seq_rm, which those models forbid.llama_set_embeddings_pre_norm
/ llama_get_embeddings_pre_norm_ith staging symbols (resolved by hand
since they are absent from the public header). This supersedes the
dev.7 note that MTP-as-draft was blocked — it is now implemented.tool/probe_mtp.dart.KvCacheType.iq4_nl — exposes GGML_TYPE_IQ4_NL, the last
upstream-supported KV cache type the binding was missing. ~4×
compression like q4_0 but better quality from a non-linear codebook.
See the README's KV-cache quantization section (incl. the symmetric
_0 integer-dot-product property and a note on fork-only TurboQuant).llama.cpp submodule bumped from gguf-v0.18.0-791-g5d56effde to tag `b9360` (sha 6b4e4bd58). 328 commits of upstream history, purely-additive C API del
llama.cpp submodule bumped from gguf-v0.18.0-791-g5d56effde to tag
b9360 (sha 6b4e4bd58). 328 commits of upstream history,
purely-additive C API delta (llama_context_type, llama_n_rs_seq,
new llama_state_seq_flags, mtmd_get_cap_from_file, plus the
context-params fields ctx_type / n_rs_seq). Bindings
regenerated; existing wrapper code unchanged.
LlamaEngine.embed(text) — pooled and per-token embeddings via
the worker isolate, with optional L2 normalization. Returns
EmbeddingResult covering both pooled (mean/cls/last/rank) and
unpooled outputs.BatchEmbedder (sync) + LlamaEngine.embedBatch(texts)
(off-thread) — embed N texts in a single decode pass by assigning
each its own sequence id; amortizes per-token compute across the
batch for RAG-style ingest. Requires embeddings: true, a pooled
pooling type, and nSeqMax >= texts.length.SpeculativeDecoder — synchronous greedy and exact stochastic
speculative decoding over a target + draft LlamaContext sharing a
vocab. Greedy output is byte-identical to plain greedy on the target;
temperature > 0 runs the min(1, p/q) accept rule with residual
resampling (distributionally identical to sampling the target), seed
for reproducibility. See example/probes/speculative_generate.dart.
(MTP-as-draft is blocked upstream — it needs the non-public
llama_set_embeddings_pre_norm; the draft-model variant works today.)ContextType + nRsSeq on ContextParams — build an MTP draft
context against an MTP-capable target model for raw-FFI use.LlamaLora + LoraBinding, with
LlamaContext.setLoraAdapters / clearLoraAdapters /
setControlVector. LoRA stack swaps, metadata accessors, aLoRA
invocation-token reads, and ReFT-style control vectors.MtmdBitmap / MtmdChunk / MtmdChunks /
MtmdCapabilities — bitmap construction (raw RGB, raw audio,
file decode, buffer decode), mtmd_input_chunks introspection
(kind, nTokens, nPos, id, text-token reads), plus a cheap
mtmd_get_cap_from_file probe.nCtxSeq, nRsSeq, effective poolingType.setThreads, setEmbeddings, setCausalAttn,
setWarmup, synchronize.memoryClear, memorySeqRm, memorySeqCp,
memorySeqKeep, memorySeqAdd, memorySeqDiv,
memorySeqPosMin/Max — covers forking, rollback, position
shifting.lastLogits, logitsAt, sampledTokenAt,
sampledProbsAt, sampledCandidatesAt, sampledLogitsAt.ContextPerf / SamplerPerf snapshots with perf() / resetPerf()
/ printPerf() on LlamaContext and Sampler. Includes
prompt/decoded tokens-per-second convenience getters.LlamaModel: isDiffusion, isHybrid, nSwa, nEmbdInp,
nEmbdOut, decoderStartToken, nClassifierOut,
classifierLabel(i), ropeType (new RopeType enum),
ropeFreqScaleTrain, metaCount + metaKeyAt / metaValueAt /
metaValue(key) / metaEntries.LlamaLibrary: supportsMmap / Mlock / Rpc,
maxParallelSequences, maxTensorBuftOverrides, timeUs,
systemInfo(), initNuma(NumaStrategy.*).SplitPath.compose / decomposePrefix for split-gguf filenames.name, seed, chainCount, chainGet(i) (borrowed),
chainRemove(i) (owned), clone(), apply(arr).captureRawStateExt / restoreRawStateExt accepting the new
StateSeqFlags (mirrors LLAMA_STATE_SEQ_FLAGS_* including
the b9360 on-device snapshot bit).tool/build_native.sh disables LLAMA_BUILD_SERVER and
LLAMA_BUILD_APP: upstream b9360's tools/server/ references
mtmd symbols missing from its own public header. We ship neither
target, so turning them off keeps --with-mtmd building cleanly.params: expose missing llama.cpp options across sampler/context/model
params: expose missing llama.cpp options across sampler/context/model
- SamplerParams: Mirostat (v1/v2), grammar (incl. lazy patterns), DRY,
XTC, dynamic temperature, adaptive-P, top-n-sigma, infill, logit
bias, shared min_keep. SamplerFactory.build now accepts model: for
vocab-dependent stages.
- ContextParams: RoPE scaling/freq, full YaRN knobs, pooling and
attention type, defrag threshold, no_perf, op_offload, swa_full,
kv_unified.
- ModelParams: split mode, main GPU, tensor split, device list, GGUF
kv overrides (int/float/bool/str), use_direct_io, use_extra_bufts,
no_host, no_alloc.
- README: link to aichat sample app.
- Bump to 0.9.0-dev.6 + CHANGELOG.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Closes the gap between the Dart binding's option surface and the underlying llama.cpp params. Purely additive — existing code keeps working with previous defaults.
SamplerParamsMirostatConfig (v1 + v2 with tau, eta, m). Terminal sampler
when enabled — replaces the dist stage.GrammarConfig — GBNF grammar plus optional lazy-trigger patterns
and trigger tokens (llama_sampler_init_grammar /
llama_sampler_init_grammar_lazy_patterns).DryConfig — DRY sampler (multiplier, base, allowed length, last-N,
seq breakers).XtcConfig — XTC sampler (probability, threshold, min keep, seed).DynamicTempConfig — dynamic temperature (temp_ext: range,
exponent).AdaptivePConfig — adaptive-P terminal sampler (target, decay,
seed).LogitBiasEntry list applied at the start of the chain.topNSigma, infill, and shared minKeep for top-p / min-p /
typical / xtc.SamplerFactory.build(params, model: ...) — model: is now
required when the chain uses grammar, DRY, infill, logit-bias, or
Mirostat v1 (anything that needs the vocab or n_ctx_train).ContextParamsRopeScalingType, PoolingType, AttentionType enums.ropeFreqBase, ropeFreqScale.yarnExtFactor, yarnAttnFactor, yarnBetaFast,
yarnBetaSlow, yarnOrigCtx.defragThreshold, noPerf, opOffload, swaFull, kvUnified.ModelParamsSplitMode enum + mainGpu + tensorSplit (allocated to
llama_max_devices() at load time).devices — list of backend device names (resolved from
LlamaBackends.list()).kvOverrides — int / float / bool / string GGUF metadata overrides
with the standard NULL-terminated array layout.useDirectIo, useExtraBufts, noHost, noAlloc.aichat
sample app as a working Flutter integration reference.progress_callback, cb_eval, abort_callback — need a
NativeCallable.listener wrapper with isolate-affinity rules.tensor_buft_overrides — needs a per-device buffer-type accessor in
BackendDevice first.llama_context_params.samplers) — still
marked [EXPERIMENTAL] upstream.Consolidates 0.9.0-dev.0 through 0.9.0-dev.5 (none of dev.0–dev.4 were published to pub.dev). The 0.2.x line is a separate package shape — see MIGRATI
Consolidates 0.9.0-dev.0 through 0.9.0-dev.5 (none of dev.0–dev.4
were published to pub.dev). The 0.2.x line is a separate package shape
— see MIGRATION.md.
LlamaEngine worker isolate is the primary public API. Streaming
token output via Stream<GenerationEvent> (sealed:
TokenEvent | ShiftEvent | DoneEvent). Cancellation via stream
subscription cancel.EngineSession (raw prompt) and EngineChat (message-history with
chat template) on top of the engine isolate.mtmd.EngineSession.saveState/loadState and
EngineChat.saveState/loadState with metadata-validated reload.llama-server-style context shift (ContextShiftPolicy.auto) gated
on engine.canShift.dart test)ios-arm64, ios-arm64-simulator, macos-arm64)arm64-v8a, two flavors: CPU+mtmd (~2 MB) and
Hexagon NPU + OpenCL + mtmd (~3.7 MB)llama_cpp.podspec.engine.devices (List<BackendDevice>),
engine.hasAccelerator, engine.primaryAcceleratorName, and the
pre-engine LlamaBackends.list(). Tells you which backends loaded
on the current device.primaryAcceleratorName priority orders by registry name (HTP →
Hexagon → Metal → CUDA → Vulkan) before type, so Snapdragon HTP wins
over OpenCL even when ggml reports both as type=gpu.ContextParams.typeK / typeV accept any
of KvCacheType.{f32, f16, bf16, q8_0, q4_0, q4_1, q5_0, q5_1}.
q8_0 halves KV memory at small quality cost; useful on 8 GB Android
devices with longer contexts.LlamaLog.captureToFile(path) /
LlamaLog.restoreStderr(). Toggleable redirect of llama.cpp/ggml
log lines for Android, where stderr is not connected to logcat.ADSP_LIBRARY_PATH. LlamaLibrary.load() reads
/proc/self/maps on Android and exports ADSP_LIBRARY_PATH so
FastRPC finds libggml-htp-v*.so skeleton libs without app-side
MethodChannel plumbing.LlamaBindings is now exported. Lets callers using the raw FFI
surface type variables / pass them around without reaching into
src/.LlamaVersion is generated at build time. Exposes the package
version, the llama.cpp submodule SHA + author date, and a runtime
systemInfo() wrapper around llama_print_system_info() (e.g.
MTL : EMBED_LIBRARY = 1 | CPU : NEON = 1 | ACCELERATE = 1 | ...).Llama god-class.LlamaParent / LlamaChild / IsolateScope (replaced by LlamaEngine).LlamaService multi-session scheduler. Mobile apps do one
conversation at a time; multi-session can be added back as a higher
layer if needed.TextChunker (RAG helper).llama_chat_apply_template instead.Q4_K_*, Q5_K_*) and I-quants (IQ*) run on OpenCL+CPU.llama-server's behaviour).ggml_metal_device_free asserts at process exit because
the worker doesn't dispose model/context. Harmless.allow freeing the active slot by switching/detaching and reselecting a fallback
example/auto_trim.dartAndroid: Added OpenCL support for GPU acceleration (#91).
mtmd context disposal._exportQwen3Jinja).llama.cpp submodule.* forgot to update version
llama.cpp 25ff6f7659f6a5c47d6a73eada5813f0495331f0
* Multimodal support - vision
Breaking change: Internal API restructuring (public API remains stable)
compatible with llama.cpp 42ae10bb
performance imporvement and bugs fix
added initial support to load lora
added static property Llama.libraryPath to set library path, in order to support linux and other platforms
Llama.libraryPath to set library path, in order to support linux and other platformsModelParams disabled options splitsMode, tensorSplit and metadataOverride
ModelParams disabled options splitsMode, tensorSplit and metadataOverrideLlamaProcessor now take context and model parameters
Nothing published for this version
refactored code to follow dart package structure
TODO: Describe initial release.
Your coding agent can read these notes before it upgrades. Set up the MCP server →