NewYour coding agent can read the release notes before it upgrades.Set up the MCP server →
Go modules · #1 by repository stars
Last release today
04 Oct 2026
Ships on a steady schedule
a new release about every 8 days
Rarely documented
notes for 14 of the last 60 stable releases
Nothing withdrawn
no release was ever pulled
3 years old
1151 releases · first in 2024
Nothing published for this version
Merge remote-tracking branch 'upstream/main' into release_v0.40.0
Merge remote-tracking branch 'upstream/main' into release_v0.40.0
mlx: match publisher tokenizer semantics
Honor pretokenizer stage order, split behavior, Unicode boundaries,
added-token normalization, and ranked BPE merges. Handle empty added
tokens and empty Metaspace input consistently.
Add shared Go/Python reference cases using published tokenizers, pulling
missing models directly and failing on errors, plus focused regressions
for configuration precedence, byte fallback, and parallel encoding.
One column per month.
Nothing published for this version
Nothing published for this version
Ollama now supports Clef and Clef Flash , Cloudflare's new open-source decision models, through /v1/systemone .
Ollama now supports Clef and Clef Flash, Cloudflare's new open-source decision models, through /v1/systemone.
Clef (27B) and Clef Flash (9B) are multimodal: requests can now include images alongside the text state, shared by all questions and scored jointly with it.
curl http://localhost:11434/v1/systemone -d '{
"model": "clef-flash",
"state": "The user took this screenshot.",
"images": ["<base64-encoded image>"],
"questions": {
"has_ollama": {"type": "noul", "instructions": "Does this image contain Ollama?"}
}
}'{
"model": "clef-flash",
"answers": {
"has_ollama": {
"type": "noul",
"noul": 0.958
}
},
"usage": {
"input_tokens": 548,
"output_tokens": 0
}
}
CAPABILITY declarations, so model creators can explicitly declare what a model can do. Declarations are preserved when creating from GGUF or safetensors, through model inheritance, and on Modelfile exportollama show and the model list now report only decision as the capability for decision models, so clients no longer offer them for general chat, tools, or thinkingFull Changelog: v0.35.0...v0.35.1
ci: fix missing build context
ci: fix missing build context (#18742)
models: add clef support
models: add clef support (#18741)
Add CAPABILITY declarations to Modelfiles and an additive capabilities field to create requests. Preserve declarations across GGUF and safetensors cre
Add CAPABILITY declarations to Modelfiles and an additive capabilities
field to create requests. Preserve declarations across GGUF and safetensors
creation, inheritance, and Modelfile export.
Require decision capability before scheduling System One requests instead
of matching Qwen architecture/renderer metadata. Retain main's GGUF-only
scoring restriction until the separate MLX runtime work lands.
Extracted from the capability foundation in 36d46a0 on system_one_mlx;
MLX scoring and manifest-list changes are intentionally separate.
Nothing published for this version
Requests containing the deprecated typical_p parameter now log a warning instead of failing.
Ollama now supports decision models through /v1/systemone, based on TypeSafe’s Jev API.
Decision models return choices, probabilities, and scores instead of text. Use them for tasks such as ticket triage, model routing, and content classification.
Available models:
ollama pull nimbleSend context and one or more questions:
curl http://localhost:11434/v1/systemone \
-H 'Content-Type: application/json' \
-d '{
"model": "nimble",
"state": "Our checkout has returned 500 errors since 9am.",
"questions": {
"label": {
"type": "choice",
"instructions": "Which label fits this ticket?",
"criteria": {
"billing": "Payments and refunds",
"bug": "Software errors",
"account": "Login and account access"
}
}
}
}'Example response:
{
"model": "nimble",
"answers": {
"label": {
"type": "choice",
"choice": "bug",
"probabilities": {
"billing": 0.0125,
"bug": 0.9781,
"account": 0.0093
},
"confidence": 0.8906
}
},
"usage": {
"input_tokens": 174,
"output_tokens": 1
}
}The API supports three question types:
choice: Select an option and return probabilities for each.noul: Return the probability that a condition is true.score: Return a score across an ordered set of criteriatypical_p parameter now log a warning instead of failing.Full Changelog: v0.34.4...v0.35.0
mlx: bound pull stall retries and let the watchdog interrupt them ( #1 …
mlx: bound pull stall retries and let the watchdog interrupt them (#1…
feat: add System One scoring API
feat: add System One scoring API (#18606)
Nothing published for this version
Nothing published for this version
Structured outputs on thinking models now apply in a single pass, making them faster and more reliable.
Full Changelog: v0.34.3...v0.34.4
We pick up schema fixes for typed dictionary values and short arrays.
We pick up schema fixes for typed dictionary values and short arrays.
Nothing published for this version
mlx: speed up Qwen 3.8 prompt processing
Use MLX's gated-delta kernel for long scans and fold dense MLP global scales into SwiGLU.
Nothing published for this version
Nothing published for this version
GET /api/show now advertises each model's thinking controls and default:
GET /api/show now advertises each model's thinking controls and default:
Available in the CLI with:
ollama show gemma4
thinking
levels false, true
default true
Available in the API with:
curl http://localhost:11434/api/show -d '{"model": "glm-5.3-flash:cloud"}'{
"thinking": {
"values": ["low", "high", "max"],
"default": "max"
}
}Also available on ollama.com directly for cloud models.
Full Changelog: v0.34.2...v0.34.3
server: allow registry cross-host redirects among allowlisted hosts (…
server: allow registry cross-host redirects among allowlisted hosts (…
api: expose model thinking levels and defaults
api: expose model thinking levels and defaults (#18473)
Nothing published for this version
Nothing published for this version
Added first-run setup when running ollama , with options to sign in or continue locally. Setup completion is shared with the desktop app on macOS and
ollama, with options to sign in or continue locally. Setup completion is shared with the desktop app on macOS and Windows.ollama://apps to open the desktop app’s Apps page directly on macOS and Windows.Full Changelog: v0.34.1...v0.34.2
cli: add first-run onboarding shared with the desktop app
cli: add first-run onboarding shared with the desktop app (#18495)
The decode loop releases MLX's pool of freed buffers every 256 generated tokens, which is also how often the KV cache grows and drops its previous, sm
The decode loop releases MLX's pool of freed buffers every 256 generated
tokens, which is also how often the KV cache grows and drops its previous,
smaller buffers. The check fires only when the token count lands exactly on
a multiple of 256. Speculative decoding emits several tokens per round, so
most rounds step over the boundary and the pool is never released. Each
growth at a long context leaves several GB of buffers that no later
allocation can reuse, so the runner's footprint keeps climbing over a long
generation until the system runs out of memory.
We now release the pool whenever a round crosses a multiple of 256 tokens,
which is what a single-token round already did. With qwen3.8:27b-mlx at a
98k-token context on a 128 GB machine, a long speculative generation
previously grew the runner past 90 GB and panicked the kernel; it now stays
flat at 30 GB.
Nothing published for this version
model is one package with three jobs: the contract between the runner and the architectures, the opened checkpoint, and building nn layers from checkp
model is one package with three jobs: the contract between the runner
and the architectures, the opened checkpoint, and building nn layers
from checkpoint tensors. Its files did not say which was which. base.go
carried the folded package's name over the interfaces and the registry,
root.go held the safetensors header scan next to Root, and quant.go
mixed the nvfp4 global-scale helpers with quant parameter resolution.
base.go becomes model.go, named for what it holds. root.go keeps Root
and Open; TensorQuantInfo and the header scan join quant.go, so
everything the checkpoint says about quantization is read and resolved
in one file. The global-scale helpers move to globalscale.go with their
tests. Root.Close, a no-op with one caller, goes. No code changes
otherwise.
llama.cpp build changes resulted in duplicate symbols between libllama and libmtmd. This moves the compat patch into libllama with exported symbols.
llama.cpp build changes resulted in duplicate symbols between libllama and libmtmd. This moves the compat patch into libllama with exported symbols.
Nothing published for this version
Deprecated typical_p : it can no longer be set when creating new models, existing GGUF models retain support.
ollama create no longer experimental. GGUF model creation now requires using llama.cpp tooling for safetensor conversion and quantization./api/tags is much faster on large model libraries (3.1 s → 294 ms cold in testing), and model capabilities are now reported consistently.typical_p: it can no longer be set when creating new models, existing GGUF models retain support.Full Changelog: v0.34.0...v0.34.1-rc1
Nothing published for this version
Drop support for creating new models with typical_p parameters, while retaining support for existing GGUF models with the setting.
Drop support for creating new models with typical_p parameters, while
retaining support for existing GGUF models with the setting.
mlx: add mlx patch to docker build context
mlx: add mlx patch to docker build context (#18440)
mlx: support ModelOpt global scales in MoE models
MLX: version bump
mlx: support ModelOpt global scales in MoE models
address comments
address comments
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Use Ollama models in ChatGPT Desktop
Ollama models can now be used directly in ChatGPT Desktop, so you can keep your existing workflow while running open models. Setup is available from the Ollama app on MacOS.
This release also improves structured output performance on Apple Silicon, adds support for OpenAI-compatible client tool search and response compaction.
Full Changelog: v0.33.3...v0.34.0
Nothing published for this version
openai: support standalone named function outputs
openai: support standalone named function outputs (#18348)
proxy: normalize namespaced commands in Full Access
proxy: normalize namespaced commands in Full Access (#18331)
openai: accept plaintext-labeled Codex agent messages
openai: accept plaintext-labeled Codex agent messages (#18329)
openai: finalize responses at the web search limit
openai: finalize responses at the web search limit (#18328)
Nothing published for this version
app: harden Codex desktop proxy handling
app: harden Codex desktop proxy handling (#18244)
Nothing published for this version
app: add Ollama to ChatGPT Desktop
app: add Ollama to ChatGPT Desktop (#18236)
Nothing published for this version
gemma4 now supports images and audio on MLX engine
Full Changelog: v0.33.2...v0.33.3
Safetensors gemma4 imports served by the MLX engine now answer image and audio chats. Images run through both vision architectures: the transformer to
Safetensors gemma4 imports served by the MLX engine now answer image
and audio chats. Images run through both vision architectures: the
transformer tower (26B, 31B, e-series) and the 12B's encoder-free
unified embedder. Audio arrives through the same intake the ollama
API already accepts for gemma4 GGUFs — WAV bytes in the images field,
OpenAI input_audio parts, and /v1/audio/transcriptions uploads — with
the e2b/e4b checkpoints running clips through their conformer audio
encoder and the 12b unified checkpoint embedding the raw waveform
directly. Clips longer than 30 seconds are split evenly into chunks
of at most 30 seconds, cut at pauses, and encoded independently.
Each modality serves only checkpoints that carry it: 26B/31B have no
audio config and reject audio input, and checkpoints with an
unrecognized vision architecture still load as text-only models and
reject image requests.
The server previously hid the vision and audio capabilities for
gemma4 safetensors because the engine served neither. Both
suppressions are removed, and existing imports start advertising the
capabilities without re-importing since import already records them.
Nothing published for this version
llama.cpp: version bump b10760
llama.cpp: version bump b10760 (#18199)
Nothing published for this version
Your coding agent can read these notes before it upgrades. Set up the MCP server →