NewYour coding agent can read the release notes before it upgrades.Set up the MCP server →
Go modules · #2960 by repository stars
Last release 9 days ago
29 Sep 2026
Release timing varies
gaps range from 8 days to 5 months
Rarely documented
notes for 14 of the last 60 stable releases
Nothing withdrawn
no release was ever pulled
2 years old
153 releases · first in 2024
fix: keep letter-suffix quants out of the fine-tune field
fix: keep letter-suffix quants out of the fine-tune field
Q2_K_XL and IQ2_XXS matched the fine-tune pattern, so the encoding came back empty. LoRA was stored there too.
Nothing published for this version
One column per month.
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
ci: configure project review rules for open-code-review
ci: configure project review rules for open-code-review
- Include test files and Markdown, which OCR's defaults would skip
- Exclude generated code and generator scaffolding
- Add path-scoped rules for tests, estimators, util helpers, the CLI,
the root parsing core, go.mod, workflows, and docs
- Emphasize boundary checking on untrusted GGUF input, following the
recent ReadString/ReadArray hardening
Signed-off-by: thxCode <thxcode0824@gmail.com>
fix: Bound the declared string length in ReadString before allocating
fix: Bound the declared string length in ReadString before allocating
ReadString allocated a buffer sized by the declared length without checking it against the file. Mirror the existing SkipReadingString / _validateCountWithRemaining guards: reject a length above math.MaxInt64, and, when the file size is known, a length beyond the bytes remaining in the file.
Adds TestReadStringRejectsOversizedLength.
refactor: Stop two parse panics, refuse a typeless projector, and cor…
refactor: Stop two parse panics, refuse a typeless projector, and cor…
…rect three graph shapes (#49)
* Return zero bytes for a tensor with an empty dimension
GGUFTensorInfo.Bytes subtracts 1 from each dimension to compute the
strides. A zero-length dimension is unsigned, so the subtraction wraps to
2^64-1 and the overflow guard refuses the multiplication, which panics the
parse. ggml_nbytes returns 0 for such a tensor before it reaches the same
arithmetic.
Return 0 when any dimension is zero.
mudler/face-detect-gguf/antelopev2.gguf holds one such tensor out of 748,
det.392 with dimensions [0]. It now parses.
* Refuse the files llama.cpp refuses to load
Two kinds of file panicked the parser or returned a confident estimate for
a model that cannot start. Both are now refused while the file is parsed,
so the caller gets an error.
general.file_type written as a string panicked in ValueNumeric, which a
caller cannot catch. llama.cpp reads the key as a uint32 and throws "key
general.file_type has wrong type". Four Serveurperso repositories publish
files on this schema: ACE-Step, MiniMax-Music3, OmniVoice and Qwen3-TTS.
A projector that names no type was reported as an mlp projector. llama.cpp
resolves the type from clip.projector_type, then from the key of the
modality it is loading, and refuses the load with "unknown projector type"
when neither names one. mys/ggml_llava-v1.5-7b carries no such key and was
estimated at 6866 MiB with no warning; the same is true of every mmproj on
the format that predates the key.
The fallback chain now also reads clip.gen.audio.projector_type, so a
projector that names its type only there is neither refused nor reported
as mlp. The invented mlp default is gone: a parsed clip file always names
its own type.
ValueNumeric's type list moves to GGUFMetadataValueType.IsNumeric so the
parse check and the accessor agree on what a number is.
* Charge RWKV7's decay projection as the two-dimensional node it is
RWKV6 and RWKV7 both name a tensor "time_mix_w2". RWKV6 reshapes it to
[ne0, ne1, 1, 5] and multiplies the five-way lerp stack, so the node is
four-dimensional and its last dimension is that fan-out. RWKV7 gives the
name to a plain decay projection whose node is [n_embd, n_tokens]. The
estimate applied RWKV6's shape to both, and RWKV7's last dimension is the
embedding length rather than a small fan-out, so the node was charged
n_embd times over: 8192 MiB of a 8216 MiB compute graph on a 1.5B model.
Size the node by architecture.
Measured on an A40 with llama-server, context 2304, one sequence, flash
attention, q8_0 key and value cache. Device total, estimate against the
figure llama.cpp reports for itself, in MiB:
rwkv7-g1c-1.5b 9710 -> 1523, really uses 1562
ARWKV-7B-Preview 33085 -> 8004, really uses 8089
rwkv-6-world-1.6b 1604 unchanged, really uses 1600
Nemotron3-Nano-4B 2636 unchanged, really uses 2636
Both RWKV6 and the other recurrent architectures are byte-identical: the
new branch is reached only by rwkv7 and arwkv7, which share the same graph
builder in llama.cpp.
* Charge gemma3's vision encoder over the grid it attends to
gemma3 pools its 64x64 patch grid 4x4 into 256 projector tokens. The
pooling runs after the vision transformer, so the encoder still attends
over all 4096 patches, and its attention term is quadratic in that count.
The estimate used the projector's output token count for both, which
understates the attention scores by the square of the pooling factor.
The estimate already carries the un-pooled position count, but only on the
path that handles a projector type it does not recognise. Set it in the
gemma3 case as well.
Measured on an A40 with llama-server, context 4096, one sequence. The
projector's own device total, estimate against llama.cpp's own figure, in
MiB:
flash attention off 831 -> 1918, really uses 1944
flash attention on 827 -> 926, really uses 933
This is also what made flash attention look free for a projector: the
switch is worth 1011 MiB on gemma3 and was modelled as 4, because the term
it halves was computed over 256 positions instead of 4096. The estimate now
moves 992 MiB across the same switch.
idefics3, llama4 and internvl reduce their patch count the same way and are
left alone, because they are not measured here. Every other projector type
is byte-identical.
* Reserve a dynamic-resolution projector for llama.cpp's warm-up image
The estimate guessed the largest image a caller might send, defaulting to
1024 pixels a side. llama.cpp reserves the projector's graph for one square
warm-up image instead, and derives its side from a token count carried in
its own source rather than from anything the file declares:
warmup_image_size = sqrt(n_tokens) * patch_size * n_merge
The Qwen vision family warms up at 46*46 tokens and pixtral at 256, so the
encoder attends over 8464 positions on the first and 256 on the second,
against the 5476 and 4096 the guess produced. The attention term is
quadratic in that count, which is why three families read low and pixtral
read about seventy times high: one wrong assumption, not four errors.
Carry llama.cpp's token counts and size the warm-up image the same way.
llama.cpp prints the size it uses, so the rule is checked against its own
figure rather than inferred from the buffer. Measured on an A40:
gemma3 896 x 896, the declared image size, unchanged
qwen2vl 1288 x 1288 = 46 * 14 * 2
qwen2.5vl 1288 x 1288 = 46 * 14 * 2
pixtral 256 x 256 = 16 * 16 * 1
The projector's own compute buffer, estimate against llama.cpp's own
reserve_compute_meta figure, context 4096, one sequence, flash attention
off then on, in MiB:
qwen2vl 1937 / 164 -> 4567 / 331, really uses 4598 / 267
qwen2.5vl 1937 / 164 -> 4840 / 605, really uses 4872 / 705
qwen3vl 360 / 46 -> 4538 / 302, really uses 4728 / 322
pixtral 1121 / 98 -> 11 / 6, really uses 16 / 16
The flash-attention-off column governs whether a model fits, and it lands
within 4 percent on the three Qwen families. Flash attention on stays
looser, and pixtral is now low by 5 MiB rather than high by 1105.
Only these four projector types are affected. Every other type reserves for
the image size it declares, which is already what the estimate uses, and
gemma3, SmolVLM, InternVL3, MiniCPM-V and ultravox were measured byte for
byte unchanged.
Nothing published for this version
Nothing published for this version
fix: resolve the sliding window layout per layer, mirror llama.cpp's …
fix: resolve the sliding window layout per layer, mirror llama.cpp's …
…per-architecture periods
The sliding window pattern was a five-entry switch (llama4, phi3, gemma2,
gemma3, cohere2) while llama.cpp interleaves windows on 21 architectures.
Every architecture missing from the switch was charged a full, context-scaled
KV cache on all blocks: gpt-oss-20b about 2x (6.0 GiB instead of about 3.0 GiB
of KV at 131072 context), gemma-3n-E4B about 6.5x.
Read <architecture>.attention.sliding_window_pattern when the file declares it
(a period as a scalar, or the per-layer use as a bool array, as Gemma 4 writes),
and fall back to a per-architecture period table mirrored from llama.cpp's
src/models/*.cpp, including the dense-first phasing the old modulo could not
express. The resolved layout lives in AttentionSlidingWindowLayers and the
estimator charges each layer by its own kind.
Also honour the layers that hold no cache of their own: the layers behind
n_layer_kv_from_start reuse leading caches (gemma3n hardcodes 20, Gemma 4
derives it from attention.shared_kv_layers), and the NextN/MTP blocks of the
hybrid family are filtered out of the main context cache. Read
attention.key_length_swa / value_length_swa so windowed layers size their
heads separately (Gemma 4).
Signed-off-by: thxCode <thxcode0824@gmail.com>
docs: readme
Signed-off-by: thxCode <thxcode0824@gmail.com>
fix: integer overflow in tensor elements/bytes calculation
fix: integer overflow in tensor elements/bytes calculation
Signed-off-by: thxCode <thxcode0824@gmail.com>
Nothing published for this version
v0.24.0 Compare # Choose a tag to compare
v0.24.0
Compare
Signed-off-by: thxCode <thxcode0824@gmail.com>
fix: rerank detection
Signed-off-by: thxCode <thxcode0824@gmail.com>
fix: --token now works with --url, add --header for custom HTTP heade…
fix: --token now works with --url, add --header for custom HTTP heade…
…rs (#16)
* fix: --token now works with --url, add --header for custom HTTP headers
* update readme
Signed-off-by: thxCode <thxcode0824@gmail.com>
refactor: support imatrix
Signed-off-by: thxCode <thxcode0824@gmail.com>
Nothing published for this version
Nothing published for this version
v0.22.0 Compare # Choose a tag to compare
v0.22.0
Compare
fix: embedding model estimation
fix: embedding model estimation
Signed-off-by: thxCode <thxcode0824@gmail.com>
v0.21.3 Compare # Choose a tag to compare
v0.21.3
Compare
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Your coding agent can read these notes before it upgrades. Set up the MCP server →