NewYour coding agent can read the release notes before it upgrades.Set up the MCP server →
MLX-VLM is a package for inference and fine-tuning of Vision Language Models (VLMs) and Omni Models (VLMs with audio and video support) on your Mac using MLX.
Last release 6 days ago
28 Sep 2026
Ships fairly regularly
a new release about every 2 weeks
Nearly every release is documented
notes for 60 of the last 60 stable releases
Nothing withdrawn
no release was ever pulled
2 years old
86 releases · first in 2024
Generate multiple images per prompt via num_images (ming_image) by @Lazarus-931 in #2353
Full Changelog: v0.7.3...v0.7.4
One column per month.
Hemmingway-1 support by @Lazarus-931 in #2330
Full Changelog: v0.7.2...v0.7.3
Downmix stereo audio before resampling in load_audio by @kiarina in #2258
Full Changelog: v0.7.1...v0.7.2
Add standalone DINOv2 model with register tokens and mask support by @eklipse2k8 in #2184
Full Changelog: v0.7.0...v0.7.1
Optimize Qwen4 n-gram shard gathers by @pyros-projects in #2034
Full Changelog: v0.6.17...v0.7.0
Optimize Qwen4 n-gram shard gathers by @pyros-projects in #2034
Full Changelog: v0.6.17...v0.7.0rc0
Fix MiniMax-M3 sparse-index compile re-trace leak (Metal 499k resource limit) by @ziomancer in #2016
Full Changelog: v0.6.16...v0.6.17
Warn when resize_shape is ignored by a custom image processor by @eptan in #1829
Full Changelog: v0.6.15...v0.6.16
Stop a batched row's answer depending on its neighbour's length by @Lazarus-931 in #1946
Full Changelog: v0.6.14...v0.6.15
Fix multimodal server prefill alignment by @Blaizzy in #1872
Full Changelog: v0.6.13...v0.6.14
Preserve APC prefix reuse for growing prepared prompts by @byoan in #1713
Full Changelog: v0.6.12...v0.6.13
Fix LFM tool call parsing and bump version to 0.6.11 by @Blaizzy in #1792
Full Changelog: v0.6.10...v0.6.12
Add and improve MLX-VLM agent skills by @Lazarus-931 in #1747
Full Changelog: v0.6.9...v0.6.10
Add Kimi K3 by @kernelpool in #1746
Full Changelog: v0.6.8...v0.6.9
Vendor standard models by @Lazarus-931 in #1702
Full Changelog: v0.6.7...v0.6.8
Support Laguna S by @Blaizzy in #1650
Full Changelog: v0.6.6...v0.6.7
vendor Gemma/Phi/GLM/Cohere text models and reuse shared MLPs by @Lazarus-931 in #1616
Full Changelog: v0.6.5...v0.6.6
Add deepseek_v3 and deepseek_v32 model types by @Lazarus-931 in #1517
Full Changelog: v0.6.4...v0.6.5
Add support for TTS and STT endpoints in the server by @lucasnewman in https://github.com/Blaizzy/mlx-vlm/pull/1358
Full Changelog: https://github.com/Blaizzy/mlx-vlm/compare/v0.6.3...v0.6.4
Fix APC splitting on multimodal inputs by @lucasnewman in https://github.com/Blaizzy/mlx-vlm/pull/1311
Full Changelog: https://github.com/Blaizzy/mlx-vlm/compare/v0.6.2...v0.6.3
Make sure to evaluate cached MRoPE parameters on model load by @lucasnewman in https://github.com/Blaizzy/mlx-vlm/pull/1272
Full Changelog: https://github.com/Blaizzy/mlx-vlm/compare/v0.6.1...v0.6.2
Move linear spec handling to model backends by @Blaizzy in https://github.com/Blaizzy/mlx-vlm/pull/1259
Full Changelog: https://github.com/Blaizzy/mlx-vlm/compare/v0.6.0...v0.6.1
fix: replace deprecated use_fast=True with backend=torchvision in load_processor by @Jonathangadeaharder in https://github.com/Blaizzy/mlx-vlm/pull/11…
finish_reason inconsistencies by @spicyneuron in https://github.com/Blaizzy/mlx-vlm/pull/1215usage and timings for server endpoints by @spicyneuron in https://github.com/Blaizzy/mlx-vlm/pull/1216not use_rht (fix masked decode under RHT) by @popfido in https://github.com/Blaizzy/mlx-vlm/pull/1244Full Changelog: https://github.com/Blaizzy/mlx-vlm/compare/v0.5.0...v0.6.0
Fix gemma4 multi-image processing for different-sized images by @Blaizzy in https://github.com/Blaizzy/mlx-vlm/pull/938
json_schema response_format support by @avbiswas in https://github.com/Blaizzy/mlx-vlm/pull/1047--max-tokens to server by @spicyneuron in https://github.com/Blaizzy/mlx-vlm/pull/1120Full Changelog: https://github.com/Blaizzy/mlx-vlm/compare/v0.4.4...v0.5.0
Fix Gemma 4 chunked prefill for KV-shared models and thinking by @Blaizzy in https://github.com/Blaizzy/mlx-vlm/pull/901
Full Changelog: https://github.com/Blaizzy/mlx-vlm/compare/v0.4.3...v0.4.4
Add SAM 3.1 with Object Multiplex and optimized realtime pipeline by @Blaizzy in https://github.com/Blaizzy/mlx-vlm/pull/880
Full Changelog: https://github.com/Blaizzy/mlx-vlm/compare/v0.4.2...v0.4.3
Ensure minimal metadata exists in model card by @pcuenca in https://github.com/Blaizzy/mlx-vlm/pull/853
Full Changelog: https://github.com/Blaizzy/mlx-vlm/compare/v0.4.1...v0.4.2
feat(server): add --model and --adapter-path flags for startup preloading by @auggie246 in https://github.com/Blaizzy/mlx-vlm/pull/811
Full Changelog: https://github.com/Blaizzy/mlx-vlm/compare/v0.4.0...v0.4.1
Fix gemma3n short prompts by @Blaizzy in https://github.com/Blaizzy/mlx-vlm/pull/751
Full Changelog: https://github.com/Blaizzy/mlx-vlm/compare/v0.3.12...v0.4.0
[MODEL] support qwen3.5 series by @JJJYmmm in https://github.com/Blaizzy/mlx-vlm/pull/722
Full Changelog: https://github.com/Blaizzy/mlx-vlm/compare/v0.3.11...v0.3.12
Refactor input embedding handling in _generate_batch function by @Blaizzy in https://github.com/Blaizzy/mlx-vlm/pull/694
Full Changelog: https://github.com/Blaizzy/mlx-vlm/compare/v0.3.10...v0.3.11
Fix qk_norm for lighton by @Blaizzy in https://github.com/Blaizzy/mlx-vlm/pull/615
Full Changelog: https://github.com/Blaizzy/mlx-vlm/compare/v0.3.9...v0.3.10
Fix qwen3_vl ValueError: Image features and image tokens do not match by @ziya32 in https://github.com/Blaizzy/mlx-vlm/pull/608
Full Changelog: https://github.com/Blaizzy/mlx-vlm/compare/v0.3.8...v0.3.9
Add chat_ui command by @Blaizzy in https://github.com/Blaizzy/mlx-vlm/pull/589
Full Changelog: https://github.com/Blaizzy/mlx-vlm/compare/v0.3.7...v0.3.8
update readme with new openai endpoints details by @mguella in https://github.com/Blaizzy/mlx-vlm/pull/585
Full Changelog: https://github.com/Blaizzy/mlx-vlm/compare/v0.3.6...v0.3.7
Fix: Input cast error by @Blaizzy in https://github.com/Blaizzy/mlx-vlm/pull/559
Full Changelog: https://github.com/Blaizzy/mlx-vlm/compare/v0.3.5...v0.3.6
Remove docs actions temporarly by @Blaizzy in https://github.com/Blaizzy/mlx-vlm/pull/536
Full Changelog: https://github.com/Blaizzy/mlx-vlm/compare/v0.3.4...v0.3.5
Add n_kv_heads property used in LM Studio to glm4v_moe.LanguageModel by @hehua2008 in https://github.com/Blaizzy/mlx-vlm/pull/472
Full Changelog: https://github.com/Blaizzy/mlx-vlm/compare/v0.3.3...v0.3.4
fix changelog task by @Blaizzy in https://github.com/Blaizzy/mlx-vlm/pull/441
Full Changelog: https://github.com/Blaizzy/mlx-vlm/compare/v0.3.2...v0.3.3
Fix energy calc in omni.py by @Blaizzy in https://github.com/Blaizzy/mlx-vlm/pull/427
pyproject.toml by @SauravMaheshkar in https://github.com/Blaizzy/mlx-vlm/pull/282Full Changelog: https://github.com/Blaizzy/mlx-vlm/compare/v0.3.1...v0.3.2
fix(chat-ui): Fix imports, blocking Chat-ui server start by @zenyr in https://github.com/Blaizzy/mlx-vlm/pull/420
Full Changelog: https://github.com/Blaizzy/mlx-vlm/compare/v0.3.0...v0.3.1
[gemma3n] Correctly scale text embeddings for quantized gemma3n conversions by @neilmehta24 in https://github.com/Blaizzy/mlx-vlm/pull/397
Full Changelog: https://github.com/Blaizzy/mlx-vlm/compare/v0.2.0...v0.3.0
Fix Gemma 3 config by @DePasqualeOrg in https://github.com/Blaizzy/mlx-vlm/pull/384
Full Changelog: https://github.com/Blaizzy/mlx-vlm/compare/v0.1.27...v0.2.0
Fix README POST request typo by @Blaizzy in https://github.com/Blaizzy/mlx-vlm/pull/358
Full Changelog: https://github.com/Blaizzy/mlx-vlm/compare/v0.1.26...v0.1.27
When converting models, default dtype to config.json's torch_dtype by @neilmehta24 in https://github.com/Blaizzy/mlx-vlm/pull/323
Full Changelog: https://github.com/Blaizzy/mlx-vlm/compare/v0.1.25...v0.1.26
Fix snapshot download by @Blaizzy in https://github.com/Blaizzy/mlx-vlm/pull/318
Full Changelog: https://github.com/Blaizzy/mlx-vlm/compare/v0.1.24...v0.1.25
Deps: Remove Scipy MLX by @Blaizzy in https://github.com/Blaizzy/mlx-vlm/pull/301
Full Changelog: https://github.com/Blaizzy/mlx-vlm/compare/v0.1.23...v0.1.24
Add internVL3 by @Blaizzy in https://github.com/Blaizzy/mlx-vlm/pull/295
Full Changelog: https://github.com/Blaizzy/mlx-vlm/compare/v0.1.22...v0.1.23
chore: fix sliding mask by @FL33TW00D in https://github.com/Blaizzy/mlx-vlm/pull/272
Full Changelog: https://github.com/Blaizzy/mlx-vlm/compare/v0.1.21...v0.1.22
fix remaining DEFAULT_TEMP by @omercelik in https://github.com/Blaizzy/mlx-vlm/pull/265
Full Changelog: https://github.com/Blaizzy/mlx-vlm/compare/v0.1.20...v0.1.21
Fix return tensors by @Blaizzy in https://github.com/Blaizzy/mlx-vlm/pull/262
Full Changelog: https://github.com/Blaizzy/mlx-vlm/compare/v0.1.19...v0.1.20
Copy all json files by @Blaizzy in https://github.com/Blaizzy/mlx-vlm/pull/249
Full Changelog: https://github.com/Blaizzy/mlx-vlm/compare/v0.1.18...v0.1.19
Fix Gemma 3 Text-only and prompt template by @Blaizzy in https://github.com/Blaizzy/mlx-vlm/pull/248
Full Changelog: https://github.com/Blaizzy/mlx-vlm/compare/v0.1.17...v0.1.18
Add support for Gemma 3 by @pcuenca @FL33TW00D and @Blaizzy in https://github.com/Blaizzy/mlx-vlm/pull/235
Full Changelog: https://github.com/Blaizzy/mlx-vlm/compare/v0.1.16...v0.1.17
Fix aya vision 32b by @Blaizzy in https://github.com/Blaizzy/mlx-vlm/pull/226
Full Changelog: https://github.com/Blaizzy/mlx-vlm/compare/v0.1.15...v0.1.16
add status badge by @Blaizzy in https://github.com/Blaizzy/mlx-vlm/pull/215
Full Changelog: https://github.com/Blaizzy/mlx-vlm/compare/v0.1.14...v0.1.15
Fix generate() call in object_detection.ipynb to follow the latest parameter order by @JoeJoe1313 in https://github.com/Blaizzy/mlx-vlm/pull/200
Full Changelog: https://github.com/Blaizzy/mlx-vlm/compare/v0.1.13...v0.1.14
Add Video Understanding Support (beta) by @Blaizzy in https://github.com/Blaizzy/mlx-vlm/pull/97
Full Changelog: https://github.com/Blaizzy/mlx-vlm/compare/v0.1.12...v0.1.13
Add support for Qwen2-5-VL by @Blaizzy in https://github.com/Blaizzy/mlx-vlm/pull/189
Full Changelog: https://github.com/Blaizzy/mlx-vlm/compare/v0.1.11...v0.1.12
Chat in CLI. by @chigkim in https://github.com/Blaizzy/mlx-vlm/pull/168
Full Changelog: https://github.com/Blaizzy/mlx-vlm/compare/v0.1.10...v0.1.11
Your coding agent can read these notes before it upgrades. Set up the MCP server →