NewYour coding agent can read the release notes before it upgrades.Set up the MCP server →
PyPI · #3890 most downloaded on PyPI
A framework for machine learning on Apple silicon.
Last release 6 days ago
29 Sep 2026
Ships fairly regularly
a new release about every 4 weeks
Some releases are documented
notes for 31 of 59 stable releases
Nothing withdrawn
no release was ever pulled
3 years old
59 releases · first in 2023
RDMA over thunderbolt with the JACCL backend (macOS >= 26.2) ( some numbers )
mx.core.load type annotation by @CC-Yeh in #2819mx.core.linspace type annotation by @CC-Yeh in #2820Numpy interfaces for masked_scatter by @CC-Yeh in #2832Full Changelog: v0.30.0...v0.30.1
One column per month.
Support for Neural Accelerators on M5 (macOS >= 26.2)
mx.depends to Python by @awni in https://github.com/ml-explore/mlx/pull/2606compile_commands.json by @andportnoy in https://github.com/ml-explore/mlx/pull/2645mx.median op by @awni in https://github.com/ml-explore/mlx/pull/2705zeros/ones_like by @CC-Yeh in https://github.com/ml-explore/mlx/pull/2726Full Changelog: https://github.com/ml-explore/mlx/compare/v0.29.0...v0.30.0
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Support for mxfp4 quantization (Metal, CPU)
mxfp4 quantization (Metal, CPU)mx.distributed supports NCCL back-end for CUDAcudaMemAdvise and cudaGraphAddDependencies for CUDA 13 by @andportnoy in https://github.com/ml-explore/mlx/pull/2525Path by @awni in https://github.com/ml-explore/mlx/pull/2543Full Changelog: https://github.com/ml-explore/mlx/compare/v0.28.0...v0.29.0
First version of fused sdpa vector for CUDA
Full Changelog: https://github.com/ml-explore/mlx/compare/v0.27.1...v0.28.0
Initial PyPi release of the CUDA back-end.
update and update_modules by @awni in https://github.com/ml-explore/mlx/pull/2239update_modules() when providing a subset by @angeloskath in https://github.com/ml-explore/mlx/pull/2308Full Changelog: https://github.com/ml-explore/mlx/compare/v0.26.0...v0.27.0
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Peaking at 23.89 TFlops on M2 Ultra benchmarks
mx.linalg.eigh and mx.linalg.eigvalshmx.nn.init.sparsemx.cumprod, mx.cumsummx.random.uniform and mx.random.bernoullimx.vmap with gather and constant outputsmx.array.__format__ with specNothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Block sparse matrix multiply speeds up MoEs by >2x
mx.bitwise_[or|and|xor], mx.[left|right]_shift, operator overloadsmx.metal.device_info to get better informed memory limitsmlx.optimizers.clip_grad_norm and mlx.utils.tree_reduce addedmx.arctan2mx.sqrt(2)Nothing published for this version
Nothing published for this version
Improvements for LLM generation
mx.async_evalmx.metal.start_capture and mx.metal.stop_capture for GPU debug/profilemx.expm1mx.stdmx.meshgridmx.random.multivariate_normalmx.cumsum (and other scans) for bfloatnn.upsample support bicubic interpolationNothing published for this version
Nothing published for this version
Perf improvements for attention ops:
mx.linalg.svd (CPU only)nn.RNN, nn.LSTM, nn.GRUFaster quantized matrix-vector multiplies
mx.fast.scaled_dot_product_attention fused opmx.fast.scaled_dot_product_attention fused opmx.arraymx.topkNothing published for this version
Default shapeless compilation for all activations
mx.compile(function, shapeless=True)mx.atleast_1d, mx.atleast_2d, mx.atleast_3dtolist with bfloat16 and float16argmax on M3Custom mx.fast.rope up to 20x faster
mx.fast subpackagemx.fast.rope up to 20x fastersafetensorsbfloat16 quantizated matrix-vector multipliesmx.fast subpackage with a fast RoPEmx.stream to set the default deviceoptimizers.step_decayoptimizers.cosine_decayopimtizers.exponential_decaySome functions are up to 10x faster (benchmarks)
mx.compile makes stuff go fast
mx.compile function transformation__abs__ overload for abs on arraysloc and scale in parameter for mx.random.normalmx.var to give inf with doff >= nelemnn.SequentialGradient checkpointing for training with mx.checkpoint
mx.checkpointmx.checkpointmx.linalg.qrmx.evalmx.diag, mx.diagonalarray.shape is a Python tupleint64 and uint64 reductionssum, prod, max, min, all, anyargmax, argmininf work, and fix mx.isinfmx.fullNaN in some binary ops
mx.logaddexp, mx.maximum, mx.minimummx.log1p with inf inputNative quantizations Q4_0, Q4_1, and Q8_0
Q4_0, Q4_1, and Q8_0Q4_0, Q4_1, and Q8_0)Module.save_weights supports safetensorsnn.init package with several commonly used neural network initializersAdafactor in nn.optimizersisinf and friends for integer typesint64, uint, and float320 inputsinf reads in gemvmx.arange crashes on NaN inputsFaster matmul: up to 2.5x faster for certain sizes, benchmarks
mx.isnan, mx.isinf, isposinf, isneginfmx.tilescatter_min and scatter_maxmx.eyemx.round to follow NumPy which rounds to evenInitial (and experimental) GGUF support
at[] syntax for scatter style operations: x.at[idx].add(y), (min, max, prod, etc)mx.array([x, y]))mx.inner, mx.outer+=, *=, -=, ...)mx.pi, mx.inf, mx.newaxis, …)cosine_similarity lossRoPE and ALiBitriretain_graphSupport for loading and saving HuggingFace's safetensor format
mlx.core.linalg sub-package with mx.linalg.norm (Frobenius, infininty, p-norms)tensordot and repeatBilinear,Identity, InstanceNormDropout2D, Dropout3DTransformer (pre/post norm, dropout)SoftSign, Softmax, HardSwish, LogSoftmaxRoPE positional encodingshinge, huber, log_coshquantize, dequantize, quantized_matmul
QuantizedLinear, ALiBi positional encodingsCore ops remainder, eye, identity
remainder, eye, identitymlx.nn
Nothing published for this version
Nothing published for this version
Nothing published for this version
Your coding agent can read these notes before it upgrades. Set up the MCP server →