NewYour coding agent can read the release notes before it upgrades.Set up the MCP server →
PyPI · #1978 most downloaded on PyPI
Library for utilization of compressed safetensors of neural network models
Last release 2 days ago
03 Oct 2026
Ships on a steady schedule
a new release about every 8 days
Most releases are documented
notes for 29 of 33 stable releases
Nothing withdrawn
no release was ever pulled
2 years old
217 releases · first in 2024
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
One column per month.
Nothing published for this version
Nothing published for this version
Nothing published for this version
Multi-GPU Checkpoint Conversion #849 , #869 , #889 : A dynamic, capacity-first multi-GPU scheduler is now built into entrypoints/convert . Conversion
entrypoints/convert. Conversion jobs are scheduled across accelerators based on measured free memory, with a profiler-based (TorchDispatchMode) memory estimator replacing heuristic multipliers.MagnitudeExpertPruner converter that structurally prunes MoE experts by magnitude, with a Qwen3-VL example. Built on the new multi-converter chaining interface (#805), which lets conversion pipelines chain multiple converters (e.g. prune → quantize) with validate/config propagation through the chain.model_compressor now supports decompression under distributed execution, with fixes for dtype mismatch (#851) and tensor-shape reconciliation in the distributed CPU/disk offload caches (#857).requires_version now requires transformers>=5.16.The entrypoints/convert pipeline can now schedule conversion work across multiple accelerators:
exec_jobs_dynamic, _pick_device, _free_bytes) was ported from llm-compressor as a reusable, protocol-agnostic building block for PTQ pipelines (#849).x3 multiplier heuristic was replaced by a TensorProfiler (TorchDispatchMode) that measures true peak memory during the converter validate pass, falling back to a 2.5x input footprint estimate with a warning if profiling fails (#889).MagnitudeExpertPruner converter (compressed_tensors.entrypoints.convert.converters.MagnitudeExpertPruner) removes low-magnitude MoE experts structurally (#834), with a Qwen3-VL expert-pruning example. New MoE pruning utilities were added under compressed_tensors/utils/moe.py.setStorage: out of bounds on tied/multimodal models)._estimate_tensor_count.from_accelerate now sets the proper execution device for container (param-less) modules when converting an accelerate device_map dispatch to compressed-tensors offloading (#884).print_offload_folder utility for inspecting offloaded model folders (#738).contiguous fix in the disk cache (#874).qstrategy is no longer inferred automatically.QuantizationArgs in this repo.QuantizedKVCache now delegates structural reads (e.g. past_key_values.layers[idx].keys) to the wrapped Cache (#747), fixing AttributeError for transformers>=5 models (e.g. DiffusionGemma) that read the KV cache structurally during calibration instead of only through update().
_round_to_fp4 Triton kernel for native dtype support, correcting wrong rounding of fp32/fp16 values near FP4 boundaries (#824); the Triton kernel is now enabled for cast_to_fp4 only (#827).clear_all_qparams behavior on load (#890).Gated module type classification (#871).QuantizedKVCache structural-read fix (see above).calculate_range memoization #855: the range tensor is now cached per (type, num_bits, device), removing a synchronous host-to-device copy per call (called ~245k times per 0.5B model under GPTQ W4A16).transformers>=5.16 is now required #848make quality added (#863); Black target version updated to py310 (#789).ImplBackend dispatch coverage and an overhead benchmark (#882); un-skipped test_pack_fp4_to_uint8_backends_match now that Triton kernels are re-enabled (#881).test_args_full now actually asserts on the parsed TransformArgs (#892); broken convert-checkpoint tests repaired (#887); frozen model loading tests skipped to work around a transformers bug (#835).g_idx initialization refactored to use torch.empty (#761); DynamicType.LOCAL constraint clarified to TENSOR_GROUP (#778).llama_1.1b examples.print_offload_folder util by @kylesayrs in #738from_accelerate by @kylesayrs in #884clear_all_qparams onload by @kylesayrs in #890Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Add SECURITY.md with vulnerability reporting policy by @arijitroy003 in https://github.com/vllm-project/compressed-tensors/pull/794
load_offloaded_model by @kylesayrs in https://github.com/vllm-project/compressed-tensors/pull/728apply_quantization_config by @kylesayrs in https://github.com/vllm-project/compressed-tensors/pull/730to_accelerate by only instantiating OffloadedWeightsLoader once by @kylesayrs in https://github.com/vllm-project/compressed-tensors/pull/752__setitem__/offload from update_offload by @kylesayrs in https://github.com/vllm-project/compressed-tensors/pull/709cast_to_fp4 by @kylesayrs in https://github.com/vllm-project/compressed-tensors/pull/774save_mtp_tensors_to_checkpoint graceful failure by @kylesayrs in https://github.com/vllm-project/compressed-tensors/pull/788ImplBackend and triton import utils by @kylesayrs in https://github.com/vllm-project/compressed-tensors/pull/804Full Changelog: https://github.com/vllm-project/compressed-tensors/compare/0.17.0...0.18.0
load_offloaded_model by @kylesayrs in #728apply_quantization_config by @kylesayrs in #730to_accelerate by only instantiating OffloadedWeightsLoader once by @kylesayrs in #752__setitem__/offload from update_offload by @kylesayrs in #709cast_to_fp4 by @kylesayrs in #774save_mtp_tensors_to_checkpoint graceful failure by @kylesayrs in #788ImplBackend and triton import utils by @kylesayrs in #804Full Changelog: 0.17.0...0.18.0
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Add Buildkite pipeline for test-check workflow on H100 duo and L4 solo by @deepak-kumar-neu in https://github.com/vllm-project/compressed-tensors/pull
load_offloaded_model by @kylesayrs in https://github.com/vllm-project/compressed-tensors/pull/728Full Changelog: https://github.com/vllm-project/compressed-tensors/compare/0.17.0...0.17.1
load_offloaded_model by @kylesayrs in #728Full Changelog: 0.17.0...0.17.1
Nothing published for this version
Nothing published for this version
Nothing published for this version
Support arbitrary 1-8 bit INT weight packing in pack-quantized by @mgoin in https://github.com/vllm-project/compressed-tensors/pull/715
Full Changelog: https://github.com/vllm-project/compressed-tensors/compare/0.16.0...0.17.0
Full Changelog: 0.16.0...0.17.0
Nothing published for this version
Nothing published for this version
Nothing published for this version
[Deprecation] [Offload] Remove deprecated offload functions by @kylesayrs in https://github.com/vllm-project/compressed-tensors/pull/685
is_source_process by @kylesayrs in https://github.com/vllm-project/compressed-tensors/pull/664get_cache_kwargs by @kylesayrs in https://github.com/vllm-project/compressed-tensors/pull/694load_offloaded_model by @kylesayrs in https://github.com/vllm-project/compressed-tensors/pull/708Full Changelog: https://github.com/vllm-project/compressed-tensors/compare/0.15.0.1...0.16.0
is_source_process by @kylesayrs in #664get_cache_kwargs by @kylesayrs in #694load_offloaded_model by @kylesayrs in #708Full Changelog: 0.15.0.1...0.16.0
Your coding agent can read these notes before it upgrades. Set up the MCP server →