NewYour coding agent can read the release notes before it upgrades.Set up the MCP server →
PyPI · #5842 most downloaded on PyPI
Neural Networks Compression Framework
Last release 17 days ago
17 Sep 2026
Ships on a steady schedule
a new release about every 2 months
Nearly every release is documented
notes for 33 of 35 stable releases
Nothing withdrawn
no release was ever pulled
6 years old
35 releases · first in 2020
One column per quarter.
(PyTorch) Replaced the deprecated export_for_training with export in TorchFX examples, documentation, and tests ( #4191 ).
GroupedMatMul support to the AWQ and Scale Estimation algorithms, enabling data-aware weight compression of grouped_mm-based MoE models (#4176).matmul -> transpose -> cos/sin structure used by models such as GPT-OSS (#4175).__getitem__ node that follows split-like operations (e.g. chunk) from the TorchFX inference graph, fixing statistic collection errors for models such as YOLO11 (#4155).KeyError in bias attribute resolution when a tensor is a model input and therefore has no parent node (#4169).ONNXEmbeddingMetatype nodes (#4144).export_for_training with export in TorchFX examples, documentation, and tests (#4191).Enabled data-aware compression methods (AWQ, GPTQ, Scale Estimation, and LoRA Correction) for all bit widths < 4, including INT2 and INT3 ( #4131 ).
GroupedMatMul operation support for data-free weight compression (#4130).safetensors >= 0.8.0 by converting NumPy scalars to arrays before saving (#4094).GroupedMatMul operation support for data-free weight compression (https://github.com/openvinotoolkit/nncf/pull/4130).safetensors >= 0.8.0 by converting NumPy scalars to arrays before saving (https://github.com/openvinotoolkit/nncf/pull/4094).Statistics, NNCFDataLoader, QuantizersCounter, and QuantizationStatistics classes, and the PyTorch patch_torch_operators, register_module, and PTInitializingDataLoader (https://github.com/openvinotoolkit/nncf/pull/4093).CompressionParameter class (https://github.com/openvinotoolkit/nncf/pull/4108).onnx to 1.22.0 (https://github.com/openvinotoolkit/nncf/pull/4101).networkx to <= 3.6.1 (https://github.com/openvinotoolkit/nncf/pull/4059).Removed deprecated CompressWeightsMode.INT8 mode. Use CompressWeightsMode.INT8_SYM or CompressWeightsMode.INT8_ASYM instead ( #4008 ).
QuantizationConstraints.from_config_dict to avoid unexpected behavior (#4046).initial_steps and scale_steps in calculate_quantization_params (#4039).transpose_a attribute support for Scale Estimation algorithm in ONNX and OpenVINO backends (#3839).to, dropout, etc.) from the inference graph in PTQ and Weight Compression algorithms (#4057).load_from_config reliability by saving module path in serialized config (#4050) (#4054).example_input parameter from load_from_config, which has been unused since 3.0.0 (#4051).OVQuantizer from NNCF; use it from the executorch package instead (#3842).CompressWeightsMode.INT8 mode. Use CompressWeightsMode.INT8_SYM or CompressWeightsMode.INT8_ASYM instead (#4008).ProgressBar class (#4044).IgnoredScope for nncf.compress_weights and nncf.prune (https://github.com/openvinotoolkit/nncf/pull/4063) (https://github.com/openvinotoolkit/nncf/pull/4064).QuantizationConstraints.from_config_dict to avoid unexpected behavior (https://github.com/openvinotoolkit/nncf/pull/4046).initial_steps and scale_steps in calculate_quantization_params (https://github.com/openvinotoolkit/nncf/pull/4039).transpose_a attribute support for Scale Estimation algorithm in ONNX and OpenVINO backends (https://github.com/openvinotoolkit/nncf/pull/3839).to, dropout, etc.) from the inference graph in PTQ and Weight Compression algorithms (https://github.com/openvinotoolkit/nncf/pull/4057).load_from_config reliability by saving module path in serialized config (https://github.com/openvinotoolkit/nncf/pull/4050) (https://github.com/openvinotoolkit/nncf/pull/4054).example_input parameter from load_from_config, which has been unused since 3.0.0 (https://github.com/openvinotoolkit/nncf/pull/4051).OVQuantizer from NNCF; use it from the executorch package instead (https://github.com/openvinotoolkit/nncf/pull/3842).CompressWeightsMode.INT8 mode. Use CompressWeightsMode.INT8_SYM or CompressWeightsMode.INT8_ASYM instead (https://github.com/openvinotoolkit/nncf/pull/4008).ProgressBar class (https://github.com/openvinotoolkit/nncf/pull/4044).onnx to 1.21.0 (https://github.com/openvinotoolkit/nncf/pull/4006).torchao to 0.17.0 (https://github.com/openvinotoolkit/nncf/pull/4013).onnxscript to 0.6.2 (https://github.com/openvinotoolkit/nncf/pull/4015).(PyTorch) Migrated to use torchao instead of deprecated torch.ao ( #3854 ).
NNCFGraph from nx.DiGraph to nx.MultiDiGraph to support models with parallel/multi-edges, enabling correct quantization of models with complex graph structures such as YOLO26 and models like a = conv(x); return a * a (#3843).f4e2m1) compression data type in Weight Compression. NVFP4 uses a constant group size of 16 with scales compressed to f8e4m3 using a second-degree scale (#3967).backup_mode parameter for FP compression formats (MXFP4, MXFP8, FP4, FP8), allowing first/last layers to be compressed with a backup FP format instead of INT8 (#3886).TopKMetatype support for TorchFX backend, enabling correct graph building for models with TopK operations such as YOLO26 (#3944).torchao instead of deprecated torch.ao (#3854).nncf.torch module to reduce startup import time (#3862).NNCFGraph from nx.DiGraph to nx.MultiDiGraph to support models with parallel/multi-edges, enabling correct quantization of models with complex graph structures such as YOLO26 and models like a = conv(x); return a * a (https://github.com/openvinotoolkit/nncf/pull/3843).f4e2m1) compression data type in Weight Compression. NVFP4 uses a constant group size of 16 with scales compressed to f8e4m3 using a second-degree scale (https://github.com/openvinotoolkit/nncf/pull/3967).backup_mode parameter for FP compression formats (MXFP4, MXFP8, FP4, FP8), allowing first/last layers to be compressed with a backup FP format instead of INT8 (https://github.com/openvinotoolkit/nncf/pull/3886).TopKMetatype support for TorchFX backend, enabling correct graph building for models with TopK operations such as YOLO26 (https://github.com/openvinotoolkit/nncf/pull/3944).torchao instead of deprecated torch.ao (https://github.com/openvinotoolkit/nncf/pull/3854).nncf.errors.ValidationError: There is no tensor with the name error (https://github.com/openvinotoolkit/nncf/pull/3988).nncf.torch module to reduce startup import time (https://github.com/openvinotoolkit/nncf/pull/3862).openvino 2026.1.0 (https://github.com/openvinotoolkit/nncf/pull/4005).torch 2.10.0 (https://github.com/openvinotoolkit/nncf/pull/3852).onnx 1.20.1 (https://github.com/openvinotoolkit/nncf/pull/3966).onnxruntime 1.24.3 (https://github.com/openvinotoolkit/nncf/pull/3977).pandas to optional dependency (https://github.com/openvinotoolkit/nncf/pull/3970).pillow dependency (https://github.com/openvinotoolkit/nncf/pull/3929).(PyTorch) Removed legacy create_compressed_model API for PyTorch backend, which was previously marked as deprecated.
Post-training Quantization:
nncf.CompressWeightsMode.CB4_F8E4M3 mode option to nncf.CompressWeightsMode.CB4.nncf.prune API function, which provides a unified interface for pruning algorithms. Currently available for PyTorch backend and supports Magnitude Pruning.nncf.build_graph API function for building NNCFGraph from a model. This API can be used to inspect and define the ignored scope.nncf.IgnoredScope.HWConfig, now using Python-style definition of hardware configuration instead of JSON files.compress_pt2e API has been introduced, enabling quantization of torch.fx.GraphModule models with the OpenVINOQuantizer. Users now can quantize their models in ExecuTorch for the OpenVINO backend via the nncf compress_pt2e employing Scale Estimation and AWQ.compress_quantize_weights_transformation() method by removing names of deleted initializers from graph inputs.Deprecations/Removals:
create_compressed_model API for PyTorch backend, which was previously marked as deprecated.NNCFNetwork: NAS, Structural Pruning, AutoML, Knowledge Distillation, Mixed-Precision Quantization, and Movement Sparsity.Requirements:
jsonschema, natsort, and pymoo from dependencies as they are no longer required.numpy to >=1.24.0, <2.5.0.Acknowledgements
Thanks for contributions from the OpenVINO developer community:
@avolkov-intel @Shehrozkashif @ruro @mostafafaheem
(TensorFlow) The TensorFlow backend is now deprecated and will be removed in future releases. It is recommended to use PyTorch analogous models for tr…
Post-training Quantization:
nncf.CompressWeightsMode.E2M1 mode option is renamed to nncf.CompressWeightsMode.MXFP4.RangeEstimatorParametersSet.HISTOGRAM through the AdvancedQuantizationParameters in nncf.quantize().nncf.CompressWeightsMode: MXFP8, FP8, and FP4. These can be used as the mode option in nncf.compress_weights() to apply the corresponding MXFP8, FP8, or FP4 precisions (experimental).QuantizeLinear nodes with constant inputs into precomputed, quantized initializers. This behavior is controlled by the COMPRESS_WEIGHTS backend parameter in nncf.quantize(), which is now enabled (True) by default.MatMul + Add subgraphs where one of the inputs to the Add operation is a constant. Previously, these cases were skipped because the MatMul operation was not recognized as having a bias, preventing the algorithm from being applied.MatMulNBits operation that previously caused graph breaks.Gemm operation when transB=1._get_smooth_quant_param_grid() method reported in #3613.MXFP4 data type.GroupSizeFallbackMode.ADJUST to automatically adjust group size for problematic layers.Compression-aware training:
nncf.strip() for StripFormat.IN_PLACE and example_input is no longer required.Deprecations/Removals:
Requirements:
(OpenVINO) Introduced new compression data types CB4_F8E4M3 and CODEBOOK. CB4_F8E4M3 is a fixed codebook with 16 fp8 values based on NF4 data type val
Post-training Quantization:
group_size_fallback_mode parameter for advanced weight compression. It controls how nodes that do not support the default group size are handled. By default (IGNORE), such nodes are skipped. With ERROR, an exception is raised if the channel size is not divisible by the group size, while ADJUST attempts to modify the group size so it becomes valid.quantize_pt2e API, including XNNPACKQuantizer and CoreMLQuantizer. Users now can quantize their models in ExecuTorch for the XNNPACK and CoreML backends via the nncf quantize_pt2e employing smooth quant, bias correction algorithms and a wide range of statistic collectors.TinyLlama/TinyLlama-1.1B-Chat-v1.0 model in ONNX format.Compression-aware training:
Deprecations/Removals:
create_compressed_model API.Requirements:
setuptools>=77 to build package.Acknowledgements
Thanks for contributions from the OpenVINO developer community:
@bopeng1234 @jpablomch
(PyTorch) The function_hook module is now the default mechanism for model tracing. It has moved out from experimental status and has been moved to the
Post-training Quantization:
TinyLlama-1.1B-Chat-v0.3 model in ONNX format using the NNCF weight compression API.BackendParameters.EXTERNAL_DATA_DIR parameter for the ONNX backend. This parameter specifies the absolute path to the directory where the model's external data files are stored. All external data files must be located in the same directory. It should be used when the model is loaded without external data using onnx.load("model.onnx", load_external_data=False), and the external data files are not in the current working directory of the process. This parameter can be omitted if the external data files are located in the current working directory of the process.transformer>4.52 by nncf.data.generate_text_data.Compression-aware training:
nncf.compress_weights API now includes a new compression_format option, nncf.CompressionFormat.FQ_LORA_NLS. A sample QAT compression pipeline with preview support is available here. Building on our previous work with absorbable LoRA adapters, this new pipeline is specifically designed for downstream tasks. In contrast, the pipeline from the previous release was tailored to enhance general accuracy through knowledge distillation using static rank settings. For a more comprehensive understanding of both approaches, please refer to "Weight-Only Quantization Aware Training with LoRA and NLS" in the "Training-Time Compression Algorithms" section of the main README in the repository.Requirements:
(PyTorch) Added support for 4-bit weight compression with AWQ and Scale Estimation data-aware methods to reduce quality loss.
narrow_range parameter, enabling more combinations of quantization configurations in the MinMax quantization algorithm.quantize_pt2e function and the transform_for_annotation method of the OpenVINOQuantizer to align with the torch.ao quantization implementation.nncf.compress_weights API now includes a new compression_format option, nncf.CompressionFormat.FQ_LORA, for this QAT method, a sample compression pipeline with preview support is available here.compressed_model.nncf.get_config was changed to nncf.torch.get_config. The documentation was updated to use the new API.Acknowledgements
Thanks for contributions from the OpenVINO developer community:
@shumaari
(TensorFlow) The nncf.tensorflow.create_compressed_model() method is now marked as deprecated. Please use the nncf.quantize() method for the quantizat…
nncf.quantize() method is now the recommended API for Quantization-Aware Training. Please refer to an example for more details about how to use a new approach.nncf.tensorflow.get_config() and nncf.tensorflow.load_from_config(). Please see the documentation for the saving/loading of a quantized model for more details.quantize_pt2e API has been introduced, enabling quantization of torch.fx.GraphModule models with the OpenVINOQuantizer and the X86InductorQuantizer quantizers. quantize_pt2e API utilizes MinMax algorithm statistic collectors, as well as SmoothQuant, BiasCorrection and FastBiasCorrection Post-Training Quantization algorithms.nncf.tensorflow.create_compressed_model() method is now marked as deprecated. Please use the nncf.quantize() method for the quantization initialization.numpy (>=1.24.0).tqdm dependency.Acknowledgements
Thanks for contributions from the OpenVINO developer community:
@rk119
@devesh-2002
(PyTorch) Fixed the get_torch_compile_wrapper function to match with the torch.compile.
get_torch_compile_wrapper function to match with the torch.compile.safetensors approach.(PyTorch) nncf.torch.create_compressed_model() function has been deprecated.
backup_mode optional parameter in nncf.compress_weights() to specify the data type for embeddings, convolutions and last linear layers during 4-bit weights compression. Available options are INT8_ASYM by default, INT8_SYM, and NONE which retains the original floating-point precision of the model weights.quantizer_propagation_rule parameter, providing fine-grained control over quantizer propagation. This advanced option is designed to improve accuracy for models where quantizers with different granularity could be merged to per-tensor, potentially affecting model accuracy.nncf.data.generate_text_data API method that utilizes LLM to generate data for further data-aware optimization. See the example for details.nncf.compress_weights() with NF4 per-channel quantization, which makes compressed LLMs more accurate and faster on NPU.statistics_path to cache and reuse statistics for nncf.compress_weights(), reducing the time required to find optimal compression configurations. See the TinyLlama example for details.torch.compile(compressed_model, backend="openvino") (see details here). Added INT8 quantization example. The list of supported features:
nncf.quantize().nncf.compress_weights().nncf.quantize_with_accuracy_control().nncf.compress_weights().nncf.compress_weights() with AWQ, Scale Estimation, LoRA and mixed-precision algorithms.nncf.compress_weights() with AWQ algorithm.networkx versions.nncf.ModelType.TRANSFORMER scheme.nncf.ModelType.TRANSFORMER scheme with GroupNorm metatype.torchvision mobilenet_v3 has been extended.nncf.quantize() method can generate inaccurate INT8 results for MobileNet models with the BiasCorrection algorithm.setup.py to pyproject.toml for the build and package configuration. It is aligned with Python packaging standards as outlined in PEP 517 and PEP 518. The installation through setup.py does not work anymore. No impact on the installation from PyPI and Conda.nncf.torch.create_compressed_model() function has been deprecated.Acknowledgements
Thanks for contributions from the OpenVINO developer community: @rk119 @zina-cs
(OpenVINO) Added support for combining GPTQ with AWQ and Scale Estimation (SE) algorithms in nncf.compress_weights() for more accurate weight compress
nncf.compress_weights() for more accurate weight compression of LLMs. Thus, the following combinations with GPTQ are now supported: AWQ+GPTQ+SE, AWQ+GPTQ, GPTQ+SE, GPTQ.lora_correction parameter of the nncf.compress_weights() API. The algorithm increases compression time and incurs a negligible model size overhead. Refer to accuracy/footprint trade-off for different int4 compression methods.nncf.compress_weights().torch.compile.Acknowledgements
Thanks for contributions from the OpenVINO developer community: @rk119
(OpenVINO, PyTorch, ONNX) Excluded comparison operators from the quantization scope for nncf.ModelType.TRANSFORMER.
nncf.ModelType.TRANSFORMER.nncf.compress_weights() method.nncf.compress_weights(). This allows apply AWQ for the wider scope of the models.nncf.CompressWeightsMode.E2M1 mode option of nncf.compress_weights() as the new MXFP4 precision (Experimental).nncf.quantize() method.torch.addmm.torch.nn.functional.scaled_dot_product_attention.nncf.IgnoredScope() functionality for models with If operation.nncf.compress_weights() to OpenVINO models.nncf.IgnoredScope().Acknowledgements
Thanks for contributions from the OpenVINO developer community: @Lars-Codes
(OpenVINO) Added Scale Estimation algorithm for 4-bit data-aware weights compression. The optional scale_estimation parameter was introduced to nncf.c
Acknowledgements
Thanks for contributions from the OpenVINO developer community: @DaniAffCH @UsingtcNower @anzr299 @AdiKsOnDev @Viditagarwal7479 @truhinnm
NNCF installation via pip install nncf[ ] option is now deprecated.
Acknowledgements
Thanks for contributions from the OpenVINO developer community: @Candyzorua @clinty @UsingtcNower @DaniAffCH
(PyTorch) Deprecated the binarization algorithm.
MatMul->Multiply->Matmul. For that awq optional parameter has been added to nncf.compress_weights() and can be used to minimize accuracy degradation of compressed models (note that this option increases the compression time).nncf.quantize_with_accuracy_control() method. Users can now perform quantization with accuracy control for onnx.ModelProto. By leveraging this feature, users can enhance the accuracy of quantized models while minimizing performance impact.nncf.quantize().nncf.AdvancedAccuracyRestorerParameters.subset_size option for the nncf.compress_weights().TargetDevice.NPU as the replacement for TargetDevice.VPU.revert_operations_to_floating_point_precision method.nncf.compress_weights() with Convolution & Embeddings compression in order to reduce memory footprint.nncf.quantize() for BERT and YOLOv5 models.nncf.quantize_with_accuracy_control() for SSD MobileNetV1 FPN model.binarization algorithm.TargetDevice.VPU was replaced by TargetDevice.NPU.NNCFNetworkInterface.get_clean_shallow_copy missed arguments.Acknowledgements
Thanks for contributions from the OpenVINO developer community: @AishwaryaDekhane @UsingtcNower @Om-Doiphode
(Common) Fixed issue with nncf.compress_weights() to avoid overflows on 32-bit Windows systems.
nncf.compress_weights() to avoid overflows on 32-bit Windows systems.nncf.compress_weights() on LLama models.nncf.quantize_with_accuracy_control pipeline with tune_hyperparams=True enabled option.nncf.compress_weights() for LLM models with the executing is_floating_point with tracing.The original nncf.CompressWeightsMode.INT8 enum value is now deprecated.
nncf.quantize signature has been changed to add mode: Optional[nncf.QuantizationMode] = None as its 3-rd argument, between the original calibration_dataset and preset arguments.nncf.common.quantization.structs.QuantizationMode has been renamed to nncf.common.quantization.structs.QuantizationSchemedataset optional parameter has been added to nncf.compress_weights() and can be used to minimize accuracy degradation of compressed models (note that this option increases the compression time).nncf.compress_weights(). The weights compression algorithm for PyTorch models is now based on tracing the model graph. The dataset parameter is now required in nncf.compress_weights() for the compression of PyTorch models.nncf.CompressWeightsMode.INT8 to nncf.CompressWeightsMode.INT8_ASYM and introduce nncf.CompressWeightsMode.INT8_SYM that can be efficiently used with dynamic 8-bit quantization of activations.
The original nncf.CompressWeightsMode.INT8 enum value is now deprecated.nncf.QuantizationMode.FP8_E4M3 and nncf.QuantizationMode.FP8_E5M2 enum values, invoked via passing one of these values as an optional mode argument to nncf.quantize. Currently, OpenVINO supports inference of FP8-quantized models in reference mode with no performance benefits and can be used for accuracy projections.nncf.quantize_with_accuracy_control() has been extended by restore_mode optional parameter to revert weights to int8 instead of the original precision.
This parameter helps to reduce the size of the quantized model and improves its performance.
By default, it's disabled and model weights are reverted to the original precision in nncf.quantize_with_accuracy_control().all_layers: Optional[bool] = None argument to nncf.compress_weights to indicate whether embeddings and last layers of the model should be compressed to a primary precision. This is relevant to 4-bit quantization only.sensitivity_metric: Optional[nncf.parameters.SensitivityMetric] = None argument to nncf.compress_weights for finer control over the sensitivity metric for assigning quantization precision to layers.
Defaults to weight quantization error if a dataset is not provided for weight compression and to maximum variance of the layers' inputs multiplied by inverted 8-bit quantization noise if a dataset is provided.
By default, the backup precision is assigned for the embeddings and last layers.gpt-2, stable-diffusion-v1-5, stable-diffusion-v2-1, opt-6.7b, falcon-7b, bloomz-7b1) are now more accurately quantized.nncf.strip(..., do_copy=True) now actually returns a deepcopy (stripped) of the model object.torch.return_type (such as torch.max).torch namespace.README.md for better readability.nncf.CompressWeightsMode.INT8 enum value is now deprecated.transformers repository is marked as deprecated and will be removed in a future release.
Developers are advised to use optimum-intel instead.…implementation of quantization algorithms. Deprecated create_compressed_model() method for Post-training Quantization.
compress_weights(…) pipeline).dump_intermediate_model parameter support for AccuracyAwareAlgorithm (quantize_with_accuracy_control(…) pipeline).quantize_with_tune_hyperparams(…) pipeline).quantize(…) pipeline and the common implementation of quantization algorithms. Deprecated create_compressed_model() method for Post-training Quantization.ModelType.Transformer scheme.QuantizationPreset.Mixed was set as the default for ModelType.Transformer scheme.ModelType.Transformer to align with the quantization scheme.compress_weights(…) pipeline) execution time for LLM's quantization, added ignored scope support.quantize_with_accuracy_control(…) pipeline).quantize(...) method can generate inaccurate int8 results for models with the BatchNormalization layer that contains biases. To get the best accuracy, use the do_constant_folding=True option during export from PyTorch to ONNX.(PyTorch) Removed deprecated NNCFNetwork.__getattr__, NNCFNetwork.get_nncf_wrapped_model methods.
CPU_SPR device type support.ModelType.Transformer.compress_weights method that provides data-free INT8 weights compression.quantize(…) pipeline (up to 4.3x speed up in total).quantize_with_accuracy_control(…) pipelilne (up to 8x speed up for 122-quantizing-model-with-accuracy-control notebook).validate_scopes parameter for NNCF configuration..strip() option to API.torch.jit.traced without calling .strip().forward instance attribute on model objects passed into create_compressed_model.torch.jit.script wrapper so that user-side handling exceptions during torch.jit.script invocation do not cause NNCF to be permanently disabled.__class__ method for ProxyModule that avoids causing error while calling .super() in forward method.NNCFNetwork.__getattr__, NNCFNetwork.get_nncf_wrapped_model methods.Official release of OpenVINO framework support.
openvino package and not on the openvino-dev package."overflow_fix" parameter (for quantize(...) & quantize_with_accuracy_control(...) methods) support & functionality. It improves accuracy for optimized model for affected devices. More details in Quantization section.quantize(...) & quantize_with_accuracy_control(...) methods.ignored_scope attribute behaviour for weights. Now, the weighted layers excludes from optimization scope correctly.nncf.quantize(...). Now, models with opset < 13 are optimized correctly in per-tensor quantization.quantize(...) method can generate inaccurate int8 results for models with the DenseNet-like architecture. Use quantize_with_accuracy_control(...) in such case.quantize(...) method can hang on models with transformer architecture when fast_bias_correction optional parameter is set to False. Don't set it to False or use quantize_with_accuracy_control(...) in such case.quantize(...) method can generate inaccurate int8 results for models with the MobileNet-like architecture on non-VNNI machines.nncf.common.utils.patcher.Patcher - this class can be used to patch methods on live PyTorch model objects with wrappers such as nncf.torch.dynamic_graph.context.no_nncf_trace when doing so in the model code is not possible (e.g. if the model comes from an external library package).nncf.api.compression.CompressionAlgorithmController class now have a .strip() method that will return the compressed model object with as many custom NNCF additions removed as possible while preserving the functioning of the model object as a compressed model.transpose/permute/getitem) for pruning node selector.nncf.set_log_file(...) can now be used to set location of the NNCF log file.torch.nn.functional.pad operation.torch.baddbmm as an alias for the matmul metatype for quantization purposes.__matmul__ magic functions to the list of patched ops (for SwinTransformer by Microsoft).Calling setup.py directly to install NNCF is deprecated and no longer guaranteed to work.
pip install nncf[openvino] will install NNCF with the required OV framework dependencies.nncf.quantize(...) function.
The parameter set of the function is the same for all frameworks - actual framework-specific implementations are being dispatched based on the type of the model object argument.transformers repo.
See description of the movement pruning involved in the JPQD for details.logging.INFO log level.transformers integration patch now allows to export to ONNX during training, and not only at the end of it.torch.nn.utils.weight_norm weights are now detected correctly."num_bn_adaptation_samples": 0 in config leading to a TypeError during quantization algo initialization.ignored_scopes only if the top-most node in data flow order matches against ignored_scopes"ignored_scopes" and "target_scopes" are now strictly checked to be matching against at least one node in the model graph instead of silently ignoring the unmatched entries.setup.py directly to install NNCF is deprecated and no longer guaranteed to work.from nncf.common.utils.logger import logger as nncf_logger is deprecated - use from nncf import nncf_logger instead.pruning_rate is renamed to pruning_level in pruning compression controllers.(ONNX) PTQ API support for ONNX.
New features
BootstrapNAS to find high-performing sub-networks from the super-network optimization.Bugfixes
ONNXGraph and MinMaxQuantization.(TensorFlow) Added TensorFlow 2.5.x support.
SubclassedConverter class was added to create NNCFGraph for the tf.Graph Keras model.TFOpLambda layer support with TFModelConverter, TFModelTransformer, and TFOpLambdaMetatype.MatMul and Conv2D to BiasAdd and Metatypes of TensorFlow operations with weights TFOpWithWeightsMetatype are added.Reshape and Linear as ReshapePruningOp and LinearPruningOp.Resnet50 and Mobilenet_v2 for the latest VPU.NNCFBatchNorm into NNCFBatchNorm1d, NNCFBatchNorm2d, NNCFBatchNorm3d.BNASTrainingController and BNASTrainingAlgorithm for BootstrapNAS to search the model's architecture.ModelProto is now converted to NNCFGraph through GraphConverter.ONNXOpMetatype and extended patterns for fusing HW config is now available.ONNXPostTrainingQuantization and MinMaxQuantization supports for ONNX.EarlyExitCompressionTrainingLoop.FakeQuantizer to make exact zeros.(PyTorch) All PyTorch operations are now NNCF-wrapped automatically.
early_exit mode.dump_checkpoints_fn callback to control the location of checkpoint saving during accuracy-aware training.softmax mode.mixed_min_max option for quantizer range initialization.determine_subtype stage of metatype assignment.torch.nn.utils.weight_norm.WeightNorm appliedRelax TensorFlow version requirements to 2.4.x
Bump target framework versions to PyTorch 1.9.1 and TensorFlow 2.4.3
self in the calls to the wrapped model refer to the wrapper NNCFNetwork object and not to the wrapped modelview operations to handle shape arguments with the torch.Tensor typeTarget version updates:
Bugfixes:
self in the calls to the wrapped model refers to the wrapper NNCFNetwork object and not to the wrapped modelview operations to handle shape arguments with the torch.Tensor typeAdded TensorFlow 2.4.2 support - NNCF can now be used to apply the compression algorithms to models originally trained in TensorFlow. NNCF with Tensor
Added TensorFlow 2.4.2 support - NNCF can now be used to apply the compression algorithms to models originally trained in TensorFlow. NNCF with TensorFlow backend supports the following features:
tf.distribute.MirroredStrategy.Added model compression samples for NNCF with TensorFlow backend:
Accuracy-aware training available for filter pruning and sparsity in order to achieve best compression results within a given accuracy drop threshold in a fully automated fashion.
Framework-specific checkpoints produced with NNCF now have NNCF-specific compression state information included, so that the exact compressed model state can be restored/loaded without having to provide the same NNCF config file that was used during the creation of the NNCF-compressed checkpoint
Common interface for compression methods for both PyTorch and TensorFlow backends (https://github.com/openvinotoolkit/nncf/tree/develop/nncf/api).
(PyTorch) Added an option to specify an effective learning rate multiplier for the trainable parameters of the compression algorithms via NNCF config, for finer control over which should tune faster - the underlying FP32 model weights or the compression parameters.
(PyTorch) Unified scales for concat operations - the per-tensor quantizers that affect the concat operations will now have identical scales so that the resulting concatenated tensor can be represented without loss of accuracy w.r.t. the concatenated subcomponents.
(TensorFlow) Algo-mixing: Added configuration files and reference checkpoints for filter-pruned + qunatized models: ResNet50@ImageNet2012(40% of filters pruned + INT8), RetinaNet@COCO2017(40% of filters pruned + INT8).
(Experimental, PyTorch) Learned Global Ranking filter pruning mechanism for better pruning ratios with less accuracy drop for a broad range of models has been implemented.
(Experimental, PyTorch) Knowledge distillation supported, ready to be used with any compression algorithm to produce an additional loss source of the compressed model against the uncompressed version
CompressionLevel has been renamed to CompressionStage"ignored_scopes" and "target_scopes" no longer allow prefix matching - use full-fledged regular expression approach via {re} if anything more than an exact match is desired.pip install nncf or python setup.py install and are assumed to be present in the user's environment; the pip's "extras" syntax must be used to install the BKC requirements, e.g. by executing pip install nncf[tf], pip install nncf[torch] or pip install nncf[tf,torch]"quantizable_subgraph_patterns" option removed from the NNCF configFixed a bug with where compressed models that were supposed to return named tuples actually returned regular tuples
Adjust Padding feature to support accurate execution of U4 on VPU - when setting "target_device" to "VPU", the training-time padding values for quanti
pip install . path of installing NNCF from a checked-out repository is now supported.with no_nncf_trace() blocks now function as expected.Added AutoQ - an AutoML-based mixed-precision initialization mode for quantization, which utilizes the power of reinforcement learning to select the b
"num_init_samples" should be used in place of "num_init_steps" in NNCF config files.domain set to org.openvinotoolkitNothing published for this version
Nothing published for this version
Models with filter pruning applied are now exportable to ONNX
"batchnorm_adaptation" config parameters in compression algorithm documentation (e.g. Quantizer.md) for instructions on how to enable it in NNCF configYour coding agent can read these notes before it upgrades. Set up the MCP server →