NewYour coding agent can read the release notes before it upgrades.Set up the MCP server →
PyPI · #1089 most downloaded on PyPI
Fast inference engine for Transformer models
Last release 1 months ago
31 Aug 2026
Release timing varies
gaps range from 3 weeks to 7 months
Nearly every release is documented
notes for 60 of the last 60 stable releases
Nothing withdrawn
no release was ever pulled
5 years old
62 releases · first in 2021
One column per quarter.
Import Python converters and specs lazily to avoid loading torch for inference-only use ( #2080 ) by @Anai-Guo
Support for Gemma4 12B dense model ( #2060 ) by @jordimas
Add support for Google T5Gemma2 ( #2058 ) by @jordimas
Gemma4 support for dense model ( #2048 ) by @jordimas
Fix Windows build ( #2007 ) @sssshhhhhh
Introduce AMD GPU support with ROCm HIP ( #1989 ) @sssshhhhhh
T5Gemma model conversion and inference ( #1962 ) by @jordimas
Fixed pkg_resources Deprecated Warning ( #1911 ) by @thawancomt
Update Intel oneAPI to version 2025.3
Note: The Ctranslate2 Python package now supports python 3.13, drop the support for python 3.8.
Note: The Ctranslate2 Python package now supports python 3.13, drop the support for python 3.8.
Note: The Ctranslate2 Python package now supports CUDNN 9 and is no longer compatible with CUDNN 8.
Note: The Ctranslate2 Python package now supports CUDNN 9 and is no longer compatible with CUDNN 8.
Removed: Flash Attention support in the Python package due to significant package size increase with minimal performance gain. Note: Flash Attention r
Removed: Flash Attention support in the Python package due to significant package size increase with minimal performance gain.
Note: Flash Attention remains supported in the C++ package with the WITH_FLASH_ATTN option.
Flash Attention may be re-added in the future if substantial improvements are made.
Note: Because of exceeding project's size on Pypi (> 20 GB), the release v4.3.0 was pushed unsuccessfully.
Note: Because of exceeding project's size on Pypi (> 20 GB), the release v4.3.0 was pushed unsuccessfully.
Note: Because of the increasing of package's size (> 100 MB), the release v4.2.0 was pushed unsuccessfully.
Note: Because of the increasing of package's size (> 100 MB), the release v4.2.0 was pushed unsuccessfully.
Read very large tensor by chunk if the size > max value of int
This major version introduces the breaking change while updating to cuda 12.
This major version introduces the breaking change while updating to cuda 12.
Support of new option offset to ignore token score of special tokens
Fix the conversion for whisper without the "alignment_heads" in the "generation_config.json"
Support "sliding window" and "chunking input" for Mistral
Minimal Support for Mistral (Loader and Rotary extension for long sequence). No sliding yet
New features
Update the Transformers converter to support more model architectures:
generate_tokensGenerator.async_generate_tokens to return an asynchronous generator compatible with asyncioWhisper::alignBinary wheels for Python 3.7 are no longer built
GenerationStepResult.hypothesis_id to identify the different hypotheses when running random sampling with num_hypotheses > 1batch_id values passed to the callback function.detach() on PyTorch tensors before getting the Numpy array in convertersConverted models now uses the same floating point precision as the original models. For example, a model saved in float16 will be converted to a float
Converted models now uses the same floating point precision as the original models. For example, a model saved in float16 will be converted to a float16 model. Before this change, the weights were casted to float32 by default.
Similarly, selecting int8 keeps non quantized weights in their original precision unless a more specific quantization type is selected:
compute_type to model instancesStorageView with additional methods and properties:
to(dtype)device_indexdevicedtypeshapeget_supported_compute_types to correctly return bfloat16 when supportedreturn_alternatives with a model using relative positionstorch<1.13Fix an error when running models with the new int8_bfloat16 computation type
int8_bfloat16 computation typegenerate_tokens is closedAdd new computation types: bfloat16 and int8_bfloat16 (require a GPU with Compute Capability 8.0 or above)
bfloat16 and int8_bfloat16 (require a GPU with Compute Capability 8.0 or above)trust_remote_code when loading the tokenizer in the Transformers converterlib64 instead of lib)Fix repeated outputs in version 3.16.0 when using include_prompt_in_result=False and a batch input with variable lengths: a typo in the code led to mi
include_prompt_in_result=False and a batch input with variable lengths: a typo in the code led to min_length being incorrectly appliedUpdate the Transformers converter to support more architectures:
sampling_topp to enable top-p (nucleus) samplingmin_length and max_length when using include_prompt_in_result=False and a batch input with variable lengths: the length constraint should only apply to the sequence after the promptFix an error when using the new static_prompt argument in the methods generate_tokens and generate_batch
static_prompt argument in the methods generate_tokens and generate_batchInitial support of encoder-only Transformer model via a new class ctranslate2.Encoder
ctranslate2.Encoderstatic_prompt to optimize the execution for models using system prompts: the model state for this prompt is cached and reused in future callsTrueconfig.jsontied-embeddings-all: falseuse_fast argument when loading Hugging Face tokenizers to use the default tokenizer for the modelUpdate the Transformers converter with new architectures:
layer_norm="rms"max_relative_positions=-1 (rotary embeddings)max_relative_positions=-2 (ALiBi)pos_ffn_activation_fn="silu"generate_tokens method to properly raise the underlying exception instead of hanging indefinitely-DBUILD_SHARED_LIBS=OFFlibctranslate2.a without using the "whole archive" flagsSupport conversion of GPT-NeoX models with the Transformers converter
end_token argument to also accept a list of tokensreturn_end_token to include the end token in the results of the methods generate_batch and translate_batch (by default the end token is removed)callback argument for the methods generate_batch and translate_batch to get early results from the decoding loopCTranslate2::ctranslate2 to facilitate the library integration in other CMake projectsAdd methods Generator.generate_tokens and Translator.generate_tokens returning a generator that yields tokens as soon as they are generated by the mod
Generator.generate_tokens and Translator.generate_tokens returning a generator that yields tokens as soon as they are generated by the model (not compatible with beam search)rotary_interleave=False in the model specification (may require to permute QK weights)Whisper.align to improve batch supportlow_cpu_mem_usage in the Transformers converter to reduce the memory usage when loading large models (requires the package accelerate)Whisper.align when num_frames // 2 <= median_filter_widthend_token or suppress_sequences contain tokens that are not in the vocabularyint8_float16The Python wheels for macOS ARM are now built with the Ruy backend to support INT8 computation. This will change the performance and results when load
auto compute type. To keep the previous behavior, set compute_type="float32".decoder_start_tokenrevision in the Transformers converter to download a specific revision of the model from the Hugging Face HubFix a synchronization issue when the model input is a CUDA storage
Select the correct device when copying a StorageView instance
StorageView instanceAdd missing device setter in Whisper.encode
Whisper.encodeAdd Generator option include_prompt_in_result (True by default)
Generator option include_prompt_in_result (True by default)Whisper.encode to only run the Whisper encoderWhisper.device and Whisper.device_indexWhisper.detect_language, Whisper.generate, and Whisper.align to accept the encoder outputGenerator.forward on GPU and the generator object is destroyed before the forward outputFix missing alignments in the Whisper.align result due to a bug in the DTW implementation
Whisper.align result due to a bug in the DTW implementationAdd method Whisper.align to return the text/audio alignment and implement word-level timestamps
Whisper.align to return the text/audio alignment and implement word-level timestampsintra_threads to 1 when loading a model on the GPU as some ops may still run on the CPUExperimental support of AVX512 in manually vectorized functions: this code path is not enabled by default but can be enabled by setting the environmen
CT2_FORCE_CPU_ISA=AVX512copy_files to copy any files from the Hugging Face model to the converted model directorymax_initial_timestamp_indexsuppress_blanksuppress_tokens--quantization float16Rename the "float" compute type to "float32" for clarity. "float" is still accepted for backward compatibility.
CT2_CUDA_TRUE_FP16_GEMM. This flag is enabled by default so that FP16 GEMMs are running in full FP16. When disabled, the compute type of FP16 GEMMs is set to FP32, which is what PyTorch and TensorFlow do by default.max_length, which could previously happen when the prompt was longer than max_length/2Build the Windows Python wheels with cuDNN to enable GPU execution of Whisper models
Whisper.is_multilingualWhisper.generateWhisper: fix an incorrect timestamp rule that prevented timestamps to be generated in pairs
Add a patience factor for beam search to continue decoding until beam_size * patience hypotheses are finished, as described in Kasai et al. 2022
beam_size * patience hypotheses are finished, as described in Kasai et al. 2022python_requires to facilitate the package installation with tools like Poetry and PDMFix incorrect vocabulary in M2M100 models after conversion with transformers>=4.24
transformers>=4.24prefix_bias_beta > 0 with beam_size == 1Support T5 models, including the variants T5v1.1 and mT5
files argument in the constructor of classes loading modelsmodels::ModelMemoryReader class--activation_scales)Add decoding option suppress_sequences to prevent specific sequences of tokens from being generated
suppress_sequences to prevent specific sequences of tokens from being generatedend_token to stop the decoding on a different token than the model EOS tokennum_hypotheses > 1return_no_speech_prob<|notimestamps|> but not othersTransformerSpec constructor to accept arbitrary encoder and decoder specifications-ffast-math which introduces unwanted side effects and enable it only for the layer norm CPU kernel where it is actually useful-DOPENMP_RUNTIME=COMPRaise a deprecation warning when reading the TranslationResult object as a list of dictionaries
Whisper.generate as it is usually not useful in a transcription loopWhisper.generate is updated from 1 to 5 to match the default value in openai/whispermin_length and no_repeat_ngram_size now penalize the logits instead of the log probs which may change some scoresTranslationResult object as a list of dictionariesctranslate2.set_log_level<|notimestamps|>return_no_speech_prob to the method Whisper.generate for the result to include the probability of the no speech token<|X.XX|>)LogitsProcessor abstract class to apply arbitrary updates to the logits during decodingWhisper: fix generate arguments that were not correctly passed to the model
generate arguments that were not correctly passed to the modelWhisper: do not implicitly add <|startoftranscript|> in generate since it is not always the first token
<|startoftranscript|> in generate since it is not always the first tokenThis major version integrates the Whisper speech recognition model published by OpenAI. It also introduces some breaking changes to remove deprecated…
This major version integrates the Whisper speech recognition model published by OpenAI. It also introduces some breaking changes to remove deprecated usages and simplify some modules.
normalize_scores: the scores are now always divided by pow(length, length_penalty) with length_penalty defaulting to 1allow_early_exit: the beam search now exits early only when no penalties are usedOpenNMTTFConverterV2 -> OpenNMTTFConverterTranslationStats -> ExecutionStatsScoringResult as a list of scores: the scores can be accessed with the attribute log_probsExecutionStats as a tupletranslate to a more specific name ct2-translatorTranslationStats -> ExecutionStatsGeneratorPool -> GeneratorTranslatorPool -> TranslatorTranslatorPool::consume_* -> Translator::translate_*TranslatorPool::consume_stream -> removedTranslatorPool::score_stream -> removedGenerator.forward_batch to get the full model output for a batch of sequencesStorageView to expose C++ methods taking or returning N-dimensional arrays: the class implements the array interface for interoperability with Numpy and PyTorchconfig.json in the model directory that contains non structual model parameters (e.g. related to the input, the vocabulary, etc.)models::ModelFactorydisable_unk, min_length, no_repeat_ngram_size, etc.ReplicaPool base class to simplify adding new classes with multiple model workersThe Linux binaries now use the GNU OpenMP runtime instead of Intel OpenMP to workaround an initialization error on systems without /dev/shm
/dev/shmTranslator.translate_iterable and Translator.score_iterable, raise an error if the input iterables don't have the same lengthIn beam search, get more candidates from the model output and replace finished hypotheses by these additional candidates
-fvisibility=hidden when building the Python modulescore_batch methods now return a list of ScoringResult instances instead of plain lists of probabilities. In most cases you should not need to update
score_batch methods now return a list of ScoringResult instances instead of plain lists of probabilities. In most cases you should not need to update your code: the result object implements the methods __len__, __iter__, and __getitem__ so that it can still be used as a list.Translator.translate_iterableTranslator.score_iterableGenerator.generate_iterableGenerator.score_iterablemin_alternative_expansion_prob to filter out unlikely alternatives in return_alternatives modeScoringResult instances from score_batch to include additional outputs. The current attributes are:
tokens: the list of tokens that were actually scored (including special tokens)log_probs: the log probability of each scored tokenscore_batch asynchronously by setting the asynchronous flagdisable_unk or use_vmap with one of the following options:
min_decoding_lengthno_repeat_ngram_sizeprefix_bias_betarepetition_penaltyFix conversion of NLLB models when tokenizer_class is missing from the configuration
tokenizer_class is missing from the configurationSupport NLLB multilingual models via the Transformers converter
microsoft/DialoGPT where EOS is used as a separatorreturn_alternatives and return_attention with a float16 modelint8Generation option no_repeat_ngram_size to prevent the repetitions of N-grams with a minimum size
no_repeat_ngram_size to prevent the repetitions of N-grams with a minimum sizereturn_alternatives mode when the target prefix is longer than max_decoding_length<pad> token when converting MarianMT models from Transformers: this token is only used to start the decoder from a zero embedding, but it is not included in the original Marian modelOpenNMTTFConverterV2.from_configFix missing final bias in some MarianMT models converted from Transformers
ctranslate2::float16_t in the public headers that is required to use some functionsSupport conversion of decoder-only Transformer models trained with OpenNMT-tf
facebook/bart-large-cnnmax_input_length after all special tokens have been added to the inputSupport Meta's OPT models via the Transformers converter
transformer_lm modelsYour coding agent can read these notes before it upgrades. Set up the MCP server →