NewYour coding agent can read the release notes before it upgrades.Set up the MCP server →
PyPI · #715 most downloaded on PyPI
Embeddings, Retrieval, and Reranking
Last release 16 days ago
18 Sep 2026
Ships fairly regularly
a new release about every 3 weeks
Nearly every release is documented
notes for 55 of the last 60 stable releases
1 version withdrawn
withdrawn after publishing
7 years old
84 releases · first in 2019
This is a small release with refreshed inference benchmarks, clearer documentation, and a few improvements to multimodal input handling.
This is a small release with refreshed inference benchmarks, clearer documentation, and a few improvements to multimodal input handling.
Install this version with
# Training + Inference
pip install sentence-transformers[train]==6.1.0
# Inference only, use one of:
pip install sentence-transformers==6.1.0
pip install sentence-transformers[onnx-gpu]==6.1.0
pip install sentence-transformers[onnx]==6.1.0
pip install sentence-transformers[openvino]==6.1.0
# Multimodal dependencies (optional):
pip install sentence-transformers[image]==6.1.0
pip install sentence-transformers[audio]==6.1.0
pip install sentence-transformers[video]==6.1.0
# Or combine as needed:
pip install sentence-transformers[train,onnx,image]==6.1.0For existing multimodal dict inputs, message content now follows each input's key order, which can change embeddings or scores. Video-frame URLs without recognizable image extensions need an explicit frame wrapper to remain one video. See #4037 for details.
docs] Document that optional imports in a custom module become load-time requirements by @abtonmoy in #3922Full Changelog: v6.0.1...v6.1.0
One column per quarter.
This patch release fixes a multi-vector loading bug: PyLate checkpoints that carry both a [Q] / [D] prefix and a text prompt lost the prefix, so they
This patch release fixes a multi-vector loading bug: PyLate checkpoints that carry both a [Q]/[D] prefix and a text prompt lost the prefix, so they were encoded without a marker they were trained with. It also grows the documented multi-vector model tables from 51 checkpoints to 80.
Install this version with
# Training + Inference
pip install sentence-transformers[train]==6.0.1
# Inference only, use one of:
pip install sentence-transformers==6.0.1
pip install sentence-transformers[onnx-gpu]==6.0.1
pip install sentence-transformers[onnx]==6.0.1
pip install sentence-transformers[openvino]==6.0.1
# Multimodal dependencies (optional):
pip install sentence-transformers[image]==6.0.1
pip install sentence-transformers[audio]==6.0.1
pip install sentence-transformers[video]==6.0.1
# Or combine as needed:
pip install sentence-transformers[train,onnx,image]==6.0.1PyLate supports a prefix as well as prompt text, and Sentence Transformers assumed that these were mutually exclusive. However, the following three checkpoints carry both a prefix and a prompt, and were trained with the prefix prepended to the prompt text:
PyLate trained these checkpoints on [CLS] [Q] search_query: ... and [CLS] [D] search_document: ..., but Sentence Transformers dropped the prefix and encoded them as [CLS] search_query: ... and [CLS] search_document: .... After the fix, the performance of the ColBERT-Zero model improved from 0.6569 NDCG@10 to 0.6824 on the NanoBEIR benchmark. To my knowledge, only these 3 checkpoints used both a prefix and a prompt, so this is the only case where the bug would have affected you.
task to routed modules in Router.preprocess (#3967)Router.forward has always passed task down to the routed module, but Router.preprocess did not, so anything a module does with the task at preprocessing time did nothing behind a Router: query_length and document_length caps, query_expansion, and the chat-template task keyword. No released checkpoint combines a Router with those settings, so this is a latent bug rather than one you are likely to have hit. It would have affected anyone building such a model themselves, with no error to indicate it.
revision and trust_remote_code notes refreshed as upstream pull requests merged (#3972).transformers v4.41.0+, where v6.0 requires PyTorch 2.2+ and transformers v5.0+. This is the text rendered as the PyPI project description.task to routed modules in Router.preprocess by @tomaarsen in #3967Full Changelog: v6.0.0...v6.0.1
Warning This is a major release with breaking changes. Upgrading from v5.x to v6.0 may require code updates. The changes marked 🚨 below are the ones m…
This major release introduces Multi-Vector Embedding models, also known as late interaction or ColBERT-style models, as a fourth model type alongside SentenceTransformer, CrossEncoder, and SparseEncoder. Going forward, you'll be able to use Sentence Transformers for training, inferencing, and interpreting Multi-Vector Embedding models.
It also modernizes the dependency floors to transformers v5, fixes a class of silent scoring bugs caused by half precision, and speeds up both training and encoding.
Install this version with
# Training + Inference
pip install sentence-transformers[train]==6.0.0
# Inference only, use one of:
pip install sentence-transformers==6.0.0
pip install sentence-transformers[onnx-gpu]==6.0.0
pip install sentence-transformers[onnx]==6.0.0
pip install sentence-transformers[openvino]==6.0.0
# Multimodal dependencies (optional):
pip install sentence-transformers[image]==6.0.0
pip install sentence-transformers[audio]==6.0.0
pip install sentence-transformers[video]==6.0.0
# Or combine as needed:
pip install sentence-transformers[train,onnx,image]==6.0.0Tip
Our Multi-Vector (Late Interaction) Embedding Models with Sentence Transformers blogpost is an excellent place to learn about multi-vector models: loading the various checkpoint formats, encoding and scoring, plugging them into a search stack, running them on page images, and keeping the index affordable.
Warning
This is a major release with breaking changes. Upgrading from v5.x to v6.0 may require code updates. The changes marked 🚨 below are the ones most likely to affect you, and the Migration Guide has the full list. If you run into issues when upgrading, feel free to open an issue.
Sentence Transformers v6.0 introduces MultiVectorEncoder, for ColBERT-style late interaction retrieval. Where a regular embedding model compresses a whole text into one vector, a multi-vector model keeps one vector per token and scores query against document with the MaxSim operator. That preserves token-level matching information that a single vector has to average away, which usually means stronger retrieval at the cost of a bigger index. It is also the state of the art for visual document retrieval, where a text query is matched against page images directly, with no OCR step in between.
Any PyLate checkpoint and any Stanford-NLP ColBERT checkpoint loads straight into it, and colpali-engine models for visual document retrieval work too, through the same familiar API you already use for dense, sparse, and reranker models.
from sentence_transformers import MultiVectorEncoder
# Download from the 🤗 Hub
model = MultiVectorEncoder("lightonai/LateOn")
query_embeddings = model.encode_query(["Which planet is known as the Red Planet?"])
document_embeddings = model.encode_document([
"Venus is often called Earth's twin because of its similar size and proximity.",
"Mars, known for its reddish appearance, is often referred to as the Red Planet.",
"Jupiter, the largest planet in our solar system, has a prominent red spot.",
"Saturn, famous for its rings, is sometimes mistaken for the Red Planet.",
])
print(query_embeddings[0].shape)
# (12, 128)
scores = model.similarity(query_embeddings, document_embeddings)
print(scores)
# tensor([[10.7942, 11.1104, 10.9743, 11.0811]])Mars wins, as it should, though notice how close the four scores are. That is normal for MaxSim: the scores often look similar, but the ranking is still exact. The blogpost explores this in more detail.
Note what you get back: a list of 2D tensors on the model device, one per input, each of shape (num_tokens, embedding_dim). Unlike dense embeddings, you cannot stack these into one rectangular tensor, because every input has its own token count. Pass convert_to_numpy=True for a list of numpy arrays instead, which is what you want once a corpus outgrows device memory.
Multi-vector models are also asymmetric: queries and documents go through different prefixes, different length caps, and different scoring masks. Unlike many dense models, where the two are interchangeable, encode_query and encode_document are required to get correct embeddings.
Scoring uses MaxSim: for each query token, take its highest similarity against any document token, then sum those maxima across the query.
$$\text{MaxSim}(Q, D) = \sum_{Q_i \in Q} \max_{D_j \in D} Q_i \cdot D_j$$
You can read the operator as a soft alignment: every query token points at the one document token that best explains it, and the score is how well the document explains the query overall. The alignment does not have to be lexical, since the token embeddings are contextualized. But when an exact match does matter to you (a product code, a surname, a function name), MaxSim has a token sitting right there to match it, where a single-vector model had to fold it into an average.
Because MaxSim sums over query tokens, its magnitude scales with the query token count, so scores are not comparable across models with different query recipes. If you want scores on a bounded scale, use similarity_fn_name="meanmaxsim", which divides by the query token count and gives you an average cosine similarity in [-1, 1].
Scoring builds a 4-dimensional intermediate of every query token against every document token, which is the largest tensor in the operation. Every scoring function takes a chunk_elements budget that bounds it, defaulting to 100 million elements (roughly 400 MB in float32), so lower it if you run out of memory. Scores and gradients are bit-identical whatever you set it to. maxsim and maxsim_pairwise also take a device, which scores one chunk at a time on that device and moves each result straight back, letting you score a corpus larger than your VRAM on the GPU. Both are reachable through similarity, which forwards any extra keyword arguments to the scoring function:
scores = model.similarity(query_embeddings, document_embeddings, chunk_elements=1_000_000, device="cuda")When training, pass the budget to the loss instead, with similarity_fct=partial(colbert_scores, chunk_elements=1_000_000). It chunks the document axis, so it composes with the loss-level score_mini_batch_size, which chunks the query axis.
lightonai/LateOn and lightonai/DenseOn were trained by LightOn on the same data with the same ModernBERT backbone and the same 149M parameters, differing only in whether they keep one vector per token or pool down to one per document. Running both over all 13 NanoBEIR datasets isolates what that choice buys:
| NanoBEIR dataset | LateOn (multi-vector, 128d) | DenseOn (dense, 768d) |
|---|---|---|
| MSMARCO | 0.7194 | 0.6517 |
| NQ | 0.7810 | 0.7511 |
| HotpotQA | 0.9295 | 0.8802 |
| FEVER | 0.9702 | 0.9612 |
| ClimateFEVER | 0.4887 | 0.4846 |
| DBPedia | 0.6836 | 0.6748 |
| QuoraRetrieval | 0.9795 | 0.9687 |
| Touche2020 | 0.5938 | 0.5673 |
| ArguAna | 0.5562 | 0.5660 |
| NFCorpus | 0.3949 | 0.3851 |
| SciFact | 0.7978 | 0.8057 |
| SCIDOCS | 0.4469 | 0.4484 |
| FiQA2018 | 0.5871 | 0.6491 |
| Mean | 0.6868 | 0.6764 |
Late interaction wins on 9 of the 13 datasets and on the mean, by roughly one NDCG point. The four it loses (ArguAna, FiQA2018, SCIDOCS, and SciFact) are the shape of the tradeoff you should expect: a real gain in retrieval quality at the same model size, paid for in index footprint, rather than a universal win on every dataset. The same pair scores 57.22 against 56.20 on the full 15-dataset BEIR, a comparable gap, so the margin is not an artifact of the small benchmark.
That footprint is the real cost. One vector per token instead of one vector per document is a lot more vectors, only partly offset by the smaller dimension. Encoding 4,874 Natural Questions passages with lightonai/LateOn produced 608,414 token vectors, an average of 124.8 per passage:
| Representation | Vectors | Dimensions | float32 index |
|---|---|---|---|
Dense, all-MiniLM-L6-v2 |
4,874 | 384 | 7.5 MB |
Dense, gte-modernbert-base |
4,874 | 768 | 15.0 MB |
Multi-vector, LateOn |
608,414 | 128 | 311.5 MB |
That is about 42x the storage of the MiniLM index. Token Pooling cuts the vector count before any of that, real late interaction indexes compress heavily (the same vectors take 88 MB as a fast-plaid PLAID index), and using a multi-vector model as a reranker over a dense first stage avoids building an index at all.
Multi-vector checkpoints have been published in several formats over the years. MultiVectorEncoder reads all of them, so loading looks the same whatever the model started life as:
from sentence_transformers import MultiVectorEncoder
# Native Sentence Transformers checkpoints. PyLate builds on the same schema,
# so any PyLate checkpoint loads identically
model = MultiVectorEncoder("lightonai/LateOn")
model = MultiVectorEncoder("mixedbread-ai/mxbai-edge-colbert-v0-17m")
model = MultiVectorEncoder("LiquidAI/LFM2-ColBERT-350M")
# Any Stanford-NLP ColBERT checkpoint, detected via the `HF_ColBERT` architecture
# marker. The inline projection weight and the recipe come from `artifact.metadata`
model = MultiVectorEncoder("colbert-ir/colbertv2.0")
model = MultiVectorEncoder("answerdotai/answerai-colbert-small-v1")
# transformers-native *ForRetrieval ports (ColPali, ColQwen2, ...)
model = MultiVectorEncoder("vidore/colqwen2-v1.0-hf")
# A bare transformer: a fresh random projection is appended, so training is required
model = MultiVectorEncoder("answerdotai/ModernBERT-base")The recipe knobs that differ per checkpoint (marker prefixes for queries and documents, length caps, whether queries are padded out with [MASK] tokens, and which tokens are skipped when scoring documents) all live in the module configs, so print(model) shows you exactly what you loaded:
model = MultiVectorEncoder("colbert-ir/colbertv2.0")
print(model)
"""
MultiVectorEncoder(
(0): Transformer({..., 'document_length': 180,
'query_expansion': {'strategy': 'fixed', 'attend': False, 'token': None, 'length': 32}})
(1): Dense({'in_features': 768, 'out_features': 128, 'bias': False, ...})
(2): MultiVectorMask({'skiplist_words': ['!', '"', '#', ...], 'skiplist_tasks': ['document'], ...})
(3): Normalize({...})
)
"""Following the design principle of the rest of the library, this behavior lives in swappable modules rather than in the model class: a Transformer producing contextualized token embeddings, a token-level Dense projecting each of them down, a MultiVectorMask deciding which tokens count during scoring, and a token-level Normalize.
These are the checkpoints we test against directly, ranked by retrieval quality. The sentence-transformers tag on the Hub is the list that stays current, and for text retrieval in particular, any PyLate or Stanford-NLP ColBERT checkpoint loads whether or not it carries the tag yet. Where a revision is listed, pass it until the pull request on that repository is merged.
Text retrieval (29 models). NanoBEIR is the mean NDCG@10 over the 13 NanoBEIR datasets, a fast proxy for English text retrieval quality. A - means the model was not evaluated on it, which is the case for the non-English models.
Visual document retrieval (22 models). These embed page images as documents and text as queries. NanoViDoRe is the equivalent proxy over the ViDoRe benchmark subsamples.
| Model | Parameters | NanoViDoRe | Notes |
|---|---|---|---|
| webAI-Official/webAI-ColVec1.1-8b | 8.4B | 0.6580 | needs trust_remote_code=True |
| webAI-Official/webAI-ColVec1.1-4b | 4.5B | 0.6520 | needs trust_remote_code=True |
| tencent/EVIE-Preview-4.5B | 4.54B | 0.6405 | - |
| TomoroAI/tomoro-colqwen3-embed-8b | 8.8B | 0.6206 | needs trust_remote_code=True |
| TomoroAI/tomoro-colqwen3-embed-4b | 4.4B | 0.6019 | needs trust_remote_code=True |
| vidore/colqwen2.5-v0.2 | 3.8B | 0.5402 | - |
| vidore/colqwen2.5-v0.1 | 3.8B | 0.5395 | - |
| vidore/colqwen-omni-v0.1 | 4.4B | 0.5309 | - |
| vidore/colpali-v1.3 | 2.9B | 0.4802 | - |
| vidore/colpali-v1.3-hf | 2.9B | 0.4793 | - |
| vidore/colpali-v1.2 | 2.9B | 0.4691 | - |
| vidore/colqwen2-v1.0 | 2.2B | 0.4685 | - |
| vidore/colqwen2-v0.1 | 2.2B | 0.4526 | - |
| vidore/colpali | 2.9B | 0.4516 | - |
| vidore/colpali-v1.1 | 2.9B | 0.4314 | - |
| vidore/colsmolvlm-v0.1 | 2.1B | 0.4054 | - |
| vidore/colpali-hard-v1.1 | 2.9B | 0.3949 | - |
| vidore/colSmol-500M | 507M | 0.3459 | - |
| vidore/colSmol-256M | 256M | 0.2673 | - |
| ModernVBERT/colmodernvbert | 252M | 0.2632 | - |
| vidore/colpali-v1.2-hf | 2.9B | - | - |
| vidore/colqwen2-v1.0-hf | 2.2B | - | - |
Note that NanoBEIR and NanoViDoRe are small benchmarks, so their scores are not a substitute for evaluating on your own data, which is always the right way to pick a model.
Late interaction is the state of the art for visual document retrieval: matching a text query against page images, with charts, tables, and layout intact, and no OCR step. This is what the ColPali family of models does, and those checkpoints run through the same API. Image documents are passed as URLs, local paths, or PIL images:
from sentence_transformers import MultiVectorEncoder
model = MultiVectorEncoder("vidore/colqwen2.5-v0.2")
queries = [
"What is the variable represented on the y-axis of the graph?",
"Total outlay is maximum in which year?",
]
images = [
"https://huggingface.co/datasets/sentence-transformers/example-documents/resolve/main/doc1.jpg",
"https://huggingface.co/datasets/sentence-transformers/example-documents/resolve/main/doc2.jpg",
"https://huggingface.co/datasets/sentence-transformers/example-documents/resolve/main/doc3.jpg",
"https://huggingface.co/datasets/sentence-transformers/example-documents/resolve/main/doc4.jpg",
]
# A page yields far more vectors than a query: one per image patch
query_embeddings = model.encode_query(queries)
document_embeddings = model.encode_document(images)
print(query_embeddings[0].shape, document_embeddings[0].shape)
# (25, 128) (755, 128)
scores = model.similarity(query_embeddings, document_embeddings)
print(scores)
# tensor([[13.8672, 12.3115, 12.1670, 11.0293],
# [ 7.2012, 14.7207, 6.9414, 6.9746]])The code is unchanged from the text case. Underneath, the processor handles the visual prompt and the image patches, and MaxSim scores query text tokens against document image patches. Page images are not the only non-text modality either: text, images, audio, and video are all accepted, and a checkpoint supports whichever of those its processor does, which model.modalities reports.
Because MaxSim is a sum of per-query-token maxima, a ranking decomposes exactly: every point of a document's score belongs to one query token and one document token. The new sentence_transformers.multi_vector_encoder.interpretability module overlays that decomposition onto the page as the standard ColPali heatmap, either aggregated over the query or one map per query token.
If the index footprint worries you, the most effective knob is to store fewer token vectors. HierarchicalTokenPooling implements the token pooling technique from Clavié, Chaffin, and Adams: it clusters each document's token vectors with Ward linkage on cosine similarity and replaces each cluster with its mean, keeping roughly 1 / pool_factor of the tokens.
from sentence_transformers import MultiVectorEncoder
from sentence_transformers.multi_vector_encoder.modules import HierarchicalTokenPooling
model = MultiVectorEncoder("lightonai/LateOn")
pooling = HierarchicalTokenPooling(pool_factor=2)
# 1. Per encode call
document_embeddings = model.encode_document(documents, token_pooling=pooling)
# 2. Standalone, on embeddings you already have saved
pooled = pooling.pool(document_embeddings)
# 3. Baked into the model, so every consumer of the checkpoint gets pooled documents
model.append(HierarchicalTokenPooling(pool_factor=2))
model.save_pretrained("my-pooled-colbert")By default, pooling applies to documents only, since queries are short and are the side you cannot afford to distort. On the Natural Questions corpus above, the reduction tracks pool_factor closely:
pool_factor |
Token vectors | Reduction | float32 index |
|---|---|---|---|
| 1 (off) | 608,414 | 1.00x | 311.5 MB |
| 2 | 305,438 | 1.99x | 156.4 MB |
| 3 | 204,407 | 2.98x | 104.7 MB |
| 4 | 153,936 | 3.95x | 78.8 MB |
The original experiments measured the retrieval cost of this on BEIR and found very little of it: 100.6% of the unpooled performance on average at pool_factor=2, and 99.0% at pool_factor=3. How much it costs on your data is corpus-specific, so measure it with an evaluator before you settle on a factor.
Introducing MultiVectorEncoder has been one of the largest updates to Sentence Transformers, introducing all of the following:
MultiVectorMask, BaseTokenPooling, HierarchicalTokenPooling, and LambdaTokenPoolingmaxsim, maxsim_pairwise, mean_maxsim, mean_maxsim_pairwise) plus 6 named ColBERT scorers and 5 XTR scoring entry pointstorch.compileSentence Transformers v6.0 requires transformers v5. The v4.x compatibility branches have been removed, which is what allows the new modality handling, chat template support, and unpadding paths to be relied upon rather than feature-detected. The floors that moved:
| Dependency | v5.7.0 | v6.0.0 |
|---|---|---|
transformers |
>=4.41.0,<6.0.0 |
>=5.0.0,<6.0.0 |
huggingface-hub |
>=0.23.0 |
>=1.3.0,<2.0.0 |
torch |
>=1.11.0 |
>=2.2 |
numpy |
>=1.20.0 |
>=1.24.0 |
scikit-learn |
>=0.22.0 |
>=1.1.0 |
typing_extensions |
>=4.5.0 |
>=4.10.0 |
datasets (train) |
>=2.0.0 |
>=2.16.0 |
accelerate (train) |
>=0.20.3 |
>=1.3.0 |
optimum-intel[openvino] |
unpinned | >=2.0.0 |
requires-python is unchanged at >=3.10. Note that multi-GPU training with streaming (IterableDataset) datasets needs accelerate>=1.13.0 in practice.
Half precision ties too many scores together to rank with. Three separate places where that mattered are now computed in float32.
Reranker scores are the big one. CrossEncoder.predict (and rank) now upcast the logits to float32 before applying the activation function. A sigmoid in bfloat16 saturates and collapses the top candidates onto a handful of tied values, which randomizes their order. Measured on cross-encoder/ettin-reranker-32m-v1 in bfloat16 over three NanoBEIR datasets with 100 candidates per query:
| Metric | v6.0.0 | v5.7.0 |
|---|---|---|
| NanoBEIR mean NDCG@10 | 0.6795 | 0.1849 |
| NanoBEIR mean MRR@10 | 0.6797 | 0.3986 |
| Unique scores over 15,040 pairs | 710 | 270 |
NanoMSMARCO NDCG@10 alone goes from 0.0965 to 0.7093. If you run a half precision reranker with the default sigmoid activation, its ranking was essentially randomized before this release. Models using activation_fn=nn.Identity() (raw logits) were unaffected, as bf16 logits keep enough relative spacing.
Similarity scores from model.similarity / similarity_pairwise and the cos_sim family are now computed in float32 for float16 and bfloat16 embeddings. With 10,000 realistic cosine scores (mean 0.7, standard deviation 0.05), float32 keeps 9,983 distinct values where float16 keeps 593 and bfloat16 keeps just 93. bfloat16 can represent only 129 distinct values in the whole of [0.5, 1.0).
MaxSim sums over query tokens, reaching magnitudes where the bfloat16 grid is 0.125 wide, so maxsim and maxsim_pairwise accumulate the per-token maxima in float32 and always return float32 scores. The 4-dimensional scoring intermediate stays in the input dtype, so this does not change peak memory.
Note that encode() output dtypes are unchanged. Only the scoring step is upcast. For CrossEncoder.predict, the returned dtype changes only with convert_to_tensor=True or convert_to_numpy=False, as the default numpy output was already float32.
Separately, the multi-vector bf16 benchmarks were re-measured under this float32 accumulation (#3924). Most of the previously reported bf16 quality drop came from the scoring accumulation rather than from the embeddings: plain bf16 now sits at 99.0% of fp32 retrieval quality (was 95.0%), and bf16 with FlashAttention-2 is indistinguishable from fp32 at 99.96% (was 97.9%).
similarity and similarity_pairwise are methods, not properties. Calls like model.similarity(embeddings1, embeddings2) work unchanged, but assigning a custom function to model.similarity is no longer supported: it now silently shadows the method where it previously raised an AttributeError. Set model.similarity_fn_name = "dot" instead, which updates both. Note also that model.similarity.__name__ is now "similarity" rather than the resolved function name, which affected loss get_config_dict() output and generated model cards. The new sentence_transformers.util.similarity_fct_name() resolves it properly and the losses use it.model.encode([{"role": "user", ...}, {"role": "assistant", ...}]) produces one embedding, where v5.x read it as a batch of two inputs. Wrap each conversation in its own list to encode a batch: model.encode([[msg1], [msg2]]). This applies to SentenceTransformer, SparseEncoder, and MultiVectorEncoder. CrossEncoder is unaffected.trust_remote_code=True (#3935). Loading a model whose modules.json references a class outside sentence_transformers executes third-party code, and a local directory no longer implies trust. This closes the bypass reported in #3801 and completes the deprecation cycle announced in v5.6 and v5.7. Unmet, it raises a ValueError naming the class and pointing at the repository or local path to inspect. Trainer checkpoint reloading (load_best_model_at_end, resume_from_checkpoint) keeps working for programmatically built models without the flag.quantize_embeddings returns a list of per-input matrices when given a list of 2D arrays, where it previously stacked them into one 3D array. Update callers that indexed the stacked array. An empty list now returns [] instead of raising, and a (0, dim) matrix returns a correctly shaped empty result.encode(pool=..., precision="int8") now quantizes once after merging the worker results, so the calibration ranges match single-process encoding. Quantized indexes built with v5.x multi-process encoding are not bit-compatible and should be regenerated. Peak memory is higher, because the full float32 matrix is materialized before quantization.CrossEncoder.rank returns Python floats (#3927) as its "score" values, where it previously returned numpy.float32 scalars or 0-dimensional tensors. The results are directly JSON serializable, matching semantic_search. convert_to_numpy and convert_to_tensor on rank are now deprecated no-ops: call predict directly if you want an array or a tensor. Beyond the cleaner output, this avoids a device synchronization per comparison when sorting, which took 212ms for 1000 CUDA scalars against 0.089ms for Python floats.Normalize moved to sentence_transformers.base.modules. Existing models load fine and silently, but a model saved by v6.0 with a Normalize module cannot be loaded by Sentence Transformers older than v6.0.SentenceTransformer checkpoint as a CrossEncoder (or any other such conversion) no longer picks up the source's prompts, default_prompt_name, similarity_fn_name, truncate_dim, or activation_fn, as those describe a model you are not loading. A reranker's default prompt being prepended to every encode call was the motivating case. Explicit keyword arguments still win. These conversions are now also logged at warning level, so they are visible at default verbosity.SimilarityFunction.possible_values() now includes "maxsim" and "meanmaxsim". Setting an unsupported similarity_fn_name on SentenceTransformer or SparseEncoder raises immediately rather than failing later, and a new SUPPORTED_SIMILARITY_FN_NAMES class attribute documents what each model type accepts.Multi-column losses now run one forward pass over merged columns (#3938). A training batch arrives as one feature dict per column (anchor, positive, negative_1, and so on), and the classic pattern runs the model once per column. The SentenceTransformer and SparseEncoder losses now pad and concatenate the like-width candidate columns into a single batch, keeping the anchor on its own forward pass since a 12-token query padded into 256-token documents costs more than it saves:
| Configuration | v5.7.0 | v6.0.0 |
|---|---|---|
| Natural Questions with 5 hard negatives | 316.2s | 250.7s (1.26x) |
| AllNLI triplets | 53.8s | 44.2s (1.22x) |
Loss trajectories match, up to dropout sampling. Losses fall back to per-column forward passes whenever the columns cannot be merged safely, for example with differing feature keys, disagreeing prompts or router tasks, or flattened Flash Attention inputs. The cached losses keep using GradCache, and AdaptiveLayerLoss opts out.
Backend benchmarks were re-measured for all four model types, with new Flash Attention columns and rewritten recommendations. For SentenceTransformer, float16 with Flash Attention and unpadding is now the fastest GPU configuration at 3.87x over float32, and ONNX on GPU is no longer recommended for short texts as float16 now beats it. For CrossEncoder, Flash Attention is explicitly not recommended, as unpadding does not apply to classification heads. For SparseEncoder, plain float16 remains the recommendation even though FA2 unpadding is now supported. See Speeding up Inference for the flowcharts.
Model authors can now record which package versions their checkpoint needs, and loading verifies them up front instead of failing in a confusing way later. Add a requirements mapping to config_sentence_transformers.json, using PEP 440 specifiers:
{
"model_type": "SentenceTransformer",
"requirements": {
"transformers": ">=5.15",
"peft": {
"specifier": ">=0.18,<0.20",
"reason": "Older versions ignore the key_mapping, which silently randomizes the adapter weights."
}
}
}Loading that model in an environment that does not satisfy it raises an ImportError listing every unmet requirement at once, with the optional reason included and a ready-to-run install command:
The model 'tomaarsen/my-model' requires:
- transformers>=5.15, but transformers==5.4.0 is installed.
- peft>=0.18,<0.20, but peft==0.17.0 is installed. Older versions ignore the key_mapping, which silently randomizes the adapter weights.
Install compatible versions with:
pip install -U "transformers>=5.15" "peft>=0.18,<0.20"
"python" and "pytorch" are understood as special names, prereleases are accepted so nightlies and .dev0 builds do not trip the check, and anything unparsable warns and is skipped rather than blocking the load. It works for all four model types. See Declaring Version Requirements for details.
Pooling(include_prompt=False) no longer corrupts repeated forward passes (#3944). The pooling module used to write its prompt-excluded mask back into features["attention_mask"], but that key is what the encoder attends over on the next forward pass, and it is where the prompt boundary is read from. Any loss that embeds the same feature dicts twice therefore got a different answer each time. AdaptiveLayerLoss is the headline victim: on a model with a 3-token prompt, two consecutive calls with identical inputs returned 2.706679 and then 10.334954, where it is now stable at 1.777709. DenoisingAutoEncoderLoss was hit from another angle, handing its decoder an all-zero cross-attention mask. If you trained AdaptiveLayerLoss on an include_prompt=False model with prompts, your results will move. As a side effect, encode(output_value=None) now reports the full mask the encoder used, matching the input_ids and token_embeddings in the same dictionary. sentence_embedding and output_value="token_embeddings" are bit-identical.TripletEvaluator and SparseTripletEvaluator embed anchors with encode_query and positives and negatives with encode_document, instead of encode for all three. This is a no-op for models without query / document prompts, but asymmetric models will report different triplet accuracy than in v5.x, since their prompts, router routes, and per-task length caps are now applied. Both also now reject unknown similarity_fn_names and unknown margin keys at construction, where a typo previously degraded silently to a missing metric or a zero margin.InformationRetrievalEvaluator breaks score ties by corpus id, making its metrics independent of corpus_chunk_size. Previously torch.topk(..., sorted=False) plus a heap comparison made tie retention depend on chunk boundaries and favor larger corpus ids. Metrics change only where exact ties exist, such as duplicate documents or quantized embeddings. Inherited by the sparse and NanoBEIR variants.DistillKLDivLoss and SparseDistillKLDivLoss gained per-side temperatures (student_temperature, teacher_temperature) following the DenseOn and LateOn recipes, plus validation that catches previously silent misuse: non-positive or non-finite temperatures, fewer than three columns (a softmax over one candidate is constant, so the loss and its gradient are identically zero), and teacher score shapes that do not match the candidate columns. A new one-time warning reports how many teacher scores underflowed to exactly zero and recommends a teacher_temperature floor.top_k must be positive, integer embeddings are upcast rather than producing an integer score grid, and a query padding mask is inferred from all-zero rows when none is given, matching the document side. XTR also now computes its Z normalizer as the paper's retrieval count rather than a positive-maxima proxy.MultipleNegativesRankingLoss rejects a NaN scale rather than accepting it.NoDuplicatesBatchSampler silently no-opping on media datasets in #3794: PIL images and torchcodec decoders stringify to a fresh object address on every access, so every row looked unique and no duplicates were ever detected. Large numpy arrays had the opposite problem, as their truncated string representation made distinct arrays collide. Values are now keyed by content. Batches change for datasets with image, audio, video, or array columns. Plain text and numeric datasets are byte-identical.xxhash 4.0 compatibility in #3928: xxh64_intdigest no longer accepts str, which crashed training with BatchSamplers.NO_DUPLICATES_HASHED (or precompute_hashes=True). Strings are now encoded before hashing. Digests are unchanged, so precomputed hashes stay valid.pixel_values and other base-model arguments being silently dropped for PEFT models in #3794: PeftModel.forward hid the wrapped model's parameters, so the forward argument allowlist was built from the wrapper. It now unwraps first.task and num_images_per_sample leaking into the transformers forward pass as unexpected keyword arguments in #3794, via an explicit denylist of Sentence Transformers internal feature keys that yields to a model actually declaring them.trainer.evaluate() for VLM losses in #3794 by no longer gating media count tracking on self.training.dataloader_persistent_workers=True that cost is paid again every epoch and every evaluation, commonly making training slower than dataloader_num_workers=0. The examples now use dataloader_num_workers=2 with peNote truncated.
The repository now has a Security Policy (#3858): please report vulnerabilities through GitHub's private vulnerability reporting, and see the policy f…
This minor version is a correctness and performance-focused release. It rebuilds all gradient-cached losses on one shared engine, fixing several silently wrong gradients and adding token-based mini-batching for up to 3.9x faster cached-loss training. It also makes model.compile() actually speed up inference, and brings a long list of fixes across embedding quantization, evaluators, hard-negative mining, community detection, and multimodal inputs.
Two changes are marked breaking (🚨): int8/uint8 embedding quantization now clips out-of-range values and floors bucket values, so int8 outputs are no longer bit-identical with earlier versions, and AdaptiveLayerLoss/Matryoshka2dLoss now weight prior-layer losses uniformly by default. There's also a forward-looking deprecation: loading models whose modules import classes from outside sentence_transformers will require trust_remote_code=True from v6.0.
Install this version with
# Training + Inference
pip install sentence-transformers[train]==5.7.0
# Inference only, use one of:
pip install sentence-transformers==5.7.0
pip install sentence-transformers[onnx-gpu]==5.7.0
pip install sentence-transformers[onnx]==5.7.0
pip install sentence-transformers[openvino]==5.7.0
# Multimodal dependencies (optional):
pip install sentence-transformers[image]==5.7.0
pip install sentence-transformers[audio]==5.7.0
pip install sentence-transformers[video]==5.7.0
# Or combine as needed:
pip install sentence-transformers[train,onnx,image]==5.7.0
The gradient-cached losses (CachedMultipleNegativesRankingLoss, CachedGISTEmbedLoss, CachedSpladeLoss, the Cross Encoder CachedMultipleNegativesRankingLoss, and MegaBatchMarginLoss) train with large batch sizes at constant memory by embedding in mini-batches and replaying them with cached gradients. Each loss carried its own diverged copy of that machinery. They are now all rebuilt on one shared engine, which fixed several bugs that silently corrupted gradients:
CachedMultipleNegativesRankingLoss on GPU: the backward pass used different dropout masks than the forward pass, silently biasing gradients for every reranker trained with dropout active on CUDA or MPS. CPU training was unaffected.CachedGISTEmbedLoss and the Cross Encoder loss backpropagate the wrong batch's gradients, because the cache was stored on the loss module. The cache now travels with each forward pass's backward hook (the .cache and .random_states loss attributes are gone as a result).Pooling(include_prompt=False) (e.g. Instructor models) mutated the attention mask in place, so the backward re-embedding of every cached loss ran with a different mask than the forward pass.MatryoshkaLoss(GISTEmbedLoss(...)): the guide model overwrote the cached embeddings, so only the largest Matryoshka dimension was actually trained.Along the way, this also fixed an autocast dtype crash in the backward pass and the trainer retaining autograd graphs between logging steps when tracking loss components.
MegaBatchMarginLoss's default mini-batched version is rebuilt on the engine as well. It crashed outright on recent releases, and underneath that, its historical implementation only applied the last mini-batch's gradients. It now trains on the full batch (results will differ, for the better), works with MatryoshkaLoss, evaluates under torch.no_grad, and raises for a third input column instead of silently ignoring it.
The headline feature is mini_batch_num_tokens, available on CachedMultipleNegativesRankingLoss, CachedMultipleNegativesSymmetricRankingLoss, CachedGISTEmbedLoss, CachedSpladeLoss, and MegaBatchMarginLoss. Instead of a fixed number of sequences per mini-batch, mini-batches are greedily packed by total non-padding token count, giving near-constant work per mini-batch on variable-length data:
from sentence_transformers import SentenceTransformer
from sentence_transformers.sentence_transformer.losses import CachedMultipleNegativesRankingLoss
model = SentenceTransformer("microsoft/mpnet-base")
loss = CachedMultipleNegativesRankingLoss(model, mini_batch_num_tokens=16384)
On the PR's Natural Questions benchmark, cached-loss training with flash attention and a tuned token budget dropped from 715 to 182 seconds (3.9x) versus the previous release, with unchanged quality. The engine also trims trailing padding from each mini-batch, which alone is worth about 26% throughput on the default padded path. The updated training efficiency documentation recommends the smallest token budget that saturates your GPU. mini_batch_size keeps working everywhere as before.
encode() and predict() previously called the model's forward() directly, bypassing nn.Module.__call__, which is where torch.compile installs its compiled path. As a result, model.compile() was silently a no-op for inference. The forward pass now runs through __call__, so compilation applies (outputs are bit-identical when not compiling).
from sentence_transformers import SentenceTransformer
model = SentenceTransformer("BAAI/bge-small-en-v1.5", model_kwargs={"torch_dtype": "bfloat16"})
model.compile(dynamic=True)
# Compilation is lazy, so warm up on representative inputs before benchmarking or serving
embeddings = model.encode(["This is an example sentence", "Each sentence is converted"])
For the largest gains at batch size 1, mode="reduce-overhead" applies CUDA graphs: measured in bf16 on an RTX 3090, roughly 3.0x on bge-small-en-v1.5 and 3.7x on modernbert-embed-large. Very small models like all-MiniLM-L6-v2 see little gain or even a slowdown, as their inference is dominated by tokenization and Python overhead, so always measure on your own model and hardware. The Speeding up Inference documentation for all three model types (SentenceTransformer, CrossEncoder, SparseEncoder) gains a torch.compile tab covering dynamic shapes, CUDA graphs, and the fixed-shape padding the latter need.
Two fixes for quantize_embeddings:
[-42, 42, -112] where [-128, 42, 127] was correct. Out-of-range values now saturate at the bounds. Bucket values are also floored before casting, which makes int8 a consistent uniform quantizer matching uint8, but shifts in-range int8 outputs in the lower half of the range down by one level out of 256. int8 quantization is therefore not bit-identical with earlier versions: re-quantize existing int8 corpora rather than mixing old and new quantized embeddings. uint8 outputs only change where they previously wrapped. Resolves #3159.precision="binary"/"ubinary" crashed with a ValueError when the embedding dimension was not a multiple of 8 (e.g. some Matryoshka-truncated dimensions), because the whole batch was packed as one flat bit array. Bits are now packed per embedding via np.packbits(..., axis=-1). Output is bit-identical for dimensions divisible by 8, like 384, 768, or 1024.AdaptiveLayerLoss and Matryoshka2dLoss train models whose embeddings remain useful when transformer layers are dropped, by also training each prior layer's output, including a KL-divergence term that distills the final layer's embeddings into the prior layers. Two changes, together resolving #3757:
layer_weighting (#3901): the prior-layer losses were hardcoded to decay as 1 / (1 + layer_idx). The new layer_weighting option accepts "uniform" (the new default, matching the 2DMSE and Starbucks papers), "log" (the ESE paper), "linear" (exactly the previous behavior), or any callable mapping a layer index to a weight. Uniform outperformed log and linear at every truncated layer count on STSb, hence the new default. Pass layer_weighting="linear" to restore the previous behavior exactly.Two silent failure modes of decoder-based models now raise actionable errors:
CrossEncoder reranker previously failed silently, as such a chat template does not know the query and document roles: both messages rendered to nothing, producing a cryptic IndexError or meaningless scores from an empty prompt. The first predict() call or training batch containing a pair now raises an error that names the model, explains why the roles are dropped, and includes a copyable Jinja template remedy. Published rerankers with working templates are unaffected (the 20 most-downloaded chat-template rerankers on the Hub all pass). Resolves #3881.processing_kwargs={"text": {"padding_side": "right"}} silently scored padding tokens. The produced attention mask itself is now verified, raising on the first batch that is actually padded on the wrong side.The repository now has a Security Policy (#3858): please report vulnerabilities through GitHub's private vulnerability reporting, and see the policy for the threat model and scope.
In the same vein, module class imports are now trust-checked (#3860). A model's modules.json references its module classes by import path. Classes from sentence_transformers itself are always fine, but a reference into any other installed package (e.g. some_library.models.CustomModule) used to be imported without any trust gate. Loading such a model from a remote, untrusted source now emits a FutureWarning, and from v6.0 it will require trust_remote_code=True, in line with the v5.6.0 deprecation of local custom code (#3807). The WordEmbeddings module's configurable tokenizer_class now goes through the same gate. Additionally, passing a file path as model_name_or_path raises a clear NotADirectoryError up front instead of potentially loading a same-named Hub repository.
Several fixes for BinaryClassificationEvaluator and ReciprocalRankFusionEvaluator, which their sparse counterparts inherit:
"euclidean" and "manhattan" in similarity_fn_names produced metrics for the opposite classifier: accuracy, F1, average precision, MCC, and the thresholds (which came out negative) were all wrong. Beyond reporting, with one of these as the first similarity function, the inverted average precision could steer metric_for_best_model/load_best_model_at_end during training. Expect euclidean/manhattan numbers to jump upward versus previous releases. Cosine and dot metrics were always correct.*_accuracy_threshold and *_f1_threshold columns, shifting every subsequent value under the wrong header, including in the default cosine-only configuration. The returned metrics and logs were unaffected, only the CSV file was misaligned.KeyError: 'cosine' (#3853): requesting multiple similarity functions without cosine (e.g. similarity_fn_names=["dot", "euclidean"]) crashed while aggregating the max_* metrics.ReciprocalRankFusionEvaluator used 0-based ranks and gave every document a score contribution from both retrievers, even from a retriever that did not return it, under-ranking documents that both retrievers agree on. The fusion now matches canonical RRF, so fused rankings and metrics shift, generally upward (the documented hybrid search example went from 32.62 to 32.95 NDCG@10).primary_metric (#3861): with a name set, ReciprocalRankFusionEvaluator prefixed its result keys but not primary_metric, so the standard results[evaluator.primary_metric] access raised KeyError, e.g. when used during training.A batch of fixes for mine_hard_negatives and the semantic_search_* helpers:
cache_folder cache key ignored the prompt arguments, so mining runs with the same texts but a different query_prompt/corpus_prompt (or prompt names) silently reused stale embeddings. The prompts are now part of the key. As a one-time side effect, caches written by older versions are recomputed after upgrading. Resolves #3870.-1 padding as candidates (#3872): the use_faiss=True path retrieved range_max + 1 candidates instead of range_max + max_positives, yielding fewer negatives than the default path on multi-positive datasets. Additionally, when the candidate window exceeded the corpus size, FAISS pads its results with index -1, which Python resolves to corpus[-1], so the last corpus document could be mined as a negative for any query. Padded slots are now disqualified.RuntimeError: selected index k out of range whenever range_max + max_positives exceeded the corpus size, easy to hit on modest corpora with default settings. It now mines as many negatives as the corpus allows, matching the FAISS path.num_negatives cannot fit in the [range_min, range_max) window used to crash mid-run with opaque tensor shape errors, and a negative range_min did not crash at all: it silently sliced the candidate window from the wrong end and mined the easiest candidates as "hard" negatives. Both now raise a clear ValueError. The 2048-candidate retrieval cap now only applies to FAISS on GPU, where it also fixes a crash by accounting for max_positives, and FAISS on CPU is no longer capped. Resolves #3903.semantic_search_faiss (#3887): with a corpus smaller than top_k, FAISS -1 padding leaked into the returned hits as {"corpus_id": -1} entries with garbage scores, could outrank real documents after rescoring, and segfaulted on an empty index with rescore=True. Result lists now only contain real documents, so they can hold fewer than top_k entries for small corpora.corpus_precision documentation (#3888): the semantic_search_faiss docs listed "int8"/"binary", but the function accepts "float32", "uint8", and "ubinary". The docs are corrected and unsupported values now raise an immediate ValueError.semantic_search_seismic (#3907): a query that matched no documents crashed the result formatting with an IndexError. Results now stay aligned with the input query order, with an empty list for no-match queries. The corpus_index type hint and docstring now describe the bare SeismicIndex the function actually accepts (the documented tuple never worked).semantic_search_qdrant (#3909): a single 1D query embedding (as returned by encode_query for one text) was iterated as if each vocabulary entry were its own query, silently running one Qdrant search per vocabulary token and returning that many result lists instead of one. It is now treated as a batch of one, and query tensors that are neither 1D nor 2D raise a clear ValueError.video_metadata alignment (#3876): in a batch mixing videos with and without metadata, the batch-level metadata list came out shorter than the batch, so metadata silently attached to the wrong videos and, with frame sampling enabled, trailing videos could be dropped entirely. Metadata is now aligned per sample. Audio batches mixing conflicting sampling_rates also now raise instead of silently processing all audio at whichever rate came last. Resolves #3874.RuntimeError (not just ImportError) when it is installed but unusable, e.g. on a broken FFmpeg setup, which made import sentence_transformers itself fail. It is now treated as an unavailable optional dependency, keeping text-only usage working. Resolves #3896.video_metadata or sampling_rate as a sibling key next to "video"/"audio" in a multimodal dict now produces an error that spells out the working nested form (e.g. {"audio": {"array": ..., "sampling_rate": 16000}}).ListMLELoss/PListMLELoss (#3827): padded list positions entered the normalizer with roughly unit mass each, so a query's loss and gradients depended on how much padding its batch happened to contain. Padded positions are now excluded before the normalizer. Retraining the documented MS MARCO recipes lifts NanoBEIR mean nDCG@10 from roughly 0.39 to 0.53 for ListMLELoss and from 0.514 to 0.525 for PListMLELoss.SparseCoSENTLoss (#3868): the default similarity_fct was the matrix-valued util.cos_sim where CoSENT needs the pairwise similarity, and broadcasting kept the loss finite while silently optimizing a different objective. The default is now util.pairwise_cos_sim, matching the docstring, CoSENTLoss, and SparseAnglELoss. If you passed similarity_fct explicitly, you were unaffected.ContrastiveLoss and OnlineContrastiveLoss expect 0/1 labels. They now emit a one-time warning when given anything else, explaining the actual behavior: ContrastiveLoss uses such labels directly as term weights, and OnlineContrastiveLoss silently drops those pairs from the loss. Resolves #3382.optimum ONNX export in #3831: model.config became a read-only property in v5.5.0, which broke optimum's ONNX export (it assigns model.config while standardizing attributes). A setter now delegates to the underlying transformers model. Resolves #3830.dataloader_persistent_workers=True in #3895: every evaluation built and prepared fresh eval DataLoaders whose persistent worker processes were never released, accumulating over long training runs until file descriptor exhaustion (Too many open files) or OOM. Prepared eval dataloaders are now cached and reused per eval dataset name, and evaluate()/get_eval_dataloader() accept a dataset name string.community_detection and reduce its memory usage in #3832: intermediate communities are stored as compact uint32 arrays (about 5x less memory), GPU results are moved to CPU once per batch, and the overlap-removal step is vectorized. Outputs are identical, and dense-community workloads measured up to 3.8x faster.community_detection candidate window on threshold ties in #3900: on the CPU path, members whose similarity equals the threshold exactly (e.g. duplicate detection with threshold=1.0 on identical or one-hot vectors) never triggered a window expansion, silently capping communities at an internal window size (often 50). Resolves #3899.select_max_active_dims in #3852: the utility zeroed non-top-k values in the caller's tensor in place. It now returns a fresh tensor, supports single 1D embeddings (previously a crash), and raises a ValueError for non-positive max_active_dims (previously returned all-zero embeddings for 0).model.similarity() and model.similarity_pairwise() with similarity_fn_name="euclidean" or "manhattan" crashed on a SparseEncoder when one side was encoded with convert_to_sparse_tensor=False (cosine and dot already handled the mix), as did pairwise_angle_sim. Mixed inputs now match the all-dense results.SentenceTransformer model cards were built from the shared base template, missing the "Additional Resources" documentation links that CrossEncoder and SparseEncoder cards already had.tokenizers as a direct dependency in #3844: it is imported directly but was only pulled in transitively via transformers, which mattered for strict resolvers and minimal environments. Resolves #3519.chore] Increment dev version by @tomaarsen in https://github.com/huggingface/sentence-transformers/pull/3824axis=-1 to np.packbits in binary/ubinary quantization by @JSap0914 in https://github.com/huggingface/sentence-transformers/pull/3825model_card] Use ST-specific model card template for ST by @tomaarsen in https://github.com/huggingface/sentence-transformers/pull/3828tests] Fix tiny-random tests with newer transformers by @tomaarsen in https://github.com/huggingface/sentence-transformers/pull/3829ci] Exclude librosa/numba/llvmlite on Python 3.13 by @tomaarsen in https://github.com/huggingface/sentence-transformers/pull/3835tests] Rename a reranker test model from v6 to v54 by @tomaarsen in https://github.com/huggingface/sentence-transformers/pull/3857security] Create a Security Policy by @tomaarsen in https://github.com/huggingface/sentence-transformers/pull/3858security] Trust-check every non-sentence-transformers module class import by @tomaarsen in https://github.com/huggingface/sentence-transformers/pull/3860tests] Skip bf16 + Windows + CPU forwards, as they can WindowsError on torch 2.13 by @tomaarsen in https://github.com/huggingface/sentence-transformers/pull/3863losses] Consolidate GradCache into one shared engine, fix silently wrong gradients, add mini_batch_num_tokens by @tomaarsen in https://github.com/huggingface/sentence-transformers/pull/3862video_metadata is passed as a sibling modality key by @tomaarsen in https://github.com/huggingface/sentence-transformers/pull/3866warn] Check causal left padding against the attention mask by @tomaarsen in https://github.com/huggingface/sentence-transformers/pull/3869fix] RRF scoring: only rank retrievers that returned a document, and use 1-based ranks by @eSVeeF in https://github.com/huggingface/sentence-transformers/pull/3883layer_weighting option to AdaptiveLayerLoss and Matryoshka2dLoss by @tomaarsen in https://github.com/huggingface/sentence-transformers/pull/3901ce] Raise when a chat template cannot carry query/document pairs by @tomaarsen in https://github.com/huggingface/sentence-transformers/pull/3902semantic_search_faiss results by @ErenAta16 in https://github.com/huggingface/sentence-transformers/pull/3887corpus_precision values documented for semantic_search_faiss by @ErenAta16 in https://github.com/huggingface/sentence-transformers/pull/3888trainer] Fix eval DataLoader worker leak with dataloader_persistent_workers by @mjun0812 in https://github.com/huggingface/sentence-transformers/pull/3895fix] Support mixed sparse/dense inputs in euclidean and manhattan similarity by @LuShadowX in https://github.com/huggingface/sentence-transformers/pull/3906Beyond the pull request authors listed above:
mini_batch_num_tokens in #3862, and @vineethsaivs surfaced the MegaBatchMarginLoss breakage (#3854) it also fixes.AdaptiveLayerLoss issues (#3757) behind #3880 and #3901.Full Changelog: https://github.com/huggingface/sentence-transformers/compare/v5.6.1...v5.7.0
This patch release fixes silently degraded embeddings for RoBERTa-family models when flash attention is requested with transformers v5, notably every
This patch release fixes silently degraded embeddings for RoBERTa-family models when flash attention is requested with transformers v5, notably every XLM-R based multilingual embedding model (BAAI/bge-m3, intfloat/multilingual-e5-large, etc.). The bug affected v5.5.0, v5.5.1, and v5.6.0.
Install this version with
# Training + Inference
pip install sentence-transformers[train]==5.6.1
# Inference only, use one of:
pip install sentence-transformers==5.6.1
pip install sentence-transformers[onnx-gpu]==5.6.1
pip install sentence-transformers[onnx]==5.6.1
pip install sentence-transformers[openvino]==5.6.1
# Multimodal dependencies (optional):
pip install sentence-transformers[image]==5.6.1
pip install sentence-transformers[audio]==5.6.1
pip install sentence-transformers[video]==5.6.1
# Or combine as needed:
pip install sentence-transformers[train,onnx,image]==5.6.1
Since v5.5.0, the Transformer module flattens text-only batches into one packed sequence when flash attention is requested, skipping all padding overhead for a notable performance improvement. The position_ids of that packed sequence restart at 0 for every text, which is correct for the vast majority of models. RoBERTa-family architectures however compute positions as padding_idx + 1 + n for the n-th token, so every token read a position embedding shifted by padding_idx + 1 (usually 2). Nothing crashes, the embeddings are just silently worse.
from sentence_transformers import SentenceTransformer
# An affected configuration: flash attention with an XLM-R based model
model = SentenceTransformer(
"BAAI/bge-m3",
model_kwargs={"attn_implementation": "flash_attention_2"},
)
Measured on BAAI/bge-m3:
| Evaluation | padded | packed, 0-based positions | packed, with this fix |
|---|---|---|---|
| stsb test Spearman | 0.8485 | 0.7239 | 0.8485 |
| NanoBEIR mean nDCG@10 | 0.6041 | 0.5414 | 0.6050 |
The quality loss recovers exactly once the offset is applied. The fix scans the loaded model's modules once for an int padding_idx stored next to a learned position_embeddings table, and offsets the packed position_ids when that pair is found. An audit of transformers finds 16 architectures with that pair (roberta, xlm_roberta, xlm_roberta_xl, camembert, roberta_prelayernorm, xmod, data2vec_text, longformer, luke, ibert, mpnet, markuplm, lilt, layoutlmv3, esm, and pp_doclayout_v2), all offset by exactly padding_idx + 1, and no 0-based or rotary architecture matches.
You are only affected if you encoded text with flash attention requested on transformers v5 with a RoBERTa-family checkpoint. The default padded path (e.g. sdpa) was never affected, and neither were MPNet models like all-mpnet-base-v2 despite mpnet appearing in the audit: transformers does not support flash attention for MPNet at all. If you did index a corpus with such a configuration, re-encode it after upgrading: pre-fix embeddings score notably worse and do not mix with post-fix embeddings.
ci] Exclude librosa/numba/llvmlite on Python 3.13 by @tomaarsen in https://github.com/huggingface/sentence-transformers/pull/3835tests] Skip bf16 + Windows + CPU forwards, as they can WindowsError on torch 2.13 by @tomaarsen in https://github.com/huggingface/sentence-transformers/pull/3863Full Changelog: https://github.com/huggingface/sentence-transformers/compare/v5.6.0...v5.6.1
There's also a forward-looking deprecation: loading local custom code without trust_remote_code=True now warns, and will require it from v6.0.
This minor version is a correctness- and robustness-focused release. It fixes a silent scoring bug for causal-LM rerankers, corrects several hard-negative mining and GIST loss edge cases, restores TSDAE on transformers v5, and adds Apple Silicon (MPS) support for the cached losses.
The headline fix affects chat-template models that read the final token position, i.e. causal-LM rerankers (like Qwen3-Reranker) and last-token-pooling embedders: when an over-long input was truncated, the chat template's trailing suffix (e.g. the assistant prefill the model scores from) was silently dropped, producing wrong scores with no error. There's also a forward-looking deprecation: loading local custom code without trust_remote_code=True now warns, and will require it from v6.0.
Install this version with
# Training + Inference
pip install sentence-transformers[train]==5.6.0
# Inference only, use one of:
pip install sentence-transformers==5.6.0
pip install sentence-transformers[onnx-gpu]==5.6.0
pip install sentence-transformers[onnx]==5.6.0
pip install sentence-transformers[openvino]==5.6.0
# Multimodal dependencies (optional):
pip install sentence-transformers[image]==5.6.0
pip install sentence-transformers[audio]==5.6.0
pip install sentence-transformers[video]==5.6.0
# Or combine as needed:
pip install sentence-transformers[train,onnx,image]==5.6.0
Chat-template models render the full conversation to a flat string before tokenizing, so when the rendered input is longer than the tokenizer's model_max_length, the tokenizer truncates it from the right and drops the template's trailing suffix: the fixed tokens a template appends after the content, e.g. a prompt, instruction, [/INST], or a trailing EOS. For models that read the final token position, this silently corrupted the result:
Qwen/Qwen3-Reranker-0.6B) score a pair from the last token's yes/no logits, andWhen the suffix was truncated away, that final position landed mid-document instead of after the prefill, so the score or embedding came from the wrong place.
Transformer.preprocess now detects when truncation drops the suffix and splices it back onto the tail of each truncated row. Because the fix lives in the shared base Transformer, it applies across SentenceTransformer, CrossEncoder, and SparseEncoder. It's enabled by default and saved to the model configuration. Pass processing_kwargs={"chat_template": {"restore_suffix": False}} to opt back into raw truncation.
A trio of correctness and scalability fixes for hard-negative mining and the GIST losses:
mine_hard_negatives(relative_margin=...) and the margin_strategy="relative" branch of GISTEmbedLoss / CachedGISTEmbedLoss used a multiplicative threshold (positive * (1 - margin)) that only behaves correctly when the positive-pair similarity is positive. When that similarity was negative, the threshold moved the wrong way and let through false negatives: candidates more similar to the anchor than the true positive. The threshold is now positive - |positive| * margin, identical to before for positive scores but correct for negative ones.gather_across_devices=True and a non-zero margin, the false-negative suppression mask protected the wrong columns on ranks beyond the first (it ignored the per-rank offset into the gathered batch), which set the true positive's logit to -inf and produced a +inf loss. The mask now accounts for the cross-rank offset, so multi-GPU GIST training stays finite.mine_hard_negatives(use_faiss=False) (the default) materialized the full (queries × corpus) similarity matrix at once, which could OOM on large corpora. It now batches over the query axis (controlled by faiss_batch_size, default 16384), bounding peak memory while producing identical results.transformers v5 (#3781)transformers v5 removed the private PreTrainedModel._tie_encoder_decoder_weights helper that DenoisingAutoEncoderLoss (TSDAE) used to tie its separate encoder and decoder. As a stopgap, v5.5 raised a RuntimeError for the default tie_encoder_decoder=True on transformers >= 5.0.0, effectively breaking TSDAE there unless you pinned an older transformers or disabled tying. TSDAE now ships its own tying routine that shares storage between encoder and decoder, so it works on both transformers <5 and >=5 with the default settings.
trust_remote_code (#3807)Sentence Transformers has historically treated any local model directory as implicitly trusted: local custom code (e.g. modeling_*.py) loaded even with trust_remote_code=False, unlike transformers. This discrepancy might be unexpected, so loading local custom code this way now emits a FutureWarning, and from v6.0 it will require trust_remote_code=True like in transformers.
Two fixes for training on Apple Silicon:
CachedMultipleNegativesRankingLoss and CachedGISTEmbedLoss crashed at construction on MPS because their RandContext used a CUDA-only RNG path. They now run on MPS with deterministic replay preserved.SparseEncoder sparsity on MPS: the legacy model.fit(..., use_amp=True) path hard-coded CUDA's AMP GradScaler / autocast, and SparseEncoder sparsity statistics called to_sparse_csr(), which is unimplemented on MPS. Both now work on Apple Silicon.bf16=True/fp16=True (with those enabled, autocast outputs float32 logits, so the common path was unaffected). Logits are now upcast to float32 in the loss.device_map placement with the device argument in #3823: loading with model_kwargs={"device_map": ...} previously placed the backbone via accelerate, then immediately moved it to the default device, defeating device_map. It now keeps the backbone in place (moving the other modules onto its device), and warns if both device and device_map are passed.{"image": img} was classified as a combined modality and rejected by models that support only that one modality (e.g. BGE-VL), with a self-contradicting error. It is now treated as the bare "image" input, which unblocks vision-retrieval benchmarks like MTEB that pass {"image": ...}.Modality 'message' is not supported). The errors are now scenario-specific and suggest what to do, such as encoding each modality separately.get_device_name in #3798: on PyTorch builds where torch.distributed is present but unavailable (some ROCm and CPU-only builds), get_device_name() crashed with AttributeError: module 'torch.distributed' has no attribute 'is_initialized'. It now checks is_available() first, across all distributed call sites.export_static_quantized_openvino_model defaulted its calibration dataset to the bare id "glue", which the stricter Hub repo-id validation now rejects. The default is now the namespaced "nyu-mll/glue".A batch of example and documentation modernization, mostly migrating example scripts off deprecated datasets script-loaders and bare ids onto maintained Hugging Face datasets so they run on datasets 4.x:
datasets 4.x in #3782 (e.g. quora, nq_open, yahoo_answers_topics).datasets in #3783.datasets in #3784.datasets in #3785.quora-duplicates labels it previously ignored.ContrastiveTensionLoss docstring in #3788, and add a low-VRAM hardware note to the efficiency docs in #3802.chore] Increment dev version by @tomaarsen in https://github.com/huggingface/sentence-transformers/pull/3775fix] Collapse single-key multimodal dicts to bare modality by @tomaarsen in https://github.com/huggingface/sentence-transformers/pull/3779ci] Pass HF read token for main CI by @tomaarsen in https://github.com/huggingface/sentence-transformers/pull/3808tests] Reuse model/dataset fixtures more in tests by @tomaarsen in https://github.com/huggingface/sentence-transformers/pull/3810ci] Avoid model2vec distill install in CI as it limits the transformers version by @tomaarsen in https://github.com/huggingface/sentence-transformers/pull/3813openvino] Fix calibration default & tests for optimum-intel 2.0 / openvino 2026 by @tomaarsen in https://github.com/huggingface/sentence-transformers/pull/3814fix] Clarify unsupported-modality error messages by @tomaarsen in https://github.com/huggingface/sentence-transformers/pull/3792MPS errors by @omkar-334 in https://github.com/huggingface/sentence-transformers/pull/3818deprecation] Warn when loading local custom code without trust_remote_code by @tomaarsen in https://github.com/huggingface/sentence-transformers/pull/3807Full Changelog: https://github.com/huggingface/sentence-transformers/compare/v5.5.1...v5.6.0
This patch release fixes a small quirk with multimodal inference when using single-key multimodal inputs like model.encode({"image": ...}).
This patch release fixes a small quirk with multimodal inference when using single-key multimodal inputs like model.encode({"image": ...}).
Install this version with
# Training + Inference
pip install sentence-transformers[train]==5.5.1
# Inference only, use one of:
pip install sentence-transformers==5.5.1
pip install sentence-transformers[onnx-gpu]==5.5.1
pip install sentence-transformers[onnx]==5.5.1
pip install sentence-transformers[openvino]==5.5.1
# Multimodal dependencies (optional):
pip install sentence-transformers[image]==5.5.1
pip install sentence-transformers[audio]==5.5.1
pip install sentence-transformers[video]==5.5.1
# Or combine as needed:
pip install sentence-transformers[train,onnx,image]==5.5.1
Previously, inference like model.encode({"image": ...}) or model.encode([{"image": ...}, ...]) would be inferred as the ("image",) modality, which differed from the inferred modality of "image" for just model.encode(my_image) or model.encode([my_image, my_image_2, ...]).
This results in confusing errors if the model doesn't have a modality_config mapping for ("image",) in addition to "image", so now a single-key multimodal dict is collapsed to the bare modality (just "image" in this example).
This affected this code:
from sentence_transformers import SentenceTransformer
model = SentenceTransformer('BAAI/BGE-VL-base', trust_remote_code=True)
embedding = model.encode({"image": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/blog/ettin-reranker/mteb_ndcg10_all-MiniLM-L6-v2.png"})
print(embedding.shape)
Which previously failed as the model only implements a path for "text", "image", and ("image", "text").
Full Changelog: https://github.com/huggingface/sentence-transformers/compare/v5.5.0...v5.5.1
This release ships the train-sentence-transformers Agent Skill, adds two new training losses, and brings a long list of robustness and correctness fix
This release ships the train-sentence-transformers Agent Skill, adds two new training losses, and brings a long list of robustness and correctness fixes.
The new train-sentence-transformers Agent Skill lets AI coding agents (Claude Code, Codex, Cursor, Gemini CLI, ...) drive end-to-end training and fine-tuning across all three model types. EmbedDistillLoss is a new embedding-level knowledge distillation loss for SentenceTransformer: it aligns a student model's embeddings with pre-computed teacher embeddings, an alternative to the score-based distillation provided by MarginMSELoss and DistillKLDivLoss. ADRMSELoss is a new listwise learning-to-rank loss for CrossEncoder from the Rank-DistiLLM paper. encode() and predict() also gain a per-call processing_kwargs override, and more.
Install this version with
# Training + Inference
pip install sentence-transformers[train]==5.5.0
# Inference only, use one of:
pip install sentence-transformers==5.5.0
pip install sentence-transformers[onnx-gpu]==5.5.0
pip install sentence-transformers[onnx]==5.5.0
pip install sentence-transformers[openvino]==5.5.0
# Multimodal dependencies (optional):
pip install sentence-transformers[image]==5.5.0
pip install sentence-transformers[audio]==5.5.0
pip install sentence-transformers[video]==5.5.0
# Or combine as needed:
pip install sentence-transformers[train,onnx,image]==5.5.0
train-sentence-transformers Agent Skill (#3752)If you use an AI coding agent (Claude Code, Codex, Cursor, Gemini CLI, OpenCode, ...), you can now install the train-sentence-transformers Agent Skill and ask your agent to fine-tune a model on your data:
hf skills add train-sentence-transformers # installs under ./.agents/skills/
hf skills add train-sentence-transformers --global # installs under ~/.agents/skills/
hf skills add train-sentence-transformers --claude # also symlinks into .claude/skills/
The skill gives the agent curated, version-aware guidance for training SentenceTransformer (bi-encoder), CrossEncoder (reranker), and SparseEncoder/SPLADE models, covering base model selection, loss and evaluator choice, hard-negative mining, distillation, LoRA, Matryoshka, multilingual training, static embeddings, plus a set of production-ready training template scripts. Then you can prompt your agent with things like:
"Train a multilingual sentence-transformer on Dutch legal pairs."
"Fine-tune a cross-encoder reranker on
(question, answer)pairs from my dataset, mine hard negatives, and push to my Hub repo.""Train a German sparse embedding model with high sparsity."
"Can you train a static embedding model on 100k code triplets?"
The skill lives in the repository under skills/train-sentence-transformers/ and is mirrored to the huggingface/skills marketplace on each release.
Introduces EmbedDistillLoss (Kim et al., 2023), an embedding-level knowledge distillation loss for SentenceTransformer. Rather than distilling teacher scores (MarginMSELoss, DistillKLDivLoss), it directly aligns the student's sentence_embedding with a pre-computed teacher embedding passed via the dataset's label column. The comparison uses a configurable distance_metric, one of "cosine" (the default), "l2", or "mse". When the student and teacher dimensions differ, pass projection_dim=<teacher_dim> to add a learnable projection from the student's embedding space into the teacher's. That projection lives on the loss rather than on the saved model, so use loss.save_projection(...) / loss.load_projection(...) to reuse it across stages (e.g. like done in Arkam et al. for Jina v5). As part of this change, MSELoss is now a thin subclass of EmbedDistillLoss with distance_metric="mse", and also gains the optional projection_dim argument.
from datasets import Dataset
from sentence_transformers import SentenceTransformer, SentenceTransformerTrainer
from sentence_transformers.sentence_transformer.losses import EmbedDistillLoss
student_model = SentenceTransformer("microsoft/mpnet-base")
teacher_model = SentenceTransformer("all-mpnet-base-v2")
train_dataset = Dataset.from_dict({
"sentence": ["It's nice weather outside today.", "He drove to work."],
})
# Pre-compute teacher embeddings once and store them as the `label` column
def add_teacher_embeddings(batch):
return {"label": teacher_model.encode(batch["sentence"]).tolist()}
train_dataset = train_dataset.map(add_teacher_embeddings, batched=True)
loss = EmbedDistillLoss(student_model, distance_metric="cosine")
# If the student and teacher dimensions differ, add a learnable projection:
# loss = EmbedDistillLoss(student_model, distance_metric="cosine", projection_dim=768)
trainer = SentenceTransformerTrainer(
model=student_model,
train_dataset=train_dataset,
loss=loss,
)
trainer.train()
See the updated model distillation examples and the loss overview for more.
Introduces ADRMSELoss (Approx Discounted Rank Mean Squared Error), a listwise learning-to-rank loss for CrossEncoder from the Rank-DistiLLM paper (Schlatt et al., ECIR 2025). It computes a differentiable approximation of each document's rank via pairwise sigmoids and minimizes the nDCG-discounted squared error against the true ranks derived from the labels. It expects listwise inputs: a (query, [doc1, ..., docN]) pair plus a [score1, ..., scoreN] label list per sample (binary or continuous labels, variable document counts allowed). It's designed for LLM-distillation reranking, where the per-document scores come from a strong LLM's ordering.
from datasets import Dataset
from sentence_transformers import CrossEncoder, CrossEncoderTrainer
from sentence_transformers.cross_encoder.losses import ADRMSELoss
model = CrossEncoder("microsoft/mpnet-base")
train_dataset = Dataset.from_dict({
"query": ["What are pandas?", "What is the capital of France?"],
"docs": [
["Pandas are a kind of bear.", "Pandas are kind of like fish."],
["The capital of France is Paris.", "Paris is the capital of France.", "Paris is quite large."],
],
"scores": [[0.95, 0.1], [0.98, 0.92, 0.2]],
})
loss = ADRMSELoss(model)
trainer = CrossEncoderTrainer(
model=model,
train_dataset=train_dataset,
loss=loss,
)
trainer.train()
There's a full MS MARCO example at training_ms_marco_adrmse.py. Note that LambdaLoss generally remains the strongest loss in the listwise family. See the Cross Encoder loss overview for guidance on picking a loss.
processing_kwargs override (#3753)SentenceTransformer.encode() / encode_query() / encode_document(), SparseEncoder.encode(), CrossEncoder.predict(), and model.preprocess() now accept a processing_kwargs argument that overrides the processor/tokenizer kwargs configured at construction time, for a single call. It has the same nested structure as the processing_kwargs constructor argument (top-level keys text, audio, image, video, common, chat_template) and is shallow-merged on top of the instance-level settings, so you can override just one setting (e.g. max_length) and leave the rest intact.
from sentence_transformers import SentenceTransformer
model = SentenceTransformer("all-MiniLM-L6-v2")
# Override processor kwargs (e.g. max_length, truncation) for this call only:
embeddings = model.encode(
["a short text", "a much longer text that you want truncated more aggressively ..."],
processing_kwargs={"text": {"max_length": 256, "truncation": True}},
)
This is especially handy for vision-language models, where you can change the image resolution per call, e.g. model.encode(images, processing_kwargs={"image": {"max_pixels": 256 * 256}}).
CrossEncoder module stacks that don't start with a Transformer, and recognize a trailing Dense(module_output_name="scores") as the scoring head, by @tomaarsen in #3742: num_labels now reads that head's out_features, and model.config / model.model return None when there's no underlying transformers model.InformationRetrievalEvaluator / NanoBEIREvaluator (or their sparse variants) was used during training, by @tomaarsen in #3741: the usage snippet then shows encode_query / encode_document, even without IR prompt names or a Router architecture.transformers version is too old to honor use_bidirectional_attention / is_causal flags in a model's config (e.g. for google/embeddinggemma-300m), rather than silently ignoring them, by @tomaarsen in #3726.pooling_mode="cls" previously returned the embedding at position 0, which is a [PAD] token for left-padded inputs (common with decoder-only models), silently producing incorrect sentence embeddings. It now uses the attention mask to find the first real token per sequence. Resolves #3208.int64-derived divisor in mean / mean_sqrt_len_tokens pooling forced the pooled output to fp32, which could crash the downstream Dense / scoring head with a dtype mismatch.DistributedDataParallel / torch.compile wrappers in AdaptiveLayerLoss (and Matryoshka2dLoss) by @tomaarsen in #3768: training with these losses under DDP or torch.compile previously crashed with TypeError: 'DistributedDataParallel' object is not subscriptable. Resolves #3170.preprocess / get_embedding_dimension on DDP-wrapped models in losses by @tomaarsen in #3746: training a CrossEncoder (or using MatryoshkaLoss) under DDP crashed with AttributeError: 'DistributedDataParallel' object has no attribute 'preprocess'.hub_strategy="every_save" / "checkpoint" / "all_checkpoints" were previously missing modules.json, config_sentence_transformers.json, README.md, and module subfolders, leaving those revisions unloadable.model_type from the archetype class on user subclasses by @tomaarsen in #3763: a plain subclass like class MyModel(SentenceTransformer): pass would silently load checkpoints via the conversion path (e.g. defaulting CLS-pooling models to mean pooling), producing wrong embeddings with no error. Resolves #3536. Note: a model previously saved through a subclass has the subclass name in its config and should be re-saved (or its config_sentence_transformers.json edited) under this fix.model.config property that delegates to the underlying transformers model's PretrainedConfig (or None if there is none) by @tomaarsen in #3764: this restores DeepSpeed ZeRO and other transformers integrations that read model.config.hidden_size, which previously crashed with AttributeError: 'SentenceTransformer' object has no attribute 'config'. Resolves #3531.file_io error handling for local paths and Hub failures by @tomaarsen in #3765: an incomplete local model path no longer raises a confusing HFValidationError (e.g. on Windows absolute paths), and transient Hub errors (auth, rate-limit, network) on critical files now propagate instead of silently falling back to a default architecture. Resolves #3370. A local directory whose name collides with a Hub repo id now takes precedence even if incomplete.Router children to load via the dynamic-module mechanism by @tomaarsen in #3749: a model whose architecture uses a Router with a repository-local custom child module class now loads with trust_remote_code=True instead of raising an ImportError.local_files_only) to the dynamic-module loader by @tomaarsen in #3766: private Hub repos with trust_remote_code=True repo-local custom modules now load on the first try instead of failing with a misleading ModuleNotFoundError. Resolves #3367.torchcodec-decoder-wrapped audio/video that appears inside a multimodal dict input (e.g. {"audio": {"array": ..., "sampling_rate": ...}, "text": ...}) so it reaches the processor correctly by @tomaarsen in #3736. Resolves #3732.ValueError from urlparse by @forhim007 in #3760: strings like "https://www.google.com)[google.com]" raised ValueError: Invalid IPv6 URL inside modality detection. They are now treated as plain text. Resolves #3758.Trainer.__init__. Such columns are now detected by modality and skipped, the stats sample is bounded (and reduced from 1000 to 100 rows), and a modality row was added to the model-card dataset stats table.ci] Reduce hub calls in tests by @tomaarsen in https://github.com/huggingface/sentence-transformers/pull/3727enh] The Qwen3 integrations are merged, no need for revision anymore by @tomaarsen in https://github.com/huggingface/sentence-transformers/pull/3729tests] Future-proof getting model keys as MODEL_MAPPING_NAMES is being removed by @tomaarsen in https://github.com/huggingface/sentence-transformers/pull/3730examples] Fix training dataset creation by @tomaarsen in https://github.com/huggingface/sentence-transformers/pull/3728feat] Add ADRMSELoss by @sky-2002 in https://github.com/huggingface/sentence-transformers/pull/3690model card] stats are computed over 100 samples, not 1000 by @tomaarsen in https://github.com/huggingface/sentence-transformers/pull/3739trainer] Push full Sentence Transformers layout from each checkpoint by @tomaarsen in https://github.com/huggingface/sentence-transformers/pull/3740model card] Set ir_model on the model card based on evaluators by @tomaarsen in https://github.com/huggingface/sentence-transformers/pull/3741feat] Allow Dense as CrossEncoder scoring head by @tomaarsen in https://github.com/huggingface/sentence-transformers/pull/3742fix] Expose preprocess/get_embedding_dimension on DDP-wrapped models in losses by @tomaarsen in https://github.com/huggingface/sentence-transformers/pull/3746fix] Allow Router children to load via dynamic-module mechanism by @tomaarsen in https://github.com/huggingface/sentence-transformers/pull/3749model_card] Fix newlines in datasets with large texts by @tomaarsen in https://github.com/huggingface/sentence-transformers/pull/3750fix] Don't upcast bf16/fp16 to fp32 in flash-attention pooling path by @tomaarsen in https://github.com/huggingface/sentence-transformers/pull/3751train-sentence-transformers by @tomaarsen in https://github.com/huggingface/sentence-transformers/pull/3752docs] Fix MTEB links + broken 'note' by @tomaarsen in https://github.com/huggingface/sentence-transformers/pull/3754examples] Modernize the MSMARCO training scripts, add MNRL + MarginMSE recipe by @tomaarsen in https://github.com/huggingface/sentence-transformers/pull/3761fix] Inherit model_type from archetype on user subclasses by @tomaarsen in https://github.com/huggingface/sentence-transformers/pull/3763fix] Delegate model.config to underlying transformers model by @tomaarsen in https://github.com/huggingface/sentence-transformers/pull/3764feat] Add EmbedDistillLoss by @yjoonjang in https://github.com/huggingface/sentence-transformers/pull/3665fix] Forward Hub auth to dynamic-module loader for private trust_remote_code models by @tomaarsen in https://github.com/huggingface/sentence-transformers/pull/3766fix] Robust file_io error handling for local paths and Hub failures by @tomaarsen in https://github.com/huggingface/sentence-transformers/pull/3765fix] Use first non-pad token for CLS pooling with left-padding by @tomaarsen in https://github.com/huggingface/sentence-transformers/pull/3767fix] Unwrap DDP/torch.compile wrappers in AdaptiveLayerLoss by @tomaarsen in https://github.com/huggingface/sentence-transformers/pull/3768docs] Use direct class imports in examples & docs (drop losses.MSELoss(...) style) by @tomaarsen in https://github.com/huggingface/sentence-transformers/pull/3770examples] Avoid LoggingHandler, silence httpx in examples by @tomaarsen in https://github.com/huggingface/sentence-transformers/pull/3771docs] Use modality-neutral terms (input, document) in loss docs & docstrings by @tomaarsen in https://github.com/huggingface/sentence-transformers/pull/3772docs] Load models in float32 in the training examples & docs by @tomaarsen in https://github.com/huggingface/sentence-transformers/pull/3773Full Changelog: https://github.com/huggingface/sentence-transformers/compare/v5.4.1...v5.5.0
This patch release allows encode() and predict() to accept 1D numpy string arrays as inputs.
This patch release allows encode() and predict() to accept 1D numpy string arrays as inputs.
Install this version with
# Training + Inference
pip install sentence-transformers[train]==5.4.1
# Inference only, use one of:
pip install sentence-transformers==5.4.1
pip install sentence-transformers[onnx-gpu]==5.4.1
pip install sentence-transformers[onnx]==5.4.1
pip install sentence-transformers[openvino]==5.4.1
# Multimodal dependencies (optional):
pip install sentence-transformers[image]==5.4.1
pip install sentence-transformers[audio]==5.4.1
pip install sentence-transformers[video]==5.4.1
# Or combine as needed:
pip install sentence-transformers[train,onnx,image]==5.4.1
encode() and predict() now correctly recognize 1D numpy string/object arrays as batches rather than singular inputs. Previously, something like model.encode(df["text"].to_numpy()) was silently treated as a single input and produced incorrect output. 1D numpy arrays with dtype.kind in ("U", "O") are now unpacked like lists, and 2D+ arrays are treated as batches of pairs (for CrossEncoder).
import numpy as np
from sentence_transformers import SentenceTransformer
model = SentenceTransformer("all-MiniLM-L6-v2")
# Previously treated as one input; now correctly encoded as 3 separate texts
embeddings = model.encode(np.array(["first", "second", "third"]))
print(embeddings.shape)
# (3, 384)
For CrossEncoder, a 1D numpy string array is still treated as a single [query, document] pair to match the existing list behavior, while a 2D array of shape (N, 2) is a batch of N pairs.
Dense (#3714)The Dense module stores its activation function as a dotted import path in its saved config (e.g. "torch.nn.modules.activation.Tanh"), which was then resolved via import_from_string whenever the module was loaded. Because any importable Python callable could be referenced, a maliciously crafted config.json on the Hub could trigger arbitrary imports at model load time.
The loader now only resolves activation functions whose import path starts with torch.. Anything else is skipped with a warning and replaced by the default activation (Tanh). To load a model with a custom (non-torch) activation function, opt in explicitly with trust_remote_code=True:
from sentence_transformers import SentenceTransformer
# Torch-provided activations load as before
model = SentenceTransformer("some/model-with-torch-activation")
# Non-torch activations now require explicit opt-in
model = SentenceTransformer("some/model-with-custom-activation", trust_remote_code=True)
This mirrors the opt-in trust model already used by transformers for custom code, and ensures untrusted model repositories cannot smuggle arbitrary imports through the Dense activation config.
tests] Fix test_trainer_prompts for SE and ST after prompt handling moved into Transformer.preprocess by @tomaarsen in https://github.com/huggingface/sentence-transformers/pull/3710chore] Increment dev version after v5.4 release by @tomaarsen in https://github.com/huggingface/sentence-transformers/pull/3711docs] No revision needed anymore for nvidia nemotron by @tomaarsen in https://github.com/huggingface/sentence-transformers/pull/3712chore] Replace evaluation_strategy with eval_strategy in a few more places by @tomaarsen in https://github.com/huggingface/sentence-transformers/pull/3713security] Only load activation functions starting with 'torch' in the Dense module by @tomaarsen in https://github.com/huggingface/sentence-transformers/pull/3714fix] Treat numpy string/object arrays as batches in encode/predict by @tomaarsen in https://github.com/huggingface/sentence-transformers/pull/3720Full Changelog: https://github.com/huggingface/sentence-transformers/compare/v5.4.0...v5.4.1
> Migration guide: Migrating from v5.x to v5.4+: covers updated import paths, renamed parameters, and other softly breaking changes with deprecation w…
This large release introduces first-class multimodal support for both SentenceTransformer and CrossEncoder, making it easy to compute embeddings and rerank across text, images, audio, and video. The CrossEncoder class has been fully modularized, allowing for generative rerankers (CausalLM-based models) via a new LogitScore module. Flash Attention 2 now automatically skips padding for text-only inputs, providing significant speedups & memory reductions, especially when input lengths vary.
Blog post: Multimodal Embedding & Reranker Models with Sentence Transformers: a walkthrough of the new multimodal capabilities with some practical examples.
Migration guide: Migrating from v5.x to v5.4+: covers updated import paths, renamed parameters, and other softly breaking changes with deprecation warnings. Note that there are no hard deprecations, all existing code should continue to work with warnings at worst.
Install this version with
# Training + Inference
pip install sentence-transformers[train]==5.4.0
# Inference only, use one of:
pip install sentence-transformers==5.4.0
pip install sentence-transformers[onnx-gpu]==5.4.0
pip install sentence-transformers[onnx]==5.4.0
pip install sentence-transformers[openvino]==5.4.0
# Multimodal dependencies (optional):
pip install sentence-transformers[image]==5.4.0
pip install sentence-transformers[audio]==5.4.0
pip install sentence-transformers[video]==5.4.0
# Or combine as needed:
pip install sentence-transformers[train,onnx,image]==5.4.0
SentenceTransformer now natively supports vision-language models (VLMs) and other multimodal architectures. You can encode and compare across text, images, audio, videos, or combinations of these, with automatic modality detection and preprocessing. Models advertise which modalities they support via the new model.modalities property and model.supports() method.
from PIL import Image
from sentence_transformers import SentenceTransformer
model = SentenceTransformer(
"Qwen/Qwen3-VL-Embedding-2B",
model_kwargs={"attn_implementation": "flash_attention_2", "torch_dtype": "bfloat16"},
processor_kwargs={"min_pixels": 28 * 28, "max_pixels": 600 * 600},
revision="refs/pr/23",
)
# Check supported modalities
print(model.modalities)
# ['text', 'image', 'video', 'message']
print(model.supports("image"))
# True
# Encode text
text_embeddings = model.encode(["A photo of a cat", "A pollinator on a flower"])
# Encode images (PIL images, file paths, or URLs all work)
image_embeddings = model.encode([
Image.open("cat.jpg"),
"https://example.com/flower.jpg",
])
# Encode mixed text+image inputs
multimodal_embeddings = model.encode([
{"text": "Describe this image", "image": Image.open("cat.jpg")},
])
# Compute cross-modal similarity
similarity = model.similarity(text_embeddings, image_embeddings)
You can also compose separate encoders for different modalities using the new Router module. Unlike the single-backbone VLM approach, Router lets you combine any existing text and image encoders and route inputs based on detected modality:
from sentence_transformers import SentenceTransformer
from sentence_transformers.sentence_transformer.modules import Dense, Pooling, Router, Transformer
# Text encoder: MiniLM with mean pooling, projected to 768 dims to match image encoder
text_encoder = Transformer("sentence-transformers/all-MiniLM-L6-v2")
text_pooling = Pooling(text_encoder.get_embedding_dimension(), pooling_mode="mean")
text_projection = Dense(text_encoder.get_embedding_dimension(), 768)
# Image encoder: SigLIP outputs pooled embeddings directly
image_encoder = Transformer("google/siglip2-base-patch16-224")
# Route inputs to the appropriate encoder based on detected modality
router = Router(
sub_modules={
"text": [text_encoder, text_pooling, text_projection],
"image": [image_encoder],
},
)
model = SentenceTransformer(modules=[router])
# Text and image inputs are automatically routed to the correct encoder
text_embeddings = model.encode(["A photo of a cat"])
image_embeddings = model.encode(["https://example.com/cat.jpg"])
similarity = model.similarity(text_embeddings, image_embeddings)
CrossEncoder now supports multimodal inputs for reranking, enabling cross-modal scoring of query-document pairs where either side can be text, images, audio, video, or mixed-modality content. This works with both generative rerankers (CausalLM-based, via the new LogitScore module) and encoder-based models. See the pretrained multimodal rerankers for models you can use right away.
from sentence_transformers import CrossEncoder
# Load a multimodal reranker
model = CrossEncoder("Qwen/Qwen3-VL-Reranker-2B", revision="refs/pr/11")
# Rank text documents against an image query (or vice versa)
results = model.rank(
query="https://example.com/product.jpg",
documents=["A red sneaker", "A blue dress", "A leather bag"],
)
Two training approaches are provided in the multimodal training examples:
CrossEncoder has been fully modularized, inheriting from BaseModel (which is a torch.nn.Sequential). You can now inspect, customize, and compose module chains, just like SentenceTransformer. See the custom models guide for full details.
from sentence_transformers import CrossEncoder
model = CrossEncoder("Qwen/Qwen3-Reranker-0.6B", revision="refs/pr/11")
print(model)
"""
CrossEncoder(
(0): Transformer({'transformer_task': 'text-generation', ...})
(1): LogitScore({'true_token_id': 9693, 'false_token_id': 2152, ...})
)
"""
Thanks to the modular architecture, generative rerankers like mixedbread-ai/mxbai-rerank-base-v2 now work out of the box. These models ship with a modules.json that configures the Transformer + LogitScore chain automatically:
from sentence_transformers import CrossEncoder
model = CrossEncoder("mixedbread-ai/mxbai-rerank-base-v2")
scores = model.predict([
("How many people live in Berlin?", "Berlin had a population of 3,520,031 in 2022."),
("How many people live in Berlin?", "Berlin is well known for its museums."),
])
# array([ 9. , -0.5], dtype=float32)
The Transformer module now supports multiple task types that determine how the underlying model is loaded and what outputs it produces:
"sequence-classification": Loads via AutoModelForSequenceClassification, returns classification logits directly."text-generation": Loads via AutoModelForCausalLM, returns raw logits from the language model head."any-to-any": Loads via AutoModelForMultimodalLM (transformers v5+), for multimodal causal LMs that accept interleaved image/text inputs."feature-extraction": Loads via AutoModel (no task-specific head), returns hidden states.Various module chains are possible now, here's some common ones:
Encoder-based (Sequence Classification): A single Transformer module with transformer_task="sequence-classification", the traditional BERT/RoBERTa approach. This was previously the only option for CrossEncoder models.
CausalLM-based (Text Generation + LogitScore): For generative rerankers (Qwen, Llama, mxbai-rerank-v2, etc.), a Transformer with transformer_task="text-generation" followed by a LogitScore module that computes logit["yes"] - logit["no"] at the last token position. For multimodal rerankers, transformer_task="any-to-any" is used instead.
Feature Extraction + Pooling + Dense: A memory-efficient alternative that uses the base model without LM head, pools the last token, and projects to a single score via a Dense layer.
When loading a model without a modules.json, CrossEncoder automatically selects the right chain: if the architecture ends with ForCausalLM, it uses text-generation + LogitScore (with "yes"/"no" tokens); otherwise it uses sequence-classification. You can also construct custom module chains explicitly:
from sentence_transformers import CrossEncoder
from sentence_transformers.cross_encoder.modules import Transformer, LogitScore
transformer = Transformer("Qwen/Qwen3-Reranker-0.6B", transformer_task="text-generation", revision="refs/pr/11")
true_id = transformer.tokenizer.convert_tokens_to_ids("1")
false_id = transformer.tokenizer.convert_tokens_to_ids("0")
model = CrossEncoder(modules=[transformer, LogitScore(true_token_id=true_id, false_token_id=false_id)])
CrossEncoder.__init__ arguments after model_name_or_path are now keyword-only.tokenizer_args/tokenizer_kwargs -> processor_kwargs (with deprecation warnings).max_length -> max_seq_length (with deprecation warning).default_activation_function -> activation_fn (with deprecation warning).When using Flash Attention 2, Sentence Transformers now automatically skips padding for text-only inputs by concatenating all sequences into a single flat tensor. This eliminates wasted computation on padding tokens and is especially beneficial when input lengths vary widely within a batch. See the efficiency docs for more details.
The feature is enabled automatically when all prerequisites are met:
transformers >= 5.0.0pip install kernels (recommended) or pip install flash-attn)attn_implementation="flash_attention_2""feature-extraction" transformer task"torch" backend and supports the "text" modalityfrom sentence_transformers import SentenceTransformer
model = SentenceTransformer(
"sentence-transformers/all-MiniLM-L6-v2",
model_kwargs={"attn_implementation": "flash_attention_2", "torch_dtype": "bfloat16"},
)
# Padding is automatically skipped for text inputs
embeddings = model.encode(["short", "a much longer sentence that would normally cause padding"])
You can manually control the behavior via the unpad_inputs property on the Transformer module:
model[0].unpad_inputs = False # Force padding (e.g. for architectures that don't support unpadded inputs)
model[0].unpad_inputs = True # Explicitly request unpadding
model[0].unpad_inputs = None # Auto-detect (default)
The following benchmark compares throughput and VRAM usage across three attention configurations using BAAI/bge-base-en-v1.5, averaged across batch sizes. Four datasets with varying text lengths are tested: stsb (avg 10 tokens), natural-questions (avg 139 tokens), imdb (avg 304 tokens), and a shuffled mix of all three.
Flash Attention 2 with input flattening always outperforms standard Flash Attention 2, while using considerably less VRAM. The gains grow with the variance in input length, with the mixed dataset with wildly varying lengths (10-500 tokens) benefitting the most.
SentenceTransformer automatically adds a Pooling module for a causal language model (e.g. Llama, Qwen), it now defaults to last-token pooling instead of mean pooling. This better matches how decoder-only models represent sequences. Existing models with an explicit pooling configuration are unaffected.TripletLoss by @tomaarsen in #3704: The Euclidean and Manhattan distance metrics in TripletLoss were missing a negation, causing them to behave as similarity metrics rather than distance metrics, which inverted the loss. Cosine distance was not affected.SparseEncoder.encode to prevent Router misrouting by @ratatouille-plat in #3695: The max_active_dims parameter was leaking into tokenize kwargs, which could cause the Router to misroute inputs.chore] Increment dev version to v5.4.0.dev0 following the v5.3.0 release by @tomaarsen in https://github.com/huggingface/sentence-transformers/pull/3688ci] Remove cache, switch to uv for CI by @tomaarsen in https://github.com/huggingface/sentence-transformers/pull/3689ci] Install model2vec with the distill extra by @tomaarsen in https://github.com/huggingface/sentence-transformers/pull/3698v5.4] Introduce cross-modality and multi-modality support; modularize CrossEncoder class by @tomaarsen in https://github.com/huggingface/sentence-transformers/pull/3554fix] Fix inverted Euclidean/Manhattan distance in TripletLoss by @tomaarsen in https://github.com/huggingface/sentence-transformers/pull/3704docs] Jina-reranker-m0 doesn't require a revision anymore by @tomaarsen in https://github.com/huggingface/sentence-transformers/pull/3705Full Changelog: https://github.com/huggingface/sentence-transformers/compare/v5.3.0...v5.4.0
Your coding agent can read these notes before it upgrades. Set up the MCP server →