NewYour coding agent can read the release notes before it upgrades.Set up the MCP server →
PyPI · #715 most downloaded on PyPI
Embeddings, Retrieval, and Reranking
Last release 16 days ago
18 Sep 2026
Ships fairly regularly
a new release about every 3 weeks
Nearly every release is documented
notes for 55 of the last 60 stable releases
1 version withdrawn
withdrawn after publishing
7 years old
84 releases · first in 2019
This minor version brings several improvements to contrastive learning: MultipleNegativesRankingLoss now supports alternative InfoNCE formulations (sy
This minor version brings several improvements to contrastive learning: MultipleNegativesRankingLoss now supports alternative InfoNCE formulations (symmetric, GTE-style) and optional hardness weighting for harder negatives. Two new losses are introduced, GlobalOrthogonalRegularizationLoss for embedding space regularization and CachedSpladeLoss for memory-efficient SPLADE training. The release also adds a faster hashed batch sampler, fixes GroupByLabelBatchSampler for triplet losses, and ensures full compatibility with the latest Transformers v5 versions.
Install this version with
# Training + Inference
pip install sentence-transformers[train]==5.3.0
# Inference only, use one of:
pip install sentence-transformers==5.3.0
pip install sentence-transformers[onnx-gpu]==5.3.0
pip install sentence-transformers[onnx]==5.3.0
pip install sentence-transformers[openvino]==5.3.0
MultipleNegativesRankingLoss received two major upgrades: support for alternative InfoNCE formulations from the literature, and optional hardness weighting to up-weight harder negatives.
MultipleNegativesRankingLoss now supports several well-known contrastive loss variants from the literature through new directions and partition_mode parameters. Previously, this loss only supported the standard forward direction (query → doc). You can now configure which similarity interactions are included in the loss:
"query_to_doc" (default): For each query, its matched document should score higher than all other documents."doc_to_query": The symmetric reverse — for each document, its matched query should score higher than all other queries."query_to_query": For each query, all other queries should score lower than its matched document."doc_to_doc": For each document, all other documents should score lower than its matched query.The partition_mode controls how scores are normalized: "joint" computes a single softmax over all directions, while "per_direction" computes a separate softmax per direction and averages the losses.
These combine to reproduce several loss formulations from the literature:
Standard InfoNCE (default, unchanged behavior):
loss = MultipleNegativesRankingLoss(model)
# equivalent to directions=("query_to_doc",), partition_mode="joint"
Symmetric InfoNCE (Günther et al. 2024) — adds the reverse direction so both queries and documents are trained to find their match:
loss = MultipleNegativesRankingLoss(
model,
directions=("query_to_doc", "doc_to_query"),
partition_mode="per_direction",
)
GTE improved contrastive loss (Li et al. 2023) — adds same-type negatives (query <-> query, doc <-> doc) for a stronger training signal, especially useful with pairs-only data:
loss = MultipleNegativesRankingLoss(
model,
directions=("query_to_doc", "query_to_query", "doc_to_query", "doc_to_doc"),
partition_mode="joint",
)
Adds optional hardness weighting to MultipleNegativesRankingLoss and CachedMultipleNegativesRankingLoss, inspired by Lan et al. 2025 (LLaVE). This up-weights harder negatives in the softmax by adding hardness_strength * stop_grad(cos_sim) to selected negative logits. The feature is off by default (hardness_mode=None), so existing behavior is unchanged.
The hardness_mode parameter controls which negatives receive the penalty:
"in_batch_negatives": Penalizes in-batch negatives only (positives and hard negatives from other samples). Works with all data formats including pairs-only."hard_negatives": Penalizes explicit hard negatives only (columns beyond the first two). Only active when hard negatives are provided."all_negatives": Penalizes both in-batch and hard negatives, leaving only the positive unpenalized.from sentence_transformers.losses import MultipleNegativesRankingLoss
loss = MultipleNegativesRankingLoss(
model,
hardness_mode="in_batch_negatives",
hardness_strength=9.0,
)
Introduces GlobalOrthogonalRegularizationLoss (Zhang et al. 2017), a regularization loss that encourages embeddings to be well-distributed in the embedding space. It penalizes two things: (1) high mean pairwise similarity across unrelated embeddings, and (2) high second moment of similarities (which indicates clustering). This loss is meant to be combined with a primary contrastive loss like MultipleNegativesRankingLoss. By wrapping both losses in a single module, you can share embeddings and only require one forward pass:
import torch
from datasets import Dataset
from torch import Tensor
from sentence_transformers import SentenceTransformer, SentenceTransformerTrainer
from sentence_transformers.losses import GlobalOrthogonalRegularizationLoss, MultipleNegativesRankingLoss
from sentence_transformers.util import cos_sim
model = SentenceTransformer("microsoft/mpnet-base")
train_dataset = Dataset.from_dict({
"anchor": ["It's nice weather outside today.", "He drove to work."],
"positive": ["It's so sunny.", "He took the car to the office."],
})
class InfoNCEGORLoss(torch.nn.Module):
def __init__(self, model: SentenceTransformer, similarity_fct=cos_sim, scale=20.0) -> None:
super().__init__()
self.model = model
self.info_nce_loss = MultipleNegativesRankingLoss(model, similarity_fct=similarity_fct, scale=scale)
self.gor_loss = GlobalOrthogonalRegularizationLoss(model, similarity_fct=similarity_fct)
def forward(self, sentence_features: list[dict[str, Tensor]], labels: Tensor | None = None) -> Tensor:
embeddings = [self.model(sentence_feature)["sentence_embedding"] for sentence_feature in sentence_features]
info_nce_loss: dict[str, Tensor] = {
"info_nce": self.info_nce_loss.compute_loss_from_embeddings(embeddings, labels)
}
gor_loss: dict[str, Tensor] = self.gor_loss.compute_loss_from_embeddings(embeddings, labels)
return {**info_nce_loss, **gor_loss}
loss = InfoNCEGORLoss(model)
trainer = SentenceTransformerTrainer(
model=model,
train_dataset=train_dataset,
loss=loss,
)
trainer.train()
Introduces CachedSpladeLoss, a gradient-cached version of SpladeLoss that enables training SPLADE models with larger batch sizes without additional GPU memory. It applies the GradCache technique at the SpladeLoss wrapper level, so both the base loss and regularizers receive pre-computed embeddings — no changes to existing base losses or regularizers are needed.
from datasets import Dataset
from sentence_transformers.sparse_encoder import SparseEncoder, SparseEncoderTrainer
from sentence_transformers.sparse_encoder.losses import CachedSpladeLoss, SparseMultipleNegativesRankingLoss
model = SparseEncoder("distilbert/distilbert-base-uncased")
train_dataset = Dataset.from_dict({
"anchor": ["It's nice weather outside today.", "He drove to work."],
"positive": ["It's so sunny.", "He took the car to the office."],
})
loss = CachedSpladeLoss(
model=model,
loss=SparseMultipleNegativesRankingLoss(model),
document_regularizer_weight=3e-5,
query_regularizer_weight=5e-5,
mini_batch_size=32,
)
trainer = SparseEncoderTrainer(model=model, train_dataset=train_dataset, loss=loss)
trainer.train()
Adds a NO_DUPLICATES_HASHED batch sampler option, which uses the existing NoDuplicatesBatchSampler with precompute_hashes=True. This pre-computes xxhash 64-bit values for each sample, providing significant speedups for large batch sizes at a small memory cost. Requires the xxhash library.
from sentence_transformers import SentenceTransformerTrainingArguments
args = SentenceTransformerTrainingArguments(
batch_sampler="NO_DUPLICATES_HASHED" # Pre-computes hashes for faster duplicate checking
)
Fixes a critical issue where GroupByLabelBatchSampler produced ~99% single-class batches, causing zero gradients with triplet losses. The sampler now uses round-robin interleaving where each label emits 2 samples per round, with the label visit order reshuffled every round. This guarantees every batch contains multiple distinct labels, each with at least 2 samples.
This release includes full compatibility updates for Transformers v5:
_nested_gather method (#3664)warmup_steps and warmup_ratio until Transformers v4 support is dropped (#3645)requests dependency with optional httpx dependency by @tomaarsen in #3618MultipleNegativesRankingLoss when num_negatives=None by @fuutot in #3636MultipleNegativesRankingLoss by @fuutot in #3641feat] Support excluding prompt tokens with pooling with left-padding tokenizer by @tomaarsen in https://github.com/huggingface/sentence-transformers/pull/3598tests] Relax the CI branches by @tomaarsen in https://github.com/huggingface/sentence-transformers/pull/3610compat] Expand test suite to full transformers v5 by @tomaarsen in https://github.com/huggingface/sentence-transformers/pull/3615deps] Replace requests dependency with optional httpx dependency by @tomaarsen in https://github.com/huggingface/sentence-transformers/pull/3618feat] Add triplets/n-tuple support to AnglE by @tomaarsen in https://github.com/huggingface/sentence-transformers/pull/3609http_get with load_dataset - wiki1m_for_simcse and STSbenchmark by @omkar-334 in https://github.com/huggingface/sentence-transformers/pull/3635tests] Use 120s HF Hub timeout for tests by @tomaarsen in https://github.com/huggingface/sentence-transformers/pull/3637MultipleNegativesRankingLoss when num_negatives=None by @fuutot in https://github.com/huggingface/sentence-transformers/pull/36362_programming_train_bi-encoder.py from v2 to v3 by @omkar-334 in https://github.com/huggingface/sentence-transformers/pull/3629train_simcse_from_file.py from v2 to v3 by @omkar-334 in https://github.com/huggingface/sentence-transformers/pull/3631ContrastiveTensionLoss and ContrastiveTensionLossInBatchNegatives by @omkar-334 in https://github.com/huggingface/sentence-transformers/pull/3639http_get with load_dataset -askubuntu and all-nli by @omkar-334 in https://github.com/huggingface/sentence-transformers/pull/3638batch_size args to CE evaluators by @omkar-334 in https://github.com/huggingface/sentence-transformers/pull/3643trec dataset and migrate training_batch_hard_trec.py from v2 to v3 by @omkar-334 in https://github.com/huggingface/sentence-transformers/pull/3624train_stsb_ct.py from v2 to v3 by @omkar-334 in https://github.com/huggingface/sentence-transformers/pull/3626train_ct_from_file.py from v2 to v3 by @omkar-334 in https://github.com/huggingface/sentence-transformers/pull/3625train_askubuntu_ct-improved.py from v2 to v3 by @omkar-334 in https://github.com/huggingface/sentence-transformers/pull/3628train_stsb_ct_improved from v2 to v3 by @omkar-334 in https://github.com/huggingface/sentence-transformers/pull/3627train_askubuntu_simcse.py from v2 to v3 by @omkar-334 in https://github.com/huggingface/sentence-transformers/pull/3630train_stsb_simcse.py from v2 to v3 by @omkar-334 in https://github.com/huggingface/sentence-transformers/pull/3648train_askubuntu_ct.py from v2 to v3 by @omkar-334 in https://github.com/huggingface/sentence-transformers/pull/3647train_ct-improved_from_file.py from v2 to v3 by @omkar-334 in https://github.com/huggingface/sentence-transformers/pull/3646DenoisingAutoEncoderLoss.py by @omkar-334 in https://github.com/huggingface/sentence-transformers/pull/3652model.fit in test files by @omkar-334 in https://github.com/huggingface/sentence-transformers/pull/3653feat] Add support for T5Gemma and T5Gemma2 models by @tomaarsen in https://github.com/huggingface/sentence-transformers/pull/3644compat] Allow for both warmup_steps and warmup_ratio until transformers v4 support is dropped by @tomaarsen in https://github.com/huggingface/sentence-transformers/pull/3645feat] Introduce GlobalOrthogonalRegularizationLoss by @tomaarsen in https://github.com/huggingface/sentence-transformers/pull/3654compat] Introduce Transformers v5.2 compatibility: trainer _nested_gather moved by @tomaarsen in https://github.com/huggingface/sentence-transformers/pull/3664perf] Speed up NoDuplicatesBatchSampler iteration (NO_DUPLICATES and NO_DUPLICATES_HASHED) by @hotchpotch in https://github.com/huggingface/sentence-transformers/pull/3658fix] GroupByLabelBatchSampler to guarantee multi-class batches for triplet losses by @MrLoh in https://github.com/huggingface/sentence-transformers/pull/3668feat] Introduce CachedSpladeLoss for memory-efficient SPLADE training by @yjoonjang in https://github.com/huggingface/sentence-transformers/pull/3670docs] Add tips for adjusting batch size to improve processing speed by @tomaarsen in https://github.com/huggingface/sentence-transformers/pull/3672docs] CE trainer: Removed IterableDataset from train and eval dataset type hints by @tomaarsen in https://github.com/huggingface/sentence-transformers/pull/3676loss] Disallow query_to_query/doc_to_doc with partition_mode="per_direction" due to negative loss by @tomaarsen in https://github.com/huggingface/sentence-transformers/pull/3677feat] Add hardness-weighted contrastive learning to losses by @yjoonjang in https://github.com/huggingface/sentence-transformers/pull/3667fix] Fix model card generation with set_transform with new column names by @tomaarsen in https://github.com/huggingface/sentence-transformers/pull/3680tests] Add slow reproduction tests for most common models by @tomaarsen in https://github.com/huggingface/sentence-transformers/pull/3681A big thanks to my repeat contributors, a lot of this release originated from your contributions. Much appreciated!
Full Changelog: https://github.com/huggingface/sentence-transformers/compare/v5.2.3...v5.3.0
One column per quarter.
This patch release introduces compatibility with Transformers v5.2.
This patch release introduces compatibility with Transformers v5.2.
Install this version with
# Training + Inference
pip install sentence-transformers[train]==5.2.3
# Inference only, use one of:
pip install sentence-transformers==5.2.3
pip install sentence-transformers[onnx-gpu]==5.2.3
pip install sentence-transformers[onnx]==5.2.3
pip install sentence-transformers[openvino]==5.2.3
Transformers v5.2 has just released, and it updated its Trainer in such a way that training with Sentence Transformers would start failing on the logging step. The https://github.com/huggingface/sentence-transformers/pull/3664 pull request has resolved this issue.
If you're not training with Sentence Transformers, then older versions of Sentence Transformers are also compatible with Transformers v5.2.
compat] Introduce Transformers v5.2 compatibility: trainer _nested_gather moved by @tomaarsen (https://github.com/huggingface/sentence-transformers/pull/3664)Full Changelog: https://github.com/huggingface/sentence-transformers/compare/v5.2.2...v5.2.3
This patch release introduces compatibility with Transformers v5.2.
Install this version with
# Training + Inference
pip install sentence-transformers[train]==5.2.3
# Inference only, use one of:
pip install sentence-transformers==5.2.3
pip install sentence-transformers[onnx-gpu]==5.2.3
pip install sentence-transformers[onnx]==5.2.3
pip install sentence-transformers[openvino]==5.2.3Transformers v5.2 has just released, and it updated its Trainer in such a way that training with Sentence Transformers would start failing on the logging step. The #3664 pull request has resolved this issue.
If you're not training with Sentence Transformers, then older versions of Sentence Transformers are also compatible with Transformers v5.2.
compat] Introduce Transformers v5.2 compatibility: trainer _nested_gather moved by @tomaarsen (#3664)Full Changelog: v5.2.2...v5.2.3
This patch release replaces mandatory requests dependency with an optional httpx dependency.
This patch release replaces mandatory requests dependency with an optional httpx dependency.
Install this version with
# Training + Inference
pip install sentence-transformers[train]==5.2.2
# Inference only, use one of:
pip install sentence-transformers==5.2.2
pip install sentence-transformers[onnx-gpu]==5.2.2
pip install sentence-transformers[onnx]==5.2.2
pip install sentence-transformers[openvino]==5.2.2
Transformers v5.0 and its required huggingface_hub versions have dropped support of requests in favor of httpx. The former was also used in sentence-transformers, but not listed explicitly as a dependency. This patch removes the use of requests in favor of httpx, although it's now optional and not automatically imported. This should also save some import time.
Importing Sentence Transformers should now not crash if requests is not installed.
deps] Replace requests dependency with optional httpx dependency by @tomaarsen (https://github.com/huggingface/sentence-transformers/pull/3618)Full Changelog: https://github.com/huggingface/sentence-transformers/compare/v5.2.1...v5.2.2
This patch release adds support for the full Transformers v5 release.
This patch release adds support for the full Transformers v5 release.
Install this version with
# Training + Inference
pip install sentence-transformers[train]==5.2.1
# Inference only, use one of:
pip install sentence-transformers==5.2.1
pip install sentence-transformers[onnx-gpu]==5.2.1
pip install sentence-transformers[onnx]==5.2.1
pip install sentence-transformers[openvino]==5.2.1
Sentence Transformers v5.2.0 already introduced support for the Transformers v5.0 release candidates, but this release is adding support for the full release. The intention is to maintain backward compatibility with v4.x. The library includes dual CI testing for both version for now, allowing users to upgrade to the newest Transformers features when ready. In future versions, Sentence Transformers may start requiring Transformers v5.0 or higher.
Full Changelog: https://github.com/huggingface/sentence-transformers/compare/v5.2.0...v5.2.1
…Transformers v5 support, Python 3.9 deprecations, and more.
This minor release introduces multi-processing for CrossEncoder (rerankers), multilingual NanoBEIR evaluators, similarity score outputs in mine_hard_negatives, Transformers v5 support, Python 3.9 deprecations, and more.
Install this version with
# Training + Inference
pip install sentence-transformers[train]==5.2.0
# Inference only, use one of:
pip install sentence-transformers==5.2.0
pip install sentence-transformers[onnx-gpu]==5.2.0
pip install sentence-transformers[onnx]==5.2.0
pip install sentence-transformers[openvino]==5.2.0
The CrossEncoder class now supports multiprocessing for faster inference on CPU and multi-GPU setups. This brings CrossEncoder functionality in line with the existing multiprocessing capabilities of SentenceTransformer models, allowing you to use multiple CPU cores or GPUs to speed up both the predict and rank methods when processing large batches of sentence pairs.
The implementation introduces these new methods, mirroring the SentenceTransformer approach:
start_multi_process_pool() - Initialize a pool of worker processesstop_multi_process_pool() - Clean up the worker poolUsage is straightforward with the new pool parameter:
from sentence_transformers.cross_encoder import CrossEncoder
def main():
model = CrossEncoder('cross-encoder/ms-marco-MiniLM-L6-v2')
# Start a pool of workers
pool = model.start_multi_process_pool()
# Use the pool for faster inference
scores = model.predict(sentence_pairs, pool=pool)
rankings = model.rank(query, documents, pool=pool)
# Clean up when done
model.stop_multi_process_pool(pool)
if __name__ == "__main__":
main()
Or simply pass a list of devices to device to have predict and rank automatically create a pool behind the scenes.
from sentence_transformers.cross_encoder import CrossEncoder
def main():
model = CrossEncoder('cross-encoder/ms-marco-MiniLM-L6-v2', device="cpu")
# Use 4 processes
scores = model.predict(sentence_pairs, device=["cpu"] * 4)
rankings = model.rank(query, documents, device=["cpu"] * 4)
if __name__ == "__main__":
main()
This enhancement is particularly beneficial for CPU-based deployments and enables multi-GPU reranking in the mine_hard_negatives function, making hard negative mining faster for large datasets.
The NanoBEIR evaluators now support custom dataset IDs, allowing for evaluation on non-English NanoBEIR collections. All three NanoBEIR evaluators (dense, sparse, and cross-encoder) support this functionality with a simple dataset_id parameter.
For example:
import logging
from pprint import pprint
from sentence_transformers import SentenceTransformer
from sentence_transformers.evaluation import NanoBEIREvaluator
logging.basicConfig(format="%(asctime)s - %(message)s", datefmt="%Y-%m-%d %H:%M:%S", level=logging.INFO)
# Load a model to evaluate
model = SentenceTransformer("google/embeddinggemma-300m")
# Use a Serbian translation of NanoBEIR
evaluator = NanoBEIREvaluator(
["msmarco", "nq"],
dataset_id="Serbian-AI-Society/NanoBEIR-sr"
)
results = evaluator(model)
print(results[evaluator.primary_metric])
pprint({key: value for key, value in results.items() if "ndcg@10" in key})
"""
{'NanoBEIR_mean_cosine_ndcg@10': 0.44754032737278326,
'NanoMSMARCO_cosine_ndcg@10': 0.4424192627754922,
'NanoNQ_cosine_ndcg@10': 0.45266139197007427}
"""
There are already supported translations for French, Arabic, German, Spanish, Italian, Portuguese, Norwegian, Swedish, Serbian, Korean, Japanese, and 22 Bharat languages in the NanoBEIR collection. Contact me (@tomaarsen) if you have found or created another translation and would like to get it added to the collection!
The mine_hard_negatives function now includes an output_scores parameter that allows you to export similarity scores alongside the mined negatives. When output_scores=False (default), these are the output formats for various output_formats:
And when output_scores=True, the format becomes:
For context, labels are binary options denoting whether the relevant pair was labeled as a positive or not, whereas scores are similarity scores from the SentenceTransformer or CrossEncoder model.
Additionally:
n-tuple-scores format has been replaced with the cleaner output_format="n-tuple" combined with output_scores=True.For example:
from sentence_transformers.util import mine_hard_negatives
from sentence_transformers import SentenceTransformer
from datasets import load_dataset
# Load a Sentence Transformer model
model = SentenceTransformer("sentence-transformers/static-retrieval-mrl-en-v1")
# Load a dataset to mine hard negatives from
dataset = load_dataset("sentence-transformers/natural-questions", split="train").select(range(10000))
print(dataset)
"""
Dataset({
features: ['query', 'answer'],
num_rows: 10000
})
"""
# Mine hard negatives into num_negatives + 3 columns:
# 'query', 'answer', 'negative_1', 'negative_2', ..., 'score'
# where 'score' is a list of similarity scores for the query-answer plus each query-negative pair.
dataset = mine_hard_negatives(
dataset=dataset,
model=model,
num_negatives=5,
sampling_strategy="top",
relative_margin=0.05,
batch_size=128,
use_faiss=True,
output_format="labeled-list",
output_scores=True,
)
"""
Negative candidates mined, preparing dataset...
Metric Positive Negative Difference
Count 10,000 49,241
Mean 0.5884 0.3909 0.2033
Median 0.6005 0.3766 0.1837
Std 0.1467 0.1050 0.1337
Min 0.0272 0.1595 0.0088
25% 0.4918 0.3127 0.0903
50% 0.6005 0.3766 0.1837
75% 0.6974 0.4558 0.2924
Max 0.9679 0.8505 0.7281
Skipped 25,451 potential negatives (4.89%) due to the relative_margin of 0.05.
Could not find enough negatives for 148 samples (1.48%). Consider adjusting the range_max and relative_margin parameters if you'd like to find more valid negatives.
"""
print(dataset)
"""
Dataset({
features: ['query', 'answer', 'scores'],
num_rows: 9852
})
"""
print(dataset[0])
{
"query": "when did richmond last play in a preliminary final",
"answer": [
"Richmond Football Club Richmond began 2017 with 5 straight wins, a feat it had not achieved since 1995. A series of close losses hampered the Tigers throughout the middle of the season, including a 5-point loss to the Western Bulldogs, 2-point loss to Fremantle, and a 3-point loss to the Giants. Richmond ended the season strongly with convincing victories over Fremantle and St Kilda in the final two rounds, elevating the club to 3rd on the ladder. Richmond's first final of the season against the Cats at the MCG attracted a record qualifying final crowd of 95,028; the Tigers won by 51 points. Having advanced to the first preliminary finals for the first time since 2001, Richmond defeated Greater Western Sydney by 36 points in front of a crowd of 94,258 to progress to the Grand Final against Adelaide, their first Grand Final appearance since 1982. The attendance was 100,021, the largest crowd to a grand final since 1986. The Crows led at quarter time and led by as many as 13, but the Tigers took over the game as it progressed and scored seven straight goals at one point. They eventually would win by 48 points – 16.12 (108) to Adelaide's 8.12 (60) – to end their 37-year flag drought.[22] Dustin Martin also became the first player to win a Premiership medal, the Brownlow Medal and the Norm Smith Medal in the same season, while Damien Hardwick was named AFL Coaches Association Coach of the Year. Richmond's jump from 13th to premiers also marked the biggest jump from one AFL season to the next.",
"2017 AFL Grand Final The 2017 AFL Grand Final was an Australian rules football game contested between the Adelaide Crows and the Richmond Tigers, held at the Melbourne Cricket Ground on 30 September 2017. It was the 121st annual grand final of the Australian Football League (formerly the Victorian Football League), staged to determine the premiers for the 2017 AFL season.[1]. Richmond defeated Adelaide by 48 points, marking the club's eleventh premiership and first since 1980. Richmond's Dustin Martin won the Norm Smith Medal as the best player on the ground. The match was attended by 100,021 people, the largest crowd since the 1986 Grand Final.",
"Raid of Richmond The Richmond Campaign was a group of British military actions against the capital of Virginia, Richmond, and the surrounding area, during the American Revolutionary War. Led by American turncoat Benedict Arnold, the Richmond Campaign is considered one of his greatest successes while serving under the British Army, and one of the most notorious actions that Arnold ever performed.",
"2001 AFL Grand Final The 2001 AFL Grand Final was an Australian rules football game contested between the Essendon Football Club and the Brisbane Lions, held at the Melbourne Cricket Ground in Melbourne on 29 September 2001. It was the 105th annual Grand Final of the Australian Football League (formerly the Victorian Football League),[1] staged to determine the premiers for the 2001 AFL season. The match, attended by 91,482 spectators, was won by Brisbane by a margin of 26 points, marking that club's first premiership victory.",
"1964 VFL Grand Final The 1964 VFL Grand Final was an Australian rules football game contested between the Collingwood Football Club and Melbourne Football Club, held at the Melbourne Cricket Ground in Melbourne on 19 September 1964. It was the 68th annual Grand Final of the Victorian Football League, staged to determine the premiers for the 1964 VFL season. The match, attended by 102,471 spectators, was won by Melbourne by a margin of 4 points, marking that club's 12th (and to date, most recent) premiership victory.",
"1998 AFL Grand Final The 1998 AFL Grand Final was an Australian rules football game contested between the Adelaide Crows and the North Melbourne Kangaroos, held at the Melbourne Cricket Ground in Melbourne on 26 September 1998. It was the 102nd annual Grand Final of the Australian Football League (formerly the Victorian Football League), staged to determine the premiers for the 1998 AFL season. The match, attended by 94,431 spectators, was won by Adelaide by a margin of 35 points marking that club's second consecutive premiership victory, and second premiership overall.",
],
"scores": [
0.5460646748542786,
0.5105829238891602,
0.4460095167160034,
0.3221113085746765,
0.3161606788635254,
0.31184709072113037,
],
}
# dataset.push_to_hub("natural-questions-hard-negatives", "labeled-list-scores")
Sentence Transformers now supports the latest Transformers v5.0 release while maintaining backward compatibility with v4.x. The library includes dual CI testing for both version for now, allowing users to upgrade to the newest Transformers features when ready. In future versions, Sentence Transformers may start requiring Transformers v5.0 or higher.
The Pillow library is now an optional dependency rather than a required one, reducing installation size for users who don't work with image-based models. Users who need image functionality can install it via pip install sentence-transformers[image] or directly with pip install pillow.
Following Python's deprecation schedule, Sentence Transformers v5.2.0 has deprecated support for Python 3.9. Users are encouraged to upgrade to Python 3.10 or newer to continue receiving updates and new features.
labels argument in the loss that's used to train (#3506).sentence-transformers[onnx] and sentence-transformers[onnx-gpu] extra's now rely on the new optimum-onnx package with optimum >= 2.0.0.tests] Loosen safetensors test rtol/atol by @tomaarsen in https://github.com/huggingface/sentence-transformers/pull/3572deprecation] Deprecate Python 3.9, upgrade ruff by @tomaarsen in https://github.com/huggingface/sentence-transformers/pull/3573fix]: correct condition for restoring layer embeddings in TransformerDecorator/AdaptiveLayerLoss by @emapco in https://github.com/huggingface/sentence-transformers/pull/3560chore] Rename master to main, update outdated URLs by @tomaarsen in https://github.com/huggingface/sentence-transformers/pull/3579tests] Increase atol/rtol from 1e-6 to 1e-5 for higher test consistency by @tomaarsen in https://github.com/huggingface/sentence-transformers/pull/3578feat] Allow transformers v5.0, add CI for transformers <v5 and >=v5 by @tomaarsen in https://github.com/huggingface/sentence-transformers/pull/3586deps] Use optimum-onnx now that both optimum-onnx and optimum-intel can use optimum==2.0.0 by @tomaarsen in https://github.com/huggingface/sentence-transformers/pull/3587An extra thanks to @Samoed, @NohTow, and @raphaelsty for engaging in valuable discussions in the pull requests, @omkar-334 for finding all kinds of open issues where possible, and @marquesafonso for working on a solid PR for multilingual NanoBEIR that we didn't end up going for.
Additionally, a big thanks to @milistu from Serbian-AI-Society, @NohTow & @raphaelsty from LightOn, @mlabonne and Fernando Fernandes Neto from LiquidAI, @lbourdois from CATIE-AQ and Arun Arumugam for creating the NanoBEIR translations that are supported out of the gate.
Full Changelog: https://github.com/huggingface/sentence-transformers/compare/v5.1.2...v5.2.0
This patch celebrates the transition of Sentence Transformers to Hugging Face, and improves model saving, loading defaults, and loss compatibilities.
This patch celebrates the transition of Sentence Transformers to Hugging Face, and improves model saving, loading defaults, and loss compatibilities.
Install this version with
# Training + Inference
pip install sentence-transformers[train]==5.1.2
# Inference only, use one of:
pip install sentence-transformers==5.1.2
pip install sentence-transformers[onnx-gpu]==5.1.2
pip install sentence-transformers[onnx]==5.1.2
pip install sentence-transformers[openvino]==5.1.2
Today, Sentence Transformers is moving from the Ubiquitous Knowledge Processing (UKP) Lab at Technische Universität Darmstadt to Hugging Face. This formalizes the existing maintenance structure, as Tom Aarsen (that's me!) from Hugging Face has been maintaining the project for the past two years. The project's development roadmap, license, support, and commitment to the community remain unchanged. Read the full announcement for more details!
<img width="1050" height="548" alt="thumbnail" src="https://github.com/user-attachments/assets/95266693-8f8f-4553-9aa3-1b8cc80afb18" />
fix] Patch Router training with Cached losses by @tomaarsen in https://github.com/UKPLab/sentence-transformers/pull/3527fix] Allow loading Dense modules not saved in fp32 by @tomaarsen in https://github.com/UKPLab/sentence-transformers/pull/3528tests] Patch Regex expected output for Python 3.9 by @tomaarsen in https://github.com/UKPLab/sentence-transformers/pull/3529fix] correct dataset link in training scripts by @thomasht86 in https://github.com/UKPLab/sentence-transformers/pull/3543fix] Allow MatryoshkaLoss with MSELoss by @tomaarsen in https://github.com/UKPLab/sentence-transformers/pull/3538docs] Format all markdown using mdformat by @tomaarsen in https://github.com/UKPLab/sentence-transformers/pull/3539Full Changelog: https://github.com/UKPLab/sentence-transformers/compare/v5.1.1...v5.1.2
This patch makes Sentence Transformers more explicit with incorrect arguments and introduces some fixes for multi-GPU processing, evaluators, and hard
This patch makes Sentence Transformers more explicit with incorrect arguments and introduces some fixes for multi-GPU processing, evaluators, and hard negatives mining.
Install this version with
# Training + Inference
pip install sentence-transformers[train]==5.1.1
# Inference only, use one of:
pip install sentence-transformers==5.1.1
pip install sentence-transformers[onnx-gpu]==5.1.1
pip install sentence-transformers[onnx]==5.1.1
pip install sentence-transformers[openvino]==5.1.1
get_model_kwargs (#3500)Some SentenceTransformer or SparseEncoder models support custom model-specific keyword arguments, such as jinaai/jina-embeddings-v4. As of this release, calling model.encode with keyword arguments that aren't used by the model will result in an error.
>>> from sentence_transformers import SentenceTransformer
>>> model = SentenceTransformer("all-MiniLM-L6-v2")
>>> model.encode("Who is Amelia Earhart?", normalize=True)
Traceback (most recent call last):
File "<stdin>", line 1, in <module>
File "[sic]/torch/utils/_contextlib.py", line 116, in decorate_context
return func(*args, **kwargs)
^^^^^^^^^^^^^^^^^^^^^
File "[sic]/SentenceTransformer.py", line 983, in encode
raise ValueError(
ValueError: SentenceTransformer.encode() has been called with additional keyword arguments that this model does not use: ['normalize']. As per SentenceTransformer.get_model_kwargs(), this model does not accept any additional keyword arguments.
Quite useful when you, for example, accidentally forget that the parameter to get normalized embeddings is normalize_embeddings. Prior to this version, this parameter would simply quietly be ignored.
To check which custom extra keyword arguments may be used for your model, you can call the new get_model_kwargs method:
>>> from sentence_transformers import SentenceTransformer, SparseEncoder
>>> SentenceTransformer("all-MiniLM-L6-v2").get_model_kwargs()
[]
>>> SentenceTransformer("jinaai/jina-embeddings-v4", trust_remote_code=True).get_model_kwargs()
['task', 'truncate_dim']
>>> SparseEncoder("opensearch-project/opensearch-neural-sparse-encoding-doc-v3-distill").get_model_kwargs()
['task']
Note: You can always pass the task parameter, it's the only model-specific parameter that will be quietly ignored. This means that you can always use model.encode(..., task="query") and model.encode(..., task="document").
batch_size being ignored in CrossEncoderRerankingEvaluator (#3497)encode, embeddings are now moved from the various devices to the CPU before being stacked into one tensor (#3488)encode_query and encode_document in mine_hard_negatives, automatically using defined "query" and "document" prompts (#3502)output_path that doesn't exist yet (#3516)mine_hard_negatives (#3504)fix] add batch size parameter to model prediction in CrossEncoderRerankingEvaluator by @emapco in https://github.com/UKPLab/sentence-transformers/pull/3497fix] Ensure multi-process embeddings are moved to CPU for concatenation by @tomaarsen in https://github.com/UKPLab/sentence-transformers/pull/3488model_card] Don't override manually provided languages in model card by @tomaarsen in https://github.com/UKPLab/sentence-transformers/pull/3501tests] Add hard negatives test showing multiple positives are correctly handled by @tomaarsen in https://github.com/UKPLab/sentence-transformers/pull/3503feat] Use encode_document and encode_query in mine_hard_negatives by @tomaarsen in https://github.com/UKPLab/sentence-transformers/pull/3502input_ids, attention_mask, token_type_ids, inputs_embeds to forward by @Samoed in https://github.com/UKPLab/sentence-transformers/pull/3509feat] add get_model_kwargs method; throw error if unused kwarg is passed by @tomaarsen in https://github.com/UKPLab/sentence-transformers/pull/3500fix] Fix the number of missing negatives in mine_hard_negatives by @tomaarsen in https://github.com/UKPLab/sentence-transformers/pull/3504Full Changelog: https://github.com/UKPLab/sentence-transformers/compare/v5.1.0...v5.1.1
This release introduces 2 new efficient computing backends for SparseEncoder embedding models: ONNX and OpenVINO + optimization & quantization, allowi
This release introduces 2 new efficient computing backends for SparseEncoder embedding models: ONNX and OpenVINO + optimization & quantization, allowing for speedups up to 2x-3x; a new "n-tuple-score" output format for hard negative mining for distillation; gathering across devices for free lunch on multi-gpu training; trackio support; MTEB documentation; any many small fixes and features.
Install this version with
# Training + Inference
pip install sentence-transformers[train]==5.1.0
# Inference only, use one of:
pip install sentence-transformers==5.1.0
pip install sentence-transformers[onnx-gpu]==5.1.0
pip install sentence-transformers[onnx]==5.1.0
pip install sentence-transformers[openvino]==5.1.0
Introducing a new backend keyword argument to the SparseEncoder initialization, allowing values of "torch" (default), "onnx", and "openvino".
These require installing sentence-transformers with specific extras:
pip install sentence-transformers[onnx-gpu]
# or ONNX for CPU only:
pip install sentence-transformers[onnx]
# or
pip install sentence-transformers[openvino]
It's as simple as:
from sentence_transformers import SparseEncoder
# Load a SparseEncoder model with the ONNX backend
model = SparseEncoder("naver/splade-v3", backend="onnx")
query = "Which planet is known as the Red Planet?"
documents = [
"Venus is often called Earth's twin because of its similar size and proximity.",
"Mars, known for its reddish appearance, is often referred to as the Red Planet.",
"Jupiter, the largest planet in our solar system, has a prominent red spot.",
"Saturn, famous for its rings, is sometimes mistaken for the Red Planet."
]
query_embeddings = model.encode_query(query)
document_embeddings = model.encode_document(documents)
print(query_embeddings.shape, document_embeddings.shape)
# torch.Size([30522]) torch.Size([4, 30522])
similarities = model.similarity(query_embeddings, document_embeddings)
print(similarities)
# tensor([[12.1450, 26.1040, 22.0025, 23.3877]])
decoded_query = model.decode(query_embeddings, top_k=5)
decoded_documents = model.decode(document_embeddings, top_k=5)
print(decoded_query)
# [('red', 3.0222), ('planet', 2.5001), ('planets', 1.9412), ('known', 1.8126), ('nasa', 0.9347)]
print(decoded_documents)
# [
# [('venus', 3.1980), ('twin', 2.7036), ('earth', 2.4310), ('twins', 2.0957), ('planet', 1.9462)],
# [('mars', 3.1443), ('planet', 2.4924), ('red', 2.4514), ('reddish', 2.2234), ('planets', 2.1976)],
# [('jupiter', 2.9604), ('red', 2.5507), ('planet', 2.3774), ('planets', 2.1641), ('spot', 2.1138)],
# [('saturn', 2.9354), ('red', 2.4548), ('planet', 2.3962), ('mistaken', 2.3361), ('cass', 2.2100)]
# ]
If you specify a backend and your model repository or directory contains an ONNX/OpenVINO model file, it will automatically be used! And if your model repository or directory doesn't have one already, an ONNX/OpenVINO model will be automatically exported. Just remember to model.push_to_hub or model.save_pretrained into the same model repository or directory to avoid having to re-export the model every time.
All keyword arguments passed via model_kwargs will be passed on to ORTModelForMaskedLM.from_pretrained or ORTModelForMaskedLM.from_pretrained. The most useful arguments are:
provider: (Only if backend="onnx") ONNX Runtime provider to use for loading the model, e.g. "CPUExecutionProvider" . See https://onnxruntime.ai/docs/execution-providers/ for possible providers. If not specified, the strongest provider (E.g. "CUDAExecutionProvider") will be used.file_name: The name of the ONNX file to load. If not specified, will default to "model.onnx" or otherwise "onnx/model.onnx" for ONNX, and "openvino_model.xml" and "openvino/openvino_model.xml" for OpenVINO. This argument is useful for specifying optimized or quantized models.export: A boolean flag specifying whether the model will be exported. If not provided, export will be set to True if the model repository or directory does not already contain an ONNX or OpenVINO model.We ran benchmarks for CPU and GPU, averaging findings across 3 datasets, and numerous batch sizes. Here are the findings:
<p float="left"> <img src="https://github.com/user-attachments/assets/5e86cf2b-6845-4cc8-a16f-98b5a1fa226d" width="45%"/> <img src="https://github.com/user-attachments/assets/fbf05540-53e5-459f-89d0-9153c3d5f1fa" width="45%"/> </p>
These findings resulted in these recommendations: <img width="825" height="610" alt="image" src="https://github.com/user-attachments/assets/a7c10e24-42f0-430b-a152-2a154f43af61" />
For GPU, you can expect 1.81x speedup with bf16 at no cost, and for CPU you can expect up to ~3x speedup at minimal cost of accuracy in our evaluation. Your mileage with the accuracy hit for quantization may vary, but it seems to remain very small.
Read the Speeding up Inference documentation for more details.
n-tuple-scores output format from mine_hard_negatives (#3430, #3481)The mine_hard_negatives utility function has been extended to support the n-tuple-scores output format, which outputs negatives into num_negatives + 3 columns:
where the 'score' is a list of scores for the query-answer plus each query-negative pair.
from sentence_transformers.util import mine_hard_negatives
from sentence_transformers import SentenceTransformer
from datasets import load_dataset
# Load a Sentence Transformer model
model = SentenceTransformer("all-MiniLM-L6-v2", device="cuda")
# Load a dataset to mine hard negatives from
dataset = load_dataset("sentence-transformers/natural-questions", split="train")
# Mine hard negatives into num_negatives + 3 columns: 'query', 'answer', 'negative_1', 'negative_2', ..., 'score'
# where 'score' is a list of scores for the query-answer plus each query-negative pair.
dataset = mine_hard_negatives(
dataset=dataset,
model=model,
num_negatives=5,
sampling_strategy="top",
batch_size=128,
use_faiss=True,
output_format="n-tuple-scores",
)
print(dataset)
print(dataset[14])
"""
{
'query': 'when did jack and the beanstalk take place',
'answer': "Jack and the Beanstalk According to researchers at the universities in Durham and Lisbon, the story originated more than 5,000 years ago, based on a widespread archaic story form which is now classified by folklorists as ATU 328 The Boy Who Stole Ogre's Treasure.[7]",
'negative_1': 'Jack and the Beanstalk "Jack and the Beanstalk" is an English fairy tale. It appeared as "The Story of Jack Spriggins and the Enchanted Bean" in 1734[1] and as Benjamin Tabart\'s moralised "The History of Jack and the Bean-Stalk" in 1807.[2] Henry Cole, publishing under pen name Felix Summerly popularised the tale in The Home Treasury (1845),[3] and Joseph Jacobs rewrote it in English Fairy Tales (1890).[4] Jacobs\' version is most commonly reprinted today and it is believed to be closer to the oral versions than Tabart\'s because it lacks the moralising.[5]',
'negative_2': 'Jack and the Beanstalk Jack climbs the beanstalk twice more. He learns of other treasures and steals them when the giant sleeps: first a goose that lays golden eggs, then a magic harp that plays by itself. The giant wakes when Jack leaves the house with the harp and chases Jack down the beanstalk. Jack calls to his mother for an axe and before the giant reaches the ground, cuts down the beanstalk, causing the giant to fall to his death.',
'negative_3': 'Jack in the Box Jack in the Box is an American fast-food restaurant chain founded February 21, 1951, by Robert O. Peterson in San Diego, California, where it is headquartered. The chain has 2,200 locations, primarily serving the West Coast of the United States and selected large urban areas in the eastern portion of the US including Texas. Food items include a variety of hamburger and cheeseburger sandwiches along with selections of internationally themed foods such as tacos and egg rolls. The company also operates the Qdoba Mexican Grill chain.[4][5]',
'negative_4': 'Jack in the Box Jack in the Box is an American fast-food restaurant chain founded February 21, 1951, by Robert O. Peterson in San Diego, California, where it is headquartered. The chain has 2,200 locations, primarily serving the West Coast of the United States and selected large urban areas in the eastern portion of the US including Texas and the Charlotte metropolitan area. The company also formerly operated the Qdoba Mexican Grill chain until Apollo Global Management bought the chain in December 2017.[4]',
'negative_5': "Jack Box Jack Box (full name Jack I. Box; or simply known as Jack) is the mascot of American restaurant chain Jack in the Box. In the advertisements, he is the founder, CEO, and ad spokesman for the chain. According to the company's web site, he has the appearance of a typical male, with the exception of his huge spherical white head, blue dot eyes, conical black pointed nose, and a curvilinear red smile. He is most of the time seen wearing his yellow clown cap, and a business suit driving a red Viper convertible.",
'score': [0.7949077486991882, 0.8010389804840088, 0.6466549634933472, 0.5222680568695068, 0.5216285586357117, 0.47328776121139526]
}
"""
This format is directly usable in various distillation losses:
Note that without applying any absolute_margin, relative_margin, max_score, etc., you can mine negatives that actually score better than your positive. With a distillation loss, this is totally fine. It will learn using the (margins between the) scores, so you don't have to worry about false negatives as much as when using e.g. MultipleNegativesRankingLoss.
This release also adds support for 1) n-tuples instead of just triplets and 2) num_negatives + 1 scores where the first score is the query-positive score for MarginMSELoss for CrossEncoder models.
Various loss functions in Sentence Transformers take advantage of so-called "in-batch negatives". With these losses, for each sample in a batch, all data for the other samples will be considered as negatives, because random inputs are likely unrelated to the sample. This pushes them further apart, resulting ideally only in higher similarity scores for inputs that really are similar.
This release introduces a new gather_across_devices parameter for each of these losses. This parameter only works in a multi-GPU setting, and will pull the other samples from other devices into the computation. In short: if you have the following setup:
mini_batch_size=16per_device_train_batch_size=128 in the SentenceTransformersTrainingArgumentsThen each device will have the memory usage corresponding to a batch size of 16, while each sample has 1 positive and 255 negatives (1 hard negative for that sample, 127 other positive values as in-batch negatives and 127 other negative values as in-batch negatives). Your global batch size will be 128 * 8 = 1024, and your learning rate should be set according to that value.
Now, if you use the exact same setup, but with gather_across_devices=True, then your setting is suddenly:
Each device will have the memory usage corresponding to a batch size of 16, while each sample has 1 positive and 2047 negatives (1 hard negative for that sample, 1023 other positive values as in-batch negatives and 1023 other negative values as in-batch negatives). Your global batch size will be 128 * 8 = 1024, and your learning rate should be set according to that value.
The difference is that the in-batch negatives will now pull from other devices too! Because a larger batch size often results in stronger models with in-batch negatives losses, this should give stronger models at almost no overhead.
Here are the results from one of my simple experiments with finetuning mpnet-base on natural-questions with 8 GPUs:
baseline:
gather_across_devices=True:
If your transformers version is high enough, and you have trackio installed (pip install trackio), then Sentence Transformers will also export logs to Trackio. It'll allow you to browse to localhost to track your experiments for free.
<img width="1911" height="922" alt="Schermafbeelding 2025-07-30 182602" src="https://github.com/user-attachments/assets/e3cbf25b-2f6c-4a48-9c1a-2cc84b24fc13" />
If you're interested in evaluating your SentenceTransformer models on common benchmarks, then MTEB is your friend. However, there wasn't yet any documentation to guide you in the right direction. This release, we added some:
prompt (#3444)Router torch initialization, resulted in issues with DataParallel and memory usage (#3454)CrossEncoder.predict is called with an empty list (#3466)docs] Fix link in README for training script name by @tomaarsen in https://github.com/UKPLab/sentence-transformers/pull/3417docs] Fix arxiv link in SpladePooling docs by @tomaarsen in https://github.com/UKPLab/sentence-transformers/pull/3418tests] Reuse models more where possible by @tomaarsen in https://github.com/UKPLab/sentence-transformers/pull/3432model card] Avoid pipe characters that mess up table formatting by @tomaarsen in https://github.com/UKPLab/sentence-transformers/pull/3429feat] Add "n-tuple-scores" output format to mine_hard_negatives function by @tomaarsen in https://github.com/UKPLab/sentence-transformers/pull/3430feat] Avoid unneeded warning when calling encode_query/document with prompt by @tomaarsen in https://github.com/UKPLab/sentence-transformers/pull/3444compat] Fix compatibility issues with datasets v4 by @tomaarsen in https://github.com/UKPLab/sentence-transformers/pull/3445prompts type with documentation by @FremyCompany in https://github.com/UKPLab/sentence-transformers/pull/3427feat] Add gather_across_devices parameter to some contrastive losses by @tomaarsen in https://github.com/UKPLab/sentence-transformers/pull/3442chore] Redistribute util.py (and its tests) to separate directory by @tomaarsen in https://github.com/UKPLab/sentence-transformers/pull/3446tests] Reduce the number of hub requests for the model card tests by @tomaarsen in https://github.com/UKPLab/sentence-transformers/pull/3447fix] cast indexing numpy int to Python int by @emapco in https://github.com/UKPLab/sentence-transformers/pull/3455fix] Fix Router torch initialization, fixes DP by @tomaarsen in https://github.com/UKPLab/sentence-transformers/pull/3454fix] Patch gather_across_devices for in-batch negatives losses by @tomaarsen in https://github.com/UKPLab/sentence-transformers/pull/3453feat] Update the trackio default project if not already defined by @tomaarsen in https://github.com/UKPLab/sentence-transformers/pull/3467docs] Fix dead link in ContrastiveLoss references by @tomaarsen in https://github.com/UKPLab/sentence-transformers/pull/3476docs] Add splade_index semantic search example by @tomaarsen in https://github.com/UKPLab/sentence-transformers/pull/3473feat] Add ONNX, OV support for SparseEncoder; refactor ONNX/OV by @tomaarsen in https://github.com/UKPLab/sentence-transformers/pull/3475fix] FIPS compatibility - use SHA256 with usedforsecurity=False in hard negatives caching by @tomaarsen in https://github.com/UKPLab/sentence-transformers/pull/3479feat] Allow n-tuples for CE MarginMSE training by @tomaarsen in https://github.com/UKPLab/sentence-transformers/pull/3481docs] Update main sbert.net page with v5.1 mention by @tomaarsen in https://github.com/UKPLab/sentence-transformers/pull/3482Also thanks to @Samoed and @KennethEnevoldsen for their reviews on the MTEB documentation, and thanks to @NohTow for the inspiration on gathering across devices.
Full Changelog: https://github.com/UKPLab/sentence-transformers/compare/v5.0.0...v5.1.0
This release consists of significant updates including the introduction of Sparse Encoder models, new methods encode_query and encode_document, multi-
This release consists of significant updates including the introduction of Sparse Encoder models, new methods encode_query and encode_document, multi-processing support in encode, the Router module for asymmetric models, custom learning rates for parameter groups, composite loss logging, and various small improvements and bug fixes.
Install this version with
# Training + Inference
pip install sentence-transformers[train]==5.0.0
# Inference only, use one of:
pip install sentence-transformers==5.0.0
pip install sentence-transformers[onnx-gpu]==5.0.0
pip install sentence-transformers[onnx]==5.0.0
pip install sentence-transformers[openvino]==5.0.0
[!TIP] Our Training and Finetuning Sparse Embedding Models with Sentence Transformers v5 blogpost is an excellent place to learn about finetuning sparse embedding models!
[!NOTE] This release is designed to be fully backwards compatible, meaning that you should be able to upgrade from older versions to v5.x without any issues. If you are running into issues when upgrading, feel free to open an issue. Also see the Migration Guide for changes that we would recommend.
The Sentence Transformers v5.0 release introduces Sparse Embedding models, also known as Sparse Encoders. These models generate high-dimensional embeddings, often with 30,000+ dimensions, where often only <1% of dimensions are non-zero. This is in contrast to the standard dense embedding models, which produce low-dimensional embeddings (e.g., 384, 768, or 1024 dimensions) where all values are non-zero.
Usually, each active dimension (i.e. the dimension with a non-zero value) in a sparse embedding corresponds to a specific token in the model's vocabulary, allowing for interpretability. This means that you can e.g. see exactly which words/tokens are important in an embedding, and that you can inspect exactly because of which words/tokens two texts are deemed similar.
Let's have a look at naver/splade-v3, a strong sparse embedding model, as an example:
from sentence_transformers import SparseEncoder
# Download from the 🤗 Hub
model = SparseEncoder("naver/splade-v3")
# Run inference
sentences = [
"The weather is lovely today.",
"It's so sunny outside!",
"He drove to the stadium.",
]
embeddings = model.encode(sentences)
print(embeddings.shape)
# (3, 30522)
# Get the similarity scores for the embeddings
similarities = model.similarity(embeddings, embeddings)
print(similarities)
# tensor([[ 32.4323, 5.8528, 0.0258],
# [ 5.8528, 26.6649, 0.0302],
# [ 0.0258, 0.0302, 24.0839]])
# Let's decode our embeddings to be able to interpret them
decoded = model.decode(embeddings, top_k=10)
for decoded, sentence in zip(decoded, sentences):
print(f"Sentence: {sentence}")
print(f"Decoded: {decoded}")
print()
Sentence: The weather is lovely today.
Decoded: [('weather', 2.754288673400879), ('today', 2.610959529876709), ('lovely', 2.431990623474121), ('currently', 1.5520408153533936), ('beautiful', 1.5046082735061646), ('cool', 1.4664798974990845), ('pretty', 0.8986214995384216), ('yesterday', 0.8603134155273438), ('nice', 0.8322536945343018), ('summer', 0.7702118158340454)]
Sentence: It's so sunny outside!
Decoded: [('outside', 2.6939032077789307), ('sunny', 2.535827398300171), ('so', 2.0600898265838623), ('out', 1.5397940874099731), ('weather', 1.1198079586029053), ('very', 0.9873268604278564), ('cool', 0.9406591057777405), ('it', 0.9026399254798889), ('summer', 0.684999406337738), ('sun', 0.6520509123802185)]
Sentence: He drove to the stadium.
Decoded: [('stadium', 2.7872302532196045), ('drove', 1.8208855390548706), ('driving', 1.6665740013122559), ('drive', 1.5565159320831299), ('he', 1.4721972942352295), ('stadiums', 1.449463129043579), ('to', 1.0441515445709229), ('car', 0.7002660632133484), ('visit', 0.5118278861045837), ('football', 0.502326250076294)]
In this example, the embeddings are 30,522-dimensional vectors, where each dimension corresponds to a token in the model's vocabulary. The decode method returned the top 10 tokens with the highest values in the embedding, allowing us to interpret which tokens contribute most to the embedding.
We can even determine the intersection or overlap between embeddings, very useful for determining why two texts are deemed similar or dissimilar:
# Let's also compute the intersection/overlap of the first two embeddings
intersection_embedding = model.intersection(embeddings[0], embeddings[1])
decoded_intersection = model.decode(intersection_embedding)
print(decoded_intersection)
Decoded: [('weather', 3.0842742919921875), ('cool', 1.379457712173462), ('summer', 0.5275946259498596), ('comfort', 0.3239051103591919), ('sally', 0.22571465373039246), ('julian', 0.14787325263023376), ('nature', 0.08582140505313873), ('beauty', 0.0588383711874485), ('mood', 0.018594780936837196), ('nathan', 0.000752730411477387)]
And if we think the embeddings are too big, we can limit the maximum number of active dimensions like so:
from sentence_transformers import SparseEncoder
# Download from the 🤗 Hub
model = SparseEncoder("naver/splade-v3") # You can also set max_active_dims here instead of encode()
# Run inference
documents = [
"UV-A light, specifically, is what mainly causes tanning, skin aging, and cataracts, UV-B causes sunburn, skin aging and skin cancer, and UV-C is the strongest, and therefore most effective at killing microorganisms. Again â\x80\x93 single words and multiple bullets.",
"Answers from Ronald Petersen, M.D. Yes, Alzheimer's disease usually worsens slowly. But its speed of progression varies, depending on a person's genetic makeup, environmental factors, age at diagnosis and other medical conditions. Still, anyone diagnosed with Alzheimer's whose symptoms seem to be progressing quickly â\x80\x94 or who experiences a sudden decline â\x80\x94 should see his or her doctor.",
"Bell's palsy and Extreme tiredness and Extreme fatigue (2 causes) Bell's palsy and Extreme tiredness and Hepatitis (2 causes) Bell's palsy and Extreme tiredness and Liver pain (2 causes) Bell's palsy and Extreme tiredness and Lymph node swelling in children (2 causes)",
]
embeddings = model.encode_document(documents, max_active_dims=64)
print(embeddings.shape)
# (3, 30522)
# Print the sparsity of the embeddings
sparsity = model.sparsity(embeddings)
print(sparsity)
# {'active_dims': 64.0, 'sparsity_ratio': 0.9979031518249132}
<details><summary>Click to see that it has minimal impact on scores</summary>
from sentence_transformers import SparseEncoder
# Download from the 🤗 Hub
model = SparseEncoder("naver/splade-v3") # You can also set max_active_dims here instead of encode()
# Run inference
queries = ["what causes aging fast"]
documents = [
"UV-A light, specifically, is what mainly causes tanning, skin aging, and cataracts, UV-B causes sunburn, skin aging and skin cancer, and UV-C is the strongest, and therefore most effective at killing microorganisms. Again â\x80\x93 single words and multiple bullets.",
"Answers from Ronald Petersen, M.D. Yes, Alzheimer's disease usually worsens slowly. But its speed of progression varies, depending on a person's genetic makeup, environmental factors, age at diagnosis and other medical conditions. Still, anyone diagnosed with Alzheimer's whose symptoms seem to be progressing quickly â\x80\x94 or who experiences a sudden decline â\x80\x94 should see his or her doctor.",
"Bell's palsy and Extreme tiredness and Extreme fatigue (2 causes) Bell's palsy and Extreme tiredness and Hepatitis (2 causes) Bell's palsy and Extreme tiredness and Liver pain (2 causes) Bell's palsy and Extreme tiredness and Lymph node swelling in children (2 causes)",
]
query_embeddings = model.encode_query(queries)
document_embeddings = model.encode_document(documents)
# Determine the sparsity
query_sparsity = model.sparsity(query_embeddings)
document_sparsity = model.sparsity(document_embeddings)
print(query_sparsity, document_sparsity)
# {'active_dims': 28.0, 'sparsity_ratio': 0.9990826289233995} {'active_dims': 174.6666717529297, 'sparsity_ratio': 0.9942773516888497}
# Calculate the similarity scores for the embeddings
similarities = model.similarity(query_embeddings, document_embeddings)
print(similarities)
# tensor([[11.3767, 10.8296, 4.3457]], device='cuda:0')
# Again with smaller max_active_dims
smaller_document_embeddings = model.encode_document(documents, max_active_dims=64)
# Determine the sparsity for the smaller document embeddings
smaller_document_sparsity = model.sparsity(smaller_document_embeddings)
print(query_sparsity, smaller_document_sparsity)
# {'active_dims': 28.0, 'sparsity_ratio': 0.9990826289233995} {'active_dims': 64.0, 'sparsity_ratio': 0.9979031518249132}
# Print the similarity scores for the smaller document embeddings
smaller_similarities = model.similarity(query_embeddings, smaller_document_embeddings)
print(smaller_similarities)
# tensor([[10.1311, 9.8360, 4.3457]], device='cuda:0')
# Very similar to the scores for the full document embeddings!
</details>
A big question is: How do sparse embedding models stack up against the “standard” dense embedding models, and what kind of performance can you expect when combining various?
For this, I ran a variation of our hybrid_search.py evaluation script, with:
Which resulted in this evaluation:
| Dense | Sparse | Reranker | NDCG@10 | MRR@10 | MAP |
|---|---|---|---|---|---|
| x | 65.33 | 57.56 | 57.97 | ||
| x | 67.34 | 59.59 | 59.98 | ||
| x | x | 72.39 | 66.99 | 67.59 | |
| x | x | 68.37 | 62.76 | 63.56 | |
| x | x | 69.02 | 63.66 | 64.44 | |
| x | x | x | 68.28 | 62.66 | 63.44 |
Here, the sparse embedding model actually already outperforms the dense one, but the real magic happens when combining the two: hybrid search. In our case, we used Reciprocal Rank Fusion to merge the two rankings.
Rerankers also help improve the performance of the dense or sparse model here, but hurt the performance of the hybrid search, as its performance is already beyond what the reranker can achieve.
[!NOTE] The naver/splade-v3-doc was trained on the MS MARCO training set, so this is in-domain performance, much like what you might expect if you finetune on your own data.
Check out the following links to get a better feel for what Sparse Encoders are, how they work, what architectures exist, how to use them, what pretrained models exist, how to finetune them, and more:
The introduction of SparseEncoder has been one of the largest updates to Sentence Transformers, introducing all of the following:
encode_query and encode_documentSentence Transformers v5.0 introduces two new core methods to the SentenceTransformer and SparseEncoder classes: encode_query and encode_document.
These methods are specialized versions of encode that differ in exactly two ways:
prompt_name or prompt is provided, it uses a predefined “query”/“document” prompt,
if available in the model’s prompts dictionary (example).task to “query”/“document”. If the model has a Router
module, it will use the “query”/“document” task type to route the input through the appropriate submodules.In short, if you use encode_query and encode_document, you can be sure that you're using the model's predefined prompts and use the correct route (if the model has multiple routes).
If you are unsure whether you should use encode, encode_query, or encode_documen),
your best bet is to use encode_query and encode_document for Information Retrieval tasks
with clear query and document/passage distinction, and use encode for all other tasks.
Note that encode is the most general method and can be used for any task, including Information
Retrieval, and that if the model was not trained with predefined prompts and/or task types, then all three methods will return identical embeddings.
See for example this snippet, which automatically uses the “query” prompt stored in the Qwen3-Embedding-0.6B model config.
from sentence_transformers import SentenceTransformer
# Load the model
model = SentenceTransformer("Qwen/Qwen3-Embedding-0.6B")
# The queries and documents to embed
queries = [
"What is the capital of China?",
"Explain gravity",
]
documents = [
"The capital of China is Beijing.",
"Gravity is a force that attracts two bodies towards each other. It gives weight to physical objects and is responsible for the movement of planets around the sun.",
]
# Encode the queries and documents
query_embeddings = model.encode_query(queries) # Equavalent to model.encode(queries, prompt_name="query")
document_embeddings = model.encode_document(documents)
# Compute the (cosine) similarity between the query and document embeddings
similarity = model.similarity(query_embeddings, document_embeddings)
print(similarity)
# tensor([[0.7646, 0.1414],
# [0.1355, 0.6000]])
encode_multi_process absorbed by encodeThe encode method (and by extension the encode_query and encode_document methods) can now be used directly for multi-processing/multi-GPU processing, instead of having to use encode_multi_process.
Previously, you had to manually start a multi-processing pool, use encode_multi_process, and stop the pool:
from sentence_transformers import SentenceTransformer
def main():
model = SentenceTransformer("all-mpnet-base-v2")
texts = ["The weather is so nice!", "It's so sunny outside.", ...]
pool = model.start_multi_process_pool(["cpu", "cpu", "cpu", "cpu"])
embeddings = model.encode_multi_process(texts, pool, chunk_size=512)
model.stop_multi_process_pool(pool)
print(embeddings.shape)
# => (4000, 768)
if __name__ == "__main__":
main()
Now you can just pass a list of devices as device to encode:
from sentence_transformers import SentenceTransformer
def main():
model = SentenceTransformer("all-mpnet-base-v2")
texts = ["The weather is so nice!", "It's so sunny outside.", ...]
embeddings = model.encode(texts, device=["cpu", "cpu", "cpu", "cpu"], chunk_size=512)
print(embeddings.shape)
# => (4000, 768)
if __name__ == "__main__":
main()
The multi-processing can be configured using these parameters:
device: If a list of devices, start multi-processing using those devices. Can be e.g. cpu, but also different GPUs.
pool: You can still use start_multi_process_pool and stop_multi_process_pool to create and stop a multi-processing pool, allowing you to reuse the pool across multiple encode calls via the pool arguments.
chunk_size: When you use multi-processing with n devices, then the inputs will be subdivided into chunks, and those chunks will be spread across the n processes. The size of the chunk can be defined here, although it’s optional. It can have a minor impact on processing speed and memory usage, but is much less important than the batch_size argument.
Documentation: Migration Guide
Documentation: SentenceTransformer.encode
The Sentence Transformers v5.0 release has refactored the Asym module into the Router module. The previous implementation wasn’t straightforward to use with the other components of the library. We’ve improved heavily on this to make the integration seamless. This module allows you to create asymmetric models that apply different modules depending on the specified route (often “query” or “document”).
Notably, you can use the task argument in model.encode to specify which route to use, and the model.encode_query and model.encode_document convenience methods automatically specify task="query" and task="document", respectively.
See for example opensearch-project/opensearch-neural-sparse-encoding-doc-v2-distill for an example of a model using a Router to specify different modules for queries vs documents. Its router_config.json specifies that the query route uses an efficient SparseStaticEmbedding module, while the document route uses the more expensive standard SPLADE modules: MLMTransformer with SpladePooling.
Usage is very straight-forward with the new encode_query and encode_document methods:
from sentence_transformers import SparseEncoder
# Download from the 🤗 Hub
model = SparseEncoder("opensearch-project/opensearch-neural-sparse-encoding-doc-v2-distill")
print(model)
# SparseEncoder(
# (0): Router(
# (query_0_SparseStaticEmbedding): SparseStaticEmbedding({'frozen': True}, dim=30522, tokenizer=DistilBertTokenizerFast)
# (document_0_MLMTransformer): MLMTransformer({'max_seq_length': 512, 'do_lower_case': False, 'architecture': 'DistilBertForMaskedLM'})
# (document_1_SpladePooling): SpladePooling({'pooling_strategy': 'max', 'activation_function': 'relu', 'word_embedding_dimension': 30522})
# )
# )
# Run inference
queries = ["what causes aging fast"]
documents = [
"UV-A light, specifically, is what mainly causes tanning, skin aging, and cataracts, UV-B causes sunburn, skin aging and skin cancer, and UV-C is the strongest, and therefore most effective at killing microorganisms. Again â\x80\x93 single words and multiple bullets.",
"Answers from Ronald Petersen, M.D. Yes, Alzheimer's disease usually worsens slowly. But its speed of progression varies, depending on a person's genetic makeup, environmental factors, age at diagnosis and other medical conditions. Still, anyone diagnosed with Alzheimer's whose symptoms seem to be progressing quickly â\x80\x94 or who experiences a sudden decline â\x80\x94 should see his or her doctor.",
"Bell's palsy and Extreme tiredness and Extreme fatigue (2 causes) Bell's palsy and Extreme tiredness and Hepatitis (2 causes) Bell's palsy and Extreme tiredness and Liver pain (2 causes) Bell's palsy and Extreme tiredness and Lymph node swelling in children (2 causes)",
]
query_embeddings = model.encode_query(queries)
document_embeddings = model.encode_document(documents)
print(query_embeddings.shape, document_embeddings.shape)
# [1, 30522] [3, 30522]
# Get the similarity scores for the embeddings
similarities = model.similarity(query_embeddings, document_embeddings)
print(similarities)
# tensor([[12.0820, 6.5648, 5.0988]])
Note that if you wish to train a model with a Router, then you must specify the router_mapping training arguments that maps dataset column names to Router routes. Then the Trainer knows which route to use for each dataset column.
Note also that any models using Asym still work as before.
Alongside introducing some new modules and refactoring the Asym module into the Router module, we also introduced two new "superclass" modules: Module and InputModule. The former is the new base class of all modules, with the latter as the base class of all modules that are also responsible for tokenization (i.e. for processing inputs).
The documentation describes which methods still need to be implemented when you subclass one of these, and also which convenience methods are available for you to use already. It should certainly simplify the creation of custom modules.
With the introduction of the Router module, it’s becoming much simpler to train a “two-tower model” where the query and document encoders differ a lot. For example, a regular Sentence Transformer for the document encoder, and a Static Embedding model for the query encoder.
In such settings, it’s worthwhile to set different learning rates for different parts of the model. Because of this, v5.0 adds a learning_rate_mapping parameter to the Training Arguments classes. This mapping consists of parameter name regular expressions to learning rates, e.g.
args = SentenceTransformerTrainingArguments(
...,
learning_rate=2e-5,
learning_rate_mapping={"StaticEmbedding.*": 1e-3},
)
Using these training arguments, the learning rate for every parameter whose name matches the regular expression is 1e-3, while all other parameters have a learning rate of 2e-5. Note that we use re.search for determining whether a parameter matches the regular expression, not match or fullmatch.
Many models are trained with just one loss, or perhaps one loss for each dataset. In those cases, all of the losses are nicely logged in both the terminal and third party logging tools (e.g. Weights & Biases, Tensorboard, etc.).
But if you’re using one loss that has multiple components, e.g. a SpladeLoss which sums the losses from FlopsLoss and a SparseMultipleNegativesRankingLoss behind the scenes, then you’re often left guessing whether the various loss components are balanced or not: perhaps one of the two is responsible for 90% of the total loss?
As of the v5.0 release, your loss classes can output dictionaries of loss components. The Trainer will sum them and train like normal, but each of the components will also be logged individually! In short, you can see the various loss components in addition to the final loss itself in your logs.
class SpladeLoss(nn.Module):
...
def forward(
self, sentence_features: Iterable[dict[str, torch.Tensor]], labels: torch.Tensor | None = None
) -> dict[str, torch.Tensor]:
# Compute embeddings using the model
embeddings = [self.model(sentence_feature)["sentence_embedding"] for sentence_feature in sentence_features]
...
return {
"base_loss": base_loss,
"document_regularizer_loss": corpus_loss * self.document_regularizer_weight,
"query_regularizer_loss": query_loss * self.query_regularizer_weight,
}
sif_coefficient, token_remove_pattern, and quantize_to parameters from Model2Vec to StaticEmbedding.from_distillation(...) (#3349)mine_hard_negatives (#3338)mine_hard_negatives (#3334)truncate_dim to encode (and encode_query, encode_document) instead of exclusively being able to set the truncate_dim when initializing the SentenceTransformer.transformers model with model.transformers_model, works for SentenceTransformer, CrossEncoder, and SparseEncoder.See our Migration Guide for more details on the changes, as well as the documentation as a whole.
docs] Point to v4.1 new docs pages in index.html by @tomaarsen in https://github.com/UKPLab/sentence-transformers/pull/3328ci] Attempt to avoid 429 Client Error in CI by @tomaarsen in https://github.com/UKPLab/sentence-transformers/pull/3342fix, cross-encoder] Propagate the gradient checkpointing to the transformer model by @tomaarsen in https://github.com/UKPLab/sentence-transformers/pull/3331tests] Update test based on M2V version by @tomaarsen in https://github.com/UKPLab/sentence-transformers/pull/3354docs] Add two useful recommendations to the docs by @tomaarsen in https://github.com/UKPLab/sentence-transformers/pull/3353refactor] Refactor module loading; introduce Module subclass by @tomaarsen in https://github.com/UKPLab/sentence-transformers/pull/3345tests] Improve robustness of model shape assertion in model2vec test by @tomaarsen in https://github.com/UKPLab/sentence-transformers/pull/3391fix] Use transformers Peft integration instead of manual get_peft_model call by @tomaarsen in https://github.com/UKPLab/sentence-transformers/pull/3405v5] Add support for Sparse Embedding models by @arthurbr11 in https://github.com/UKPLab/sentence-transformers/pull/3401docs] Fix formatting of docstring arguments in SpladeRegularizerWeightSchedulerCallback by @tomaarsen in https://github.com/UKPLab/sentence-transformers/pull/3408fix] Update .gitignore by @arthurbr11 in https://github.com/UKPLab/sentence-transformers/pull/3409fix] Remove hub_kwargs in SparseStaticEmbedding.from_json in favor of more explicit kwargs by @tomaarsen in https://github.com/UKPLab/sentence-transformers/pull/3407docs] Update collections links by @arthurbr11 in https://github.com/UKPLab/sentence-transformers/pull/3410I especially want to thank the following teams and individuals for their contributions to this release, small and large, in no particular order:
Apologies if I forgot anyone. And finally a big thanks to Arthur Bresnu, who led a lot of the work on this release. I wouldn't have been able to introduce Sparse Encoders in this fashion, in this timeline, without his excellent work.
Full Changelog: https://github.com/UKPLab/sentence-transformers/compare/v4.1.0...v5.0.0
This PR softly deprecates the margin option in `mine_hard_negatives` in favor of absolute_margin and relative_margin. In short:
This release introduces 2 new efficient computing backends for CrossEncoder (reranker) models: ONNX and OpenVINO + optimization & quantization, allowing for speedups up to 2x-3x; improved hard negatives mining strategies, and minor improvements.
Install this version with
# Training + Inference
pip install sentence-transformers[train]==4.1.0
# Inference only, use one of:
pip install sentence-transformers==4.1.0
pip install sentence-transformers[onnx-gpu]==4.1.0
pip install sentence-transformers[onnx]==4.1.0
pip install sentence-transformers[openvino]==4.1.0
Introducing a new backend keyword argument to the CrossEncoder initialization, allowing values of "torch" (default), "onnx", and "openvino".
These require installing sentence-transformers with specific extras:
pip install sentence-transformers[onnx-gpu]
# or ONNX for CPU only:
pip install sentence-transformers[onnx]
# or
pip install sentence-transformers[openvino]
It's as simple as:
from sentence_transformers import CrossEncoder
model = CrossEncoder("cross-encoder/ms-marco-MiniLM-L6-v2", backend="onnx")
query = "Which planet is known as the Red Planet?"
passages = [
"Venus is often called Earth's twin because of its similar size and proximity.",
"Mars, known for its reddish appearance, is often referred to as the Red Planet.",
"Jupiter, the largest planet in our solar system, has a prominent red spot.",
"Saturn, famous for its rings, is sometimes mistaken for the Red Planet."
]
scores = model.predict([(query, passage) for passage in passages])
print(scores)
If you specify a backend and your model repository or directory contains an ONNX/OpenVINO model file, it will automatically be used! And if your model repository or directory doesn't have one already, an ONNX/OpenVINO model will be automatically exported. Just remember to model.push_to_hub or model.save_pretrained into the same model repository or directory to avoid having to re-export the model every time.
All keyword arguments passed via model_kwargs will be passed on to ORTModelForSequenceClassification.from_pretrained or OVModelForSequenceClassification.from_pretrained. The most useful arguments are:
provider: (Only if backend="onnx") ONNX Runtime provider to use for loading the model, e.g. "CPUExecutionProvider" . See https://onnxruntime.ai/docs/execution-providers/ for possible providers. If not specified, the strongest provider (E.g. "CUDAExecutionProvider") will be used.file_name: The name of the ONNX file to load. If not specified, will default to "model.onnx" or otherwise "onnx/model.onnx" for ONNX, and "openvino_model.xml" and "openvino/openvino_model.xml" for OpenVINO. This argument is useful for specifying optimized or quantized models.export: A boolean flag specifying whether the model will be exported. If not provided, export will be set to True if the model repository or directory does not already contain an ONNX or OpenVINO model.For example:
from sentence_transformers import CrossEncoder
model = CrossEncoder(
"cross-encoder/ms-marco-MiniLM-L6-v2",
backend="onnx",
model_kwargs={
"file_name": "model_O3.onnx",
"provider": "CPUExecutionProvider",
}
)
query = "Which planet is known as the Red Planet?"
passages = [
"Venus is often called Earth's twin because of its similar size and proximity.",
"Mars, known for its reddish appearance, is often referred to as the Red Planet.",
"Jupiter, the largest planet in our solar system, has a prominent red spot.",
"Saturn, famous for its rings, is sometimes mistaken for the Red Planet."
]
scores = model.predict([(query, passage) for passage in passages])
print(scores)
We ran benchmarks for CPU and GPU, averaging findings across 4 models of various sizes, 3 datasets, and numerous batch sizes. Here are the findings:
<p float="left"> <img src="https://github.com/user-attachments/assets/eda6bb28-e997-4f95-993d-a0d1783c6c37" width="45%" /> <img src="https://github.com/user-attachments/assets/c27e6933-9af6-4d3c-aa4b-66b44956ff85" width="45%" /> </p>
These findings resulted in these recommendations:
For GPU, you can expect 1.88x speedup with fp16 at no cost, and for CPU you can expect ~3x speedup at no cost of accuracy in our evaluation. Your mileage with the accuracy hit for quantization may vary, but it seems to remain very small.
Read the Speeding up Inference documentation for more details.
<details><summary>ONNX & OpenVINO Optimization and Quantization</summary>
In addition to exporting default ONNX and OpenVINO models, you can also use one of the helper methods for optimizing and quantizing ONNX models:
export_optimized_onnx_model: This function uses Optimum to implement several optimizations in the ONNX model, ranging from basic optimizations to approximations and mixed precision. Read about the 4 default options here. This function accepts:
model A SentenceTransformer or CrossEncoder model loaded with backend="onnx".optimization_config: "O1", "O2", "O3", or "O4" from 🤗 Optimum or a custom OptimizationConfig instance.model_name_or_path: The directory or model repository where the optimized model will be saved.push_to_hub: Whether the push the exported model to the hub with model_name_or_path as the repository name. If False, the model will be saved in the directory specified with model_name_or_path.create_pr: If push_to_hub, then this denotes whether a pull request is created rather than pushing the model directly to the repository. Very useful for optimizing models of repositories that you don't have write access to.file_suffix: The suffix to add to the optimized model file name. Will use the optimization_config string or "optimized" if not set.The usage is like this:
from sentence_transformers import SentenceTransformer, export_optimized_onnx_model
onnx_model = CrossEncoder("cross-encoder/ms-marco-MiniLM-L6-v2", backend="onnx")
export_optimized_onnx_model(
model=onnx_model,
optimization_config="O4",
model_name_or_path="cross-encoder/ms-marco-MiniLM-L6-v2",
push_to_hub=True,
create_pr=True,
)
After which you can load the model with:
from sentence_transformers import CrossEncoder
pull_request_nr = 2 # TODO: Update this to the number of your pull request
model = CrossEncoder(
"cross-encoder/ms-marco-MiniLM-L6-v2",
backend="onnx",
model_kwargs={"file_name": "onnx/model_O4.onnx"},
revision=f"refs/pr/{pull_request_nr}"
)
or when it gets merged:
from sentence_transformers import CrossEncoder
model = CrossEncoder(
"cross-encoder/ms-marco-MiniLM-L6-v2",
backend="onnx",
model_kwargs={"file_name": "onnx/model_O4.onnx"},
)
export_dynamic_quantized_onnx_model: This function uses Optimum to quantize the ONNX model to int8, also allowing for hardware-specific optimizations. This results in impressive speedups for CPUs. In my findings, each of the default quantization configuration options gave approximately the same performance improvements. This function accepts
model A SentenceTransformer or CrossEncoder model loaded with backend="onnx".quantization_config: "arm64", "avx2", "avx512", or "avx512_vnni" representing quantization configurations from AutoQuantizationConfig, or an QuantizationConfig instance.model_name_or_path: The directory or model repository where the optimized model will be saved.push_to_hub: Whether the push the exported model to the hub with model_name_or_path as the repository name. If False, the model will be saved in the directory specified with model_name_or_path.create_pr: If push_to_hub, then this denotes whether a pull request is created rather than pushing the model directly to the repository. Very useful for quantizing models of repositories that you don't have write access to.file_suffix: The suffix to add to the optimized model file name. Will use the quantization_config string or e.g. "int8_quantized" if not set.The usage is like this:
from sentence_transformers import CrossEncoder, export_dynamic_quantized_onnx_model
model = CrossEncoder("cross-encoder/ms-marco-MiniLM-L6-v2", backend="onnx")
export_dynamic_quantized_onnx_model(
model,
"avx512_vnni",
"sentence-transformers/cross-encoder/ms-marco-MiniLM-L6-v2",
push_to_hub=True,
create_pr=True,
)
After which you can load the model with:
from sentence_transformers import CrossEncoder
pull_request_nr = 2 # TODO: Update this to the number of your pull request
model = CrossEncoder(
"cross-encoder/ms-marco-MiniLM-L6-v2",
backend="onnx",
model_kwargs={"file_name": "onnx/model_qint8_avx512_vnni.onnx"},
revision=f"refs/pr/{pull_request_nr}",
)
or when it gets merged:
from sentence_transformers import CrossEncoder
model = CrossEncoder(
"cross-encoder/ms-marco-MiniLM-L6-v2",
backend="onnx",
model_kwargs={"file_name": "onnx/model_qint8_avx512_vnni.onnx"},
)
OpenVINO models can be quantized to int8 precision using Optimum Intel to speed up inference. To do this, you can use the export_static_quantized_openvino_model() function, which saves the quantized model in a directory or model repository that you specify. Post-Training Static Quantization expects:
model: a Sentence Transformer or Cross Encoder model loaded with the OpenVINO backend.quantization_config: (Optional) The quantization configuration. This parameter accepts either: None for the default 8-bit quantization, a dictionary representing quantization configurations, or an OVQuantizationConfig instance.model_name_or_path: a path to save the quantized model file, or the repository name if you want to push it to the Hugging Face Hub.dataset_name: (Optional) The name of the dataset to load for calibration. If not specified, defaults to sst2 subset from the glue dataset.dataset_config_name: (Optional) The specific configuration of the dataset to load.dataset_split: (Optional) The split of the dataset to load (e.g., ‘train’, ‘test’).column_name: (Optional) The column name in the dataset to use for calibration.push_to_hub: (Optional) a boolean to push the quantized model to the Hugging Face Hub.create_pr: (Optional) a boolean to create a pull request when pushing to the Hugging Face Hub. Useful when you don’t have write access to the repository.file_suffix: (Optional) a string to append to the model name when saving it. If not specified, "qint8_quantized" will be used.The usage is like this:
from sentence_transformers import CrossEncoder, export_static_quantized_openvino_model
model = CrossEncoder("cross-encoder/ms-marco-MiniLM-L6-v2", backend="openvino")
export_static_quantized_openvino_model(
model,
quantization_config=None,
model_name_or_path="cross-encoder/ms-marco-MiniLM-L6-v2",
push_to_hub=True,
create_pr=True,
)
After which you can load the model with:
from sentence_transformers import CrossEncoder
pull_request_nr = 2 # TODO: Update this to the number of your pull request
model = CrossEncoder(
"cross-encoder/ms-marco-MiniLM-L6-v2",
backend="openvino",
model_kwargs={"file_name": "openvino/openvino_model_qint8_quantized.xml"},
revision=f"refs/pr/{pull_request_nr}"
)
or when it gets merged:
from sentence_transformers import CrossEncoder
model = CrossEncoder(
"cross-encoder/ms-marco-MiniLM-L6-v2",
backend="openvino",
model_kwargs={"file_name": "openvino/openvino_model_qint8_quantized.xml"},
)
Read the Speeding up Inference documentation for more details.
</details>
This PR softly deprecates the margin option in mine_hard_negatives in favor of absolute_margin and relative_margin. In short:
absolute_margin: Discards negative candidates whose anchor_negative_similarity score is greater than or equal to anchor_positive_similarity - absolute_margin. With an absolute_margin of 0.1 and an anchor-positive similarity of 0.86, the maximum anchor-negative similarity for that anchor (e.g. query) is 0.76.relative_margin: Discards negative candidates whose anchor_negative_similarity score is greater than or equal to anchor_positive_similarity * (1 - relative_margin). With a relative_margin of 0.05 and an anchor-positive similarity of 0.86, the maximum anchor-negative similarity for that anchor (e.g. query) is 0.817 (i.e. 95% of the anchor-positive similarity).This means that we now support the recommended hard negatives mining strategy from the excellent NV-Retriever paper, a.k.a. the TopK-PercPos (95%) strategy:
from sentence_transformers.util import mine_hard_negatives
...
dataset = mine_hard_negatives(
dataset=dataset,
model=model,
relative_margin=0.05, # 0.05 means that the negative is at most 95% as similar to the anchor as the positive
num_negatives=num_negatives, # 10 or less is recommended
sampling_strategy="top", # "top" means that we sample the top candidates as negatives
batch_size=batch_size, # Adjust as needed
use_faiss=True, # Optional: Use faiss/faiss-gpu for faster similarity search
)
margin and margin_strategy to GISTEmbedLoss and CachedGISTEmbedLoss (#3299, #3323)activation_function=None in Dense module (#3316)all_layer_embeddings outputs are determined (#3320)SentenceTransformer.encode if prompts are provided and output_value=None (#3327)docs] Update a removed article with a new source by @lakshminarasimmanv in https://github.com/UKPLab/sentence-transformers/pull/3309typing] Fix typing for CrossEncoder.to by @tomaarsen in https://github.com/UKPLab/sentence-transformers/pull/3324feat] hard neg mining: deprecate margin in favor of absolute_margin & relative margin by @tomaarsen in https://github.com/UKPLab/sentence-transformers/pull/3321fix] Use return_dict=True in Transformer; improve how all_layer_embeddings are determined by @tomaarsen in https://github.com/UKPLab/sentence-transformers/pull/3320fix] Avoid error if prompts & output_value=None by @tomaarsen in https://github.com/UKPLab/sentence-transformers/pull/3327backend] Add ONNX & OpenVINO support for Cross Encoder (reranker) models by @tomaarsen in https://github.com/UKPLab/sentence-transformers/pull/3319Full Changelog: https://github.com/UKPLab/sentence-transformers/compare/v4.0.2...v4.1.0
This patch release updates some logic for maximum sequence lengths, typing issues, FSDP training, and distributed training device placement.
This patch release updates some logic for maximum sequence lengths, typing issues, FSDP training, and distributed training device placement.
Install this version with
# Training + Inference
pip install sentence-transformers[train]==4.0.2
# Inference only, use one of:
pip install sentence-transformers==4.0.2
pip install sentence-transformers[onnx-gpu]==4.0.2
pip install sentence-transformers[onnx]==4.0.2
pip install sentence-transformers[openvino]==4.0.2
When loading CrossEncoder models, we now rely on the minimum of the tokenizer model_max_length and the config max_position_embeddings (if they exist), rather than only relying on the latter if it exists. This previously resulted in the maximum sequence length of BAAI/bge-reranker-base being 514, whereas it can only handle sequences up to 512 tokens.
from sentence_transformers import CrossEncoder
model = CrossEncoder("BAAI/bge-reranker-base")
print(model.max_length)
# => 512
# The texts for which to predict similarity scores
query = "How many people live in Berlin?"
passages = [
"Berlin had a population of 3,520,031 registered inhabitants in an area of 891.82 square kilometers.",
"In 2013 around 600,000 Berliners were registered in one of the more than 2,300 sport and fitness clubs.",
]
scores = model.predict([(query, passage) for passage in passages])
print(scores)
# => [0.99953485 0.01062613]
# Or test really long inputs to ensure that there's no crash:
score = model.predict([["one " * 1000, "two " * 1000]])
print(score)
# => [0.95482624]
Note that you can use the activation_fn option with torch.nn.Identity() to avoid the default Sigmoid that maps everything to [0, 1]:
from sentence_transformers import CrossEncoder
import torch
model = CrossEncoder("BAAI/bge-reranker-base", activation_fn=torch.nn.Identity())
# The texts for which to predict similarity scores
query = "How many people live in Berlin?"
passages = [
"Berlin had a population of 3,520,031 registered inhabitants in an area of 891.82 square kilometers.",
"In 2013 around 600,000 Berliners were registered in one of the more than 2,300 sport and fitness clubs.",
]
scores = model.predict([(query, passage) for passage in passages])
print(scores)
# => [ 7.672551 -4.5337563]
By default, in a distributed training setup with multiple CUDA devices, the model is now placed on the CUDA device corresponding with that local rank. This should lower the VRAM usage on GPU 0 when performing distributed training.
SentenceTransformer class outside of the encode method. In v4.0.1, it was possible to no longer get help from your IDE for e.g. model.similarity, for example. (#3297)loss class instance when required for FSDP training. (#3295)docs] Resolve more broken links throughout the docs by @tomaarsen in https://github.com/UKPLab/sentence-transformers/pull/3294docs] Fix some broken docs redirects by @tomaarsen in https://github.com/UKPLab/sentence-transformers/pull/3296typing] Move encode typings back to .py from .pyi by @tomaarsen in https://github.com/UKPLab/sentence-transformers/pull/3297fix] Avoid "Only if model is wrapped" check which is faulty for FSDP by @tomaarsen in https://github.com/UKPLab/sentence-transformers/pull/3295cross-encoder] Set the tokenizer model_max_length to the min. of model_max_length & max_pos_embeds by @tomaarsen in https://github.com/UKPLab/sentence-transformers/pull/3304ci] Attempt to fix CI by @tomaarsen in https://github.com/UKPLab/sentence-transformers/pull/3305get_device_name for distributed training by @uminaty in https://github.com/UKPLab/sentence-transformers/pull/3303docs] Add missing docstring for push_to_hub by @tomaarsen in https://github.com/UKPLab/sentence-transformers/pull/3306docs] Specify that exported ONNX/OpenVINO models don't include pooling/normalization by @tomaarsen in https://github.com/UKPLab/sentence-transformers/pull/3307Full Changelog: https://github.com/UKPLab/sentence-transformers/compare/v4.0.1...v4.0.2
…using the `CrossEncoder.fit` method. Rather than deprecating this method, starting from v4.0, this method will use the `CrossEncoderTrainer` behind th…
This release consists of a major refactor that overhauls the reranker a.k.a. Cross Encoder training approach (introducing multi-gpu training, bf16, loss logging, callbacks, and much more), including all new Training Overview, Loss Overview, API Reference docs, training examples and more!
Install this version with
# Training + Inference
pip install sentence-transformers[train]==4.0.1
# Inference only, use one of:
pip install sentence-transformers==4.0.1
pip install sentence-transformers[onnx-gpu]==4.0.1
pip install sentence-transformers[onnx]==4.0.1
pip install sentence-transformers[openvino]==4.0.1
[!TIP] My Training and Finetuning Reranker Models with Sentence Transformers v4 blogpost is an excellent place to learn 1) why finetuning rerankers makes sense and 2) how you can do it, too!
The v4.0 release centers around this huge modernization of the training approach for CrossEncoder models, following v3.0 which introduced the same for SentenceTransformer models. Whereas training before v4.0 used to be all about InputExample, DataLoader and model.fit, the new training approach relies on 5 components. You can learn more about these components in our Training and Finetuning Embedding Models with Sentence Transformers v4 blogpost. Additionally, you can read the new Training Overview, check out the Training Examples, or read this summary:
Dataset or DatasetDict. This class is much more suited for sharing & efficient modifications than lists/DataLoaders of InputExample instances. A Dataset can contain multiple text columns that will be fed in order to the corresponding loss function. So, if the loss expects (anchor, positive, negative) triplets, then your dataset should also have 3 columns. The names of these columns are irrelevant. If there is a "label" or "score" column, it is treated separately, and used as the labels during training.
A DatasetDict can be used to train with multiple datasets at once, e.g.:DatasetDict({
natural_questions: Dataset({
features: ['anchor', 'positive'],
num_rows: 392702
})
gooaq: Dataset({
features: ['anchor', 'positive', 'negative'],
num_rows: 549367
})
stsb: Dataset({
features: ['sentence1', 'sentence2', 'label'],
num_rows: 5749
})
})
When a DatasetDict is used, the loss parameter to the CrossEncoderTrainer must also be a dictionary with these dataset keys, e.g.:{
'natural_questions': CachedMultipleNegativesRankingLoss(...),
'gooaq': CachedMultipleNegativesRankingLoss(...),
'stsb': BinaryCrossEntropyLoss(...),
}
SentenceEvaluator instance. Unlike before, models can now be evaluated both on an evaluation dataset with some loss function and/or a SentenceEvaluator instance.CrossEncoderTrainer instance based on the transformers Trainer. This instance can be initialized with a CrossEncoder model, a CrossEncoderTrainingArguments class, a SentenceEvaluator, a training and evaluation Dataset/DatasetDict and a loss function/dict of loss functions. Most of these parameters are optional. Once provided, all you have to do is call trainer.train().Some of the major features that are now implemented include:
This script is a minimal example (no evaluator, no training arguments) of training mpnet-base on a part of the sentence-transformers/hotpotqa dataset using BinaryCrossEntropyLoss:
from datasets import load_dataset
from sentence_transformers import CrossEncoder, CrossEncoderTrainer
from sentence_transformers.cross_encoder.losses import BinaryCrossEntropyLoss
# 1. Define the model. Either from scratch of by loading a pre-trained model
model = CrossEncoder("microsoft/mpnet-base")
# 2. Load a dataset to finetune on
dataset = load_dataset("sentence-transformers/hotpotqa", "triplet", split="train")
def triplet_to_labeled_pair(batch):
anchors = batch["anchor"]
positives = batch["positive"]
negatives = batch["negative"]
return {
"sentence_A": anchors * 2,
"sentence_B": positives + negatives,
"labels": [1] * len(positives) + [0] * len(negatives),
}
dataset = dataset.map(triplet_to_labeled_pair, batched=True, remove_columns=dataset.column_names)
train_dataset = dataset.select(range(10_000))
eval_dataset = dataset.select(range(10_000, 11_000))
# 3. Define a loss function
loss = BinaryCrossEntropyLoss(model)
# 4. Create a trainer & train
trainer = CrossEncoderTrainer(
model=model,
train_dataset=train_dataset,
eval_dataset=eval_dataset,
loss=loss,
)
trainer.train()
# 5. Save the trained model
model.save_pretrained("models/mpnet-base-hotpotqa")
# model.push_to_hub("mpnet-base-hotpotqa")
Additionally, trained models now automatically produce extensive model cards. Each of the following models were trained using some script from the Training Examples, and the model cards were not edited manually whatsoever:
Prior to the Sentence Transformer v4 release, all reranker models would be trained using the CrossEncoder.fit method. Rather than deprecating this method, starting from v4.0, this method will use the CrossEncoderTrainer behind the scenes. This means that your old training code should still work, and should even be upgraded with the new features such as multi-gpu training, loss logging, etc. That said, the new training approach is much more powerful, so it is recommended to write new training scripts using the new approach.
To help you out, all of the Cross Encoder (a.k.a. reranker) training scripts were updated to use the new Trainer-based approach.
Finetuning reranker models on your data is very valuable. Consider for example these 2 models that I finetuned on 100k samples from the GooAQ dataset in 30 minutes and 1 hour, respectively. After finetuning, my models heavily outperformed general-purpose reranker models, even though GooAQ is a very generic dataset/domain!
Read my Training and Finetuning Reranker Models with Sentence Transformers v4 blogpost for many more details on these models and how they were trained.
show_progress_bar for the InformationRetrievalEvaluator (#3227)SubsetRandomSampler with RandomSampler in the default batch sampler, should result in reduced memory usage and increased training speed! (#3261)SentenceTransformer.fit (#3269)model.max_seq_length for CLIP models (#2969)MatryoshkaLoss with n_dims_per_step and an unsorted matryoshka_dims crashing (#3203)GISTEmbedLoss failing with some base models whose tokenizers don't have the vocab attribute (#3219, #3226)Asym-based SentenceTransformer models (#3220, #3244)numpy or torch (#3277)The v4.0.0 version did not include the model_card_template.md in the package, this has been resolved in v4.0.1 via ba1260d58c8804e97989e5af06ac90a0f4be8594.
docs] Resolve broken URL due to weird & behaviour in pretrained ST models by @tomaarsen in https://github.com/UKPLab/sentence-transformers/pull/3213docs] Update incorrect name: pairwise_similarity -> similarity_pairwise by @tomaarsen in https://github.com/UKPLab/sentence-transformers/pull/3224fix] Use .get_vocab() instead of .vocab for checking tokenizer vocabulary by @tomaarsen in https://github.com/UKPLab/sentence-transformers/pull/3226feat] Add progress bar support for corpus in IR Evaluator by @tomaarsen in https://github.com/UKPLab/sentence-transformers/pull/3227fix] Fix Syntax issue; move 'as fIn' to after the if-else in STSDataReader by @tomaarsen in https://github.com/UKPLab/sentence-transformers/pull/3235typing] Fix the type hints in CGISTEmbedLoss by @tomaarsen in https://github.com/UKPLab/sentence-transformers/pull/3272v4] CrossEncoder Training refactor - MultiGPU, loss logging, bf16, etc. by @tomaarsen in https://github.com/UKPLab/sentence-transformers/pull/3222A special shoutout to @milistu for contributing the LambdaLoss & ListNetLoss and @yjoonjang for contributing the ListMLELoss, PListMLELoss, and RankNetLoss. Much appreciated, you really helped improve this release!
Full Changelog: https://github.com/UKPLab/sentence-transformers/compare/v3.4.1...v4.0.1
Nothing published for this version
This release introduces a convenient compatibility with Model2Vec models, and fixes a bug that caused an outgoing request even when using a local mode
This release introduces a convenient compatibility with Model2Vec models, and fixes a bug that caused an outgoing request even when using a local model.
Install this version with
# Training + Inference
pip install sentence-transformers[train]==3.4.1
# Inference only, use one of:
pip install sentence-transformers==3.4.1
pip install sentence-transformers[onnx-gpu]==3.4.1
pip install sentence-transformers[onnx]==3.4.1
pip install sentence-transformers[openvino]==3.4.1
This release introduces support to load an efficient Model2Vec embedding model directly in Sentence Transformers:
from sentence_transformers import SentenceTransformer
# Download from the 🤗 Hub
model = SentenceTransformer(
"minishlab/potion-base-8M",
device="cpu",
)
# Run inference
sentences = [
'Gadofosveset-enhanced MR angiography of carotid arteries: does steady-state imaging improve accuracy of first-pass imaging?',
'To evaluate the diagnostic accuracy of gadofosveset-enhanced magnetic resonance (MR) angiography in the assessment of carotid artery stenosis, with digital subtraction angiography (DSA) as the reference standard, and to determine the value of reading first-pass, steady-state, and "combined" (first-pass plus steady-state) MR angiograms.',
'In a longitudinal study we investigated in vivo alterations of CVO during neuroinflammation, applying Gadofluorine M- (Gf) enhanced magnetic resonance imaging (MRI) in experimental autoimmune encephalomyelitis, an animal model of multiple sclerosis. SJL/J mice were monitored by Gadopentate dimeglumine- (Gd-DTPA) and Gf-enhanced MRI after adoptive transfer of proteolipid-protein-specific T cells. Mean Gf intensity ratios were calculated individually for different CVO and correlated to the clinical disease course. Subsequently, the tissue distribution of fluorescence-labeled Gf as well as the extent of cellular inflammation was assessed in corresponding histological slices.',
]
embeddings = model.encode(sentences)
print(embeddings.shape)
# [3, 1024]
# Get the similarity scores for the embeddings
similarities = model.similarity(embeddings[0], embeddings[1:])
print(similarities)
# tensor([[0.8085, 0.4884]])
<details><summary>Previously, loading a Model2Vec model required you to load a StaticEmbedding module.</summary>
from sentence_transformers import SentenceTransformer
from sentence_transformers.models import StaticEmbedding
# Download from the 🤗 Hub
module = StaticEmbedding.from_model2vec("minishlab/potion-base-8M")
model = SentenceTransformer(modules=[module], device="cpu")
# Run inference
sentences = [
'Gadofosveset-enhanced MR angiography of carotid arteries: does steady-state imaging improve accuracy of first-pass imaging?',
'To evaluate the diagnostic accuracy of gadofosveset-enhanced magnetic resonance (MR) angiography in the assessment of carotid artery stenosis, with digital subtraction angiography (DSA) as the reference standard, and to determine the value of reading first-pass, steady-state, and "combined" (first-pass plus steady-state) MR angiograms.',
'In a longitudinal study we investigated in vivo alterations of CVO during neuroinflammation, applying Gadofluorine M- (Gf) enhanced magnetic resonance imaging (MRI) in experimental autoimmune encephalomyelitis, an animal model of multiple sclerosis. SJL/J mice were monitored by Gadopentate dimeglumine- (Gd-DTPA) and Gf-enhanced MRI after adoptive transfer of proteolipid-protein-specific T cells. Mean Gf intensity ratios were calculated individually for different CVO and correlated to the clinical disease course. Subsequently, the tissue distribution of fluorescence-labeled Gf as well as the extent of cellular inflammation was assessed in corresponding histological slices.',
]
embeddings = model.encode(sentences)
print(embeddings.shape)
# [3, 1024]
# Get the similarity scores for the embeddings
similarities = model.similarity(embeddings[0], embeddings[1:])
print(similarities)
# tensor([[0.8085, 0.4884]])
</details>
Model2Vec was the inspiration of the recent Static Embedding work; all of these models can be used to approach the performance of normal transformer-based embedding models at a fraction of the latency. For example, both Model2Vec and Static Embedding models are ~25x faster than tiny embedding models on a GPU and ~400x faster than those models on a CPU.
local_files_only=True still triggered a request to Hugging Face for the model card metadata; this has been resolved in (#3202).StaticEmbedding.__init__ by @altescy in https://github.com/UKPLab/sentence-transformers/pull/3196integration] Work towards full model2vec integration by @tomaarsen in https://github.com/UKPLab/sentence-transformers/pull/3182set_base_model when local_files_only=True by @Davidyz in https://github.com/UKPLab/sentence-transformers/pull/3202Full Changelog: https://github.com/UKPLab/sentence-transformers/compare/v3.4.0...v3.4.1
[fix] Fix breaking change in PyLate when loading modules by @tomaarsen in https://github.com/UKPLab/sentence-transformers/pull/3110
This release resolves a memory leak when deleting a model & trainer, adds compatibility between the Cached... losses and the Matryoshka loss modifier, resolves numerous bugs, and adds several small features.
Install this version with
# Training + Inference
pip install sentence-transformers[train]==3.4.0
# Inference only, use one of:
pip install sentence-transformers==3.4.0
pip install sentence-transformers[onnx-gpu]==3.4.0
pip install sentence-transformers[onnx]==3.4.0
pip install sentence-transformers[openvino]==3.4.0
It is now possible to combine the strong Cached losses (CachedMultipleNegativesRankingLoss, CachedGISTEmbedLoss, CachedMultipleNegativesSymmetricRankingLoss) with the Matryoshka loss modifier:
from sentence_transformers import SentenceTransformer, SentenceTransformerTrainer, losses
from datasets import Dataset
model = SentenceTransformer("microsoft/mpnet-base")
train_dataset = Dataset.from_dict({
"anchor": ["It's nice weather outside today.", "He drove to work."],
"positive": ["It's so sunny.", "He took the car to the office."],
})
loss = losses.CachedMultipleNegativesRankingLoss(model, mini_batch_size=16)
loss = losses.MatryoshkaLoss(model, loss, [768, 512, 256, 128, 64])
trainer = SentenceTransformerTrainer(
model=model,
train_dataset=train_dataset,
loss=loss,
)
trainer.train()
See for example tomaarsen/mpnet-base-gooaq-cmnrl-mrl which was trained with CachedMultipleNegativesRankingLoss (CMNRL) with the Matryoshka loss modifier (MRL).
Due to a circular dependency in the SentenceTransformerTrainer -> SentenceTransformer -> SentenceTransformerModelCardData -> SentenceTransformerTrainer, deleting the trainer and model still doesn't clear them up via garbage disposal. I've moved a lot of components around, and now SentenceTransformerModelCardData does not need to store the SentenceTransformerTrainer, breaking the cycle.
We ran the seed optimization script (which frequently creates and deletes models and trainers):
16332MiB / 24576MiB
8222MiB / 24576MiB
margin parameter to the TripletEvaluator in #2862.mine_hard_negatives in #2967.no_duplicates Batch Sampler (#3069). This has been resolved in #3073model.fit() training with write_csv on an evaluator would crash (#3062). This has been resolved in #3066.np.float instead of float (#3075). This has been resolved in #3076 and #3096.revision or cache_dir when loading a PEFT Adapter model (#3061). This has been resolved in #3079 and #3174.model.to (#3078). This has been resolved in #3104.kwargs keys were not saved in modules.json correctly, e.g. relevant for jina-embeddings-v3 (#3111). This has been resolved in #3112.HfArgumentParser(SentenceTransformerTrainingArguments) would crash due to prompts typing (#3090). This has been resolved in #3178.training] Pass steps/epoch/output_path to Evaluator during training by @tomaarsen in https://github.com/UKPLab/sentence-transformers/pull/3066examples] Update the quantization script by @tomaarsen in https://github.com/UKPLab/sentence-transformers/pull/3070fix] Fix different batches per epoch in NoDuplicatesBatchSampler by @tomaarsen in https://github.com/UKPLab/sentence-transformers/pull/3073docs] Add links to backend-export in Speeding up Inference by @tomaarsen in https://github.com/UKPLab/sentence-transformers/pull/3071model_card] Keep the model card readable even with many datasets by @tomaarsen in https://github.com/UKPLab/sentence-transformers/pull/3088docs] Add NanoBEIR to the Training Overview evaluators by @tomaarsen in https://github.com/UKPLab/sentence-transformers/pull/3089docs] Update from Sphinx==3.5.4 to 8.1.3, recommonmark -> myst-parser by @tomaarsen in https://github.com/UKPLab/sentence-transformers/pull/3099docs] List 'prompts' as a key training argument by @tomaarsen in https://github.com/UKPLab/sentence-transformers/pull/3101enhancement] Make MultipleNegativesRankingLoss easier to understand by @tomaarsen in https://github.com/UKPLab/sentence-transformers/pull/3100fix] Fix breaking change in PyLate when loading modules by @tomaarsen in https://github.com/UKPLab/sentence-transformers/pull/3110typo] Add missing space between sentences in error message by @tomaarsen in https://github.com/UKPLab/sentence-transformers/pull/3125fix] Save custom module kwargs if specified by @tomaarsen in https://github.com/UKPLab/sentence-transformers/pull/3112memory] Avoid storing trainer in ModelCardCallback and SentenceTransformerModelCardData by @tomaarsen in https://github.com/UKPLab/sentence-transformers/pull/3144docs] Update the Static Embedding example snippet by @tomaarsen in https://github.com/UKPLab/sentence-transformers/pull/3177fix] Use HfArgumentParser-compatible typing for prompts by @tomaarsen in https://github.com/UKPLab/sentence-transformers/pull/3178docs] Add PEFT documentation + training example by @tomaarsen in https://github.com/UKPLab/sentence-transformers/pull/3180tests] Make TripletEvaluator test more consistent by @tomaarsen in https://github.com/UKPLab/sentence-transformers/pull/3183deprecation] Clarify that datasets and readers are deprecated since v3 by @tomaarsen in https://github.com/UKPLab/sentence-transformers/pull/3184An explicit thanks to @JINO-ROHIT who has made a large amount of contributions in this release.
Full Changelog: https://github.com/UKPLab/sentence-transformers/compare/v3.3.1...v3.4.0
This patch release fixes a small issue with loading private models from Hugging Face using the token argument.
This patch release fixes a small issue with loading private models from Hugging Face using the token argument.
Install this version with
# Training + Inference
pip install sentence-transformers[train]==3.3.1
# Inference only, use one of:
pip install sentence-transformers==3.3.1
pip install sentence-transformers[onnx-gpu]==3.3.1
pip install sentence-transformers[onnx]==3.3.1
pip install sentence-transformers[openvino]==3.3.1
If you're loading model under this scenario:
HF_TOKEN environment variable via huggingface-cli login or some other approach.token argument to SentenceTransformer to load the model.Then you may have encountered a crash in v3.3.0. This should be resolved now.
docs] Fix the prompt link to the training script by @tomaarsen in https://github.com/UKPLab/sentence-transformers/pull/3060Full Changelog: https://github.com/UKPLab/sentence-transformers/compare/v3.3.0...v3.3.1
…v4.46.0 compatibility, and Python 3.8 deprecation.
4x speedup for CPU with OpenVINO int8 static quantization, training with prompts for a free performance boost, convenient evaluation on NanoBEIR: a subset of a strong Information Retrieval benchmark, PEFT compatibility by easily adding/loading adapters, Transformers v4.46.0 compatibility, and Python 3.8 deprecation.
Install this version with:
# Training + Inference
pip install sentence-transformers[train]==3.3.0
# Inference only, use one of:
pip install sentence-transformers==3.3.0
pip install sentence-transformers[onnx-gpu]==3.3.0
pip install sentence-transformers[onnx]==3.3.0
pip install sentence-transformers[openvino]==3.3.0
We introduce int8 static quantization using OpenVINO, a highly performant solution that outperforms all other current backends by a mile, at a minimal loss in performance. Here are the updated benchmarks:
<p align="center"> <img src="https://github.com/user-attachments/assets/96f96da5-b65e-4293-8b67-d47430aa5fae" width="50%" /> </p>
from sentence_transformers import SentenceTransformer, export_static_quantized_openvino_model
# 1. Load a model with the OpenVINO backend
model = SentenceTransformer("all-MiniLM-L6-v2", backend="openvino")
# 2. Quantize the model to int8, push the model to https://huggingface.co/sentence-transformers/all-MiniLM-L6-v2
# as a pull request:
export_static_quantized_openvino_model(
model,
quantization_config=None,
model_name_or_path="sentence-transformers/all-MiniLM-L6-v2",
push_to_hub=True,
create_pr=True,
)
You can immediately use the model, even before it's merged, by using the revision argument:
from sentence_transformers import SentenceTransformer
pull_request_nr = 2 # TODO: Update this to the number of your pull request
model = SentenceTransformer(
"all-MiniLM-L6-v2",
backend="openvino",
model_kwargs={"file_name": "openvino_model_qint8_quantized.xml"},
revision=f"refs/pr/{pull_request_nr}"
)
And once it's merged:
from sentence_transformers import SentenceTransformer
model = SentenceTransformer(
"all-MiniLM-L6-v2",
backend="openvino",
model_kwargs={"file_name": "openvino/openvino_model_qint8_quantized.xml"},
)
You can also quantize a model and save it locally:
from sentence_transformers import SentenceTransformer, export_static_quantized_openvino_model
from optimum.intel import OVQuantizationConfig
model = SentenceTransformer("all-mpnet-base-v2", backend="openvino")
model.save_pretrained("path/to/all-mpnet-base-v2-local")
quantization_config = OVQuantizationConfig() # <- You can update settings here
export_static_quantized_openvino_model(model, quantization_config, "path/to/all-mpnet-base-v2-local")
And after quantizing, you can load it like so:
from sentence_transformers import SentenceTransformer
model = SentenceTransformer(
"path/to/all-mpnet-base-v2-local",
backend="openvino",
model_kwargs={"file_name": "openvino_model_qint8_quantized.xml"},
)
All original Sentence Transformer models already have these new openvino_model_qint8_quantized.xml files, so you can load them without exporting directly! I would recommend making pull requests for other models on Hugging Face that you'd like to see quantized.
Learn more about how to Speed up Inference in the documentation: https://sbert.net/docs/sentence_transformer/usage/efficiency.html
Many modern embedding models are trained with “instructions” or “prompts” following the INSTRUCTOR paper. These prompts are strings, prefixed to each text to be embedded, allowing the model to distinguish between different types of text.
For example, the mixedbread-ai/mxbai-embed-large-v1 model was trained with Represent this sentence for searching relevant passages: as the prompt for all queries. This prompt is stored in the model configuration under the prompt name "query", so users can specify that prompt_name in model.encode:
from sentence_transformers import SentenceTransformer
model = SentenceTransformer("mixedbread-ai/mxbai-embed-large-v1")
query_embedding = model.encode("What are Pandas?", prompt_name="query")
# or
# query_embedding = model.encode("What are Pandas?", prompt="Represent this sentence for searching relevant passages: ")
document_embeddings = model.encode([
"Pandas is a software library written for the Python programming language for data manipulation and analysis.",
"Pandas are a species of bear native to South Central China. They are also known as the giant panda or simply panda.",
"Koala bears are not actually bears, they are marsupials native to Australia.",
])
similarity = model.similarity(query_embedding, document_embeddings)
print(similarity)
# => tensor([[0.7594, 0.7560, 0.4674]])
Various papers (INSTRUCTOR, BGE) show that including prompts or instructions both during training and inference results in stronger performance. As of this release, it's now possible to easily train with prompts in Sentence Transformers with just one extra training argument: prompts. There are 4 accepted formats for it:
str: A single prompt to use for all columns in all datasets. For example:args = SentenceTransformerTrainingArguments(
...,
prompts="text: ",
...,
)
Dict[str, str]: A dictionary mapping column names to prompts, applied to all datasets. For example:args = SentenceTransformerTrainingArguments(
...,
prompts={
"query": "query: ",
"answer": "document: ",
},
...,
)
Dict[str, str]: A dictionary mapping dataset names to prompts. This should only be used if your training/evaluation/test datasets are a DatasetDict or a dictionary of Dataset. For example:args = SentenceTransformerTrainingArguments(
...,
prompts={
"stsb": "Represent this text for semantic similarity search: ",
"nq": "Represent this text for retrieval: ",
},
...,
)
Dict[str, Dict[str, str]]: A dictionary mapping dataset names to dictionaries mapping column names to prompts. This should only be used if your training/evaluation/test datasets are a DatasetDict or a dictionary of Dataset. For example:args = SentenceTransformerTrainingArguments(
...,
prompts={
"stsb": {
"sentence1": "sts: ",
"sentence2": "sts: ",
},
"nq": {
"query": "query: ",
"document": "document: ",
},
},
...,
)
I've trained models with and without prompts for 2 base models: mpnet-base and bert-base-uncased:
For both base models, the model with prompts consistently outperformed the baseline model. After training, the models with prompts resulted in a 0.66% and 0.90% relative improvement on NDCG@10 at no extra cost.
mpnet-base tests |
bert-base-uncased tests |
|---|---|
This update introduced a new simple NanoBEIREvaluator, evaluating your model against NanoBEIR: a collection of subsets of the 13 BEIR datasets. BEIR corresponds to the retrieval tab of MTEB, and is commonly seen as a valuable indicator of general-purpose information retrieval performance.
With the NanoBEIREvaluator, you can easily evaluate your models on a much faster benchmark that should give similar insights in performance as BEIR. You can use it like so:
from sentence_transformers.evaluation import NanoBEIREvaluator
from sentence_transformers import SentenceTransformer
import logging
# Optional, but nice to get human-readable results in the terminal
logging.basicConfig(
format="%(asctime)s - %(message)s", datefmt="%Y-%m-%d %H:%M:%S", level=logging.INFO
)
# 1. Load a model
model = SentenceTransformer("all-mpnet-base-v2", backend="onnx")
# 2. Initialize the evaluator
evaluator = NanoBEIREvaluator()
# 3. Call the evaluator to get a dictionary of metric names to values
results = evaluator(model)
"""
NanoBEIR Evaluation of the model on ['climatefever', 'dbpedia', 'fever', 'fiqa2018', 'hotpotqa', 'msmarco', 'nfcorpus', 'nq', 'quoraretrieval', 'scidocs', 'arguana', 'scifact', 'touche2020'] dataset:
Evaluating NanoClimateFEVER
Information Retrieval Evaluation of the model on the NanoClimateFEVER dataset:
Queries: 50
Corpus: 3408
Score-Function: cosine
Accuracy@1: 24.00%
Accuracy@3: 36.00%
Accuracy@5: 44.00%
Accuracy@10: 66.00%
Precision@1: 24.00%
Precision@3: 14.00%
Precision@5: 10.40%
Precision@10: 9.00%
Recall@1: 9.50%
Recall@3: 17.33%
Recall@5: 22.90%
Recall@10: 36.07%
MRR@10: 0.3311
NDCG@10: 0.2618
MAP@100: 0.1982
Evaluating NanoDBPedia
Information Retrieval Evaluation of the model on the NanoDBPedia dataset:
Queries: 50
Corpus: 6045
Score-Function: cosine
Accuracy@1: 66.00%
Accuracy@3: 88.00%
Accuracy@5: 88.00%
Accuracy@10: 88.00%
Precision@1: 66.00%
Precision@3: 58.00%
Precision@5: 52.00%
Precision@10: 43.60%
Recall@1: 6.87%
Recall@3: 14.70%
Recall@5: 20.30%
Recall@10: 27.62%
MRR@10: 0.7533
NDCG@10: 0.5384
MAP@100: 0.3796
Evaluating NanoFEVER
Information Retrieval Evaluation of the model on the NanoFEVER dataset:
Queries: 50
Corpus: 4996
... (truncated for brevity)
Aggregated for Score Function: cosine
Accuracy@1: 52.87%
Accuracy@3: 71.35%
Accuracy@5: 78.45%
Accuracy@10: 85.07%
Precision@1: 52.87%
Recall@1: 30.28%
Precision@3: 33.78%
Recall@3: 47.93%
Precision@5: 26.23%
Recall@5: 55.04%
Precision@10: 18.07%
Recall@10: 62.54%
MRR@10: 0.6334
NDCG@10: 0.5758
"""
# 4. Print the results
print(evaluator.primary_metric)
# => "NanoBEIR_mean_cosine_ndcg@10"
print(results[evaluator.primary_metric])
# => 0.5758124378869705
<details><summary>Advanced Usage</summary>
You can also specify a subset of datasets, and you can specify query and/or corpus prompts, if your model uses them. For example:
import logging
from sentence_transformers import SentenceTransformer
from sentence_transformers.evaluation import NanoBEIREvaluator
# Optional, but nice to get human-readable results in the terminal
logging.basicConfig(
format="%(asctime)s - %(message)s", datefmt="%Y-%m-%d %H:%M:%S", level=logging.INFO
)
model = SentenceTransformer('intfloat/multilingual-e5-large-instruct')
datasets = ["QuoraRetrieval", "MSMARCO"]
query_prompts = {
"QuoraRetrieval": "Instruct: Given a question, retrieve questions that are semantically equivalent to the given question\\nQuery: ",
"MSMARCO": "Instruct: Given a web search query, retrieve relevant passages that answer the query\\nQuery: "
}
evaluator = NanoBEIREvaluator(
dataset_names=datasets,
query_prompts=query_prompts,
)
results = evaluator(model)
'''
NanoBEIR Evaluation of the model on ['QuoraRetrieval', 'MSMARCO'] dataset:
Evaluating NanoQuoraRetrieval
Information Retrieval Evaluation of the model on the NanoQuoraRetrieval dataset:
Queries: 50
Corpus: 5046
Score-Function: cosine
Accuracy@1: 92.00%
Accuracy@3: 98.00%
Accuracy@5: 100.00%
Accuracy@10: 100.00%
Precision@1: 92.00%
Precision@3: 40.67%
Precision@5: 26.00%
Precision@10: 14.00%
Recall@1: 81.73%
Recall@3: 94.20%
Recall@5: 97.93%
Recall@10: 100.00%
MRR@10: 0.9540
NDCG@10: 0.9597
MAP@100: 0.9395
Evaluating NanoMSMARCO
Information Retrieval Evaluation of the model on the NanoMSMARCO dataset:
Queries: 50
Corpus: 5043
Score-Function: cosine
Accuracy@1: 40.00%
Accuracy@3: 74.00%
Accuracy@5: 78.00%
Accuracy@10: 88.00%
Precision@1: 40.00%
Precision@3: 24.67%
Precision@5: 15.60%
Precision@10: 8.80%
Recall@1: 40.00%
Recall@3: 74.00%
Recall@5: 78.00%
Recall@10: 88.00%
MRR@10: 0.5849
NDCG@10: 0.6572
MAP@100: 0.5892
Average Queries: 50.0
Average Corpus: 5044.5
Aggregated for Score Function: cosine
Accuracy@1: 66.00%
Accuracy@3: 86.00%
Accuracy@5: 89.00%
Accuracy@10: 94.00%
Precision@1: 66.00%
Recall@1: 60.87%
Precision@3: 32.67%
Recall@3: 84.10%
Precision@5: 20.80%
Recall@5: 87.97%
Precision@10: 11.40%
Recall@10: 94.00%
MRR@10: 0.7694
NDCG@10: 0.8085
'''
print(evaluator.primary_metric)
# => "NanoBEIR_mean_cosine_ndcg@10"
print(results[evaluator.primary_metric])
# => 0.8084508771660436
</details>
NanoBEIREvaluatorSentence Transformers has been integrated much more closely with PEFT. Notably, we introduce new methods:
These methods allow you to add new PEFT adapters or load pretrained ones, for example:
from sentence_transformers import SentenceTransformer
# 1. Load a model to finetune with 2. (Optional) model card data
model = SentenceTransformer(
"all-MiniLM-L6-v2",
model_card_data=SentenceTransformerModelCardData(
language="en",
license="apache-2.0",
model_name="all-MiniLM-L6-v2 adapter finetuned on GooAQ pairs",
),
)
# 2. Create a LoRA adapter for the model & add it
peft_config = LoraConfig(
task_type=TaskType.FEATURE_EXTRACTION,
inference_mode=False,
r=8,
lora_alpha=32,
lora_dropout=0.1,
)
model.add_adapter(peft_config)
# Proceed as usual... See https://sbert.net/docs/sentence_transformer/training_overview.html
Given sentence-transformers-testing/stsb-bert-tiny-lora as a small adapter model (the adapter_model.safetensors file is only 33.8kB!) on top of sentence-transformers-testing/stsb-bert-tiny-safetensors, you can either load this adapter directly:
from sentence_transformers import SentenceTransformer
model = SentenceTransformer("sentence-transformers-testing/stsb-bert-tiny-lora")
embeddings = model.encode(["This is an example sentence", "Each sentence is converted"])
print(embeddings.shape)
# (2, 128)
Or you can load the original model and load the adapter into it:
from sentence_transformers import SentenceTransformer
model = SentenceTransformer("sentence-transformers-testing/stsb-bert-tiny-safetensors")
model.load_adapter("sentence-transformers-testing/stsb-bert-tiny-lora")
embeddings = model.encode(["This is an example sentence", "Each sentence is converted"])
print(embeddings.shape)
# (2, 128)
The recent transformers v4.46.0 update introduced a few changes that were incompatible with Sentence Transformers. For example:
num_items_in_batch argument to the compute_loss method in the TrainerValueError if eval_dataset is None while eval_strategy is not "no" (this should be possible in Sentence Transformers, as we accept evaluating with just an evaluator as well)These issues and deprecation warnings have been resolved.
Given that Python 3.8 has now reached it's end of life, Sentence Transformers will no longer support it.
peft] If AutoModel is wrapped with PEFT for prompt learning, then extend the attention mask by @tomaarsen in https://github.com/UKPLab/sentence-transformers/pull/3000integration] Add support for Transformers v4.46.0 by @tomaarsen in https://github.com/UKPLab/sentence-transformers/pull/3026datasets must be install to fit a model by @h4c5 in https://github.com/UKPLab/sentence-transformers/pull/3020feat] Integrate NanoBeIR datasets; use model.similarity by default in evaluators by @ArthurCamara in https://github.com/UKPLab/sentence-transformers/pull/2966fix] Avoid passing eval_dataset=None to transformers due to >=v4.46.0 crash by @tomaarsen in https://github.com/UKPLab/sentence-transformers/pull/3035docs] Update the dated example in the NanoBEIREvaluator by @tomaarsen in https://github.com/UKPLab/sentence-transformers/pull/3034deprecate] Drop Python 3.8 support due to EOL by @tomaarsen in https://github.com/UKPLab/sentence-transformers/pull/3033tests] Remove evaluation_steps from model.fit test without evaluator by @tomaarsen in https://github.com/UKPLab/sentence-transformers/pull/3037fix] Fix loading pre-exported OV/ONNX model if export=False by @tomaarsen in https://github.com/UKPLab/sentence-transformers/pull/3036chore] If Transformers 4.46.0, use processing_class instead of tokenizer when saving by @tomaarsen in https://github.com/UKPLab/sentence-transformers/pull/3038docs] Add some missing docs for include_prompt in Pooling by @tomaarsen in https://github.com/UKPLab/sentence-transformers/pull/3042feat] Trainer with prompts and prompt masking by @ArthurCamara in https://github.com/UKPLab/sentence-transformers/pull/2964enh] Add Support for multiple adapters on Transformers-based models by @carlesonielfa in https://github.com/UKPLab/sentence-transformers/pull/3046 & https://github.com/UKPLab/sentence-transformers/pull/2993Big thanks to @ArthurCamara for leading the work on both 1) training with prompts and 2) NanoBEIR.
Full Changelog: https://github.com/UKPLab/sentence-transformers/compare/v3.2.1...v3.3.0
This patch release fixes some small bugs, such as related to loading CLIP models, automatic model card generation issues, and ensuring compatibility w
This patch release fixes some small bugs, such as related to loading CLIP models, automatic model card generation issues, and ensuring compatibility with third party libraries.
Install this version with
# Training + Inference
pip install sentence-transformers[train]==3.2.1
# Inference only, use one of:
pip install sentence-transformers==3.2.1
pip install sentence-transformers[onnx-gpu]==3.2.1
pip install sentence-transformers[onnx]==3.2.1
pip install sentence-transformers[openvino]==3.2.1
In v3.2.0, a non-Transformer based model (e.g. CLIP) would not load correctly if the model was saved in the root of the model repository/directory. This has been resolved in #3007.
StaticEmbedding-based model is finetuned with incompatible lossesThe following losses are not compatible with StaticEmbedding-based models:
An error is now thrown when one of these are used with a StaticEmbedding-based model. I recommend using MultipleNegativesRankingLoss to finetune these models, e.g. as in https://huggingface.co/tomaarsen/static-bert-uncased-gooaq.
Note: to get good performance, you must use much higher learning rates than otherwise. In my experiments, 2e-1 worked well.
output_hidden_statesFor example, this script used to fail, but passes now:
from sentence_transformers import SentenceTransformer
model = SentenceTransformer(
"distiluse-base-multilingual-cased",
backend="onnx",
model_kwargs={"provider": "CPUExecutionProvider"},
)
sentences = ["This is an example sentence", "Each sentence is converted"]
embeddings = model.encode(sentences)
print(embeddings.shape)
docs] Update the training snippets for some losses that should use the v3 Trainer by @tomaarsen in https://github.com/UKPLab/sentence-transformers/pull/2987enh] Throw error if StaticEmbedding-based model is trained with incompatible loss by @tomaarsen in https://github.com/UKPLab/sentence-transformers/pull/2990fix] Fix semantic_search_usearch with 'binary' by @tomaarsen in https://github.com/UKPLab/sentence-transformers/pull/2989model cards] Prevent crash on generating widgets if dataset column is empty by @tomaarsen in https://github.com/UKPLab/sentence-transformers/pull/2997warn] Throw a warning if compute_metrics is set, as it's not used by @tomaarsen in https://github.com/UKPLab/sentence-transformers/pull/3002fix] Prevent IndexError if output_hidden_states & ONNX by @tomaarsen in https://github.com/UKPLab/sentence-transformers/pull/3008Full Changelog: https://github.com/UKPLab/sentence-transformers/compare/v3.2.0...v3.2.1
This release introduces 2 new efficient computing backends for SentenceTransformer models: ONNX and OpenVINO + optimization & quantization, allowing f
This release introduces 2 new efficient computing backends for SentenceTransformer models: ONNX and OpenVINO + optimization & quantization, allowing for speedups up to 2x-3x; static embeddings via Model2Vec allowing for lightning-fast models (i.e., 50x-500x speedups) at a ~10%-20% performance cost; and various small improvements and fixes.
Install this version with
# Training + Inference
pip install sentence-transformers[train]==3.2.0
# Inference only, use one of:
pip install sentence-transformers==3.2.0
pip install sentence-transformers[onnx-gpu]==3.2.0
pip install sentence-transformers[onnx]==3.2.0
pip install sentence-transformers[openvino]==3.2.0
Introducing a new backend keyword argument to the SentenceTransformer initialization, allowing values of "torch" (default), "onnx", and "openvino".
These come with new installations:
pip install sentence-transformers[onnx-gpu]
# or ONNX for CPU only:
pip install sentence-transformers[onnx]
# or
pip install sentence-transformers[openvino]
It's as simple as:
from sentence_transformers import SentenceTransformer
model = SentenceTransformer("all-MiniLM-L6-v2", backend="onnx")
sentences = ["This is an example sentence", "Each sentence is converted"]
embeddings = model.encode(sentences)
If you specify a backend and your model repository or directory contains an ONNX/OpenVINO model file, it will automatically be used! And if your model repository or directory doesn't have one already, an ONNX/OpenVINO model will be automatically exported. Just remember to model.push_to_hub or model.save_pretrained into the same model repository or directory to avoid having to re-export the model every time.
All keyword arguments passed via model_kwargs will be passed on to ORTModel.from_pretrained or OVBaseModel.from_pretrained. The most useful arguments are:
provider: (Only if backend="onnx") ONNX Runtime provider to use for loading the model, e.g. "CPUExecutionProvider" . See https://onnxruntime.ai/docs/execution-providers/ for possible providers. If not specified, the strongest provider (E.g. "CUDAExecutionProvider") will be used.file_name: The name of the ONNX file to load. If not specified, will default to "model.onnx" or otherwise "onnx/model.onnx" for ONNX, and "openvino_model.xml" and "openvino/openvino_model.xml" for OpenVINO. This argument is useful for specifying optimized or quantized models.export: A boolean flag specifying whether the model will be exported. If not provided, export will be set to True if the model repository or directory does not already contain an ONNX or OpenVINO model.For example:
from sentence_transformers import SentenceTransformer
model = SentenceTransformer(
"all-MiniLM-L6-v2",
backend="onnx",
model_kwargs={
"file_name": "model_O3.onnx",
"provider": "CPUExecutionProvider",
}
)
sentences = ["This is an example sentence", "Each sentence is converted"]
embeddings = model.encode(sentences)
We ran benchmarks for CPU and GPU, averaging findings across 4 models of various sizes, 3 datasets, and numerous batch sizes. Here are the findings:
<p float="left"> <img src="https://github.com/user-attachments/assets/d3f423ff-ad4e-4c91-9beb-8217a062a61d" width="45%" /> <img src="https://github.com/user-attachments/assets/3b9ae402-1127-4152-a925-70c3d626b27d" width="45%" /> </p>
These findings resulted in these recommendations:
For GPU, you can expect 2x speedup with fp16 at no cost, and for CPU you can expect ~2.5x speedup at a cost of 0.4% accuracy.
<details><summary>ONNX Optimization and Quantization</summary>
In addition to exporting default ONNX and OpenVINO models, we also introduce 2 helper methods for optimizing and quantizing ONNX models:
export_optimized_onnx_model: This function uses Optimum to implement several optimizations in the ONNX model, ranging from basic optimizations to approximations and mixed precision. Read about the 4 default options here. This function accepts:
model A SentenceTransformer model loaded with backend="onnx".optimization_config: "O1", "O2", "O3", or "O4" from 🤗 Optimum or a custom OptimizationConfig instance.model_name_or_path: The directory or model repository where the optimized model will be saved.push_to_hub: Whether the push the exported model to the hub with model_name_or_path as the repository name. If False, the model will be saved in the directory specified with model_name_or_path.create_pr: If push_to_hub, then this denotes whether a pull request is created rather than pushing the model directly to the repository. Very useful for optimizing models of repositories that you don't have write access to.file_suffix: The suffix to add to the optimized model file name. Will use the optimization_config string or "optimized" if not set.The usage is like this:
from sentence_transformers import SentenceTransformer, export_optimized_onnx_model
onnx_model = SentenceTransformer("BAAI/bge-large-en-v1.5", backend="onnx")
export_optimized_onnx_model(
model=onnx_model,
optimization_config="O4",
model_name_or_path="BAAI/bge-large-en-v1.5",
push_to_hub=True,
create_pr=True,
)
After which you can load the model with:
from sentence_transformers import SentenceTransformer
pull_request_nr = 2 # TODO: Update this to the number of your pull request
model = SentenceTransformer(
"BAAI/bge-large-en-v1.5",
backend="onnx",
model_kwargs={"file_name": "onnx/model_O4.onnx"},
revision=f"refs/pr/{pull_request_nr}"
)
or when it gets merged:
from sentence_transformers import SentenceTransformer
model = SentenceTransformer(
"BAAI/bge-large-en-v1.5",
backend="onnx",
model_kwargs={"file_name": "onnx/model_O4.onnx"},
)
export_dynamic_quantized_onnx_model: This function uses Optimum to quantize the ONNX model to int8, also allowing for hardware-specific optimizations. This results in impressive speedups for CPUs. In my findings, each of the default quantization configuration options gave approximately the same performance improvements. This function accepts
model A SentenceTransformer model loaded with backend="onnx".quantization_config: "arm64", "avx2", "avx512", or "avx512_vnni" representing quantization configurations from AutoQuantizationConfig, or an QuantizationConfig instance.model_name_or_path: The directory or model repository where the optimized model will be saved.push_to_hub: Whether the push the exported model to the hub with model_name_or_path as the repository name. If False, the model will be saved in the directory specified with model_name_or_path.create_pr: If push_to_hub, then this denotes whether a pull request is created rather than pushing the model directly to the repository. Very useful for quantizing models of repositories that you don't have write access to.file_suffix: The suffix to add to the optimized model file name. Will use the quantization_config string or e.g. "int8_quantized" if not set.The usage is like this:
from sentence_transformers import SentenceTransformer, export_quantized_onnx_model
onnx_model = SentenceTransformer("BAAI/bge-large-en-v1.5", backend="onnx")
export_quantized_onnx_model(
model=onnx_model,
quantization_config="avx512",
model_name_or_path="BAAI/bge-large-en-v1.5",
push_to_hub=True,
create_pr=True,
)
After which you can load the model with:
from sentence_transformers import SentenceTransformer
pull_request_nr = 2 # TODO: Update this to the number of your pull request
model = SentenceTransformer(
"BAAI/bge-large-en-v1.5",
backend="onnx",
model_kwargs={"file_name": "onnx/model_qint8_avx512.onnx"},
revision=f"refs/pr/{pull_request_nr}"
)
or when it gets merged:
from sentence_transformers import SentenceTransformer
model = SentenceTransformer(
"BAAI/bge-large-en-v1.5",
backend="onnx",
model_kwargs={"file_name": "onnx/model_qint8_avx512.onnx"},
)
</details>
If ONNX or OpenVINO isn't fast enough for you yet, then perhaps you'll enjoy Static Embeddings. These embeddings are a bit akin to GLoVe or Word2vec, i.e. they're bags of token embeddings that are summed together to create text embeddings, allowing for lightning-fast embeddings that don't require any neural networks.
However, these Static Embeddings are created in different ways. For example:
Distillation via the Model2Vec technique. This projects allows you to distill any Sentence Transformer model into Static Embeddings. For example, distilling BAAI/bge-base-en-v1.5 resulted in a Static Embeddings Sentence Transformer model that reaches 87.5% of the performance of all-MiniLM-L6-v2 on MTEB (+ PEARL & WordSim) and 97.4% of the performance of all-MiniLM-L6-v2 on various classification benchmarks. You can initialize Static Embeddings via Model2Vec in two ways:
from_model2vec: You can load one of the pretrained Model2Vec models:# note: `pip install model2vec` is needed, but not for inference
from sentence_transformers import SentenceTransformer
from sentence_transformers.models import StaticEmbedding
# Initialize a Sentence Transformer model with a static embedding from a pretrained model2vec model
static_embedding = StaticEmbedding.from_model2vec("minishlab/M2V_multilingual_output")
model = SentenceTransformer(modules=[static_embedding])
# Encode some texts
queries = ["What is the capital of France?", "How many people live in the Netherlands?"]
documents = ["Paris is the capital of France", "The Netherlands has 17 million inhabitants"]
query_embeddings = model.encode(queries)
document_embeddings = model.encode(documents)
# Compute similarities
scores = model.similarity(query_embeddings, document_embeddings)
print(scores)
"""
tensor([[0.8170, 0.3843],
[0.3929, 0.5818]])
"""
from_distillation: You can use the name of any Sentence Transformer model alongside some parameters (See this docs for more information) to perform the distillation yourself, without needing any dataset. On my device, this takes ~4s on a GPU and ~2 minutes on a CPU:# note: `pip install model2vec` is needed, but not for inference
from sentence_transformers import SentenceTransformer
from sentence_transformers.models import StaticEmbedding
# Initialize a Sentence Transformer model with a static embedding by distilling via model2vec
static_embedding = StaticEmbedding.from_distillation(
"mixedbread-ai/mxbai-embed-large-v1",
device="cuda",
pca_dims=256,
apply_zipf=True,
)
model = SentenceTransformer(modules=[static_embedding])
# Encode some texts
queries = ["What is the capital of France?", "How many people live in the Netherlands?"]
documents = ["Paris is the capital of France", "The Netherlands has 17 million inhabitants"]
query_embeddings = model.encode(queries)
document_embeddings = model.encode(documents)
# Compute similarities
scores = model.similarity(query_embeddings, document_embeddings)
print(scores)
"""
tensor([[0.8430, 0.3271],
[0.3213, 0.5861]])
"""
Random initialization: Although this initialization needs finetuning, finetuning a Sentence Transformers model backed by StaticEmbedding is extremely fast. For example, I was able to finetune tomaarsen/static-bert-uncased-gooaq with MatryoshkaLoss & MultipleNegativesRankingLoss on the entire (3 million pairs) gooaq dataset in just 7 minutes. This model reaches a NDCG@10 of 79.33 on a hold-out set of 10k samples from gooaq, whereas e.g. BAAI/bge-base-en-v1.5 reaches 85.01 NDCG@10. In short, only 6.6% less performance for a model that's about 500x faster. That's not a typo: I can compute embeddings for about 14000 stsb sentences from per second on CPU, compared to about ~24 with BAAI/bge-base-en-v1.5, a.k.a. 625x faster.
[!NOTE] You can
save_pretrainedand load these models like any other Sentence Transformer models, theStaticEmbeddinginitialization is only necessary when you're creating a new model.
- Creation:
from sentence_transformers import SentenceTransformer from sentence_transformers.models import StaticEmbedding # Initialize a Sentence Transformer model with a static embedding from a pretrained model2vec model static_embedding = StaticEmbedding.from_distillation( "mixedbread-ai/mxbai-embed-large-v1", device="cuda", pca_dims=256, apply_zipf=True, ) model = SentenceTransformer(modules=[static_embedding]) model.save_pretrained("static-mxbai-embed-large-v1") # or # model.push_to_hub("tomaarsen/static-mxbai-embed-large-v1")- Inference:
from sentence_transformers import SentenceTransformer # Initialize a Sentence Transformer model with a static embedding model = SentenceTransformer("static-mxbai-embed-large-v1") model.encode([...])
InformationRetrievalEvaluator now accepts query_prompt, query_prompt_name, corpus_prompt, and corpus_prompt_name arguments, useful if your model requires specific prompts for queries and/or documents for the best performance. (#2951)mine_hard_negatives function now accepts anchor_column_name and positive_column_name for specifying which dataset columns will be used. If not specified, the first two columns are used, respectively. Additionally, the min_score parameter is added, ensuring that all mined negatives have a similarity score of at least min_score according to the chosen SentenceTransformer or CrossEncoder model. (#2977)CachedGISTEmbedLoss has been improved to support multiple negatives per sample, i.e. the loss now accepts data in the (anchor, positive, negative_1, …, negative_n) format. It is the third loss to support this format (see docs):fix] Only save first module in root if "save_in_root" is specified. by @tomaarsen in https://github.com/UKPLab/sentence-transformers/pull/2957feat] Add query prompts to Information Retrieval Evaluator by @ArthurCamara in https://github.com/UKPLab/sentence-transformers/pull/2951model cards] Keep evaluation order in training logs if there's multiple evaluators by @tomaarsen in https://github.com/UKPLab/sentence-transformers/pull/2963CrossEncoder.rank by @it176131 in https://github.com/UKPLab/sentence-transformers/pull/2947feat] Add lightning-fast StaticEmbedding module based on model2vec by @tomaarsen in https://github.com/UKPLab/sentence-transformers/pull/2961feat] Add ONNX and OpenVINO backends by @helena-intel and @tomaarsen in https://github.com/UKPLab/sentence-transformers/pull/2712Special thanks to @echarlaix for making the new backends possible due to some last-minute changes in optimum and optimum-intel.
Full Changelog: https://github.com/UKPLab/sentence-transformers/compare/v3.1.1...v3.2.0
This patch release fixes hard negatives mining for models that don't automatically normalize their embeddings and it lifts the numpy<2 restriction tha
This patch release fixes hard negatives mining for models that don't automatically normalize their embeddings and it lifts the numpy<2 restriction that was previously required.
Install this version with
# Full installation:
pip install sentence-transformers[train]==3.1.1
# Inference only:
pip install sentence-transformers==3.1.1
The mine_hard_negatives utility introduced in the previous release would fail if use_faiss=True & the model does not automatically normalize its embeddings. This release patches that, allowing the utility to work with all Sentence Transformer models:
from sentence_transformers.util import mine_hard_negatives
from sentence_transformers import SentenceTransformer
from datasets import load_dataset
# Load a Sentence Transformer model
model = SentenceTransformer("mixedbread-ai/mxbai-embed-large-v1").bfloat16()
# Load a dataset to mine hard negatives from
dataset = load_dataset("sentence-transformers/natural-questions", split="train[:10000]")
print(dataset)
"""
Dataset({
features: ['query', 'answer'],
num_rows: 10000
})
"""
# Mine hard negatives
dataset = mine_hard_negatives(
dataset=dataset,
model=model,
range_min=10,
range_max=50,
max_score=0.8,
margin=0.1,
num_negatives=5,
sampling_strategy="random",
batch_size=128,
use_faiss=True,
)
'''
Batches: 100%|██████████████████████████████████████████████████████████████████████████████████████████████████████████████████| 75/75 [00:21<00:00, 3.51it/s]
Batches: 100%|██████████████████████████████████████████████████████████████████████████████████████████████████████████████████| 79/79 [00:03<00:00, 25.77it/s]
Querying FAISS index: 100%|███████████████████████████████████████████████████████████████████████████████████████████████████████| 1/1 [00:00<00:00, 3.98it/s]
Metric Positive Negative Difference
Count 10,000 47,711
Mean 0.7600 0.5376 0.2299
Median 0.7673 0.5379 0.2274
Std 0.0658 0.0387 0.0629
Min 0.3858 0.3732 0.1044
25% 0.7219 0.5129 0.1833
50% 0.7673 0.5379 0.2274
75% 0.8058 0.5617 0.2724
Max 0.9341 0.7024 0.4780
Skipped 48770 potential negatives (9.56%) due to the margin of 0.1.
Could not find enough negatives for 2289 samples (4.58%). Consider adjusting the range_max, range_min, margin and max_score parameters if you'd like to find more valid negatives.
'''
print(dataset)
'''
Dataset({
features: ['query', 'answer', 'negative'],
num_rows: 47711
})
'''
print(dataset[0])
'''
{
'query': 'where is the us navy base in japan located',
'answer': 'United States Fleet Activities Yokosuka The United States Fleet Activities Yokosuka (横須賀海 軍施設, Yokosuka kaigunshisetsu) or Commander Fleet Activities Yokosuka (司令官艦隊活動横須賀, Shirei-kan kantai katsudō Yokosuka) is a United States Navy base in Yokosuka, Japan. Its mission is to maintain and operate base facilities for the logistic, recreational, administrative support and service of the U.S. Naval Forces Japan, Seventh Fleet and other operating forces assigned in the Western Pacific. CFAY is the largest strategically important U.S. naval installation in the western Pacific.[1] As of August 2013[update], it was commanded by Captain David Glenister.',
'negative': "2011 Tōhoku earthquake and tsunami The earthquake took place at 14:46 JST (UTC 05:46) around 67\xa0km (42\xa0mi) from the nearest point on Japan's coastline, and initial estimates indicated the tsunami would have taken 10 to 30\xa0minutes to reach the areas first affected, and then areas farther north and south based on the geography of the coastline.[127][128] Just over an hour after the earthquake at 15:55 JST, a tsunami was observed flooding Sendai Airport, which is located near the coast of Miyagi Prefecture,[129][130] with waves sweeping away cars and planes and flooding various buildings as they traveled inland.[131][132] The impact of the tsunami in and around Sendai Airport was filmed by an NHK News helicopter, showing a number of vehicles on local roads trying to escape the approaching wave and being engulfed by it.[133] A 4-metre-high (13\xa0ft) tsunami hit Iwate Prefecture.[134] Wakabayashi Ward in Sendai was also particularly hard hit.[135] At least 101 designated tsunami evacuation sites were hit by the wave.[136]"
}
'''
dataset.push_to_hub("natural-questions-hard-negatives", "triplet")
Thanks to @omarnj-lab for pointing out the bug to me.
The v3.1.0 Sentence Transformers release required numpy<2 to prevent crashes on Windows. However, various third-parties (e.g. scipy) have now been recompiled & released, allowing the Windows tests to pass again.
If you experience the following snippet:
A module that was compiled using NumPy 1.x cannot be run in NumPy 2.0.0 as it may crash. To support both 1.x and 2.x versions of NumPy, modules must be compiled with NumPy 2.0. Some module may need to rebuild instead e.g. with 'pybind11>=2.12'. If you are a user of the module, the easiest solution will be to downgrade to 'numpy<2' or try to upgrade the affected module. We expect that some modules will need time to support NumPy 2.
Then consider 1) upgrading the dependency from which the error occurred or 2) downgrading numpy to below v2:
pip install -U numpy<2
Thanks to @kozlek for pointing this out to me and helping getting it resolved.
deps] Attempt to remove numpy restrictions by @tomaarsen in https://github.com/UKPLab/sentence-transformers/pull/2937metadata] Extend pyproject.toml metadata by @tomaarsen in https://github.com/UKPLab/sentence-transformers/pull/2943fix] Ensure that the embeddings from hard negative mining are normalized by @tomaarsen in https://github.com/UKPLab/sentence-transformers/pull/2944Full Changelog: https://github.com/UKPLab/sentence-transformers/compare/v3.1.0...v3.1.1
[deprecation] Push deprecation cycle for use_auth_token to v4 by @tomaarsen in https://github.com/UKPLab/sentence-transformers/pull/2926
This release introduces a hard negatives mining utility to get better models out of your data, a new strong loss function for symmetric tasks, training with streaming datasets to avoid having to store datasets fully on disk, custom modules to allow for more creativity from model authors, and many bug fixes, small additions and documentation improvements.
Install this version with
# Full installation:
pip install sentence-transformers[train]==3.1.0
# Inference only:
pip install sentence-transformers==3.1.0
[!WARNING] Due to incompatibilities with Windows, we have set
numpy<2in the Sentence Transformers requirements. If you're not on Windows, you can still installnumpy>=2and everything should work as expected.
Hard negatives are texts that are rather similar to some anchor text (e.g. a question), but are not the correct match. For example:
These negatives are more difficult for a model to distinguish from the correct answer, leading to a stronger training signal and a stronger overall model when used with one of the Loss Functions that accepts (anchor, positive, negative) pairs such as the one above.
This release introduces a utility function called mine_hard_negatives that allows you to mine for these hard negatives given a (anchor, positive) dataset (and optionally a corpus of negative candidate texts).
It boasts the following features to give you fine-grained control over the similarity of the mined negatives relative to the anchor:
margin of the true similarity between anchor and positive.max_score.2 + num_negatives-tuples.from sentence_transformers.util import mine_hard_negatives
from sentence_transformers import SentenceTransformer
from datasets import load_dataset
# Load a Sentence Transformer model
model = SentenceTransformer("all-MiniLM-L6-v2")
# Load a dataset to mine hard negatives from
dataset = load_dataset("sentence-transformers/natural-questions", split="train")
print(dataset)
"""
Dataset({
features: ['query', 'answer'],
num_rows: 100231
})
"""
# Mine hard negatives
dataset = mine_hard_negatives(
dataset=dataset,
model=model,
range_min=10,
range_max=50,
max_score=0.8,
margin=0.1,
num_negatives=5,
sampling_strategy="random",
batch_size=128,
use_faiss=True,
)
'''
Batches: 100%|████████████████████████████████████████████████████████████████████████████████████████████████████████████████| 588/588 [00:33<00:00, 17.37it/s]
Batches: 100%|███████████████████████████████████████████████████████████████████████████████████████████████████████████████| 784/784 [00:07<00:00, 101.55it/s]
Querying FAISS index: 100%|███████████████████████████████████████████████████████████████████████████████████████████████████████| 7/7 [00:07<00:00, 1.06s/it]
Metric Positive Negative Difference
Count 100,231 460,725 460,725
Mean 0.6866 0.4133 0.2917
Median 0.7010 0.4059 0.2873
Std 0.1125 0.0673 0.1006
Min 0.0303 0.1638 0.1029
25% 0.6221 0.3649 0.2112
50% 0.7010 0.4059 0.2873
75% 0.7667 0.4561 0.3647
Max 0.9584 0.7362 0.7073
Skipped 882722 potential negatives (17.27%) due to the margin of 0.1.
Skipped 27 potential negatives (0.00%) due to the maximum score of 0.8.
Could not find enough negatives for 40430 samples (8.07%). Consider adjusting the range_max, range_min, margin and max_score parameters if you'd like to find more valid negatives.
'''
print(dataset)
'''
Dataset({
features: ['query', 'answer', 'negative'],
num_rows: 460725
})
'''
print(dataset[0])
'''
{
'query': 'the first person to use the word geography was',
'answer': 'History of geography The history of geography includes many histories of geography which have differed over time and between different cultural and political groups. In more recent developments, geography has become a distinct academic discipline. \'Geography\' derives from the Greek γεωγραφία – geographia,[1] a literal translation of which would be "to describe or write about the Earth". The first person to use the word "geography" was Eratosthenes (276–194 BC). However, there is evidence for recognizable practices of geography, such as cartography (or map-making) prior to the use of the term geography.',
'negative': 'Terminology of the British Isles The word "Great" means "larger", in comparison with Brittany in modern-day France. One historical term for the peninsula in France that largely corresponds to the modern French province is Lesser or Little Britain. That region was settled by many British immigrants during the period of Anglo-Saxon migration into Britain, and named "Little Britain" by them. The French term "Bretagne" now refers to the French "Little Britain", not to the British "Great Britain", which in French is called Grande-Bretagne. In classical times, the Graeco-Roman geographer Ptolemy in his Almagest also called the larger island megale Brettania (great Britain). At that time, it was in contrast to the smaller island of Ireland, which he called mikra Brettania (little Britain).[62] In his later work Geography, Ptolemy refers to Great Britain as Albion and to Ireland as Iwernia. These "new" names were likely to have been the native names for the islands at the time. The earlier names, in contrast, were likely to have been coined before direct contact with local peoples was made.[63]'
}
'''
dataset.push_to_hub("natural-questions-hard-negatives", "triplet")
This dataset can immediately be used in conjunction with MultipleNegativesRankingLoss, likely resulting in a stronger model than if you had just used the natural-questions dataset outright.
Here are some example datasets that I created using this new function:
Big thanks to @ChrisGeishauser and @ArthurCamara for assisting with this feature.
Let's break this down:
The v3.1 Sentence Transformers release now introduces a new loss: CachedMultipleNegativesSymmetricRankingLoss (CMNSRL), which combines both of the previous adaptations. The result is a loss adept at symmetric training tasks for which you can pick an arbitrarily large batch size. It is likely the strongest loss for Semantic Textual Similarity (STS) tasks in Sentence Transformers now. Big thanks to @madhavthaker1 for working to include it.
The v3.1 release introduces support for training with datasets.IterableDataset (Differences between Dataset and IterableDataset docs). This means that you can train without first downloading the full dataset to disk. For example:
from datasets import load_dataset
# Load a streaming dataset to finetune on
train_dataset = load_dataset("sentence-transformers/gooaq", split="train", streaming=True)
# IterableDataset({
# features: ['question', 'answer'],
# n_shards: 2
# })
or
from datasets import IterableDataset, Value, Features
def dataset_generator_fn():
# Gather, fetch, load, or generate data here
for ... in ...:
yield ...
train_dataset = IterableDataset.from_generator(dataset_generator_fn)
train_dataset = train_dataset.cast(Features({'question': Value(dtype='string', id=None), 'answer': Value(dtype='string', id=None)}))
(Read more about Dataset features here)
For a full example of training with a streaming dataset, consider this script:
import logging
from datasets import load_dataset
from sentence_transformers import (
SentenceTransformer,
SentenceTransformerTrainer,
SentenceTransformerTrainingArguments,
SentenceTransformerModelCardData,
)
from sentence_transformers.losses import MultipleNegativesRankingLoss
from sentence_transformers.training_args import BatchSamplers
logging.basicConfig(
format="%(asctime)s - %(message)s", datefmt="%Y-%m-%d %H:%M:%S", level=logging.INFO
)
# 1. Load a model to finetune with 2. (Optional) model card data
model = SentenceTransformer(
"microsoft/mpnet-base",
model_card_data=SentenceTransformerModelCardData(
language="en",
license="apache-2.0",
model_name="MPNet base trained on GooAQ pairs",
),
)
name = "mpnet-base-gooaq-streaming"
# 2. Load a streaming dataset to finetune on
train_dataset = load_dataset("sentence-transformers/gooaq", split="train", streaming=True)
# 3. Define a loss function
loss = MultipleNegativesRankingLoss(model)
# 4. (Optional) Specify training arguments
train_batch_size = 64
args = SentenceTransformerTrainingArguments(
# Required parameter:
output_dir=f"models/{name}",
# Optional training parameters:
num_train_epochs=1,
per_device_train_batch_size=train_batch_size,
learning_rate=2e-5,
warmup_ratio=0.1,
fp16=False, # Set to False if you get an error that your GPU can't run on FP16
bf16=True, # Set to True if you have a GPU that supports BF16
batch_sampler=BatchSamplers.NO_DUPLICATES, # MultipleNegativesRankingLoss benefits from no duplicate samples in a batch
# Optional tracking/debugging parameters:
save_strategy="steps",
save_steps=100,
save_total_limit=2,
logging_steps=250,
logging_first_step=True,
run_name=name, # Will be used in W&B if `wandb` is installed
)
# 5. Create a trainer & train
trainer = SentenceTransformerTrainer(
model=model,
args=args,
train_dataset=train_dataset,
loss=loss,
)
trainer.train()
# 6. Save the trained model
model.save_pretrained(f"models/{name}/final")
# 7. (Optional) Push it to the Hugging Face Hub
model.push_to_hub(name)
Sentence Transformer models consist of several modules that are executed sequentially. Most models consist of a Transformer module, a Pooling module, and perhaps a Dense and/or Normalize module. However, as of the v3.1 release, model authors can create their own modules by writing some custom modeling code. This code can be uploaded to the Hugging Face Hub alongside the model itself, after which users can load the model like normal.
This allows for authors to replace the Transformer module with one that includes model-specific quirks, or replace the Pooling module with an all-new pooling method. This even allows for multi-modal models as authors can customize the preprocessing of the first module.
jinaai/jina-clip-v1 is the first model to take advantage of this new feature, allowing you to encode both texts and images (via paths to local images or URLs) due to their custom preprocessing. Try it out yourself:
from sentence_transformers import SentenceTransformer
# Load the model; must use trust_remote_code=True to run the custom module
model = SentenceTransformer("jinaai/jina-clip-v1", trust_remote_code=True)
# Texts and images of blue and red cats to embed
sentences = ['A blue cat', 'A red cat']
image_urls = [
'https://i.pinimg.com/600x315/21/48/7e/21487e8e0970dd366dafaed6ab25d8d8.jpg',
'https://i.pinimg.com/736x/c9/f2/3e/c9f23e212529f13f19bad5602d84b78b.jpg'
]
# Embed the texts and images like normal
text_embeddings = model.encode(sentences)
image_embeddings = model.encode(image_urls)
# Compute similarity between text embeddings:
print(model.similarity(text_embeddings[0], text_embeddings[1]))
# tensor([[✅0.5636]])
# or cross-modal text and image embeddings:
print(model.similarity(text_embeddings, image_embeddings))
# tensor([[✅0.2906, ❌0.0569],
# [❌0.1277, ✅0.2916]]
Additionally, model authors can take advantage of keyword argument passthrough. By updating the modules.json file to include a list of kwargs, e.g.:
[
{
"idx": 0,
"name": "0",
"path": "",
"type": "custom_transformer.CustomTransformer",
"kwargs": ["task_type"]
},
...
]
then if a user provides the task_type keyword argument in model.encode, this value will be propagated to the forward of the custom module(s). This way, users can specify some custom functionality on the fly during inference time (as well as during load time via the model_kwargs option when initializing a SentenceTransformer model).
numpy<2.0.0 due to issues with torch and numpy interoperability on Windows.transformers version to 4.38.0 & huggingface-hub to 0.19.3 to prevent a training crash related to the prefetch_factor optionshow_progress_bar to encode_multi_process (#2762)revision to push_to_hub (#2902)cache_dir and config_args to CrossEncoder (#2784)GISTEmbedLoss with DataParallel (DP) and DataDistributedParallel (DDP) (#2772)GroupByLabelBatchSampler resulting in some data not being used in training (#2788)datasets directory exists locally (#2859)Matryoshka2dLoss not importing correctly (#2907)dataloader_drop_last=True (#2877)torch_compile=True not working in the SentenceTransformersTrainingArguments: should now work for faster training (#2884)SoftmaxLoss performing worse since v3.0 as a Linear layer was ignored by the optimizer (#2881)trainer.train(resume_from_checkpoint="...") with custom models (i.e. trust_remote_code) (#2918)model_kwargs={"torch_dtype": torch.float16} with models that use Dense layers (#2889)versions] Increment transformers/hf-hub versions to prevent training crash by @tomaarsen in https://github.com/UKPLab/sentence-transformers/pull/2757fix] Prevent crash when encoding empty list by @tomaarsen in https://github.com/UKPLab/sentence-transformers/pull/2759feat] Add show_progress_bar to encode_multi_process by @tomaarsen in https://github.com/UKPLab/sentence-transformers/pull/2762fix] Fix retokenization on DDP/DP with GIST losses by @tomaarsen in https://github.com/UKPLab/sentence-transformers/pull/2775Warning (W) by @fpgmaas in https://github.com/UKPLab/sentence-transformers/pull/2789feat] Add hard negatives mining utility by @tomaarsen in https://github.com/UKPLab/sentence-transformers/pull/2768chore] Clean-up .gitignore by @fpgmaas in https://github.com/UKPLab/sentence-transformers/pull/2799setup.py and setup.cfg to pyproject.toml by @fpgmaas in https://github.com/UKPLab/sentence-transformers/pull/2786pytest-cov and add test coverage command to the Makefile by @fpgmaas in https://github.com/UKPLab/sentence-transformers/pull/2794pytest config to pyproject.toml and remove pytest.ini by @fpgmaas in https://github.com/UKPLab/sentence-transformers/pull/2819fix] Fix packages discovery in pyproject.toml by @fpgmaas in https://github.com/UKPLab/sentence-transformers/pull/2825ruff pre-commit hook. by @fpgmaas in https://github.com/UKPLab/sentence-transformers/pull/2826chore] Enable isort with ruff by @fpgmaas in https://github.com/UKPLab/sentence-transformers/pull/2828chore] Enable ruff rules UP006 and UP007 to improve type hints. by @fpgmaas in https://github.com/UKPLab/sentence-transformers/pull/2830chore] Enable ruff's pypgrade (UP) ruleset by @fpgmaas in https://github.com/UKPLab/sentence-transformers/pull/2834is_<library>_available functions by @leblancfg in https://github.com/UKPLab/sentence-transformers/pull/2859style] Replace Huggingface with Hugging Face by @tomaarsen in https://github.com/UKPLab/sentence-transformers/pull/2905ci] Attempt to fix CI disk space issues by @tomaarsen in https://github.com/UKPLab/sentence-transformers/pull/2906docs] Fix typo and broken links in documentation by @ZiyiXia in https://github.com/UKPLab/sentence-transformers/pull/2861chore] Add unittests for InformationRetrievalEvaluator by @fpgmaas in https://github.com/UKPLab/sentence-transformers/pull/2838fix] Safely continue if ProportionalBatchSampler sub-batch sampler throws StopIteration by @tomaarsen in https://github.com/UKPLab/sentence-transformers/pull/2877fix] Fix torch_compile=True by always inserting a wrapped model into the loss by @tomaarsen in https://github.com/UKPLab/sentence-transformers/pull/2884fix] Fix SoftmaxLoss by initializing the optimizer over the loss(es) rather than the model by @tomaarsen in https://github.com/UKPLab/sentence-transformers/pull/2881fix] Fix trainer.train(resume_from_checkpoint="...") with custom models (i.e. trust_remote_code) by @tomaarsen in https://github.com/UKPLab/sentence-transformers/pull/2918docs] Heavily extend sampler documentation by @tomaarsen in https://github.com/UKPLab/sentence-transformers/pull/2921feat] Add support for streaming datasets by @tomaarsen in https://github.com/UKPLab/sentence-transformers/pull/2792fix] Change eval dataloader to use eval_batch_size by @akashd-2 in https://github.com/UKPLab/sentence-transformers/pull/2847feat] Add cache_dir support to CrossEncoder by @RoyBA in https://github.com/UKPLab/sentence-transformers/pull/2784deprecation] Push deprecation cycle for use_auth_token to v4 by @tomaarsen in https://github.com/UKPLab/sentence-transformers/pull/2926security] Load weights only with torch.load & pytorch_model.bin by @tomaarsen in https://github.com/UKPLab/sentence-transformers/pull/2927feat] Allow loading custom modules; encode kwargs passthrough to modules by @tomaarsen in https://github.com/UKPLab/sentence-transformers/pull/2773fix] Add dtype cast for modules other than Transformer by @ir2718 in https://github.com/UKPLab/sentence-transformers/pull/2889docs] Move losses up in the package reference; they're more important by @tomaarsen in https://github.com/UKPLab/sentence-transformers/pull/2929feat] Add column order warnings to the data collator by @tomaarsen in https://github.com/UKPLab/sentence-transformers/pull/2928Big thanks to @fpgmaas for the large number of valuable contributions surrounding tests, CI, config files, and overall project health.
Full Changelog: https://github.com/UKPLab/sentence-transformers/compare/v3.0.1...v3.1.0
This patch release introduces some improvements for the SentenceTransformerTrainer, as well as some updates for the automatic model card generation. I
This patch release introduces some improvements for the SentenceTransformerTrainer, as well as some updates for the automatic model card generation. It also patches some minor evaluator bugs and a bug with MatryoshkaLoss. Lastly, every single Sentence Transformer model can now be saved and loaded with the safer model.safetensors files.
Install this version with
# Full installation:
pip install sentence-transformers[train]==3.0.1
# Inference only:
pip install sentence-transformers==3.0.1
push_to_hub=True Training Argument, also implement trainer.push_to_hub(...) (#2718)This patch release improves on the automatically generated model cards in several ways:
generated_from_trainer tag is now also added (#2710)...
# Download from the 🤗 Hub
model = SentenceTransformer("sentence_transformers_model_id")
...
This now gets replaced with your new model ID on Hugging Face (#2714)
SequentialEvaluator would be ignored in the scores calculation (#2700)print_wrong_matches=True (#1894)primary_metric in InformationRetrievalEvaluator (#2701)SequentialEvaluator (#2717)MatryoshkaLoss crash if the first dimension is not the biggest (#2719)pytorch_model.bin anymore (#2722)fix] Always override the originally saved version in the ST config by @tomaarsen in https://github.com/UKPLab/sentence-transformers/pull/2709model cards] Also include HF datasets in the model card metadata by @tomaarsen in https://github.com/UKPLab/sentence-transformers/pull/2711model cards] Improve the widget example selection: not based on embeddings, better for QA by @tomaarsen in https://github.com/UKPLab/sentence-transformers/pull/2713model cards] Replace 'sentence_transformers_model_id' from reused model if possible by @tomaarsen in https://github.com/UKPLab/sentence-transformers/pull/2714feat] Allow passing a list of evaluators to the Trainer by @tomaarsen in https://github.com/UKPLab/sentence-transformers/pull/2716fix] Fix gradient checkpointing to allow for much lower memory usage by @tomaarsen in https://github.com/UKPLab/sentence-transformers/pull/2717fix] Implement create_model_card on the Trainer, allowing args.push_to_hub=True by @tomaarsen in https://github.com/UKPLab/sentence-transformers/pull/2718fix] Fix MatryoshkaLoss crash if the first dimension is not the biggest by @tomaarsen in https://github.com/UKPLab/sentence-transformers/pull/2719typing] Improve typing for many functions & add py.typed to satisfy mypy by @tomaarsen in https://github.com/UKPLab/sentence-transformers/pull/2724fix] Fix edge case with evaluator being None by @tomaarsen in https://github.com/UKPLab/sentence-transformers/pull/2726simplify] Set can_return_loss=True globally, instead of via the data collator by @tomaarsen in https://github.com/UKPLab/sentence-transformers/pull/2727feat] Integrate safetensors with Dense, etc. modules too. by @tomaarsen in https://github.com/UKPLab/sentence-transformers/pull/2722model cards] Specify the exact dataset size as a tag, will be bucketized by HF by @tomaarsen in https://github.com/UKPLab/sentence-transformers/pull/2728Full Changelog: https://github.com/UKPLab/sentence-transformers/compare/v3.0.0...v3.0.1
…`SentenceTransformer.fit` method. Rather than deprecating this method, starting from v3.0, this method will use the `SentenceTransformerTrainer` behin…
This release consists of a major refactor that overhauls the training approach (introducing multi-gpu training, bf16, loss logging, callbacks, and much more), adds convenient similarity and similarity_pairwise methods, adds extra keyword arguments, introduces Hyperparameter Optimization, and includes a massive reformatting and release of 50+ datasets for training embedding models. In total, this is the largest Sentence Transformers update since the project was first created.
Install this version with
# Full installation:
pip install sentence-transformers[train]==3.0.0
# Inference only:
pip install sentence-transformers==3.0.0
The v3.0 release centers around this huge modernization of the training approach for SentenceTransformer models. Whereas training before v3.0 used to be all about InputExample, DataLoader and model.fit, the new training approach relies on 5 new components. You can learn more about these components in our Training and Finetuning Embedding Models with Sentence Transformers v3 blogpost. Additionally, you can read the new Training Overview, check out the Training Examples, or read this summary:
Dataset or DatasetDict. This class is much more suited for sharing & efficient modifications than lists/DataLoaders of InputExample instances. A Dataset can contain multiple text columns that will be fed in order to the corresponding loss function. So, if the loss expects (anchor, positive, negative) triplets, then your dataset should also have 3 columns. The names of these columns are irrelevant. If there is a "label" or "score" column, it is treated separately, and used as the labels during training.
A DatasetDict can be used to train with multiple datasets at once, e.g.:DatasetDict({
multi_nli: Dataset({
features: ['premise', 'hypothesis', 'label'],
num_rows: 392702
})
snli: Dataset({
features: ['snli_premise', 'hypothesis', 'label'],
num_rows: 549367
})
stsb: Dataset({
features: ['sentence1', 'sentence2', 'label'],
num_rows: 5749
})
})
When a DatasetDict is used, the loss parameter to the SentenceTransformerTrainer must also be a dictionary with these dataset keys, e.g.:{
'multi_nli': SoftmaxLoss(...),
'snli': SoftmaxLoss(...),
'stsb': CosineSimilarityLoss(...),
}
SentenceEvaluator instance. Unlike before, models can now be evaluated both on an evaluation dataset with some loss function and/or a SentenceEvaluator instance.SentenceTransformersTrainer instance based on the transformers Trainer. This instance is provided with a SentenceTransformer model, a SentenceTransformerTrainingArguments class, a SentenceEvaluator, a training and evaluation Dataset/DatasetDict and a loss function/dict of loss functions. Most of these parameters are optional. Once provided, all you have to do is call trainer.train().Some of the major features that are now implemented include:
This script is a minimal example (no evaluator, no training arguments) of training mpnet-base on a part of the all-nli dataset using MultipleNegativesRankingLoss:
from datasets import load_dataset
from sentence_transformers import SentenceTransformer, SentenceTransformerTrainer
from sentence_transformers.losses import MultipleNegativesRankingLoss
# 1. Load a model to finetune
model = SentenceTransformer("microsoft/mpnet-base")
# 2. Load a dataset to finetune on
dataset = load_dataset("sentence-transformers/all-nli", "triplet")
train_dataset = dataset["train"].select(range(10_000))
eval_dataset = dataset["dev"].select(range(1_000))
# 3. Define a loss function
loss = MultipleNegativesRankingLoss(model)
# 4. Create a trainer & train
trainer = SentenceTransformerTrainer(
model=model,
train_dataset=train_dataset,
eval_dataset=eval_dataset,
loss=loss,
)
trainer.train()
# 5. Save the trained model
model.save_pretrained("models/mpnet-base-all-nli")
Additionally, trained models now automatically produce extensive model cards. Each of the following models were trained using some script from the Training Examples, and the model cards were not edited manually whatsoever:
Prior to the Sentence Transformer v3 release, all models would be trained using the SentenceTransformer.fit method. Rather than deprecating this method, starting from v3.0, this method will use the SentenceTransformerTrainer behind the scenes. This means that your old training code should still work, and should even be upgraded with the new features such as multi-gpu training, loss logging, etc. That said, the new training approach is much more powerful, so it is recommended to write new training scripts using the new approach.
Many of the old training scripts were updated to use the new Trainer-based approach, but not all have been updated yet. We accept help via Pull Requests to assist in updating the scripts.
Sentence Transformers v3.0 introduces two new useful methods:
and one property:
These can be used to calculate the similarity between embeddings, and to specify which similarity function should be used, for example:
>>> from sentence_transformers import SentenceTransformer
>>> model = SentenceTransformer("all-mpnet-base-v2")
>>> sentences = [
... "The weather is so nice!",
... "It's so sunny outside.",
... "He's driving to the movie theater.",
... "She's going to the cinema.",
... ]
>>> embeddings = model.encode(sentences, normalize_embeddings=True)
>>> model.similarity(embeddings, embeddings)
tensor([[1.0000, 0.7235, 0.0290, 0.1309],
[0.7235, 1.0000, 0.0613, 0.1129],
[0.0290, 0.0613, 1.0000, 0.5027],
[0.1309, 0.1129, 0.5027, 1.0000]])
>>> model.similarity_fn_name
"cosine"
>>> model.similarity_fn_name = "euclidean"
>>> model.similarity(embeddings, embeddings)
tensor([[-0.0000, -0.7437, -1.3935, -1.3184],
[-0.7437, -0.0000, -1.3702, -1.3320],
[-1.3935, -1.3702, -0.0000, -0.9973],
[-1.3184, -1.3320, -0.9973, -0.0000]])
Additionally, you can compute the similarity between pairs of embeddings, resulting in a 1-dimensional vector of similarities rather than a 2-dimensional matrix:
>>> model = SentenceTransformer("all-mpnet-base-v2")
>>> sentences = [
... "The weather is so nice!",
... "It's so sunny outside.",
... "He's driving to the movie theater.",
... "She's going to the cinema.",
... ]
>>> embeddings = model.encode(sentences, normalize_embeddings=True)
>>> model.similarity_pairwise(embeddings[::2], embeddings[1::2])
tensor([0.7235, 0.5027])
>>> model.similarity_fn_name
"cosine"
>>> model.similarity_fn_name = "euclidean"
>>> model.similarity_pairwise(embeddings[::2], embeddings[1::2])
tensor([-0.7437, -0.9973])
The similarity_fn_name can now be specified via the SentenceTransformer like so:
from sentence_transformers import SentenceTransformer
model = SentenceTransformer("sentence-transformers/multi-qa-mpnet-base-dot-v1", similarity_fn_name="dot")
Valid options include "cosine" (default), "dot", "euclidean", "manhattan". The chosen similarity_fn_name will also be saved into the model configuration, and loaded automatically. For example, the msmarco-distilbert-dot-v5 model was trained to work best with dot, so we've configured it to use that similarity_fn_name in its configuration:
>>> from sentence_transformers import SentenceTransformer
>>> model = SentenceTransformer("sentence-transformers/msmarco-distilbert-dot-v5")
>>> model.similarity_fn_name
'dot'
Big thanks to @ir2718 for helping set up this major feature.
model_kwargs, tokenizer_kwargs, and config_kwargs to SentenceTransformer (#2578)To those familiar with the internals of Sentence Transformers, you might know that internally, we call AutoModel.from_pretrained, AutoTokenizer.from_pretrained and AutoConfig.from_pretrained from transformers.
Each of these are rather powerful, and they are constantly improved with new features. For example, the AutoModel keyword arguments include:
torch_dtype - this allows you to immediately load a model in bfloat16 or float16 (or "auto", i.e. whatever the model was stored in), which can speed up inference a lot.quantization_configattn_implementation - all models support "eager", but some also support the much faster "fa2" (Flash Attention 2) and "sdpa" (Scaled Dot Product Attention).These options allow for speeding up the model inference. Additionally, via AutoConfig you can update the model configuration, e.g. updating the dropout probability during training, and with AutoTokenizer you can disable the fast Rust-based tokenizer if you're having issues with it via use_fast=False.
Due to how useful these options can be, the following arguments are added to SentenceTransformer:
model_kwargs for AutoModel.from_pretrained keyword argumentstokenizer_kwargs for AutoTokenizer.from_pretrained keyword argumentsconfig_kwargs for AutoConfig.from_pretrained keyword argumentsYou can use it like so:
from sentence_transformers import SentenceTransformer
import torch
model = SentenceTransformer(
"mixedbread-ai/mxbai-embed-large-v1",
model_kwargs={"torch_dtype": torch.bfloat16, "attn_implementation": "sdpa"},
config_kwargs={"hidden_dropout_prob": 0.3},
)
embeddings = model.encode(["He drove his yellow car to the beach.", "He played football with his friends."])
print(embeddings.shape)
Big thanks to @satyamk7054 for starting this work.
Sentence Transformers v3.0 introduces Hyperparameter Optimization (HPO) by extending the transformers HPO support. We recommend reading the all new Hyperparameter Optimization for many more details.
Alongside Sentence Transformers v3.0, we reformat and release 50+ useful datasets in our Embedding Model Datasets Collection on Hugging Face. These can be used with at least one loss function in Sentence Transformers v3.0 out of the box. We recommend browsing through these to see if there are datasets akin to your use cases - training a model on them might just produce large gains on your task(s).
The MSELoss now accepts multiple text columns for each label (where each label is a target/gold embedding), rather than only accepting one text column. This is extremely powerful for following the excellent Multilingual Models strategy to convert a monolingual model into a multilingual one. You can now conveniently train both English and (identical but translated) non-English texts to represent the same embedding (that was generated by a powerful English embedding model).
local_files_only argument to SentenceTransformer & CrossEncoder (#2603)You can now initialize a SentenceTransformer and CrossEncoder with local_files_only. If True, then it will not try and download a model from Hugging Face, it will only look in the local filesystem for the model or try and load it from a cache.
Thanks @debanjum for this change.
v3] Training refactor - MultiGPU, loss logging, bf16, etc. by @tomaarsen in https://github.com/UKPLab/sentence-transformers/pull/2449v3] Add similarity and similarity_pairwise methods to Sentence Transformers by @tomaarsen in https://github.com/UKPLab/sentence-transformers/pull/2615v3] Fix various model card errors by @tomaarsen in https://github.com/UKPLab/sentence-transformers/pull/2616v3] Fix trainer compute_loss when evaluating/predicting if the loss updated the inputs in-place by @tomaarsen in https://github.com/UKPLab/sentence-transformers/pull/2617v3] Never return None in infer_datasets, could result in crash by @tomaarsen in https://github.com/UKPLab/sentence-transformers/pull/2620v3] Trainer: Implement resume from checkpoint support by @tomaarsen in https://github.com/UKPLab/sentence-transformers/pull/2621trust_remote_code to CrossEncoder.tokenizer by @michaelfeil in https://github.com/UKPLab/sentence-transformers/pull/2623v3] Update example scripts to the new v3 training format by @tomaarsen in https://github.com/UKPLab/sentence-transformers/pull/2622v3] Remove "return_outputs" as it's not strictly necessary. Avoids OOM & speeds up training by @tomaarsen in https://github.com/UKPLab/sentence-transformers/pull/2633v3] Fix crash from inferring the dataset_id from a local dataset by @tomaarsen in https://github.com/UKPLab/sentence-transformers/pull/2636v3] Fix multilingual conversion script; extend MSELoss to multi-column by @tomaarsen in https://github.com/UKPLab/sentence-transformers/pull/2641v3] Update evaluation scripts to use HF Datasets by @tomaarsen in https://github.com/UKPLab/sentence-transformers/pull/2642b1 quantization for USearch by @ashvardanian in https://github.com/UKPLab/sentence-transformers/pull/2644v3] Fix resume_from_checkpoint by also updating the loss model by @tomaarsen in https://github.com/UKPLab/sentence-transformers/pull/2648v3] Fix backwards pass on MSELoss due to in-place update by @tomaarsen in https://github.com/UKPLab/sentence-transformers/pull/2647v3] Simplify load_from_checkpoint using load_state_dict by @tomaarsen in https://github.com/UKPLab/sentence-transformers/pull/2650v3] Use torch.arange instead of torch.tensor(range(...)) by @tomaarsen in https://github.com/UKPLab/sentence-transformers/pull/2651v3] Resolve inplace modification error in DDP by @tomaarsen in https://github.com/UKPLab/sentence-transformers/pull/2654v3] Add hyperparameter optimization support by letting loss be a Callable that accepts a model by @tomaarsen in https://github.com/UKPLab/sentence-transformers/pull/2655v3] Add tag hinting at the number of training samples by @tomaarsen in https://github.com/UKPLab/sentence-transformers/pull/2660v3] For the Cached losses; ignore gradients if grad is disabled (e.g. eval) by @tomaarsen in https://github.com/UKPLab/sentence-transformers/pull/2668docs] Rewrite the https://sbert.net documentation by @tomaarsen in https://github.com/UKPLab/sentence-transformers/pull/2632v3] Chore - include import sorting in ruff by @tomaarsen in https://github.com/UKPLab/sentence-transformers/pull/2672v3] Prevent warning with 'model.fit' with transformers >= 4.41.0 due to evaluation_strategy by @tomaarsen in https://github.com/UKPLab/sentence-transformers/pull/2673v3] Add various useful Sphinx packages (copy code, link to code, nicer tabs) by @tomaarsen in https://github.com/UKPLab/sentence-transformers/pull/2674v3] Make the "primary_metric" for evaluators a bit more robust by @tomaarsen in https://github.com/UKPLab/sentence-transformers/pull/2675v3] Set broadcast_buffers = False when training with DDP by @tomaarsen in https://github.com/UKPLab/sentence-transformers/pull/2663v3] Warn about using DP instead of DDP + set dataloader_drop_last with DDP by @tomaarsen in https://github.com/UKPLab/sentence-transformers/pull/2677v3] Add warning that Evaluators only run on 1 GPU when multi-GPU training by @tomaarsen in https://github.com/UKPLab/sentence-transformers/pull/2678v3] Move training dependencies into a "train" extra by @tomaarsen in https://github.com/UKPLab/sentence-transformers/pull/2676v3] Docs: update references to the API reference by @tomaarsen in https://github.com/UKPLab/sentence-transformers/pull/2679v3] Add "dataset_size:" to the tag denoting the number of training samples by @tomaarsen in https://github.com/UKPLab/sentence-transformers/pull/2680A special shoutout to @Jakobhenningjensen, @smerrill, @b5y, @ScottishFold007, @pszemraj, @bwanglzu, @igorkurinnyi, for experimenting with the v3.0 release prior to release and @matthewfranglen for the initial work on the training refactor back in October of 2022 in #1733.
cc @AlexJonesNLP as I know you are interested in this release!
Full Changelog: https://github.com/UKPLab/sentence-transformers/compare/v2.7.0...v3.0.0
This release introduces a new promising loss function, easier inference for Matryoshka models, new functionality for CrossEncoders and Inference on In
This release introduces a new promising loss function, easier inference for Matryoshka models, new functionality for CrossEncoders and Inference on Intel Gaudi2, along much more.
Install this version with
pip install sentence-transformers==2.7.0
For a number of years, MultipleNegativesRankingLoss (also known as SimCSE, InfoNCE, in-batch negatives loss) has been the state of the art in embedding model training. Notably, this loss function performs better with a larger batch size.
Recently, various improvements have been introduced:
CachedMultipleNegativesRankingLoss was introduced, which allows you to pick much higher batch sizes (e.g. 65536) with constant memory.GISTEmbedLoss takes a guide model to guide the in-batch negative sample selection. This prevents false negatives, resulting in a stronger training signal.Now, @JacksonCakes has combined these two approaches to produce the best of both worlds: CachedGISTEmbedLoss. This loss function allows for high batch sizes with constant memory usage, while also using a guide model to assist with the in-batch negative sample selection.
As can be seen in our Loss Overview, this model should be used with (anchor, positive) pairs or (anchor, positive, negative) triplets, much like MultipleNegativesRankingLoss, CachedMultipleNegativesRankingLoss, and GISTEmbedLoss. In short, any example using those loss functions can be updated to use CachedGISTEmbedLoss! Feel free to experiment, e.g. with this training script.
Sentence Transformers v2.4.0 introduced Matryoshka models: models whose embeddings are still useful after truncation. Since then, many useful Matryoshka models have been trained.
As of this release, the truncation for these Matryoshka embedding models can be done automatically via a new truncate_dim constructor argument:
from sentence_transformers import SentenceTransformer
from sentence_transformers.util import cos_sim
matryoshka_dim = 64
model = SentenceTransformer("nomic-ai/nomic-embed-text-v1.5", trust_remote_code=True, truncate_dim=matryoshka_dim)
embeddings = model.encode(
[
"search_query: What is TSNE?",
"search_document: t-distributed stochastic neighbor embedding (t-SNE) is a statistical method for visualizing high-dimensional data by giving each datapoint a location in a two or three-dimensional map.",
"search_document: Amelia Mary Earhart was an American aviation pioneer and writer.",
]
)
print(embeddings.shape)
# => [3, 64]
similarities = cos_sim(embeddings[0], embeddings[1:])
# => tensor([[0.7839, 0.4933]])
Extra information:
Alongside easier inference with Matryoshka models, evaluating them is now also much easier. You can also pass truncate_dim to any Evaluator. This way you can easily check the performance of any Sentence Transformer model at various truncated dimensions (even if the model was not trained with MatryoshkaLoss!)
from sentence_transformers.evaluation import EmbeddingSimilarityEvaluator
from sentence_transformers import SentenceTransformer
import datasets
model = SentenceTransformer("tomaarsen/mpnet-base-nli-matryoshka")
stsb = datasets.load_dataset("mteb/stsbenchmark-sts", split="test")
for dim in [768, 512, 256, 128, 64, 32, 16, 8, 4]:
evaluator = EmbeddingSimilarityEvaluator(
stsb["sentence1"],
stsb["sentence2"],
[score / 5 for score in stsb["score"]],
name=f"sts-test-{dim}",
truncate_dim=dim,
)
print(f"dim={dim:<3}: {evaluator(model) * 100:.2f} Spearman Correlation")
dim=768: 86.81 Spearman Correlation
dim=512: 86.76 Spearman Correlation
dim=256: 86.66 Spearman Correlation
dim=128: 86.20 Spearman Correlation
dim=64 : 85.40 Spearman Correlation
dim=32 : 82.42 Spearman Correlation
dim=16 : 79.31 Spearman Correlation
dim=8 : 72.82 Spearman Correlation
dim=4 : 63.44 Spearman Correlation
Here are some example training scripts that use this new truncate_dim option to assist with training Matryoshka models:
This release improves the support for CrossEncoder reranker models.
push_to_hub (#2524)You can now push trained CrossEncoder models to the 🤗 Hugging Face Hub!
from sentence_transformers import CrossEncoder
...
model = CrossEncoder("distilroberta-base")
# Train the model
model.fit(
train_dataloader=train_dataloader,
evaluator=evaluator,
epochs=num_epochs,
warmup_steps=warmup_steps,
)
model.push_to_hub("tomaarsen/distilroberta-base-stsb-cross-encoder")
CrossEncoder.push_to_hubtrust_remote_code for custom models (#2595)You can now load custom models from the Hugging Face Hub, i.e. models that have custom modelling code that require trust_remote_code to load.
from sentence_transformers import CrossEncoder
# Note: this model does not require `trust_remote_code=True` - there are currently no models that require it yet.
model = CrossEncoder("cross-encoder/ms-marco-MiniLM-L-6-v2", trust_remote_code=True)
# We want to compute the similarity between the query sentence
query = "A man is eating pasta."
# With all sentences in the corpus
corpus = [
"A man is eating food.",
"A man is eating a piece of bread.",
"The girl is carrying a baby.",
"A man is riding a horse.",
"A woman is playing violin.",
"Two men pushed carts through the woods.",
"A man is riding a white horse on an enclosed ground.",
"A monkey is playing drums.",
"A cheetah is running behind its prey.",
]
# We rank all sentences in the corpus for the query
ranks = model.rank(query, corpus)
# Print the scores
print("Query:", query)
for rank in ranks:
print(f"{rank['score']:.2f}\t{corpus[rank['corpus_id']]}")
CrossEncoderFrom this release onwards, you will be able to perform inference on Intel Gaudi2 accelerators. No modifications are needed, as the library will automatically detect the hpu device and configure the model accordingly. Thanks to Intel Habana for the support here.
docs] Add simple Makefile for building docs by @tomaarsen in https://github.com/UKPLab/sentence-transformers/pull/2566examples] Add Matryoshka evaluation plot by @kddubey in https://github.com/UKPLab/sentence-transformers/pull/2564push_to_hub to CrossEncoder by @imvladikon in https://github.com/UKPLab/sentence-transformers/pull/2524requirements] Set minimum transformers version to 4.34.0 for is_nltk_available by @tomaarsen in https://github.com/UKPLab/sentence-transformers/pull/2574docs] Update link: retrieve_rerank_simple_wikipedia.py -> .ipynb by @tomaarsen in https://github.com/UKPLab/sentence-transformers/pull/2580feat] Add truncation support by @kddubey in https://github.com/UKPLab/sentence-transformers/pull/2573examples] Add model upload for training_nli_v3 with GISTEmbedLoss by @tomaarsen in https://github.com/UKPLab/sentence-transformers/pull/2584fix] Matryoshka training always patch original forward, and check matryoshka_dims by @kddubey in https://github.com/UKPLab/sentence-transformers/pull/2593docs] Fix search bar on sbert.net by @tomaarsen in https://github.com/UKPLab/sentence-transformers/pull/2597clip] Prevent warning with padding when tokenizing for CLIP by @tomaarsen in https://github.com/UKPLab/sentence-transformers/pull/2599I especially want to thank @JacksonCakes for their excellent CachedGISTEmbedLoss PR and @kddubey for their wonderful PRs surrounding Matryoshka models and general repository housekeeping.
Full Changelog: https://github.com/UKPLab/sentence-transformers/compare/v2.6.1...v2.7.0
This is a patch release to fix a bug in `semantic_search_faiss` and `semantic_search_usearch` that caused the scores to not correspond to the returned
This is a patch release to fix a bug in semantic_search_faiss and semantic_search_usearch that caused the scores to not correspond to the returned corpus indices. Additionally, you can now evaluate embedding models after quantizing their embeddings.
You can now pass precision to the EmbeddingSimilarityEvaluator to evaluate the performance after quantization:
from sentence_transformers import SentenceTransformer
from sentence_transformers.evaluation import EmbeddingSimilarityEvaluator, SimilarityFunction
import datasets
model = SentenceTransformer("all-mpnet-base-v2")
stsb = datasets.load_dataset("mteb/stsbenchmark-sts", split="test")
print("Spearman correlation based on Cosine Similarity on the STS Benchmark test set:")
for precision in ["float32", "uint8", "int8", "ubinary", "binary"]:
evaluator = EmbeddingSimilarityEvaluator(
stsb["sentence1"],
stsb["sentence2"],
[score / 5 for score in stsb["score"]],
main_similarity=SimilarityFunction.COSINE,
name="sts-test",
precision=precision,
)
print(precision, evaluator(model))
Spearman correlation based on Cosine Similarity on the STS Benchmark test set:
float32 0.8342190421330611
uint8 0.8260094846238505
int8 0.8312754408857808
ubinary 0.8244338431442343
binary 0.8244338431442343
Full Changelog: https://github.com/UKPLab/sentence-transformers/compare/v2.6.0...v2.6.1
[deprecation] Deprecate save_to_hub in favor of push_to_hub; add safe_serialization support to push_to_hub by @tomaarsen in https://github.com/UKPLab/…
This release brings embedding quantization: a way to heavily speed up retrieval & other tasks, and a new powerful loss function: GISTEmbedLoss.
Install this version with
pip install sentence-transformers==2.6.0
Embeddings may be challenging to scale up, which leads to expensive solutions and high latencies. However, there is a new approach to counter this problem; it entails reducing the size of each of the individual values in the embedding: Quantization. Experiments on quantization have shown that we can maintain a large amount of performance while significantly speeding up computation and saving on memory, storage, and costs.
To be specific, using binary quantization may result in retaining 96% of the retrieval performance, while speeding up retrieval by 25x and saving on memory & disk space with 32x. Do not underestimate this approach! Read more about Embedding Quantization in our extensive blogpost.
Two forms of quantization exist at this time: binary and scalar (int8). These quantize embedding values from float32 into binary and int8, respectively. For Binary quantization, you can use the following snippet:
from sentence_transformers import SentenceTransformer
from sentence_transformers.quantization import quantize_embeddings
# 1. Load an embedding model
model = SentenceTransformer("mixedbread-ai/mxbai-embed-large-v1")
# 2a. Encode some text using "binary" quantization
binary_embeddings = model.encode(
["I am driving to the lake.", "It is a beautiful day."],
precision="binary",
)
# 2b. or, encode some text without quantization & apply quantization afterwards
embeddings = model.encode(["I am driving to the lake.", "It is a beautiful day."])
binary_embeddings = quantize_embeddings(embeddings, precision="binary")
References:
GISTEmbedLoss, as introduced in Solatorio (2024), is a guided variant of the more standard in-batch negatives (MultipleNegativesRankingLoss) loss. Both loss functions are provided with a list of (anchor, positive) pairs, but while MultipleNegativesRankingLoss uses anchor_i and positive_i as positive pair and all positive_j with i != j as negative pairs, GISTEmbedLoss uses a second model to guide the in-batch negative sample selection.
This can be very useful, because it is plausible that anchor_i and positive_j are actually quite semantically similar. In this case, GISTEmbedLoss would not consider them a negative pair, while MultipleNegativesRankingLoss would. When finetuning MPNet-base on the AllNLI dataset, these are the Spearman correlation based on cosine similarity using the STS Benchmark dev set (higher is better):
The blue line is MultipleNegativesRankingLoss, whereas the grey line is GISTEmbedLoss with the small all-MiniLM-L6-v2 as the guide model. Note that all-MiniLM-L6-v2 by itself does not reach 88 Spearman correlation on this dataset, so this is really the effect of two models (mpnet-base and all-MiniLM-L6-v2) reaching a performance that they could not reach separately.
save_to_hub DeprecationMost codebases that allow for pushing models to the Hugging Face Hub adopt a push_to_hub method instead of a save_to_hub method, and now Sentence Transformers will follow that convention. The push_to_hub method will now be the recommended approach, although save_to_hub will continue to exist for the time being: it will simply call push_to_hub internally.
from sentence_transformers import SentenceTransformer
model = SentenceTransformer("all-mpnet-base-v2")
...
# Train the model
model.fit(
train_objectives=[(train_dataloader, train_loss)],
evaluator=dev_evaluator,
epochs=num_epochs,
evaluation_steps=1000,
warmup_steps=warmup_steps,
)
# Push the model to Hugging Face
model.push_to_hub("tomaarsen/mpnet-base-nli-stsb")
feat] Add 'get_config_dict' method to GISTEmbedLoss for better model cards by @tomaarsen in https://github.com/UKPLab/sentence-transformers/pull/2543deprecation] Deprecate save_to_hub in favor of push_to_hub; add safe_serialization support to push_to_hub by @tomaarsen in https://github.com/UKPLab/sentence-transformers/pull/2544docs] Update return docstring of encode_multi_process by @tomaarsen in https://github.com/UKPLab/sentence-transformers/pull/2548feat] Add binary & scalar embedding quantization support to Sentence Transformers by @tomaarsen in https://github.com/UKPLab/sentence-transformers/pull/2549Full Changelog: https://github.com/UKPLab/sentence-transformers/compare/v2.5.1...v2.6.0
This is a patch release to fix a bug in CrossEncoder.rank that caused the last value to be discarded when using the default top_k=-1.
This is a patch release to fix a bug in CrossEncoder.rank that caused the last value to be discarded when using the default top_k=-1.
CrossEncoder.rank patch:from sentence_transformers.cross_encoder import CrossEncoder
# Pre-trained cross encoder
model = CrossEncoder("cross-encoder/stsb-distilroberta-base")
# We want to compute the similarity between the query sentence
query = "A man is eating pasta."
# With all sentences in the corpus
corpus = [
"A man is eating food.",
"A man is eating a piece of bread.",
"The girl is carrying a baby.",
"A man is riding a horse.",
"A woman is playing violin.",
"Two men pushed carts through the woods.",
"A man is riding a white horse on an enclosed ground.",
"A monkey is playing drums.",
"A cheetah is running behind its prey.",
]
# We rank all sentences in the corpus for the query
ranks = model.rank(query, corpus)
# Print the scores
print("Query:", query)
for rank in ranks:
print(f"{rank['score']:.2f}\t{corpus[rank['corpus_id']]}")
Query: A man is eating pasta.
0.67 A man is eating food.
0.34 A man is eating a piece of bread.
0.08 A man is riding a horse.
0.07 A man is riding a white horse on an enclosed ground.
0.01 The girl is carrying a baby.
0.01 Two men pushed carts through the woods.
0.01 A monkey is playing drums.
0.01 A woman is playing violin.
0.01 A cheetah is running behind its prey.
Previously, the lowest score document would be removed from the output.
examples] Update model repo_id in 2dMatryoshka example by @tomaarsen in https://github.com/UKPLab/sentence-transformers/pull/2515feat] Add get_config_dict to new Matryoshka2dLoss & AdaptiveLayerLoss by @tomaarsen in https://github.com/UKPLab/sentence-transformers/pull/2516chore] Update to ruff 0.3.0; update ruff.toml by @tomaarsen in https://github.com/UKPLab/sentence-transformers/pull/2517example] Don't always normalize the embeddings in clustering example by @tomaarsen in https://github.com/UKPLab/sentence-transformers/pull/2520top_k by @xenova in https://github.com/UKPLab/sentence-transformers/pull/2518Full Changelog: https://github.com/UKPLab/sentence-transformers/compare/v2.5.0...v2.5.1
This release brings two new loss functions, a new way to (re)rank with CrossEncoder models, and more fixes
This release brings two new loss functions, a new way to (re)rank with CrossEncoder models, and more fixes
Install this version with
pip install sentence-transformers==2.5.0
Embedding models are often encoder models with numerous layers, such as 12 (e.g. all-mpnet-base-v2) or 6 (e.g. all-MiniLM-L6-v2). To get embeddings, every single one of these layers must be traversed. 2D Matryoshka Sentence Embeddings (2DMSE) revisits this concept by proposing an approach to train embedding models that will perform well when only using a selection of all layers. This results in faster inference speeds at relatively low performance costs.
For example, using Sentence Transformers, you can train an Adaptive Layer model that can be sped up by 2x at a 15% reduction in performance, or 5x on GPU & 10x on CPU for a 20% reduction in performance. The 2DMSE paper highlights scenarios where this is superior to using a smaller model.
Training with Adaptive Layer support is quite elementary: rather than applying some loss function on only the last layer, we also apply that same loss function on the pooled embeddings from previous layers. Additionally, we employ a KL-divergence loss that aims to make the embeddings of the non-last layers match that of the last layer. This can be seen as a fascinating approach of knowledge distillation, but with the last layer as the teacher model and the prior layers as the student models.
For example, with the 12-layer microsoft/mpnet-base, it will now be trained such that the model produces meaningful embeddings after each of the 12 layers.
from sentence_transformers import SentenceTransformer
from sentence_transformers.losses import CoSENTLoss, AdaptiveLayerLoss
model = SentenceTransformer("microsoft/mpnet-base")
base_loss = CoSENTLoss(model=model)
loss = AdaptiveLayerLoss(model=model, loss=base_loss)
Additionally, this can be combined with the MatryoshkaLoss such that the resulting model can be reduced both in the number of layers, but also in the size of the output dimensions. See also the Matryoshka Embeddings for more information on reducing output dimensions. In Sentence Transformers, the combination of these two losses is called Matryoshka2dLoss, and a shorthand is provided for simpler training.
from sentence_transformers import SentenceTransformer
from sentence_transformers.losses import CoSENTLoss, Matryoshka2dLoss
model = SentenceTransformer("microsoft/mpnet-base")
base_loss = CoSENTLoss(model=model)
loss = Matryoshka2dLoss(model=model, loss=base_loss, matryoshka_dims=[768, 512, 256, 128, 64])
<details><summary>Performance Results</summary>
Let's look at the performance that we may be able to expect from an Adaptive Layer embedding model versus a regular embedding model. For this experiment, I have trained two models:
MultipleNegativesRankingLoss rather than AdaptiveLayerLoss on top of MultipleNegativesRankingLoss. I also use microsoft/mpnet-base as the base model.Both of these models were trained on the AllNLI dataset, which is a concatenation of the SNLI and MultiNLI datasets. I have evaluated these models on the STSBenchmark test set using multiple different embedding dimensions. The results are plotted in the following figure:
The first figure shows that the Adaptive Layer model stays much more performant when reducing the number of layers in the model. This is also clearly shown in the second figure, which displays that 80% of the performance is preserved when the number of layers is reduced all the way to 1.
Lastly, the third figure shows the expected speedup ratio for GPU & CPU devices in my tests. As you can see, removing half of the layers results in roughly a 2x speedup, at a cost of ~15% performance on STSB (~86 -> ~75 Spearman correlation). When removing even more layers, the performance benefit gets larger for CPUs, and between 5x and 10x speedups are very feasible with a 20% loss in performance.
</details>
<details><summary>Inference</summary>
After a model has been trained using the Adaptive Layer loss, you can then truncate the model layers to your desired layer count. Note that this requires doing a bit of surgery on the model itself, and each model is structured a bit differently, so the steps are slightly different depending on the model.
First of all, we will load the model & access the underlying transformers model like so:
from sentence_transformers import SentenceTransformer
model = SentenceTransformer("tomaarsen/mpnet-base-nli-adaptive-layer")
# We can access the underlying model with `model[0].auto_model`
print(model[0].auto_model)
MPNetModel(
(embeddings): MPNetEmbeddings(
(word_embeddings): Embedding(30527, 768, padding_idx=1)
(position_embeddings): Embedding(514, 768, padding_idx=1)
(LayerNorm): LayerNorm((768,), eps=1e-05, elementwise_affine=True)
(dropout): Dropout(p=0.1, inplace=False)
)
(encoder): MPNetEncoder(
(layer): ModuleList(
(0-11): 12 x MPNetLayer(
(attention): MPNetAttention(
(attn): MPNetSelfAttention(
(q): Linear(in_features=768, out_features=768, bias=True)
(k): Linear(in_features=768, out_features=768, bias=True)
(v): Linear(in_features=768, out_features=768, bias=True)
(o): Linear(in_features=768, out_features=768, bias=True)
(dropout): Dropout(p=0.1, inplace=False)
)
(LayerNorm): LayerNorm((768,), eps=1e-05, elementwise_affine=True)
(dropout): Dropout(p=0.1, inplace=False)
)
(intermediate): MPNetIntermediate(
(dense): Linear(in_features=768, out_features=3072, bias=True)
(intermediate_act_fn): GELUActivation()
)
(output): MPNetOutput(
(dense): Linear(in_features=3072, out_features=768, bias=True)
(LayerNorm): LayerNorm((768,), eps=1e-05, elementwise_affine=True)
(dropout): Dropout(p=0.1, inplace=False)
)
)
)
(relative_attention_bias): Embedding(32, 12)
)
(pooler): MPNetPooler(
(dense): Linear(in_features=768, out_features=768, bias=True)
(activation): Tanh()
)
)
This output will differ depending on the model. We will look for the repeated layers in the encoder. For this MPNet model, this is stored under model[0].auto_model.encoder.layer. Then we can slice the model to only keep the first few layers to speed up the model:
new_num_layers = 3
model[0].auto_model.encoder.layer = model[0].auto_model.encoder.layer[:new_num_layers]
Then we can run inference with it using <a href="https://sbert.net/docs/package_reference/SentenceTransformer.html#sentence_transformers.SentenceTransformer.encode"><code>SentenceTransformers.encode</code></a>.
from sentence_transformers import SentenceTransformer
from sentence_transformers.util import cos_sim
model = SentenceTransformer("tomaarsen/mpnet-base-nli-adaptive-layer")
new_num_layers = 3
model[0].auto_model.encoder.layer = model[0].auto_model.encoder.layer[:new_num_layers]
embeddings = model.encode(
[
"The weather is so nice!",
"It's so sunny outside!",
"He drove to the stadium.",
]
)
# Similarity of the first sentence with the other two
similarities = cos_sim(embeddings[0], embeddings[1:])
# => tensor([[0.7761, 0.1655]])
# compared to tensor([[ 0.7547, -0.0162]]) for the full model
As you can see, the similarity between the related sentences is much higher than the unrelated sentence, despite only using 3 layers. Feel free to copy this script locally, modify the new_num_layers, and observe the difference in similarities.
</details>
Extra information:
Example training scripts:
CrossEncoder models are often even better than biencoder (SentenceTransformer) models, as the model can compare two texts using the attention mechanism, unlike biencoders. However, they are more computationally expensive as well. They are commonly used for reranking the top retrieval results of a biencoder model. As of this release, that should now be more convenient!
We now support a rank method, which allows you to rank a bunch of documents given a query:
from sentence_transformers.cross_encoder import CrossEncoder
# Pre-trained cross encoder
model = CrossEncoder("cross-encoder/stsb-distilroberta-base")
# We want to compute the similarity between the query sentence
query = "A man is eating pasta."
# With all sentences in the corpus
corpus = [
"A man is eating food.",
"A man is eating a piece of bread.",
"The girl is carrying a baby.",
"A man is riding a horse.",
"A woman is playing violin.",
"Two men pushed carts through the woods.",
"A man is riding a white horse on an enclosed ground.",
"A monkey is playing drums.",
"A cheetah is running behind its prey.",
]
# We rank all sentences in the corpus for the query
ranks = model.rank(query, corpus)
# Print the scores
print("Query:", query)
for rank in ranks:
print(f"{rank['score']:.2f}\t{corpus[rank['corpus_id']]}")
0.67 A man is eating food.
0.34 A man is eating a piece of bread.
0.08 A man is riding a horse.
0.07 A man is riding a white horse on an enclosed ground.
0.01 The girl is carrying a baby.
0.01 Two men pushed carts through the woods.
0.01 A monkey is playing drums.
0.01 A woman is playing violin.
Extra information:
Semantic Textual Similarity example by @alvarobartt in https://github.com/UKPLab/sentence-transformers/pull/2511loss] Add AdaptiveLayerLoss; 2d Matryoshka loss modifiers by @tomaarsen in https://github.com/UKPLab/sentence-transformers/pull/2506I especially want to thank @SeanLee97 and @fkdosilovic for their valuable contributions in this release.
Full Changelog: https://github.com/UKPLab/sentence-transformers/compare/v2.4.0...v2.5.0
This release introduces numerous notable features that are well worth learning about!
This release introduces numerous notable features that are well worth learning about!
Install this version with
pip install sentence-transformers==2.4.0
Dense embedding models typically produce embeddings with a fixed size, such as 768 or 1024. All further computations (clustering, classification, semantic search, retrieval, reranking, etc.) must then be done on these full embeddings. Matryoshka Representation Learning revisits this idea, and proposes a solution to train embedding models whose embeddings are still useful after truncation to much smaller sizes. This allows for considerably faster (bulk) processing.
Training using Matryoshka Representation Learning (MRL) is quite elementary: rather than applying some loss function on only the full-size embeddings, we also apply that same loss function on truncated portions of the embeddings. For example, if a model has an embedding dimension of 768 by default, it can now be trained on 768, 512, 256, 128, 64 and 32. Each of these losses will be added together, optionally with some weight:
from sentence_transformers import SentenceTransformer
from sentence_transformers.losses import CoSENTLoss, MatryoshkaLoss
model = SentenceTransformer("microsoft/mpnet-base")
base_loss = CoSENTLoss(model=model)
loss = MatryoshkaLoss(model=model, loss=base_loss, matryoshka_dims=[768, 512, 256, 128, 64])
<details><summary>Inference</summary>
After a model has been trained using a Matryoshka loss, you can then run inference with it using <a href="https://sbert.net/docs/package_reference/SentenceTransformer.html#sentence_transformers.SentenceTransformer.encode"><code>SentenceTransformers.encode</code></a>. You must then truncate the resulting embeddings, and it is recommended to renormalize the embeddings.
from sentence_transformers import SentenceTransformer
from sentence_transformers.util import cos_sim
import torch.nn.functional as F
model = SentenceTransformer("nomic-ai/nomic-embed-text-v1.5", trust_remote_code=True)
matryoshka_dim = 64
embeddings = model.encode(
[
"search_query: What is TSNE?",
"search_document: t-distributed stochastic neighbor embedding (t-SNE) is a statistical method for visualizing high-dimensional data by giving each datapoint a location in a two or three-dimensional map.",
"search_document: Amelia Mary Earhart was an American aviation pioneer and writer.",
]
)
embeddings = embeddings[..., :matryoshka_dim] # Shrink the embedding dimensions
similarities = cos_sim(embeddings[0], embeddings[1:])
# => tensor([[0.7839, 0.4933]])
As you can see, the similarity between the search query and the correct document is much higher than that of an unrelated document, despite the very small matryoshka dimension applied. Feel free to copy this script locally, modify the matryoshka_dim, and observe the difference in similarities.
Note: Despite the embeddings being smaller, training and inference of a Matryoshka model is not faster, not more memory-efficient, and not smaller. Only the processing and storage of the resulting embeddings will be faster and cheaper.
</details>
Extra information:
Example training scripts:
CoSENTLoss was introduced by Jianlin Su, 2022 as a drop-in replacement of CosineSimilarityLoss. Experiments have shown that it produces a stronger learning signal than CosineSimilarityLoss.
from sentence_transformers import SentenceTransformer, losses
from sentence_transformers.readers import InputExample
model = SentenceTransformer('bert-base-uncased')
train_examples = [
InputExample(texts=['My first sentence', 'My second sentence'], label=1.0),
InputExample(texts=['My third sentence', 'Unrelated sentence'], label=0.3)
]
train_dataloader = DataLoader(train_examples, shuffle=True, batch_size=train_batch_size)
train_loss = losses.CoSENTLoss(model=model)
You can update training_stsbenchmark.py by replacing CosineSimilarityLoss with CoSENTLoss & you can observe the improved performance.
AnglELoss was introduced in Li and Li, 2023. It is an adaptation of the CoSENTLoss, and also acts as a strong drop-in replacement of CosineSimilarityLoss. Compared to CoSENTLoss, AnglELoss uses a different similarity function which aims to avoid vanishing gradients.
Like with CoSENTLoss, you can use it just like CosineSimilarityLoss.
from sentence_transformers import SentenceTransformer, losses
from sentence_transformers.readers import InputExample
model = SentenceTransformer('bert-base-uncased')
train_examples = [
InputExample(texts=['My first sentence', 'My second sentence'], label=1.0),
InputExample(texts=['My third sentence', 'Unrelated sentence'], label=0.3)
]
train_dataloader = DataLoader(train_examples, shuffle=True, batch_size=train_batch_size)
train_loss = losses.AnglELoss(model=model)
You can update training_stsbenchmark.py by replacing CosineSimilarityLoss with AnglELoss & you can observe the improved performance.
Some models require using specific text prompts to achieve optimal performance. For example, with intfloat/multilingual-e5-large you should prefix all queries with query: and all passages with passage: . Another example is BAAI/bge-large-en-v1.5, which performs best for retrieval when the input texts are prefixed with Represent this sentence for searching relevant passages: .
Sentence Transformer models can now be initialized with prompts and default_prompt_name parameters:
prompts is an optional argument that accepts a dictionary of prompts with prompt names to prompt texts. The prompt will be prepended to the input text during inference. For example,model = SentenceTransformer(
"intfloat/multilingual-e5-large",
prompts={
"classification": "Classify the following text: ",
"retrieval": "Retrieve semantically similar text: ",
"clustering": "Identify the topic or theme based on the text: ",
},
)
# or
model.prompts = {
"classification": "Classify the following text: ",
"retrieval": "Retrieve semantically similar text: ",
"clustering": "Identify the topic or theme based on the text: ",
}
default_prompt_name is an optional argument that determines the default prompt to be used. It has to correspond with a prompt name from prompts. If None, then no prompt is used by default. For example,model = SentenceTransformer(
"intfloat/multilingual-e5-large",
prompts={
"classification": "Classify the following text: ",
"retrieval": "Retrieve semantically similar text: ",
"clustering": "Identify the topic or theme based on the text: ",
},
default_prompt_name="retrieval",
)
# or
model.default_prompt_name="retrieval"
Both of these parameters can also be specified in the config_sentence_transformers.json file of a saved model. That way, you won't have to specify these options manually when loading. When you save a Sentence Transformer model, these options will be automatically saved as well.
During inference, prompts can be applied in a few different ways. All of these scenarios result in identical texts being embedded:
prompt option in SentenceTransformer.encode:embeddings = model.encode("How to bake a strawberry cake", prompt="Retrieve semantically similar text: ")
prompt_name option in SentenceTransformer.encode by relying on the prompts loaded from a) initialization or b) the model config.embeddings = model.encode("How to bake a strawberry cake", prompt_name="retrieval")
prompt nor prompt_name are specified in SentenceTransformer.encode, then the prompt specified by default_prompt_name will be applied. If it is None, then no prompt will be applied.embeddings = model.encode("How to bake a strawberry cake")
Some INSTRUCTOR models, such as hkunlp/instructor-large, are natively supported in Sentence Transformers. These models are special, as they are trained with instructions in mind. Notably, the primary difference between normal Sentence Transformer models and Instructor models is that the latter do not include the instructions themselves in the pooling step.
The following models work out of the box:
You can use these models like so:
from sentence_transformers import SentenceTransformer
model = SentenceTransformer("hkunlp/instructor-large")
embeddings = model.encode(
[
"Dynamical Scalar Degree of Freedom in Horava-Lifshitz Gravity",
"Comparison of Atmospheric Neutrino Flux Calculations at Low Energies",
"Fermion Bags in the Massive Gross-Neveu Model",
"QCD corrections to Associated t-tbar-H production at the Tevatron",
],
prompt="Represent the Medicine sentence for clustering: ",
)
print(embeddings.shape)
# => (4, 768)
<details><summary>Information Retrieval usage</summary>
from sentence_transformers import SentenceTransformer
from sentence_transformers.util import cos_sim
model = SentenceTransformer("hkunlp/instructor-large")
query = "where is the food stored in a yam plant"
query_instruction = (
"Represent the Wikipedia question for retrieving supporting documents: "
)
corpus = [
'Yams are perennial herbaceous vines native to Africa, Asia, and the Americas and cultivated for the consumption of their starchy tubers in many temperate and tropical regions. The tubers themselves, also called "yams", come in a variety of forms owing to numerous cultivars and related species.',
"The disparate impact theory is especially controversial under the Fair Housing Act because the Act regulates many activities relating to housing, insurance, and mortgage loans—and some scholars have argued that the theory's use under the Fair Housing Act, combined with extensions of the Community Reinvestment Act, contributed to rise of sub-prime lending and the crash of the U.S. housing market and ensuing global economic recession",
"Disparate impact in United States labor law refers to practices in employment, housing, and other areas that adversely affect one group of people of a protected characteristic more than another, even though rules applied by employers or landlords are formally neutral. Although the protected classes vary by statute, most federal civil rights laws protect based on race, color, religion, national origin, and sex as protected traits, and some laws include disability status and other traits as well.",
]
corpus_instruction = "Represent the Wikipedia document for retrieval: "
query_embedding = model.encode(query, prompt=query_instruction)
corpus_embeddings = model.encode(corpus, prompt=corpus_instruction)
similarities = cos_sim(query_embedding, corpus_embeddings)
print(similarities)
# => tensor([[0.8835, 0.7037, 0.6970]])
</details>
All other Instructor models either 1) will not load as they refer to InstructorEmbedding in their modules.json or 2) require calling model.set_pooling_include_prompt(include_prompt=False) after loading.
Sentence Transformers now no longer uses sentencepiece or nltk as mandatory dependencies. This should make Sentence Transformers 1) lighter to install, 2) quicker to import and 3) less likely to result in dependency issues.
The documentation has been upgraded with a Loss Overview: https://sbert.net/docs/training/loss_overview.html This section contains tables with loss functions and their required data formats. This should help you narrow down which loss functions might suit your use cases:
Additionally, each loss function now has extended documentation, including references, requirements, relations to other loss functions, inputs, a code snippet with example usage, and/or links to other documentation/examples using that loss function.
hotfix] Don't require loading files for Normalize by @tomaarsen in https://github.com/UKPLab/sentence-transformers/pull/2460fix] Avoid sets and dicts in BinaryClassificationEvaluator when sentences are not hashable by @tomaarsen in https://github.com/UKPLab/sentence-transformers/pull/2462docs] Refactor & improve the loss documentation by @ir2718 in https://github.com/UKPLab/sentence-transformers/pull/2447deps] Remove the sentencepiece & nltk dependencies by @tomaarsen in https://github.com/UKPLab/sentence-transformers/pull/2476ci] On Ubuntu CI runner, use temporary directories as cache folders for some models by @tomaarsen in https://github.com/UKPLab/sentence-transformers/pull/2481docs] Slight improvements to docs phrasing by @tomaarsen in https://github.com/UKPLab/sentence-transformers/pull/2486feat] Add Matryoshka loss + examples + docs by @tomaarsen in https://github.com/UKPLab/sentence-transformers/pull/2485feat] Add prompt templates by @tomaarsen in https://github.com/UKPLab/sentence-transformers/pull/2477docs] Move loss overview to "main" documentation by @tomaarsen in https://github.com/UKPLab/sentence-transformers/pull/2496feat] Allow saving a model to the Hub without providing a user + Upload Matryoshka models after training by @tomaarsen in https://github.com/UKPLab/sentence-transformers/pull/2497docs] Address some small mistakes by @tomaarsen in https://github.com/UKPLab/sentence-transformers/pull/2498I especially want to thank @ir2718, @johneckberg & @SeanLee97 for their valuable contributions in this release, and @fkdosilovic and @milistu for their valuable improvements to the CrossEncoder.
Full Changelog: https://github.com/UKPLab/sentence-transformers/compare/v2.3.1...v2.4.0
This releases patches a niche bug when loading a Sentence Transformer model which:
This releases patches a niche bug when loading a Sentence Transformer model which:
Normalize module as specified in modules.jsonThis only occurs when a model with Normalize is downloaded from the Hugging Face hub and then later used locally.
See #2458 and #2459 for more details.
Full Changelog: https://github.com/UKPLab/sentence-transformers/compare/v2.3.0...v2.3.1
Sentence Transformers has deprecated Python 3.7 following its end of security support. Additionally, various dependencies have been updated to prevent…
This release focuses on various bug fixes & improvements to keep up with adjacent works like transformers and huggingface_hub. These are the key changes in the release:
Prior to Sentence Transformers v2.3.0, saving models to the Hugging Face Hub may have resulted in various errors depending on the versions of the dependencies. Sentence Transformers v2.3.0 introduces a refactor to save_to_hub to resolve these issues.
from sentence_transformers import SentenceTransformer
model = SentenceTransformer("sentence-transformers/all-MiniLM-L6-v2")
...
model.save_to_hub("tomaarsen/all-MiniLM-L6-v2-quora")
pytorch_model.bin: 100%|█████████████████████████████████████████████████████████████████████████████████████████████████| 90.9M/90.9M [00:06<00:00, 13.7MB/s]
Upload 1 LFS files: 100%|███████████████████████████████████████████████████████████████████████████████████████████████████████| 1/1 [00:07<00:00, 7.11s/it]
Recently, transformers has shifted towards using safetensors files as their primary model file formats. Additionally, various other file formats are commonly used, such as PyTorch (pytorch_model.bin), Rust (rust_model.ot), Tensorflow (tf_model.h5) and ONNX (model.onnx).
Prior to Sentence Transformers v2.3.0, almost all files of a repository would be downloaded, even if theye are not strictly required. Since v2.3.0, only the strictly required files will be downloaded. For example, when loading sentence-transformers/all-MiniLM-L6-v2 which has its model weights in three formats (pytorch_model.bin, rust_model.ot, tf_model.h5), only pytorch_model.bin will be downloaded. Additionally, when downloading intfloat/multilingual-e5-small with two formats (model.safetensors, pytorch_model.bin), only model.safetensors will be downloaded.
from sentence_transformers import SentenceTransformer
model = SentenceTransformer("sentence-transformers/all-MiniLM-L6-v2")
Downloading modules.json: 100%|██████████████████████████████████████████████████████████████████████████████████████████████████████| 349/349 [00:00<?, ?B/s]
Downloading (…)ce_transformers.json: 100%|███████████████████████████████████████████████████████████████████████████████████████████| 116/116 [00:00<?, ?B/s]
Downloading README.md: 100%|█████████████████████████████████████████████████████████████████████████████████████████████████████| 10.6k/10.6k [00:00<?, ?B/s]
Downloading (…)nce_bert_config.json: 100%|█████████████████████████████████████████████████████████████████████████████████████████| 53.0/53.0 [00:00<?, ?B/s]
Downloading config.json: 100%|███████████████████████████████████████████████████████████████████████████████████████████████████████| 612/612 [00:00<?, ?B/s]
Downloading pytorch_model.bin: 100%|█████████████████████████████████████████████████████████████████████████████████████| 90.9M/90.9M [00:06<00:00, 15.0MB/s]
Downloading tokenizer_config.json: 100%|█████████████████████████████████████████████████████████████████████████████████████████████| 350/350 [00:00<?, ?B/s]
Downloading vocab.txt: 100%|███████████████████████████████████████████████████████████████████████████████████████████████| 232k/232k [00:00<00:00, 1.37MB/s]
Downloading tokenizer.json: 100%|██████████████████████████████████████████████████████████████████████████████████████████| 466k/466k [00:00<00:00, 4.61MB/s]
Downloading (…)cial_tokens_map.json: 100%|███████████████████████████████████████████████████████████████████████████████████████████| 112/112 [00:00<?, ?B/s]
Downloading 1_Pooling/config.json: 100%|█████████████████████████████████████████████████████████████████████████████████████████████| 190/190 [00:00<?, ?B/s]
[!NOTE]
This release updates the default cache location from~/.cache/torch/sentence_transformersto the default cache location oftransformers, i.e.~/.cache/huggingface. You can still specify custom cache locations via theSENTENCE_TRANSFORMERS_HOMEenvironment variable or thecache_folderargument. Additionally, by supporting newer versions of various dependencies (e.g.huggingface_hub), the cache format changed. A consequence is that the old cached models cannot be used in v2.3.0 onwards, and those models need to be redownloaded. Once redownloaded, an airgapped machine can load the model like normal despite having no internet access.
This release brings models with custom code to Sentence Transformers through trust_remote_code, such as jinaai/jina-embeddings-v2-base-en.
from sentence_transformers import SentenceTransformer
from sentence_transformers.util import cos_sim
model = SentenceTransformer("jinaai/jina-embeddings-v2-base-en", trust_remote_code=True)
embeddings = model.encode(['How is the weather today?', 'What is the current weather like today?'])
print(cos_sim(embeddings[0], embeddings[1]))
# => tensor([[0.9341]])
If an embedding model is ever updated, it would invalidate all of the embeddings that you have created with the prior version of that model. We promise to never update the weights of any sentence-transformers/... model, but we cannot offer this guarantee for models by the community.
That is why this version introduces a revision keyword, allowing you to specify exactly which revision or branch you'd like to load:
from sentence_transformers import SentenceTransformer
model = SentenceTransformer("BAAI/bge-small-en-v1.5", revision="982532469af0dff5df8e70b38075b0940e863662")
# or a branch:
model = SentenceTransformer("BAAI/bge-small-en-v1.5", revision="main")
use_auth_token, use token instead (#2376)Following updates from transformers & huggingface_hub, Sentence Transformers now recommends that you use the token argument to provide your Hugging Face authentication token to download private models.
from sentence_transformers import SentenceTransformer
# new:
model = SentenceTransformer("tomaarsen/all-mpnet-base-v2", token="hf_...")
# old, still works, but throws a warning to upgrade to "token"
model = SentenceTransformer("tomaarsen/all-mpnet-base-v2", use_auth_token="hf_...")
[!NOTE] The recommended way to include your Hugging Face authentication token is to run
huggingface-cli login& paste your User Access Token from your Hugging Face Settings. See these docs for more information. Then, you don't have to include thetokenargument at all; it'll be automatically read from your filesystem.
Prior to this release, SentenceTransformers.device would not always correspond to the device on which embeddings were computed, or on which a model gets trained. This release brings a few fixes:
SentenceTransformers.device now always corresponds to the device that the model is on, and on which it will do its computations.SentenceTransformers.to(...), SentenceTransformers.cpu(), SentenceTransformers.cuda(), etc. will now work as expected, rather than being ignored.MultipleNegativesRankingLoss (MNRL) is a powerful loss function that is commonly applied to train embedding models. It uses in-batch negative sampling to produce a large number of negative pairs, allowing the model to receive a training signal to push the embeddings of this pair apart. It is commonly shown that a larger batch size results in better performing models (Qu et al., 2021, Li et al., 2023), but a larger batch size requires more VRAM in practice.
To counteract that, @kwang2049 has implemented a slightly modified GradCache technique that is able to separate the batch computation into mini-batches without any reduction in training quality. This allows the common practitioner to train with competitive batch sizes, e.g. 65536! The downside is that training with Cached MNRL (CMNRL) is roughly 2 to 2.4 times slower than using normal MNRL.
CachedMultipleNegativesRankingLoss is a drop-in replacement for MultipleNegativesRankingLoss, but with a new mini_batch_size argument. I recommend trying out CMNRL with a large batch size and a fairly small mini_batch_size - the larger mini batch size that will fit into memory.
from sentence_transformers import SentenceTransformer, losses, InputExample
from torch.utils.data import DataLoader
model = SentenceTransformer("distilbert-base-uncased")
train_examples = [
InputExample(texts=['Anchor 1', 'Positive 1']),
InputExample(texts=['Anchor 2', 'Positive 2']),
]
train_dataloader = DataLoader(train_examples, shuffle=True, batch_size=1024) # Here we can try much larger batch sizes!
train_loss = losses.CachedMultipleNegativesRankingLoss(model=model, mini_batch_size = 32)
model.fit([(train_dataloader, train_loss)], ...)
This release updates the community_detection function in various ways. Notably:
show_progress_bar option has been added (#1879)In the below graph, master refers to Sentence Transformers v2.2.2 and refactor refers to v2.3.0. On GPU, the computation time was heavily reduced.
Sentence Transformers has deprecated Python 3.7 following its end of security support. Additionally, various dependencies have been updated to prevent functionality from breaking. In particular:
torch >= 1.11.0transformers>= 4.32.0huggingface_hub>=0.15.1Lastly, torchvision has been removed as a dependency.
See the following for a list of release highlights:
community_detection from running forever by @nreimers (d8982c9f0d44f8a3c41579fa64c603eca029649b)to from getting ignored, replace ._target_device with .device by @tomaarsen (#2351)normalize_embeddings support to multi-process encoding by @tomaarsen (#2377)save_to_hub, remote git dependency, add token argument by @tomaarsen (#2376)transformers>=4.32.0 and huggingface_hub>=0.15.1 by @tomaarsen (#2376)community_detection by @dyaaalbakour (#2277)token and trust_remote_code to tokenizer_args too by @tomaarsen (#2411)cache_folder nor SENTENCE_TRANSFORMERS_HOME are set, use HF default cache by @tomaarsen (#2412)@k at the end of csv file name for RerankingEvaluator by @milistu (#2427)huggingface_hub dropped support in version 0.5.0 for Python 3.6
huggingface_hub dropped support in version 0.5.0 for Python 3.6
This release fixes the issue so that huggingface_hub with version 0.4.0 and Python 3.6 can still be used.
Version 0.8.1 of huggingface_hub introduces several changes that resulted in errors and warnings. This version of sentence-transformers fixes these is
Version 0.8.1 of huggingface_hub introduces several changes that resulted in errors and warnings. This version of sentence-transformers fixes these issues.
Further, several improvements have been added / merged:
util.community_detection was improved: 1) It works in a batched mode to save memory, 2) Overlapping clusters are no longer dropped but removed by overlapping items, 3) The parameter init_max_size was removed and replaced by a heuristic to estimate the max size of clustersYou can now use the encoder from T5 to learn text embeddings. You can use it like any other transformer model: `python from sentence_transformers impo
You can now use the encoder from T5 to learn text embeddings. You can use it like any other transformer model:
from sentence_transformers import SentenceTransformer, models
word_embedding_model = models.Transformer('t5-base', max_seq_length=256)
pooling_model = models.Pooling(word_embedding_model.get_word_embedding_dimension())
model = SentenceTransformer(modules=[word_embedding_model, pooling_model])
See T5-Benchmark results - the T5 encoder is not the best model for learning text embeddings models. It requires quite a lot of training data and training steps. Other models perform much better, at least in the given experiment with 560k training triplets.
The models from the papers Sentence-T5: Scalable sentence encoders from pre-trained text-to-text models and Large Dual Encoders Are Generalizable Retrievers have been added:
For benchmark results, see https://seb.sbert.net
Thanks to #1406 you can now load private models from the hub:
model = SentenceTransformer("your-username/your-model", use_auth_token=True)
This is a smaller release with some new features
This is a smaller release with some new features
MarginMSELoss is a great method to train embeddings model with the help of a cross-encoder model. The details are explained here: MSMARCO - MarginMSE Training
You pass your training data in the format:
InputExample(texts=[query, positive, negative], label=cross_encoder.predict([query, positive])-cross_encoder.predict([query, negative])
MultipleNegativesRankingLoss computes the loss just in one way: Find the correct answer for a given question.
MultipleNegativesSymmetricRankingLoss also computes the loss in the other direction: Find the correct question for a given answer.
The CLIPModel is now based on the transformers model.
You can still load it like this:
model = SentenceTransformer('clip-ViT-B-32')
Older SentenceTransformers versions are now longer able to load and use the 'clip-ViT-B-32' model.
PR #1116 checks if you have all files in your local cache or if there are added files on the hub. If this is the case, it will automatically download them.
SentenceTransformers.encode() can return all valuesWhen you set output_value=None for the encode method, all values (token_ids, token_embeddings, sentence_embedding) will be returned.
There should be no breaking changes. Old models can still be loaded from disc. However, if you use one of the provided pre-trained models, it will be…
All pre-trained models are now hosted on the Huggingface Models hub.
Our pre-trained models can be found here: https://huggingface.co/sentence-transformers
But you can easily share your own sentence-transformer model on the hub and have other people easily access it. Simple upload the folder and have people load it via:
model = SentenceTransformer('[your_username]/[model_name]')
For more information, see: Sentence Transformers in the Hugging Face Hub
There should be no breaking changes. Old models can still be loaded from disc. However, if you use one of the provided pre-trained models, it will be downloaded again in version 2 of sentence transformers as the cache path has slightly changed.
You can filter the hub for sentence-transformers models: https://huggingface.co/models?filter=sentence-transformers
Add the sentence-transformers tag to you model card so that others can find your model.
A widget was added to sentence-transformers models on the hub that lets you interact directly on the models website: https://huggingface.co/sentence-transformers/paraphrase-MiniLM-L6-v2
Further, models can now be used with the Accelerated Inference API: Send you sentences to the API and get back the embeddings from the respective model.
A new method was added to the SentenceTransformer class: save_to_hub.
Provide the model name and the model is saved on the hub.
Here you find the explanation from transformers how the hub works: Model sharing and uploading
When you save a model with save or save_to_hub, a README.md (also known as model card) is automatically generated with basic information about the respective SentenceTransformer model.
Final release of version 1: Makes v1 of sentence-transformers forward compatible with models from version 2 of sentence-transformers.
Final release of version 1: Makes v1 of sentence-transformers forward compatible with models from version 2 of sentence-transformers.
New methods integrated to train sentence embedding models without labeled data. See Unsupervised Learning for an overview of all existent methods.
New methods integrated to train sentence embedding models without labeled data. See Unsupervised Learning for an overview of all existent methods.
New methods:
New NLI & STS models: Following the Paraphrase Data training example we published new models trained on NLI and NLI+STS data. Training code is available: training_nli_v2.py.
| Model-Name | STSb-test performance |
|---|---|
| Previous best models | |
| nli-bert-large | 79.19 |
| stsb-roberta-large | 86.39 |
| New v2 models | |
| nli-mpnet-base-v2 | 86.53 |
| stsb-mpnet-base-v2 | 88.57 |
New MS MARCO model for Semantic Search: Hofstätter et al. optimized the training procedure on the MS MARCO dataset. The resulting model is integrated as msmarco-distilbert-base-tas-b and improves the performance on the MS MARCO dataset from 33.13 to 34.43 MRR@10
SentenceTransformer.fit() Checkpoints: The fit() method now allows to save checkpoints during the training at a fixed number of steps. More infomodels.Pooling() as string:pooling_model = models.Pooling(word_embedding_model.get_word_embedding_dimension(), pooling_mode='mean')
Valid values are mean/max/cls.Nothing published for this version
This release integrates methods that allows to learn sentence embeddings without having labeled data:
This release integrates methods that allows to learn sentence embeddings without having labeled data:
default_activation_function, that is applied on-top of the output logits generated by the class.It was not possible to fine-tune and save the CLIPModel. This release fixes it. CLIPModel can now be saved like any other model by calling model.save(
It was not possible to fine-tune and save the CLIPModel. This release fixes it. CLIPModel can now be saved like any other model by calling model.save(path)
v1.0.3 - Patch for util.paraphrase_mining method
v1.0.3 - Patch for util.paraphrase_mining method
v1.0.2 - Patch for CLIPModel, new Image Examples
v1.0.2 - Patch for CLIPModel, new Image Examples
Nothing published for this version
This release brings many new improvements and new features. Also, the version number scheme is updated. Now we use the format x.y.z with x: for major
This release brings many new improvements and new features. Also, the version number scheme is updated. Now we use the format x.y.z with x: for major releases, y: smaller releases with new features, z: bugfixes
You can now encode text and images in the same vector space using the OpenAI CLIP Model. You can use the model like this:
from sentence_transformers import SentenceTransformer, util
from PIL import Image
#Load CLIP model
model = SentenceTransformer('clip-ViT-B-32')
#Encode an image:
img_emb = model.encode(Image.open('two_dogs_in_snow.jpg'))
#Encode text descriptions
text_emb = model.encode(['Two dogs in the snow', 'A cat on a table', 'A picture of London at night'])
#Compute cosine similarities
cos_scores = util.cos_sim(img_emb, text_emb)
print(cos_scores)
More Information IPython Demo Colab Demo
Examples how to train the CLIP model on your data will be added soon.
util.dot_score computes the dot product of two embedding matrices. util.normalize_embeddings will normalize embeddings to unit lengthSentenceTransformer.encode method: normalize_embeddings if set to true, it will normalize embeddings to unit length. In that case the faster util.dot_score can be used instead of util.cos_sim to compute cosine similarity scores.models.Transformer(do_lower_case=True) when creating a new SentenceTransformer, then all input will be lower cased.output_value='token_embeddings' is definedLabelAccuracyEvaluatorencode(sent, convert_to_tensor=True). They now stay on the GPUNothing published for this version
Nothing published for this version
Faster tokenization speed: Using batched tokenization for training & inference - Now, all sentences in a batch are tokenized simoultanously.
Refactored Tokenization
SentencesDataset no longer needed for training. You can pass your train examples directly to the DataLoader:train_examples = [InputExample(texts=['My first sentence', 'My second sentence'], label=0.8),
InputExample(texts=['Another pair', 'Unrelated sentence'], label=0.3)]
train_dataloader = DataLoader(train_examples, shuffle=True, batch_size=16)
InputExample objects instead of tokenized textsSentenceLabelDataset has been updated to new tokenization flow: It returns always two or more InputExamples with the same labelAsymmetric Models
Add new models.Asym class that allows different encoding of sentences based on some tag (e.g. query vs paragraph). Minimal example:
word_embedding_model = models.Transformer(base_model, max_seq_length=250)
pooling_model = models.Pooling(word_embedding_model.get_word_embedding_dimension())
d1 = models.Dense(word_embedding_model.get_word_embedding_dimension(), 256, bias=False, activation_function=nn.Identity())
d2 = models.Dense(word_embedding_model.get_word_embedding_dimension(), 256, bias=False, activation_function=nn.Identity())
asym_model = models.Asym({'QRY': [d1], 'DOC': [d2]})
model = SentenceTransformer(modules=[word_embedding_model, pooling_model, asym_model])
##Your input examples have to look like this:
inp_example = InputExample(texts=[{'QRY': 'your query'}, {'DOC': 'your document text'}], label=1)
##Encoding (Note: Mixed inputs are not allowed)
model.encode([{'QRY': 'your query1'}, {'QRY': 'your query2'}])
Inputs that have the key 'QRY' will be passed through the d1 dense layer, while inputs with they key 'DOC' through the d2 dense layer.
More documentation on how to design asymmetric models will follow soon.
New Namespace & Models for Cross-Encoder Cross-Encoder are now hosted at https://huggingface.co/cross-encoder. Also, new pre-trained models have been added for: NLI & QNLI.
Logging
Log messages now use a custom logger from logging thanks to PR #623. This allows you which log messages you want to see from which components.
Unit tests A lot more unit tests have been added, which test the different components of the framework.
Updated the dependencies so that it works with Huggingface Transformers version 4. Sentence-Transformers still works with huggingface transformers ver
This release only include some smaller updates:
This release only include some smaller updates:
SentenceTransformer.fit() method - Parameter output_path_ignore_not_empty deprecated. No longer checks that target folder must be empty
Smaller changes:
Nothing published for this version
Nothing published for this version
Upgrade transformers dependency, transformers 3.1.0, 3.2.0 and 3.3.1 are working
Minor changes:
Hugginface Transformers version 3.1.0 had a breaking change with previous version 3.0.2
Hugginface Transformers version 3.1.0 had a breaking change with previous version 3.0.2
This release fixes the issue so that Sentence-Transformers is compatible with Huggingface Transformers 3.1.0. Note, that this and future version will not be compatible with transformers < 3.1.0.
Nothing published for this version
The old FP16 training code in model.fit() was replaced by using Pytorch 1.6.0 automatic mixed precision (AMP). When setting model.fit(use_amp=True), A
model.fit(use_amp=True), AMP will be used. On suitable GPUs, this leads to a significant speed-up while requiring less memory.The documentation is substantially improved and can be found at: www.SBERT.net - Feedback welcome
num_workers to a positive integer in your DataLoader, tokenization will happen in a background thread. This substantially increases the start-up time for training.model.encode() uses also a PyTorch DataSet + DataLoader. If you set num_workers to a positive integer, tokenization will happen in the background leading to faster encoding speed for large corpora.Breaking changes:
Multi-process tokenization (Linux only) for the model encode function. Significant speed-up when encoding large sets
This is a minor release. There should be no breaking changes.
This is a minor release. There should be no breaking changes.
Your coding agent can read these notes before it upgrades. Set up the MCP server →