NewYour coding agent can read the release notes before it upgrades.Set up the MCP server →
PyPI · #2804 most downloaded on PyPI
Massive Text Embedding Benchmark
Last release today
01 Oct 2026
Ships fairly regularly
a new release about every 1 weeks
Nearly every release is documented
notes for 58 of the last 60 stable releases
Nothing withdrawn
no release was ever pulled
4 years old
811 releases · first in 2022
One column per quarter.
fix: remove n_jobs-1 from logistic regression
n_jobs-1 from logistic regression (#4211)It is currently ignored and gives a future warning.
closes #4210 (52ec861)
model: add the colbert zero serie (#4206)
ColBERT-Zero serie
Add CodeSearchNet to training data
Exact n_parameters + memory usage + embed_dim
Factorize nomic embed training datasets definition
Factorize citations (5b8131f)
fix: Code leaderboard is failing
fix code leaderboard (dec66d6)
model: Add nomic-ai/nomic-embed-multimodal-7b dense embedding model (#4186)
feat: Add nomic-ai/nomic-embed-multimodal-7b dense embedding model
Add BiQwen2_5Wrapper and ModelMeta for nomic-embed-multimodal-7b, a dense (single-vector) multimodal embedding model for visual document retrieval using cosine similarity.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
PEFT adapter repo has no config.json or model.safetensors, so _from_hub cannot extract n_embedding_parameters.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Remove custom similarity method from BiQwen2_5Wrapper to use the processor's built-in scoring functionality, following the established pattern used by other ColPali models.
Addresses review feedback in PR #4186.
Co-Authored-By: Claude Sonnet 4 <noreply@anthropic.com>
Remove custom get_image_embeddings and get_text_embeddings methods to use the inherited ColPaliEngineWrapper implementations. For dense embedding models with fixed-size vectors, the parent class methods (extend + pad_sequence) produce equivalent results to the custom implementation (append + torch.cat).
Also remove unused tqdm import.
Addresses second review feedback in PR #4186.
Co-Authored-By: Claude Sonnet 4 <noreply@anthropic.com>
Co-Authored-By: Claude Sonnet 4 <noreply@anthropic.com>
Co-Authored-By: Claude Sonnet 4 <noreply@anthropic.com>
fix nomic
add 3b revision and lint
add fixes
rem
cleanup
Update mteb/models/model_implementations/nomic_multimodal.py
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>
add embedding parameters
add base revision
remove exceptions from test
added training data
lint
Update mteb/models/model_implementations/nomic_multimodal.py
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Co-authored-by: Your Name <you@example.com>
Co-authored-by: Kenneth Enevoldsen <kennethcenevoldsen@gmail.com>
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com> (b398ea7)
dataset: Update vidore tasks to include the OCR'd text (#4191)
dataset: Add OCR adaption of vidore tasks
Add beta tag to the new tasks
add nuclear and telecom
Apply suggestions from code review
Update mteb/tasks/retrieval/multilingual/vidore3_bench_retrieval.py
updated description and added version
add superseeded from
fix imports
fix init
fix private test
add some tasks statics
add nuclear
Co-authored-by: Roman Solomatin <36135455+Samoed@users.noreply.github.com> (9a6b98f)
final reupload from mmteb (7037bfe)
fix dataset transform signature (`009209a`)
fix: fleurs loading (#4197)
fix fleurs
fix dataset transform signature (009209a)
Start video (#4148)
start video integration
start video integration
upd task structure
upd batched input
upd video input type
combine video and audio to dict
use only one video per time
remove __main__
remove PostProcessingCollator (34d060c)
model: Add Perplexity pplx-embed-v1 models (0.6B and 4B) (#4189)
model: Add Perplexity pplx-embed-v1 models (0.6B and 4B)
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com> (1c90bfd)
model: LateOn-Code models definition
fix vn models (a80fae9)
model: LateOn-Code models definition (#4175)
First draft of LateOn code models definition
Fix reference for LateOn-Code
Fix reference LateOn code edge pretrain
Add memory_usage_mb (and embed_dim)
fix lint
Add training datasets
Co-authored-by: Roman Solomatin <36135455+Samoed@users.noreply.github.com> (4345e63)
model: Vietnamese model for VN-MTEB (#4187)
[ADD] Vietnamese model for VN-MTEB
[ADD] Vietnamese model for VN-MTEB (rename variable) (3ab29f4)
fix: remove select column from dataloader
fix: remove select column from dataloader (#4185)
remove select column
fix sts (7477902)
fix: Add repr to benchmarks to avoid excessive prints
fix: Add repr to benchmarks to avoid excessive prints (#4180)
fix: don't specify keyerror when it is a keyerror
fix: Add repr for benchmarks
It now looks like:
Benchmark(name='BEIR', desciption='BEIR is a heterogeneous benchmark containing diver..., tasks=[...] (#15), ...)
8b6cd0c)fix: add hit_rate to retrieval metrics and remove cv_recall
fix: add hit_rate to retrieval metrics and remove cv_recall (#4142)
add hit_rate
upd lotte metric
fix tests
fix name
remove cv_recall
upd lotte comment (3efc32f)
fix: move mteb(multilingual, v2) datasets to the mteb org
move datasets to mteb (55d6103)
fix: prompt type for non-retrieval types (#4174)
fix prompt names
fix pyproject
rollback pyproject.toml (e453c15)
Update zeroshot classification template path for pyproject.toml (#4171)
update path for zeroshot classification templates
fix path for zeroshot classification template (4a305d8)
feat: save and load experiments results
feat: save and load experiments results (#4071)
start experiments
sanitize experiment name
sanitize inputs
add methods to load results
fix to python test
fix loading with model meta
fix test path
apply suggestions
don't serilize classes
Apply suggestions from code review
Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com>
add different loading strategies
align implementation
align implementation
upd docstring
update enum
don't change logic for modelmeta
fix typing
try to run with -n auto
always work with metadata copy
simplify test
fix corner cases and add doc
add one more test
fix when model loading from meta with experiment
remove additional typos from pyppproject
rename to experiment_kwargs
fixes after review
rename baseline everywhere
add model_name_with_experiment
Update mteb/models/model_implementations/random_baseline.py
Co-authored-by: Kenneth Enevoldsen <kennethcenevoldsen@gmail.com>
Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com>
Co-authored-by: Kenneth Enevoldsen <kennethcenevoldsen@gmail.com> (ed4f476)
fix: Update _calculate_memory_usage_mb to optionally fetch_from_hf'
fix: Update _calculate_memory_usage_mb to optionally fetch_from_hf' (#4156)
Update _calculate_memory_usage_mb to optionally fetch_from_hf
Apply changes at other places (8f8d8bc)
Change leaderboard healthcheck schedule to every 12 hours (#4165) (469611c)
fix eval from existing (#4166)
fix eval (5e57dd2)
fix: Deprecate GritLM Wrapper and use Sentence Transformers
leaderboard health check (c9c623e)
feat: Integrate eval results (#4114)
start integration of eval results
fix models
Update mteb/abstasks/_eval/eval_model.py
Co-authored-by: Kenneth Enevoldsen <kennethcenevoldsen@gmail.com>
working version
add method to model meta for uploading
fix tests
test file dump in tests
add what's new
fix dumping
rename bm25
upd what's new
fix notes https://github.com/embeddings-benchmark/mteb/issues/4150
fix test
raname baseline models
rename models in tests
fix naming
Co-authored-by: Kenneth Enevoldsen <kennethcenevoldsen@gmail.com> (d892fd7)
fix: Deprecate GritLM Wrapper and use Sentence Transformers (#4085)
Depreceate GritLM Wrapper and use Sentence Transformers
Add include_prompt in InstructSentenceTransformerModel
Update Wrapper to InstructSentenceTransformerModel for remaining models
apply_instruction_to_passages=False in loader_kwargs
update docstring
Add depreciation warning
Added warning to both class and function (338af33)
fix: languages filtering scores (#4145)
fix languages filtering scores (3a2416d)
docs: add cache/search backends to doc
docs: add cache/search backends to doc (#4143)
add cache backend to doc
Update docs/advanced_usage/cache_embeddings.md
Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com>
Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com>
Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com> (25c9a2b)
fix: Refactor pylate models such that it does not follow the Encoder interface (#4151)
make multivec not comatible with most tasks
fix
remove numpy (b0e5f5a)
fix: pylate reranking (#4152)
fix pylate reranking (0b11c69)
23856d1)ci: raise leaderboard docker workflow timeout to 12min
ci: raise leaderboard docker workflow timeout to 12min (#4153)
feat: enhance leaderboard Docker timeout debugging
This enables clear identification of which initialization phase causes timeouts, with granular timing data for each step.
Co-Authored-By: Claude Sonnet 4 <noreply@anthropic.com>
This reverts commit cba77a5e9781d05efa949371f2d48d21931d03be.
Co-Authored-By: Claude Sonnet 4 <noreply@anthropic.com>
Co-authored-by: Claude Sonnet 4 <noreply@anthropic.com> (8edfd07)
fix: Metadata for nvidia colembed models (#4157)
fix metadata
fix metadata (017024e)
Fix Error in Performance Per Model Size Plot in MIEB (#4149)
Fix Error in Performance Plot in MIEB
added as customizable column (71fca40)
# 2.8.8 (2026-02-23) ## Fix * fix: num_proc for reuploading (#4146) fix num_proc (`007ab9e`)
Update mteb/models/model_implementations/pylate_models.py
fix: lotte wrong main metric (#4128)
fix lotte
fix function name (c997546)
fix: pylate model prompts (#4120)
fix pylate model prompts
fix log message
simplify a bit
fix imoprt
Update mteb/models/model_implementations/pylate_models.py
Co-authored-by: Antoine Chaffin <38869395+NohTow@users.noreply.github.com>
Co-authored-by: Antoine Chaffin <38869395+NohTow@users.noreply.github.com> (dc3aa46)
docs: fix incorrect example in docs for MAEB
change default metric (`b2cb4c4`)
fix: add performance over time (#4134)
add performance over time
change default metric (b2cb4c4)
dataset: Added Ukr Toxicity classification dataset (#4108)
Added Ukr Toxicity classification dataset
Update mteb/tasks/classification/ukr/ukr_toxicity_classification.py
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>
format citiatons and lint
add descriptive statistics
fix citations
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>
Co-authored-by: Kenneth Enevoldsen <kennethcenevoldsen@gmail.com> (7a18af4)
Add msmarco-minilm-v2 series (#4132)
Add msmarco-minilm-v2 series
fix metadata "n_embedding_param" (7f1fa2b)
fix: Change transformers version to 5.0.0 for nemotron-colembed-v2 models and update uv.lock
fix: Change transformers version to 5.0.0 for nemotron-colembed-v2 models and update uv.lock (#4125)
change transformers version to 5.0.0 for nemotron-colembed and update lockfile
update model revisions (fc929e0)
Fill up n_embedding_parameters value in ModelMeta (#4050)
Fill up n_embedding_parameters value in ModelMeta
fix naming
add more results and fix tests
Remmove private models from _HISTORIC_MODELS
fix tests
Address comments
Add parameter for vlm2vec models
Add fetch_from_hf parameter
apply suggestions
apply suggestions
fx typechecking
correctly apply suggestion
make lint
fix test
fix documentation
Reuse Loaded json (72bde5c)
add method for creating collection for benchmark (#4115)
add method for pushing
simplify
add comment to description
simplify imports
customize collection name (5f12691)
fix: Pair classification convert to dict
fix pair classification (88fb54e)
updating to correct revision for miniac-embed (#4116)
Adds ManiacLabs/miniac-embed. Sentence Transformers–compatible; 1024-d output, cosine similarity.
New file: mteb/models/model_implementations/maniac_labs.py
mteb.get_model(model_name, revision) andmteb.get_model_meta(model_name, revision)Co-authored-by: Cursor <cursoragent@cursor.com>
Addresses reviewer feedback from @ayush1298
Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: JD Pruett <85522589+jdpruett44@users.noreply.github.com>
Co-authored-by: Cursor <cursoragent@cursor.com> (3a7566f)
fix: refactor from_* function in model metas
fix: refactor from_* function in model metas (#4101)
fix: refactor from_* function in model metas
minor fixes based on review
fixes from review
format
reworked computation of embedding size
minor fixes
fixed failed tests and typecheck
Update mteb/models/model_meta.py
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com> (c916c0b)
add: jina-embeddings-v5-text family (#4102)
add: jina-v5
fix: linting error
fix: missing values
fix: remove useless code
add: citation
Update mteb/models/model_implementations/jina_models.py
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>
fix: use different mapping strategy
fix implementation
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>
Co-authored-by: Roman Solomatin <36135455+Samoed@users.noreply.github.com> (f607bae)
model: Add the model ManiacLabs/miniac-embed (#4096)
Adds ManiacLabs/miniac-embed. Sentence Transformers–compatible; 1024-d output, cosine similarity.
New file: mteb/models/model_implementations/maniac_labs.py
mteb.get_model(model_name, revision) andmteb.get_model_meta(model_name, revision)Co-authored-by: Cursor <cursoragent@cursor.com>
Addresses reviewer feedback from @ayush1298
Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: JD Pruett <85522589+jdpruett44@users.noreply.github.com>
Co-authored-by: Cursor <cursoragent@cursor.com> (e64ed80)
model: Add NIFE models (#4058)
model: Add NIFE models
closes #3586
@stephantul can I ask you to review the metadata
ran it using:
import mteb
model = mteb.get_model("stephantulkens/NIFE-mxbai-embed-large-v1")
# dummy small tasks
task1 = mteb.get_task("LccSentimentClassification")
task2 = mteb.get_task("TwitterHjerneRetrieval")
results = mteb.evaluate(model, [task1, task2], encode_kwargs={"device": "cpu"}) # to prevent MPS error
which uses the encode function for the classification task Is that intended?
updated based on review
updated to new models and metadata
fixes from review
fixed revision (7f24538)
fix: Remove duplicate citations and add test to prevent it going forward
fix: Remove duplicate citations and add test to prevent it going forward (#4032)
test: add test to detect duplicate citations
quality
move changes to task file
fix falsepositives
update citations
add models and benchmarks
search close titles
fix maeb
Co-authored-by: Roman Solomatin <36135455+Samoed@users.noreply.github.com> (377f1b4)
benchmark: add 6 new VisRAG retrieval tasks and corresponding stats (#4059)
dataset: add 6 new VisRAG retrieval tasks and corresponding stats
fix a linter error
dataset: introduce VisRAG Retrieval Benchmark
fix: metadata update of VisRag
Update benchmark metadata
Update VisRAG datasets metadata, including one-line description and the domains
Update slideVQA domain
Add Aliases for VisRAG
Fix bibtex format
Update dataset metadata to point to the mteb versions
Remove redundant data loading
Co-authored-by: Kenneth Enevoldsen <kennethcenevoldsen@gmail.com> (7a9e653)
fix: Remove MAEB+ and MAEB(extended) from leaderboard and add "beta" to all MAEB (#4103)
fix: Remove MAEB+ and MAEB(extended)
related to #3470
We could consider keeping the two benchmarks (would still need to be beta as paper is review) so they could change.
related to #3470
Currently implemented it as keeping the two temporary benchmarks. We could consider removing them as well (I am unsure how much of a burden it is for us to maintain them), but I would probably not add them to the leaderboard.
All of these changes should be backward compatible
docs: Added whatsnew
implement fixes
update description to explain beta status (77ac52b)
fix: correct reference link for MIRACLVisionRetrieval task (#4092) (`d3e9b06`)
d3e9b06)docs: Improved adding a benchmark docs
docs: Improved adding a benchmark docs (#4087)
docs: Improved adding a benchmark docs
expanded the explanation of provide more information about the process.
Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com>
Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com>
fix
fix
minor heading change
Apply suggestions from code review
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>
Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com>
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com> (8149d4a)
fix: constrain the transformers library version for jina-clip (#4061)
fix: constrain the transformers library version for jina-clip to avoid compatibility issue
add require package
ad to conflicts
try to run again
upd lock
tmp
try
fix: pylate dependency on outdated version of transformers
Co-authored-by: Roman Solomatin <36135455+Samoed@users.noreply.github.com> (eee82cf)
dataset: Add MTEB(spa) Spanish language benchmark (#4053)
dataset: Add MTEB(spa) Spanish language benchmark
Define MTEB(spa, v1) benchmark grouping 23 existing Spanish tasks across 6 task types: Classification (8), Clustering (3), PairClassification (2), Reranking (1), Retrieval (5), and STS (4).
fix: Replace MIRACLRetrieval with HardNegatives.v2 per review
fix: Remove tasks with known issues, add contact, reduce to 16 tasks
Apply suggestion from @KennethEnevoldsen
Co-authored-by: Clemente <clemente@Clementes-MacBook-Pro.local>
Co-authored-by: Kenneth Enevoldsen <kenevoldsen@pm.me> (2507bec)
Move MIEB datasets to mteb HuggingFace org (#4070)
Move 15 MIEB datasets to mteb HuggingFace org
Update dataset paths and revisions for tasks that now use datasets forked to the mteb org:
Part of #4049
Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>
transfer mrbench
Move 8 isaacchung MIEB datasets to mteb HuggingFace org
Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>
Datasets moved:
Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>
Datasets moved:
Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>
Datasets moved:
Note: wds_imagenet1k failed due to storage limits.
Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>
Datasets migrated:
Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>
Co-authored-by: Claude Opus 4.5 <noreply@anthropic.com>
Co-authored-by: Roman Solomatin <36135455+Samoed@users.noreply.github.com> (2ef04f8)
model: add voyage-4-nano (#4086)
model: add voyage-4-nano model implementation
Apply suggestion from @Samoed
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>
Co-authored-by: Kenneth Enevoldsen <kenevoldsen@pm.me>
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com> (2ce07c4)
fix: Remove task performance by type tab when there is only one type
docs: Outline for adding a task documentation
docs: Outline for adding a task documentation (#4082)
docs: Outline for adding a task documentation
This is a suggested structure, PR is just to get feedback before I finish it up.
fixes #4077
upd docs
install dependencies in ci
add example with retrieval
filled out the missing segments
lint and format
Apply suggestions from code review
Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com>
fix numerating and indent
add missing imports
fix links
add full example for retrieval dataset
Co-authored-by: Roman Solomatin <36135455+Samoed@users.noreply.github.com>
Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com> (42b8058)
docs: Improve docstring for some of the main abstasks (#4083)
docs: fix AbsTaskClassification docstring formatting and improve docstrings for some of the main tasks
format (50bd0fa)
Add Performance per language Tab to more benchmarks (4ca1922)
dataset: add 'law-ir_ko' dataset for IR task (#4052)
law_ir_ko
Update mteb/tasks/retrieval/kor/law_ir_ko.py
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>
law_ir_ko info revision
description
metadata-info rev
metadata-info rev
Update mteb/tasks/retrieval/kor/law_ir_ko.py
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>
statistics(), reference
format citation
author & howpublished rev
make lint
description rev
Update mteb/tasks/retrieval/kor/law_ir_ko.py
Co-authored-by: Kenneth Enevoldsen <kenevoldsen@pm.me>
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>
Co-authored-by: Roman Solomatin <36135455+Samoed@users.noreply.github.com>
Co-authored-by: Kenneth Enevoldsen <kenevoldsen@pm.me> (1cdc662)
Add ModelMeta for geoffsee/auto-g-embed-st (81540a2)
Add MetaCLIP 2 model integration (#4065)
Add MetaCLIP 2 model integration
Add support for facebook/metaclip-2-mt5-worldwide-b32, a multilingual vision-language model using mT5 tokenizer for worldwide language support.
Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>
Set n_embedding_parameters to 128,057,344 (mT5 vocab size 250,112 × embed_dim 512) to fix test_n_embedding_parameters test failure.
Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>
Co-authored-by: Claude Opus 4.5 <noreply@anthropic.com>
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com> (4280dd9)
fix: filter corrupted image in Birdsnap
fix: filter corrupted image in Birdsnap (#4068)
fix: filter corrupted image in Birdsnap and drop unused splits in zero-shot tasks
Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>
In transformers 5.x, get_text_features and get_image_features return BaseModelOutputWithPooling instead of a tensor directly. Extract the pooler_output when needed.
Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>
Fixes mypy type errors where self.dataset could be None when accessing .keys() and deleting splits in dataset_transform method.
Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>
Co-authored-by: Claude Opus 4.5 <noreply@anthropic.com> (6c00506)
Backfill missing metadata for historic datasets (#4063)
Backfill missing metadata for historic datasets
Fill in missing TaskMetadata fields for ~90 historic datasets as described in issue #2502. This includes:
Fields filled include: date, domains, task_subtypes, license, annotations_creators, dialect, sample_creation, and bibtex_citation.
The _HISTORIC_DATASETS list is reduced from ~90 entries to just 4 aggregate tasks whose metadata computation has a separate issue (the compute* methods return None for single-valued fields).
Closes #2502
Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>
Add StrURL to the return type and set type annotation to match the license field type (Licenses | StrURL | None).
Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>
Co-authored-by: Claude Opus 4.5 <noreply@anthropic.com> (b5fb471)
fix Remove the hardcoded batch_size=1 when generating text and image embeddings for Nemotron-Colembed-v2 models
fix: mscoco (#4062)
fix mscoco
fix jina clip (c2d1bfe)
remove hardcoded batch_size 1 (1682b2f)
update nemotron v2 citation (0be1df3)
fix docs deploy command (#4044) (`acb3d8c`)
fix: Fill in embedding and total parameters in ModelMeta
fix: Fill in embedding and total parameters in ModelMeta (#4031)
Filling Embedding/Total Parameters in ModelMeta
Add parameter for other models
Add parameters for more models
Added exact value for n_parameters
Fix tests
set n_embedding_parameters to None
Add results of some more models
Add tests
Add _HISTORIC_MODELS list in test
Update tests/test_models/test_model_meta.py
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>
fix tests
correct tests
fix _HISTORIC_MODELS list
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com> (bc6e6cb)
dataset: Add ERESS reranking task (#3991)
dataset: Add ERESS reranking task
fix: align dataset_transform signature with base class
fix: dataset reuploaded, custom transformation removed
fix: rev updated with title + text combination
description moved away from docstring
Update mteb/tasks/reranking/eng/ecommerce_product_relevance_reranking.py
Co-authored-by: Kenneth Enevoldsen <kenevoldsen@pm.me> (fe67f8e)
Clean up docs to prepare for adding the changelog. By adding missing links and removing references to documentation that does not exist
docs: Added changelog (#3741)
docs: Added changelog
I think going forward we can just update this as well go.
minor fix
added autogenerated changelog
rename
add autogenerated workflows
updates
update
update (2082d3e)
fix: backfilling historic tasks (#4034)
fix: backfilling historic tasks
addresses #2502
back citation, date and task subtypes where only those are missing
Update mteb/tasks/pair_classification/pol/polish_pc.py
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com> (e542519)
fix: Add Optional dependencies for NemotronColEmbed models as extras
fix: Add Optional dependencies for NemotronColEmbed models as extras (#4036)
fix: Move nemotron-colembed models to separate module
Signed-off-by: Oliver Holworthy <1216955+oliverholworthy@users.noreply.github.com>
Signed-off-by: Oliver Holworthy <1216955+oliverholworthy@users.noreply.github.com>
Signed-off-by: Oliver Holworthy <1216955+oliverholworthy@users.noreply.github.com>
Signed-off-by: Oliver Holworthy <1216955+oliverholworthy@users.noreply.github.com>
Signed-off-by: Oliver Holworthy <1216955+oliverholworthy@users.noreply.github.com>
Signed-off-by: Oliver Holworthy <1216955+oliverholworthy@users.noreply.github.com>
Signed-off-by: Oliver Holworthy <1216955+oliverholworthy@users.noreply.github.com>
Signed-off-by: Oliver Holworthy <1216955+oliverholworthy@users.noreply.github.com>
Signed-off-by: Oliver Holworthy <1216955+oliverholworthy@users.noreply.github.com>
Signed-off-by: Oliver Holworthy <1216955+oliverholworthy@users.noreply.github.com>
Signed-off-by: Oliver Holworthy <1216955+oliverholworthy@users.noreply.github.com>
Signed-off-by: Oliver Holworthy <1216955+oliverholworthy@users.noreply.github.com>
Signed-off-by: Oliver Holworthy <1216955+oliverholworthy@users.noreply.github.com>
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>
Signed-off-by: Oliver Holworthy <1216955+oliverholworthy@users.noreply.github.com>
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com> (2cc8dc2)
fix: avoid supplying trust_remote_code twice to `OpenSearch-AI/Ops-Co…
fix: avoid supplying trust_remote_code twice to OpenSearch-AI/Ops-Colqwen3-4B
related to: #4040 (9aa5ae5)
fix: change np.bool to np.bool_
np.bool to np.bool_ (#4041)change np.bool (53de8c1)
model: add the model of boom (#4022)
add the model of boom
add the model of boom
rename the file from boom_models.py to ict_time_and_querit_models.py, match MTEB task names, and remove the #
rename the file from boom_models.py to ict_time_and_querit_models.py, match MTEB task names, and remove the #
lint
dirctly use the InstructSentenceTransformerModel class
Remove unused boom_4b_v1_loader function
Removed commented-out loader function for BOOM_4B_v1 model.
Removed the commented code
lint again
Removed the commented code
lint again
Remove the description comment at the beginning
Remove the description comment at the beginning and update the n_parameters
Add adapted_from field to model metadata
add the description of some instruction templates
reformatted the ict_time_and_querit_models.py
update the model revision
Co-authored-by: zhanghengran <zhanghengran@baidu.com>
Co-authored-by: Roman Solomatin <36135455+Samoed@users.noreply.github.com>
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com> (62d354a)
fix: set num proc to None by default
None by default (#4038)set num proc to none by default (b3a51c6)
fix: Improve array typing by also specifying dtype
fix: Improve array typing by also specifying dtype (#4018)
Correct array typing
Added arrat typing in docs
fix typechecking
Address comments
fix typechecking
simplify typechecking
fix typecheck
Co-authored-by: Roman Solomatin <36135455+Samoed@users.noreply.github.com> (3a40adc)
fix: Update ViDoRe3 to rank based on the mean (#4033)
Update ViDoRe3 benchmark and leaderboard
revert uv.lock (b616950)
This is technically breaking so I added a deprecation.
fix: remove undocumented models or models with no model card (#4030)
fix: don't fill-missing metedata by default
This include two changes:
compute_missing to fill_missing, it seems like this change was only applied to get_model_meta, but not throughout the stack. This is technically breaking so I added a deprecation.from_hub, I believe this was unintentionally missedfixes #4027
set fill_missing=False for from_{model}
propegate deprecation
add framework default
fix: remove undocumented models or models with no model card
Also removed scripts/generate_metadata.py
fixes #3746
compute_missing > compute_metadata
clean up tests to match new functionality
format
lint (dbc7334)
This is technically breaking so I added a deprecation.
We want to support people using pre-commit, but we don't want to force it. (b45ac3a)
fix: don't fill-missing metedata by default (#4029)
fix: don't fill-missing metedata by default
This include two changes:
compute_missing to fill_missing, it seems like this change was only applied to get_model_meta, but not throughout the stack. This is technically breaking so I added a deprecation.from_hub, I believe this was unintentionally missedfixes #4027
set fill_missing=False for from_{model}
propegate deprecation
add framework default (d06184f)
fix: Make mteb.get_model compatible with CrossEncoders
fix: Make mteb.get_model compatible with CrossEncoders (#3988)
Made mteg.get_model compatible with CrossEncoders and SparseEncoders
update loader for sparseEncoder
fix import
Simplify structure
Add model_type to sparseEncoder models
remove detection logic of sparsencoder
Add tests and documentation
simplified tests
updated docs
fix docs
fix
fix grammar
Update docs/usage/defining_the_model.md
Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com>
Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com>
Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com>
Co-authored-by: Kenneth Enevoldsen <kennethcenevoldsen@gmail.com>
Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com> (8cacef7)
Add to_disk and from_disk to ModelResults (#3973)
Add to_disk and from_disk to ModelResults
fix signature
change import to high-level
Changed model_validate to model_validate_json
fix tests
remove duplicate path (4a33277)
Fix support for datsets 4.5 with pandas 3 (#3983)
fix test
fix: sanitize type for label during array conversion
lint
revert typo fix
Co-authored-by: Kenneth Enevoldsen <kennethcenevoldsen@gmail.com> (8876169)
rename bm25s to baseline/bm25s (#4007)
rename bm25s to baseline/bm25s
Update mteb/models/get_model_meta.py
Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com>
remove logger message
rename Human to baseline/Human
Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com> (3bab8d9)
fix: simplify dependencies (#4017) (`7f07463`)
7f07463)build image on leaderboard refresh (b9ff905)
model: added Querit/Querit (#3996)
querit_models_add
Querit_Models_Change
Update
format revise
add future
format revise
format revise
last format revison
last last revise
last last last revison
revise
revise
change the instruction
last revison
revise
revise
revise
Co-authored-by: Roman Solomatin <36135455+Samoed@users.noreply.github.com> (5ca3387)
Adding nvidia/nemotron-colembed models (#3941)
Adding nvidia/nemotron-colembed models
add colembed 4b, 8b model meta
fix colembed-3b-v2 model name
update revision for colembed 3b
update revisions
Update mteb/models/model_implementations/nvidia_llama_nemoretriever_colemb.py
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>
Co-authored-by: Roman Solomatin <36135455+Samoed@users.noreply.github.com> (47d59ad)
model: added nomic-ai/nomic-embed-code (#4006)
Add model metadata for nomic-embed-code
Added new model metadata for 'nomic-embed-code'
fix nomic_embed_code
lint
Co-authored-by: Roman Solomatin <36135455+Samoed@users.noreply.github.com> (88e72c6)
model: Adding Ops-Colqwen3 models (#3987)
Create ops_colqwen3_models.py
Refactor OpsColQwen3 model and processor classes
Update model revision in ops_colqwen3_models.py
Remove calculate_probs method and fix model name
Removed the calculate_probs method and updated model name.
format
fix ds name
Co-authored-by: Roman Solomatin <36135455+Samoed@users.noreply.github.com> (7c9bab2)
fix dataset transform (`78687f7`)
fix: BedrockModel initialization arguments
BedrockModel initialization arguments (#3999)fix: add model_name arg to BedrockModel init to prevent multiple values for model_id (0b9de9b)
fix: NomicWrapper get_prompt_name call
fix: add kwargs to pub chem load data
add kwargs to pub chem load data (aeb22cd)
fix: Filled active_parameter_overiride for GritLM/GritLM-8x7B nomic-ai/nomic-embed-text-v2-moe
fix: Filled active_parameter_overiride for GritLM/GritLM-8x7B nomic-ai/nomic-embed-text-v2-moe (#3967)
Filled active_parameter_overiride for ritLM/GritLM-8x7B and nomic-ai/nomic-embed-text-v2-moe
add correct parameters for nomic-ai/nomic-embed-text-v2-moe (dbd4287)
Update mteb/results/task_result.py
fix: leaderboard Nan handling (#3965)
fix leaderboard
fix loading aggregated tasks
Update mteb/results/task_result.py
Co-authored-by: Kenneth Enevoldsen <kennethcenevoldsen@gmail.com>
Co-authored-by: Kenneth Enevoldsen <kennethcenevoldsen@gmail.com> (a369c26)
fix: Add fill_missing parameter in get_model_meta (#3801)
Add compute missing parameter in get_model_meta
fix logs
fix
fix from comments
apply suggestion
fix method
add test and fix logic
address comments
rename compute_missing to fill_missing
Co-authored-by: Roman Solomatin <36135455+Samoed@users.noreply.github.com> (90536d4)
fix: Ensure that retrieval tasks only evaluate on specified subsets instead of all (#3946)
fix dataset loading
update logging
add test (8186392)
# 2.7.6 (2026-01-20) ## Fix * fix: saving aggregated tasks (#3915) fix saving (`ced5f71`)
fix: use num_proc for dataset processing
fix: use num_proc for dataset processing (#3832)
add typehint for encode kwargs
remove num_proc
start adding num_proc
remove all num proc
fix import
add num proc to transform
add to push to hub
use num proc in vidore v2
move num proc to evaluate
pass num proc everywhere
fix tests
fix pylate
fix image text pair
fix num workers
add kwargs to load_data (daf2b6f)
fix: Update metadata to include active number of parameter to ModelMeta
fix: Update metadata to include active number of parameter to ModelMeta (#3837)
Add active parameter column on LB
update ModelMeta with parameters
update ModelMeta of models
Delete parameter_update_results.csv
fix test
fix tests
delete script
rename for consistency
convert active_parameter to property
rename and fix property
update embedding parameters for model2vec models
remove duplicate loading of models
fix
lintter
fix
remove separate method for embedding parameter calculation
fix embedding calculation to pass typecheck
lintter
fix checking
rename active parameters
upd docstring
fix tests
remove n_active_parameters_override from ModelMeta of all models
lintter
rename file instead of merging main
fix tests
correct tests
Delete model total and active parameters - model_parameters.csv
Co-authored-by: Roman Solomatin <36135455+Samoed@users.noreply.github.com> (a45359e)
refactor: split BRIGHT benchmark into individual subset tasks (#3285)
refactor: split BRIGHT benchmark into individual subset tasks
readd bright
readd bright subset tasks
feat: add descriptive stats for BRIGHT subsets retrieval tasks
feat: add top_ranked for excluded_ids handling
change main score to recall@1 for long version
improve BRIGHT task descriptions
add prompts to BRIGHT retrieval tasks
refactor: BRIGHT(v1.1)
calculate descriptive stats for BRIGHTLongRetrieval
update prompts
normalize names in prompts
don't filter tasks
remove filter_queries_without_positives and update revision
don't create top ranked if not necessary
get back naucs
fix instructions
add warning
fix import
Co-authored-by: Roman Solomatin <36135455+Samoed@users.noreply.github.com> (2c9b9e9)
fix: temporarily remove private column from RTEB
fix: temporarily remove private column from RTEB (#3932)
fix: temporarily remove private column from RTEB
Link is still missing the note as I am waiting for @isaac-chung and @Samoed to confirm the write-up.
fixes #3902
added issue link
fix remove mean (Task)
lint
merge in fixes to remove_private (#3940)
fix: exclude private tasks from Borda rank calculation in RTEB
Co-authored-by: bflhc <kunka.xgw@gmail.com>
Co-authored-by: bflhc <kunka.xgw@gmail.com> (b968433)
Add tests verifying preloaded data is preserved.
Co-authored-by: Daniel Svonava <daniel@superlinked.com> (1c5d9c6)
refactor: Activate TC (#3800)
activate tc
activate TC
small import fix
fix imports
fix imports
fix pil import
fix benchmark result validation
full benchmark fix
update
fix unpack imports
upd vllm type (16e0211)
fix: expose ResultCache directly as mteb.ResultCache
fix vllm link (d045d53)
fix: expose ResultCache directly as mteb.ResultCache (#3912)
fix: expose ResultCache directly as mteb.ResultCache
fixes #3910
docs: Update docs usage of ResultCache (3103f97)
fix: computation of results with missing scores (#3874)
fix computation of results with missing scores
fix test
change 0 to nan
change 0 to nan
remove fill_missing_scores (d60e916)
model: add pixie_models (#3938)
model: add pixie_models
Apply lint formatting (8b54f0e)
model: mixedbread-ai/mxbai-edge-colbert-v0-32m and mixedbread-ai/mxbai-edge-colbert-v0-17m (#3931)
Add model: mixedbread-ai/mxbai-edge-colbert-v0-32m and mixedbread-ai/mxbai-edge-colbert-v0-17m
Lintter
Add quotes
Update dataset name
Apply suggestions from code review
Update mixedbread_ai_models.py
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com> (5de2194)
fix: Minor logging fixes by activate LOG rule
LOG rule (#3820)activate logger rule (65313c9)
model: Adding voyage-4 model (#3927)
Adding voyage-4 model
Adding voyage-4 model configs (b80da30)
Update references and citations for ViDoRe V3 benchmark (#3930)
fix: Update references and citations for ViDoRe V3 benchmark
foramat citation
format again
Co-authored-by: Roman Solomatin <36135455+Samoed@users.noreply.github.com> (e7d077e)
model: add nemotron rerank (#3750)
add nemotron rerank
move to nvidia models
removed extra params
Apply suggestions from code review
Co-authored-by: Isaac Chung <chungisaac1217@gmail.com>
remove or
add docstring
Update mteb/models/model_implementations/nvidia_models.py
Co-authored-by: Yauhen Babakhin <ybabakhin@nvidia.com>
Co-authored-by: Isaac Chung <chungisaac1217@gmail.com>
Co-authored-by: Yauhen Babakhin <ybabakhin@nvidia.com> (330a601)
dataset: Add EuroPIRQRetrieval dataset (#3924)
dataset: Add EuroPIRQRetrieval dataset
Removed unnecessary load dataset functions (7966e06)
dataset: add ChemRxivRetrieval task to ChemTEB benchmark (#3923)
dataset: add ChemRxivRetrieval task to ChemTEB benchmark
fix: add descriptive statistics
feat: add ChemTEB v1.1 with ChemRxivRetrieval task
fix: chemteb v1.1 alias (86359fd)
docs: Resolve problems with missing documentation links
docs: Resolve problems with missing documentation links (#3834)
resolve problems with missing documentation links
split into files (0d277cd)
feat: Add vLLM support (#3794)
init
init
Signed-off-by: wang.yuqi <noooop@126.com>
Signed-off-by: wang.yuqi <noooop@126.com>
Signed-off-by: wang.yuqi <noooop@126.com>
Signed-off-by: wang.yuqi <noooop@126.com>
Signed-off-by: wang.yuqi <noooop@126.com>
Signed-off-by: wang.yuqi <noooop@126.com>
Signed-off-by: wang.yuqi <yuqi.wang@daocloud.io>
Signed-off-by: wang.yuqi <noooop@126.com>
Signed-off-by: wang.yuqi <noooop@126.com>
Co-authored-by: Kenneth Enevoldsen <kenevoldsen@pm.me>
Signed-off-by: wang.yuqi <yuqi.wang@daocloud.io>
Signed-off-by: wang.yuqi <yuqi.wang@daocloud.io>
Signed-off-by: wang.yuqi <yuqi.wang@daocloud.io>
Signed-off-by: wang.yuqi <yuqi.wang@daocloud.io>
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>
Signed-off-by: wang.yuqi <yuqi.wang@daocloud.io>
Signed-off-by: wang.yuqi <yuqi.wang@daocloud.io>
fix typing
move type ignore
doc upd
add test
Update Makefile
add support for prompts
add support for prompts
Signed-off-by: wang.yuqi <yuqi.wang@daocloud.io>
Signed-off-by: wang.yuqi <yuqi.wang@daocloud.io>
fix typehints
update pyproject
update pyproject
update pyproject
The pooling + dp fails to run.
fix uv lock
fix docs
simplify conflicts
upd lock
upd lock
Update docs/advanced_usage/vllm_wrapper.md
Co-authored-by: Kenneth Enevoldsen <kenevoldsen@pm.me>
Co-authored-by: Kenneth Enevoldsen <kenevoldsen@pm.me>
Co-authored-by: Kenneth Enevoldsen <kenevoldsen@pm.me>
Co-authored-by: Kenneth Enevoldsen <kenevoldsen@pm.me>
Co-authored-by: Kenneth Enevoldsen <kenevoldsen@pm.me>
Co-authored-by: Kenneth Enevoldsen <kenevoldsen@pm.me>
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>
Signed-off-by: wang.yuqi <noooop@126.com>
Signed-off-by: wang.yuqi <noooop@126.com>
Signed-off-by: wang.yuqi <noooop@126.com>
Signed-off-by: wang.yuqi <yuqi.wang@daocloud.io>
Co-authored-by: Kenneth Enevoldsen <kenevoldsen@pm.me>
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>
Co-authored-by: Roman Solomatin <36135455+Samoed@users.noreply.github.com> (4d568a8)
model: Update the nemo retriever reversions to avoid error when loading the model (#3925)
Update the nemo retriever versions to fix the crash issue with visual_config
Update mteb/models/model_implementations/nvidia_llama_nemoretriever_colemb.py
Update mteb/models/model_implementations/nvidia_llama_nemoretriever_colemb.py
Co-authored-by: Kenneth Enevoldsen <kenevoldsen@pm.me> (dcd31fa)
model: Adding voyage-4-large, voyage-4 and voyage-4-lite (#3885)
Adding voyage-4-large and voyage-4-lite
Adding voyage-4-large and voyage-4-lite
Adding voyage-4
Reverting voyage-4 (as the tokenizer is not yet available publicly)
added superseeded_by
Co-authored-by: Kenneth Enevoldsen <kennethcenevoldsen@gmail.com> (64ce6ba)
trigger on dependencies change (`227ce95`)
fix: KoVidore2EnergyRetrieval revision fix (#3913) (`0221b94`)
0221b94)add model: bflhc/Octen-Embedding-0.6B (#3906) (c5c481f)
model: mixedbread-ai/mxbai-rerank-large-v1 (#3905)
Add model: mixedbread-ai/mxbai-rerank-large-v1
apply suggestions
Added xsmall and base version of reranker models
lintter (cb29623)
Add typehint for encode kwargs (#3831)
add typehint for encode kwargs
remove num_proc
remove all num proc
fix import
fix docstrings (a417426)
add dataset: KoViDoRe(v2) (#3876)
add dataset: KoViDoRe v2
fix citation format
add direct loading
lint format
delete benchmark language view
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com> (24f3f85)
model: add missing sentence transformers and jina models
fix: nv embed version (#3715)
fix nv embed wrapper
try to fix
fix sbert version (d3020f3)
test: Add HF Space Dockerfile using pre-built leaderboard image
fix: Simplify conflicts (#3875)
simplify conflicts
add lock
remove torch (1bc3fba)
test: Add HF Space Dockerfile using pre-built leaderboard image (#3838)
Add HF Space Dockerfile using pre-built leaderboard image
Adds a lightweight Dockerfile for HuggingFace Space deployment that uses the pre-built ghcr.io/embeddings-benchmark/mteb/leaderboard image as base. Also adds a workflow to test the Dockerfile.
🤖 Generated with Claude Code
Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>
Delete .github/workflows/hf_space_docker.yml
test: Add CI workflow for HF Space Dockerfile validation
Co-authored-by: Claude Opus 4.5 <noreply@anthropic.com>
Co-authored-by: MTEB Agent <agent@example.com> (b905e27)
update bflhc/Octen-Embedding-8B revision (1e78793)
dataset: Vietnamese VN-MTEB TVPLRetrieval, NanoClimateFEVER-VN, NanoFEVER-VN, NanoDBPedia-VN, NanoNQ-VN, NanoHotpotQA-VN, NanoMSMARCO-VN (#3810)
[ADD] Vietnamese VN-MTEB TVPLRetrieval, NanoClimateFEVER-VN, NanoFEVER-VN, NanoDBPedia-VN, NanoNQ-VN, NanoHotpotQA-VN, NanoMSMARCO-VN
[UPDATE] descriptive stats
[UPDATE] bibtext
[UPDATE] dataset path
[UPDATE] nano db pedia retrieval
[UPDATE] size dataset from 1M corpus to 100k
[ADD] add note about what's different in nano version
[ADD] TVPLRetrieval description (adb5b42)
docs: Fix docs build strict mode errors
docs: Fix docs build strict mode errors (#3809)
fix: resolve mkdocs strict mode errors
fix: remove duplicate line in installation.md
build: add --strict flag to mkdocs build
fix: resolve invalid BibTeX keys in task citations
feat: filter BibTeX warnings in strict docs build
Update docs/usage/defining_the_model.md
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>
fix: dynamic mkdocs path discovery for CI
fix: improve docs build script with clear warning counts
fix: resolve 6 real docs build warnings
fix: remove broken PylateSearchEncoder reference
Remove unused build scripts
docs: wrap multimodal example as code snippet
fix: export SklearnModelProtocol for docs API
docs: add API reference for SklearnModelProtocol
fix: remove SklearnModelProtocol export to avoid circular import
feat: add SklearnModelProtocol docs with lazy import
fix: convert Sphinx cross-references to MkDocs syntax for proper linking
Convert :class: Sphinx syntax to [Text][module.path] MkDocs syntax to ensure cross-references are properly clickable in the generated documentation.
🤖 Generated with Claude Code
Co-Authored-By: Claude Sonnet 4 <noreply@anthropic.com>
chore: remove SklearnModelProtocol docs to avoid circular import
chore: remove SklearnModelProtocol export from _evaluators
style: fix indentation in _evaluators/init.py
fix: enable mkdocs build --strict without warnings
🤖 Generated with Claude Code
Co-Authored-By: Claude Sonnet 4 <noreply@anthropic.com>
fix: rename evaluator to multilabel_classifier to avoid type conflict
fix: use evaluator_model instead of evaluator to avoid type conflict
🤖 Generated with Claude Code
Co-Authored-By: Claude Sonnet 4 <noreply@anthropic.com>
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>
Co-authored-by: Claude Sonnet 4 <noreply@anthropic.com> (3723d27)
fix: Extend framework annotations for ModelMeta (#3819)
Update framework and filter based on them
update ModelMeta of models
update ModelMeta of models
update ModelMeta of models
update ModelMeta of models
update ModelMeta of models
update ModelMeta of models
update ModelMeta
add csv
update ModelMeta
added framework to ModelMeta
update ModelMeta of models
update ModelMeta of models
update framework in ModelMeta of models
update framework
update framework
update framework in ModelMeta
fix tests
Add models
fix tests
add tags extraction in from_hub()
fix typecheck
apply suggestions
apply suggestions
keep only static method
delete csv and script (d033c24)
fix dataset generation tags (#3835) (bf2627a)
model: Add SauerkrautLM-ColPali visual document retrieval models (#3804)
model: Add SauerkrautLM-ColPali visual document retrieval models
Add inference code and requirements for SauerkrautLM-ColPali visual document retrieval models.
These are multi-vector embedding models based on the ColPali architecture:
All models produce 128-dimensional embeddings per text/image token and use MaxSim (late interaction) for retrieval scoring.
Model checkpoints:
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>
fix: Update release_date to 2025-12-20
fix: address review comments - remove partial, add adapted_from and training_datasets
Update mteb/models/model_implementations/slm_models.py
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>
fix: import COLPALI_CITATION from colpali_models and add model_type
add training datasets
fix: remove section headers and use PyPI package instead of Git URL
fix: resolve merge conflicts and remove section headers
fix: use COLPALI_TRAINING_DATA for training_datasets
fix: use exact n_parameters and memory_usage_mb values from HuggingFace
don't build 3.14
lint
Co-authored-by: David Golchinfar <d.golchin@web.de>
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>
Co-authored-by: Roman Solomatin <36135455+Samoed@users.noreply.github.com> (44e9b20)
fix: Add leaderboard docker workflow
fix: Add leaderboard docker workflow (#3828)
Add GitHub workflow to test leaderboard Dockerfile
🤖 Generated with Claude Code
Co-Authored-By: Claude Sonnet 4 <noreply@anthropic.com>
🤖 Generated with Claude Code
Co-Authored-By: Claude Sonnet 4 <noreply@anthropic.com>
🤖 Generated with Claude Code
Co-Authored-By: Claude Sonnet 4 <noreply@anthropic.com>
🤖 Generated with Claude Code
Co-Authored-By: Claude Sonnet 4 <noreply@anthropic.com>
🤖 Generated with Claude Code
Co-Authored-By: Claude Sonnet 4 <noreply@anthropic.com>
🤖 Generated with Claude Code
Co-Authored-By: Claude Sonnet 4 <noreply@anthropic.com>
🤖 Generated with Claude Code
Co-Authored-By: Claude Sonnet 4 <noreply@anthropic.com>
fix: remove redundant typer dependency from leaderboard extras
fix: only push Docker images from main branch
fix: use COPY instead of git clone in Dockerfile
Co-authored-by: Claude Sonnet 4 <noreply@anthropic.com> (c7c04e5)
fix: Allow passing device to model
fix: Allow passing device to model (#3812)
Allow passing device to model
revert incorrect modification and fix typeerror
add device to get_model and address comments
Correct CDEWrapper (17ef363)
fix: remove redundant pip install uv commands from Makefile
ci: Switch CI to use uv (#3702)
use uv to all make commands
read the docs a bit more...
try out system flag
fix: remove redundant pip install uv commands from Makefile
Removes duplicate uv installations that were conflicting with the properly configured uv from astral-sh/setup-uv GitHub Action. The GitHub Action already installs and configures uv correctly, so the Makefile pip installs were overwriting this configuration and causing "No system Python installation found" errors.
🤖 Generated with Claude Code
Co-Authored-By: Claude Sonnet 4 <noreply@anthropic.com>
The astral-sh/setup-uv GitHub Action configures uv to manage its own Python installations, not to use system Python. The --system flag was causing "No system Python installation found" errors because uv expects to use its managed Python environment.
🤖 Generated with Claude Code
Co-Authored-By: Claude Sonnet 4 <noreply@anthropic.com>
This aligns with uv's recommended project workflow and should resolve the CI environment issues we were experiencing.
🤖 Generated with Claude Code
Co-Authored-By: Claude Sonnet 4 <noreply@anthropic.com>
Updated workflows:
🤖 Generated with Claude Code
Co-Authored-By: Claude Sonnet 4 <noreply@anthropic.com>
🤖 Generated with Claude Code
Co-Authored-By: Claude Sonnet 4 <noreply@anthropic.com>
remove 3.14
try out 3.14 again with python_full_version
specify torch version for pylate dep
try to skip colpali
try split torch
Add --no-sync flag and group/extra flags to uv run commands
Address review comments from PR #3702:
Add --no-sync to all uv run commands in Makefile for:
Add appropriate group/extra flags to uv run commands:
Update CI workflows to use --no-sync and appropriate groups:
These changes improve performance while maintaining compatibility for contributors who prefer using pip directly.
🤖 Generated with Claude Code
Co-Authored-By: Claude Sonnet 4.5 <noreply@anthropic.com>
try removing install block
add back install block
remove install block in doc CI without --no-sync
add uv lock file
replace install-for-test with just install
install pre-commit with uv
fix doc workflow
address review comments
remove no-sync from run-leaderboard make command
remove --no-sync from selected make commands
update typechecking
fix type checking
sync to install
fix tests
test pre-commit setup
remove test file
fix: separate install and install-for-tests with uv commands
fix: add leaderboard extra to typecheck command for gradio imports
fix: add faiss-cpu extra to test targets
fix: update CI workflows for uv dependency management
docs: update all documentation for uv migration
This provides users and contributors with modern, fast uv tooling while maintaining backward compatibility with existing pip workflows.
🤖 Generated with Claude Code
Co-Authored-By: Claude Sonnet 4 <noreply@anthropic.com>
Co-authored-by: Claude Sonnet 4 <noreply@anthropic.com>
Co-authored-by: Roman Solomatin <36135455+Samoed@users.noreply.github.com> (75743f1)
043ea38)Your coding agent can read these notes before it upgrades. Set up the MCP server →