NewYour coding agent can read the release notes before it upgrades.Set up the MCP server →
PyPI · #2804 most downloaded on PyPI
Massive Text Embedding Benchmark
Last release today
01 Oct 2026
Ships fairly regularly
a new release about every 1 weeks
Nearly every release is documented
notes for 60 of the last 60 stable releases
Nothing withdrawn
no release was ever pulled
4 years old
811 releases · first in 2022
One column per quarter.
fix: results mteb version parse
ci: revert leaderboard refresh trusted publishing
revert leaderboard refresh (14cf6d6)
ci: Setup hf trusted publishing (#4804)
setup hf trusted publishing
remove comment (a1f0a62)
infer modalities (2c17992)
model: Update VultronRetrieverPrime-Qwen3.5-8B metadata to reflect new repository path (e091c0d)
Late-interaction (ColBERT MaxSim) visual document retriever: ColQwen3.5, dim 320, 8.4B params, Apache-2.0, 6 languages. Reuses the existing ColQwen3_5Wrapper. Official ViDoRe scores V1 0.9208 / V2 0.6818 / V3 0.6472 (results PR to follow).
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> (99a830c)
Update Querit model implementation: Supplementing the citation information of the Querit
Co-authored-by: zhongyunfei <zhongyunfei@baidu.com> (cfce4f2)
model: Add VIRTUE multimodal embedding models (Sony VIRTUE-2B/7B-SCaR) (#4822)
Add VIRTUE multimodal embedding models (Sony VIRTUE-2B/7B-SCaR)
address review feedback (41f6c7e)
Fix Querit model implementation: Supplementing the base model information of the Querit/Querit-4B (#4819)
Fix Querit model implementation: Supplementing the base model information of the Querit/Querit-4B
Fix Querit model implementation: Supplementing the base model information of the Querit/Querit-4B
Co-authored-by: zhongyunfei <zhongyunfei@baidu.com> (240da5b)
mveb: fix and unify domain tags across all 50 source datasets (#4738)
mveb: fix and unify domain tags across all 50 source datasets
The MVEB+ video task set had inconsistent and partially-wrong domains
tags. Issues fixed:
All 50 unique source datasets across 184 video tasks now have consistent, non-empty domain tags. Verified by re-importing every task: 184 tasks load cleanly.
Tags use only the existing TaskDomain Literal vocabulary in task_metadata.py; no new domains added.
Adds 5 video content domains to TaskDomain (Activity, Instructional, Egocentric, Nature, Animation) and re-tags datasets that were mislabeled or under-characterized, so the domain set actually reflects benchmark content:
Scene is now reserved for genuine visual-scene content (WorldSense). All 184 video tasks load; every domain validates against TaskDomain.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> (343df1a)
Update Querit model implementation: 4B version of Querit-Reranker newly open-sourced (#4808)
Update Querit model implementation: 4B version of Querit-Reranker newly open-sourced.
Update Querit model implementation: 4B version of Querit-Reranker newly open-sourced.
Update Querit model implementation: 4B version of Querit-Reranker newly open-sourced.
Apply suggestions from code review
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>
Co-authored-by: zhongyunfei <zhongyunfei@baidu.com>
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com> (4976113)
Rename number_texts_intersect_with_train to samples_in_train (#4809) (e76f291)
polish MVEB leaderboard names + icons (#4803)
benchmarks: polish MVEB leaderboard names + icons
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> (5039c18)
ci: Update healthcheck for new leaderboard
fix: Support scikit-learn 1.9 in ZeroShotClassification
scikit-learn 1.9 raises "ValueError: Mix of label input types" when classification metrics receive string y_true with numeric y_pred. Zeroshot predictions are always integer indices into the candidate labels, so string dataset labels are now mapped to their candidate index before scoring. Unmappable string labels raise a clear error instead of silently scoring 0.0, which is what scikit-learn < 1.9 did.
Removes the <1.9.0 pin introduced as a stopgap in #4783.
Fixes #4784
Co-authored-by: Kenneth Enevoldsen <kennethcenevoldsen@gmail.com> (5e8f0e1)
fix: normalize benchmark definitions
fix: normalize benchmark definitions (#4792)
fix: remove *extended and MAEB+
These are mostly for reproducibility so we have moved them to their respectives script repos.
https://github.com/embeddings-benchmark/maeb-paper/pull/3 https://github.com/embeddings-benchmark/mveb-paper/pull/1
remove script to maeb
fix imports
fix: normalize benchamrk definitions
Normalize to:
7071da1)model: NightOwl-CodeEmbedding (#4791)
model: add NightOwl CodeEmbedding metadata
fix: remove programming language codes from model metadata
fix: update memory usage for NightOwl CodeEmbedding model
fix: update NightOwl CodeEmbedding model metadata
fix: update revision for NightOwl CodeEmbedding model (b691769)
fix: remove *extended and MAEB+
fix: remove *extended and MAEB+ (#4786)
fix: remove *extended and MAEB+
These are mostly for reproducibility so we have moved them to their respectives script repos.
https://github.com/embeddings-benchmark/maeb-paper/pull/3 https://github.com/embeddings-benchmark/mveb-paper/pull/1
remove script to maeb
fix imports (f795558)
Merge multimodal sentence transformers (#4785)
merge multimodal
add warning & replace usages (d5430cc)
model: Add video support to qwen3-vl embedding (#4699)
add video to qwen3
upd implementation
upd revision
fix MultimodalInstructSentenceTransformerModel
disable double sampling
move imports inside
simpliffy wrapper
fix typing (beee210)
rollback changes in deprecated_evaluator
ci: Update actions (#4780)
update github
remove remove
upd
remove free space
rename step
fix conda wargning
add concurrency to lb
remove concurrences (52a6f73)
feat: Add evaluation runtime for indexing and retrieval (#4639)
Add evaluation runtime for indexing and retrieval
change timer from maintaining state to passed as an attribute
add timer argument to load_data
add timer argument to evaluate
removed timer from kwargs
update plots for split/subsets
fix tests
fix typing errors
typing errors
change typing
correct typing
apply changes from review
update to handle overwritten load_data
changes from review
added evaluation phases merging logic and fix typecheck
added * seprator at all places in load_data
changes from review
added split/subset at all places
change split/subset in other functions as well
small typecheck update
changes from review
reordering
implement for clustering task
implement for classification task
remove logger statement from clustering
implement for pair classification task
remove logger statement for pair classification task
implement for bitext mining task
implement for STS task
implement for summarization task
implement for sklearn evaluator
fix lintter
Delete .specstory/history/2026-04-23_09-55Z-testing-mechanism-for-new-datasets.md
minor changes from review
fix typecheck
update stacklevel=2
add TimingStack as default argument
fix evaluators tests
update deprecate_evaluator
minor fix
modified implementation to handle override task
added utlity function in TaskResult to plot timings
Added docs
changes from review
added commnts
simplify plot calling and update docs with example
changed phases naming format in plots
add tests and minor changes
update to handle indexing and searching phase
rollback changes in deprecated_evaluator
move import to top level and minor fix in tests
fix import
update tests
change Scoring to aggregate level in Classification task
remove unwanted file
fix lintter and typecheck errors after merge
revert changes in other classification task
changes from copilot review and add new test
changes from review
update condition
make lint
updated docs
Co-authored-by: Roman Solomatin <36135455+Samoed@users.noreply.github.com> (6a3e816)
fix: Update lock and remove python limit fo pylate and colbert_engine
fix: Update lock and remove python limit fo pylate and colbert_engine (#4783)
update lock
fix typing
fix typing
fix tests
remove torchcodec from missing imports
fix zero shot
pin skelearn (d93a2c0)
fix: Load aggregated results from task results
fix: Load aggregated results from task results (#4719)
Load aggregated results from task results
changes from copilot review
changes from review and fix typecheck
simplify condition
minor fixed and updated tests
minor fixes
remove try-except
added comments
update tests
fix test
changes from review (2a87342)
fix: Update BidirLM models patched for transformers 5.x and addition of LongEmbed benchmark instructions
Update BidirLM models patched for transformers 5.x
The BidirLM checkpoints were patched to run on transformers>=5.0 (their main branch); this updates the reference implementation accordingly:
f21e2fc)benchmarks: add MVEB benchmark suite + leaderboard Video menu
fix: PC multimodal support (#4645)
fix PC multimodal
fix modality get
fix tasks
fix category
fix multimodal pair tasks
typecheck
fix typing
upd docs
change to mapping
change to mapping (a0e219c)
benchmarks: add MVEB benchmark suite + leaderboard Video menu (#4763)
benchmarks: add MVEB benchmark suite + leaderboard Video menu
Adds the MVEB (Massive Video Embedding Benchmark) benchmark objects to main so the leaderboard and get_benchmark() can resolve them. The underlying tasks are already on main; this adds only the curated benchmark groupings and their registration.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Addresses review:
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
Co-authored-by: Roman Solomatin <36135455+Samoed@users.noreply.github.com> (b5e03a4)
# 2.14.5 (2026-06-06) ## Fix * fix: Make project releasablle (#4777) fix pyproject (`e0b4aa2`)
fix: Don't specify hf_subsets for aggregated tasks
The eval_langs are computed by the object itself
Added tests to ensure this going forward.
Fixes #4473 (ada5a81)
refactor: move generate_model_card to ModelMeta method (#4209)
refactor: move generate_model_card to ModelMeta method
Moves the generate_model_card logic from the CLI module into ModelMeta as an instance method, fixing a mutable default argument bug along the way.
refactor: add push_eval_results CLI wrapper
refactor: revert push_eval_results CLI wrapper
Revert "refactor: revert push_eval_results CLI wrapper"
This reverts commit 3f4a9a2574cad7e728f6e32e5d3a048746a85104.
refactor: call push_eval_results from generate_model_card
refactor: remove standalone push_eval_results CLI wrapper
limit args
simplify pushing
rename function
fix typing
update gold
remove release date
Co-authored-by: Copilot <copilot@github.com>
Co-authored-by: Roman Solomatin <36135455+Samoed@users.noreply.github.com>
Co-authored-by: Kenneth Enevoldsen <kennethcenevoldsen@gmail.com>
Co-authored-by: Copilot <copilot@github.com> (5246403)
Update codefuse_models.py (9ae31ec)
fix: Allow kwargs in push_dataset_to_hub to allow pushing a private dataset
fix: Allow kwargs in push_dataset_to_hub to allow pushing a private dataset (#4508)
fix: Allow kwargs in push_dataset_to_hub to allow pushing a private dataset
Fixes #4506
Co-authored-by: Copilot <copilot@github.com>
Co-authored-by: Copilot <copilot@github.com>
Co-authored-by: Roman Solomatin <36135455+Samoed@users.noreply.github.com> (93db557)
fix: rework gemini using asyncio (#4507)
fix: rework gemini using asyncio
As gemini 2 only allows one sample at a time (see #4490) I refactored it to use async calls instead. We could also use the batch API, but then it could take 24 hours for a batch (probably wont).
The new script got 0.47 for 'AILAStatutes', with the previous script obtaining 0.47.
I also refactored the different google models into three seperate scripts. No code was changed in the refactor.
Fixes #4490
Co-authored-by: Copilot <copilot@github.com>
added batch size
format and fix tests
Co-authored-by: Copilot <copilot@github.com>
Co-authored-by: Copilot <copilot@github.com> (9834ff4)
fix: Added Git Actions using Command Pattern
fix: Added Git Actions using Command Pattern (#4329)
Added Git Actions using Command Pattern
lintter
Added better error logging in case of rollback
added init.py and github group
update docstring of function based on copilot comment
update docstring of PushToFork
fix dependencies
add in Makefile install and fix import
fix GithubException import
fix install command in Makefile
apply changes from review
update description of CopyResultsAction
added github in install-for-tests
make module private
make folder private
Fix undo in case of failure
remove CopyResultsAction
fix imports
added init.py
fix import
fix lintter
remove comments
Added pytest.importorskip
fix lint
fix lint
Remove monkeypatch from tests and update tests
fix default branch in test
fix test and cleanup
Co-authored-by: Copilot <copilot@github.com>
apply changes from review
setting email and username in config only when not set
Co-authored-by: Roman Solomatin <36135455+Samoed@users.noreply.github.com>
Co-authored-by: Copilot <copilot@github.com> (b2dfda9)
Add VideoCollator/AudioCollator for proper video/audio support,
remove qwen_omni_utils dependency, add L2 normalization, and
expose fps/max_frames/num_frames/max_audio_length params. (eae2080)
b804411)fix ci (baa25a9)
[MVEB] Add video statistics (#4456)
start video statistics
add some more video processing
more changes
simplify statistics calculation
refactor updates
update mock video tasks
fix typing
make all tasks support multimodality
add minimal version for av
fix sts
fix tests
update skip rule
activate video statistics check
add missing tasks
move statistics functions
update column format
upd lock
add resolutions
add to ignore newer datasets (6d862aa)
add remaining v2a and a2v tasks for video retrieval datasets (#4504) (0294578)
add v2a and a2v tasks for valid video retrieval datasets (#4494)
add v2a and a2v tasks to MSR-VTT
Made-with: Cursor
Made-with: Cursor
Made-with: Cursor
Made-with: Cursor
Made-with: Cursor
Made-with: Cursor
Required for v2a and a2v retrieval tasks across video datasets (DiDeMo, YouCook2, Shot2Story20K, VALOR-32K, VATEX, MSR-VTT).
Made-with: Cursor (6119afb)
Fix an incorrect retrieval example in docs (#4496)
Fix an incorrect retrieval example in docs
Incorrect assignment of the data to self.dataset. It is None, should be reinitialized
Co-authored-by: Roman Solomatin <36135455+Samoed@users.noreply.github.com> (cf4a7bf)
f46321b)[MVEB] Move collators to separate file (#4482)
move collators
make parameters by name
fix last import (25a0991)
[MVEB] Adding HMDB51 Task (Clustering) (#4488) (8d381dc)
[MVEB] Omni Embed Nemotron (#4388)
fix: Reclassify SIBFLEURS as AudioClassification instead of AudioMultilabelClassification
fix: Move SIBFLEURS descriptive stats to AudioClassification
add: NVIDIA Omni-Embed-Nemotron-3B model wrapper
fix: convert video frame tensors to PIL images for qwen_omni_utils
fix: add default system prompt to suppress Qwen audio output warning
fix: correct text max_length to 204800 per model card
refactor: remove qwen_omni_utils dependency from omni-embed-nemotron wrapper
fix: address PR review - lint, vidore training datasets, lowercase vars
refactor: use SentenceTransformerMultimodalEncoderWrapper instead of custom wrapper
refactor: add processing_kwargs for video/audio in ST wrapper
fix: support ST multimodal wrapper for omni-embed-nemotron
RAVDESSAVClustering v_measure: 0.0794 MSRVTTV2T ndcg_at_10: 0.3196
change placement
refactor: filter non-modality keys in ST multimodal wrapper instead of per-task
refactor: move audio unwrap to ST base wrapper, simplify omni-embed-nemotron
if "audio" in batch).
Works around sentence-transformers#3732.With the merge of upstream/main (PR #4356), ModelMeta now supports extra_requirements_groups. Declare multimodal_sbert so the runtime enforces the pyproject pins at load time.
fix: simplify multimodal wrapper docstring
fix tests
partly fix typing
add: num_frames param to OmniEmbedNemotronWrapper
add: e5-omni-3B and e5-omni-7B model implementations
fix: use hyphenated extra name multimodal-sbert for pip compatibility
fix: normalize multimodal_sbert to multimodal-sbert in pyproject.toml
remove: e5-omni models and results per reviewer request
Moving e5-omni to a separate PR to not block this one.
Co-authored-by: Roman Solomatin <36135455+Samoed@users.noreply.github.com> (4a319fa)
add mteb/VGGSound_AV_RETRIEVAL dataset (#4479)
add mteb/VGGSound_AV_RETRIEVAL dataset
update to video captions
refine VGGSound-AV: guard av_caption map and improve descriptions/prompts
Made-with: Cursor (29c2a27)
Apparently pipy normalizes package name which gives a probleM that only appear when install from pipy:
Apparently pipy normalizes package name which gives a probleM that only appear when install from pipy:
ValueError: Unknown extras group(s) for mteb: ['google_genai']. Available: ['ark', 'audio', 'blip2', 'bm25s', 'codecarbon', 'cohere', 'colpali-engine', 'colqwen3', 'eager-embed', 'embeddinggemma', 'faiss-cpu', 'flagembedding', 'flash-attention', 'google-genai', 'gritlm', 'image', 'jina', 'jina-clip', 'jina-v4', 'leaderboard', 'llama-embed-nemotron', 'llama-nemotron-colembed-vl', 'llama-nemotron-embed-vl-1b-v2', 'llm2vec', 'mctct', 'model2vec', 'msclap', 'muq', 'nemotron-colembed-vl-v2', 'nomic', 'open-clip-torch', 'openai', 'peft', 'pylate', 'qwen-omni-utils', 'qwen-vl', 'sauerkrautlm-colpali', 'siglip', 'speechbrain', 'timm', 'torch-vggish-yamnet', 'vertexai', 'video', 'vllm', 'voyage-v', 'voyageai', 'wav2clip', 'xet', 'xformers', 'youtu']
pep: https://peps.python.org/pep-0685/
initially proposed a fix here, but discovered that it was added: https://github.com/embeddings-benchmark/mteb/pull/4384
in this commit: https://github.com/embeddings-benchmark/mteb/pull/4384/changes/45f1419b6516ab0aa4114696a1acb6d1919ffcec
This PR just bumps the version to release the fix.
(cc @isaac-chung)
Co-authored-by: Kenneth <kennethenevoldsen@gmail.com> (f3f242d)
add mteb/AudioCaps_AV dataset (#4478)
add mteb/AudioCaps_AV dataset
Add import for AVMemeExam retrieval classes
fix license, domains, date, and va2t/t2va prompts for AudioCaps-AV
Made-with: Cursor
AudioCaps_AV has a single caption field containing human-written audio scene descriptions (what is heard, not seen). Descriptions now clarify the cross-modal nature of v2t/t2v and that va2t/t2va are audio-focused. Prompts updated to match: 'sounds in', 'audio description', 'what is heard'.
Made-with: Cursor
This reverts commit 9139801a1fe2cf015f4e00be07e9aa128750fd6e.
This reverts commit 97a9e4904fdb6b468439dfe683cd8ab1f12a3660.
Co-authored-by: AdnanElAssadi56 <115242814+AdnanElAssadi56@users.noreply.github.com> (0b71456)
[MVEB] Add Qwen Omni Video Support (#4412)
fix: Reclassify SIBFLEURS as AudioClassification instead of AudioMultilabelClassification
fix: Move SIBFLEURS descriptive stats to AudioClassification
feat: add video support to Qwen Omni models and remove qwen_omni_utils dependency
Add video modality to QwenOmniWrapper by handling video frames directly via the processor instead of going through qwen_omni_utils.process_mm_info, which only supports URL/path loading and crashes on pre-loaded tensors. Video frames are pre-resized via smart_resize to match the model's expected pixel range. Also fixes dtype mismatch for Qwen3 models and updates max_audio_length to 300s to match the model's preprocessor config (n_samples).
Add video modality to QwenOmniWrapper by handling video frames directly via the processor instead of going through qwen_omni_utils.process_mm_info, which only supports URL/path loading and crashes on pre-loaded tensors. Video frames are pre-resized via smart_resize to match the model's expected pixel range. Also fixes dtype mismatch for Qwen3 models and updates max_audio_length to 300s to match the model's preprocessor config (n_samples).
fix: remove unused extra_requirements_groups for qwen_omni_utils
fix: remove unnecessary dtype cast per reviewer feedback (857b425)
add image to LCO embed model (#4384)
add image and video to LCO embed model
Fix extras group comparison to be underscore/hyphen insensitive
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com> (fb64050)
add method for benchmark card creation
fix: HF benchmark result (#4344)
init benchmark eval results
add get score to benchmark
update scoring
add method for benchmark card creation
fix typing (990c1cf)
[MVEB] Add fps implementation to Video Sampling (#4441)
fix: Reclassify SIBFLEURS as AudioClassification instead of AudioMultilabelClassification
fix: Move SIBFLEURS descriptive stats to AudioClassification
refactor: add FPS-based video frame sampling to collator
FramesCollator and VideoCollator now support two modes:
Existing callers (PE-AV, random baseline) switched to num_frames to preserve their current fixed-sample behavior.
Both models now use the default FPS-based mode (fps=2.0, max_frames=256) instead of fixed num_frames. This gives duration-proportional frame coverage across videos of different lengths.
Allow fps, max_frames, num_frames, and max_samples to be configured via the PE-AV wrapper constructor instead of being hardcoded. Defaults to fps=2.0 matching the standard video understanding rate.
fix: address PR review - rename max_frames to max_fps_frames, raise on conflicting args
fix: rename max_fps_frames back to max_frames, clarify docstrings
fix: pass fps=None, num_frames=16 to 16-frame PE-AV variants
The *-16-frame checkpoints were trained with fixed 16-frame uniform sampling (processor config has do_sample_frames=true, num_frames=16). Without explicit loader_kwargs, the collator used the default fps=2.0, producing ~40 frames on typical clips that the processor then re-sampled down to 16 — a distribution shift from training. Setting num_frames=16 makes the collator do the sampling directly, and the processor's built-in sample becomes an identity no-op.
fix: clarify fps docstrings - downsamples only, no upsampling (9363ea7)
Don't display license links in the documentation (#4465)
Fixes #4461 (46582d9)
[MVEB] Adding UCF101 Task (Clustering) (#4454) (b8b3722)
leaderboard: add MTEB(spa, v1) to Language-specific section (#4217)
Add MTEB(spa, v1) to leaderboard language-specific menu
Co-authored-by: Clemente <clemente@Clementes-MacBook-Pro.local> (e5521a6)
Add VALOR-32K retrieval tasks (#4453)
Add VALOR-32K retrieval tasks (v2t, t2v, va2t, t2va)
Adds four bidirectional multimodal retrieval tasks for the VALOR-32K dataset (mteb/VALOR-32K), a vision-audio-language benchmark with 3,491 test samples.
Made-with: Cursor
Made-with: Cursor (792f61f)
fix: drop unused modality columns in dataloader for cross-modal tasks
fix: drop unused modality columns in dataloader for cross-modal tasks (#4440)
fix: handle None text/image in multimodal retrieval tasks
Cross-modal retrieval tasks (CIRRIT2IRetrieval, NIGHTSI2IRetrieval, Fashion200kI2TRetrieval, VisualNewsI2TRetrieval) have corpus/query items where text or image can be None for single-modality entries.
Closes #4436
Normalize None text to "" in _combine_queries_with_instruction_text, matching the existing pattern in _corpus_to_dict. Revert random_baseline and collation changes as they're no longer needed.
Cross-modal retrieval tasks have None values for modalities not used by that side of the retrieval (e.g. text=None in image-only corpus for it2i tasks). Instead of adding None-guards throughout the collate function and models, drop columns for modalities not needed for the current prompt type in _prepare_dataset. The task category (e.g. it2i) already encodes which modalities each side needs.
Closes #4436
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com> (eaf2c9d)
Apply suggestions from code review
fix: remove columns with none (#4446)
remove columns with none
Apply suggestions from code review
Co-authored-by: Kenneth Enevoldsen <kenevoldsen@pm.me>
Co-authored-by: Kenneth Enevoldsen <kennethcenevoldsen@gmail.com>
Co-authored-by: Kenneth Enevoldsen <kenevoldsen@pm.me> (63c0af6)
[MVEB] Add VGGSound audio-visual classification tasks
Add VGGSoundVAClassification (video+audio, va2c) and VGGSoundVClassification (video-only, v2c) for the VGGSound audio-visual dataset (Chen et al., ICASSP 2020). Dataset contains 9,888 test clips across 308 sound classes from YouTube videos. Audio is the primary signal in the original task; the v2c variant serves as a video-only baseline. Uses 5-fold cross-validation since the released split only contains test. Follows the standard MVEB classification task structure. Addresses part of #4130 (MVEB Overview - Classification).
Co-authored-by: Yashwanth Devavarapu <yashwanthdevavarapu@Yashwanths-MacBook-Pro.local> (bf113b4)
add mteb/Shot2Story20K_test dataset (b597f37)
Add YouCook2_val retrieval tasks (#4432)
Add YouCook2_val retrieval tasks (V2T, T2V, A2T, T2A)
Made-with: Cursor
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>
add stats
update
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com> (f0e5ab6)
Add VATEX retrieval tasks (#4433)
Add VATEX_test_1k retrieval tasks (V2T, T2V, A2T, T2A)
Made-with: Cursor
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com> (f93dc75)
model: add BidirLM/BidirLM-Omni-2.5B-Embedding (#4370)
feat: add BidirLM/BidirLM-Omni-2.5B-Embedding model implementation
fix: address reviewer comments on BidirLM-Omni-2.5B-Embedding:
refactor: update to sentence transformers 5.4 and rely on encode function to get embedding
Apply suggestions from code review
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>
Feat: instruct with prompt args chat template
Fix: rely on EncodeKwargs for encoder function
Feat: improve readability
Refactor: Change how modality are passed to encode
Fix: lint error
Refactor: args encode
Refactor: Import from Bidir
comments update
Simplify get instruction (ne need for _lookup_prompt stripped)
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com> (9e5c592)
update query filtering (`65025bb`)
update query filtering (65025bb)
[MVEB] Add SomethingSomethingV2 video classification task (#4434)
[MVEB] Add SomethingSomethingV2 video classification task
fix: correct bibtex authors for SomethingSomethingV2
Co-authored-by: zach <zacharie@example.com> (f5775fc)
fix: KeyError on aggregated tasks with eval_langs
Fix KeyError on aggregated tasks with dict eval_langs
When aggregated tasks (e.g. VisualSTS17Multilingual) have eval_langs as a dict, hf_subsets_to_langscripts lacks a "default" key. The aggregated score uses "default" as subset, causing a KeyError in TaskResult.from_task_results. Fall back to collecting all languages from the mapping when the subset key is missing.
Closes #4437
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com> (5fc2867)
Remove skip_first for jamaltartistit (#4435) (16ba72a)
Fix: apply skip_first_result when computing hit_rate metric (#4427)
Fix skip_first_result not applied to hit_rate metric
lint
Co-authored-by: Rakshitha Ireddi <rakshithaireddi@Rakshithas-MacBook-Pro.local>
Co-authored-by: Roman Solomatin <36135455+Samoed@users.noreply.github.com> (8570b74)
[MVEB] Adding MUSIC-AVQA Task (Clustering) (#4426)
[MVEB] Adding MUSIC-AVQA Task (Clustering)
simplify description (04f2f4b)
[MVEB] Add Breakfast video classification task (#4431)
[MVEB] Add Breakfast video classification task\n\nAdd BreakfastClassification task for the Breakfast Actions dataset (Kuehne et al., CVPR 2014). The dataset contains 433 videos of 10 breakfast-related activities recorded in 18 kitchens. Uses 5-fold cross-validation since the dataset only has a test split.\n\nRandom baseline accuracy: 0.1247 (near-random for 10 classes).\n\nAddresses part of #4130 (MVEB Overview - Classification).
lint
Co-authored-by: Yashwanth Devavarapu <yashwanthdevavarapu@Yashwanths-MacBook-Pro.local>
Co-authored-by: Roman Solomatin <36135455+Samoed@users.noreply.github.com> (98d682c)
Add ActivityNet_Captions_val2 video retrieval tasks (V2T and T2V) (#4429)
Add ActivityNet_Captions_val2 video retrieval tasks (V2T and T2V)
Made-with: Cursor
Made-with: Cursor
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>
Made-with: Cursor
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com> (7539339)
add mteb/DiDeMo dataset (#4425)
add mteb/MSVD dataset
add mteb/DiDeMo dataset
add comb
update
Consolidate DiDeMo retrieval tasks into a single file
Merge 4 separate DiDeMo task files into didemo_retrieval.py with a shared _load_didemo helper, reducing duplication while preserving all task names and metadata.
Made-with: Cursor
Made-with: Cursor (2e00e5b)
Add TUNA-Bench_1K video retrieval tasks (V2T and T2V)
Made-with: Cursor (b14dcc5)
fix compatibility with newer vllm versions (5ef64c3)
ci: add workflow to auto-update leaderboard model list
Adds a standalone script that generates the model list from scratch and a CI workflow that pushes it to the HF leaderboard space weekly, on model file changes, or via manual dispatch.
Closes #4316 (18e8e63)
fix: Add required_dependencies to model meta (#4356)
add required_dependencies to model meta
add extra group name
add to model to python
update handling dependencies
fix deps
fix test
remove usage of requires_package
remove image/audio dependencies
fixes after merge
add deprecated function
fix test
skip check for baseline
fix test
update lock
optionally check torchaudio in test (e2e7174)
remove video folder (011bbf5)
add mteb/MSVD dataset (#4413) (6427ea5)
Update dataset cardv2 (#4420)
update dataset card
fix cardv2 (a91046e)
Update dataset card (#4419)
update dataset card (43d1b21)
tests: Add test to ensure coverage of reference models (#4216)
Reference models tests
Reference models tests
Reference models tests
fix: address PR review comments for reference model tests
Check isinstance(task, AbsTaskRetrieval) instead of string comparison with task.metadata.type, so reranking and instruction retrieval tasks are correctly included for retrieval-only models like bm25s.
Return zero confidence scores when sim_scores list is empty, which can happen when BM25 returns no results for a query in reranking tasks.
Address Kenneth's review comments:
Task-level filtering (_is_text_only_task, RETRIEVAL_ONLY_MODELS) already handles model-task compatibility. No need to exclude entire benchmarks — non-text tasks within multimodal benchmarks are skipped automatically.
Simplify _get_target_benchmarks to use display_on_leaderboard=True, which now correctly reflects the actual leaderboard (fixed in #4288). Remove benchmark_selector imports and exclusion list — task-level filtering handles model-task compatibility.
Address Samoed's review: use Benchmark objects in parametrize instead of looking up by name twice.
speedup test
fix issue with aggregate
fix: address review - reuse _check_model_modalities, trim workflow triggers
fix: restore TARGET_BENCHMARKS definition, remove stale _get_target_benchmarks call
fix: inline modality check to avoid private import, filter image-only tasks
fix: use strict modality subset check to exclude image/multimodal tasks
fix: restore RETRIEVAL_ONLY_MODELS for BM25 task filtering
fix: add mteb/benchmarks/** to workflow triggers
Co-authored-by: Roman Solomatin <36135455+Samoed@users.noreply.github.com> (4a90b28)
[MVEB] Adding WorldSense1Min Task (Clustering) (#4393)
[MVEB] Adding WorldSense1Min Task (Clustering)
remove local test
Update mteb/tasks/init.py
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>
removing stats
moving video clustering tasks to clustering
uncomment Video task
add results
update license
remove results
Co-authored-by: wissam-KH <wissam.siblini@komodohealth.com>
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com> (f46cb7b)
[MVEB] Adding AVE-Dataset Task (Clustering) (#4416)
[MVEB] Adding AVE-Dataset Task (Clustering)
uncomment video clustering task
remove results (61e7f3f)
tests: add regression test for double loading (#4407)
add regression test (e946e1e)
add HMDB51 dataset (#4398)
add HMDB51 dataset
update
Update mteb/abstasks/task_metadata.py
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com> (22bc680)
fix: handle Transformers v5 BaseModelOutputWithPooling return types i…
fix: handle Transformers v5 BaseModelOutputWithPooling return types i… (#4328)
fix: handle Transformers v5 BaseModelOutputWithPooling return types
Transformers v5 changed get_text_features, get_image_features, and get_audio_features to return BaseModelOutputWithPooling instead of plain tensors. This caused AttributeError when tensor operations like .norm() were applied directly to the output.
Added isinstance(output, BaseModelOutputWithPooling) checks to extract pooler_output when needed, maintaining backward compatibility with Transformers v4 tensor returns.
Affected model wrappers:
Closes #4081
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
Co-authored-by: Roman Solomatin <36135455+Samoed@users.noreply.github.com> (367d554)
fix retrieval dataset loading (316fca3)
docs: Update adding dataset checklist
docs: Update adding dataset checklist (#4394)
docs: Update adding dataset checklist
fix the checklist to make it less text-specific
add score reproduction to description
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com> (dc58c76)
fix: Auto add base model to ModelMeta (#4395)
fetch source model from hub
fix tests
check if model has model card attr (ed1833a)
model: Add Google Gemini embedding 2 (#4247)
Adding Google Gemini embedding 2 model
feat: add per-task prompt mapping and multimodal support for Gemini Embedding 2
The google-genai SDK's embed_content doesn't handle the "google/" prefix format. Strip it in the constructor like Voyage does.
Retry up to 10 times with exponential backoff (60s, 120s, 240s... up to 600s) when hitting API quota limits. Essential for large multilingual benchmarks like MIRACL.
fix: replace print with logger.warning for lint compliance
fix: handle audio+text interleaved input and note MRL embed_dim support
fix: use MRL embed_dim list and remove duplicate logger
fix unused param
Co-authored-by: Kenneth Enevoldsen <kennethcenevoldsen@gmail.com> (633e41c)
Move Kinetics400 out of video and add zeroshot version (#4383)
distinguish between AV and V tasks
move out of video folder and add zeroshot version
fix task type
update task metadata based on discussions
fix mveb task type mapping
fix: Add is_beta to task metadata (#4392)
fix: Add is_beta to task metadata
todo:
add test and updates metadata
format
re-enable tests for beta datasets
format
feat: comment out MVEB task types without existing tasks
VideoClustering, VideoPairClassification, and VideoCentricQA are defined in task_metadata but have no corresponding task implementations yet, causing create_available_tasks.py to fail. Comment them out until tasks are added. Also regenerate available_tasks docs and add qwen_omni_utils optional dependency.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Change assertion to <= so task types that only have beta tasks don't break the docs generation. Use .get() with continue to skip task types with no non-beta tasks.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Skip Kinetics400 video tasks in test_all_metadata_is_filled_and_valid until descriptive stats are added. Regenerate available_tasks docs.
revert: restore docs/overview/available_tasks to main
revert: remove all generated available_tasks changes from branch
Co-authored-by: Kenneth Enevoldsen <kennethcenevoldsen@gmail.com>
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com> (6ec3f40)
Add nicher92/saga-embed_v1 to MTEB models (#4371)
Add nicher92/saga-embed_v1 to MTEB models
Update training_datasets in ModelMeta
fix: fixed naming
Replace custom SagaModel class with standard SentenceTransformerEncoderWrapper and model_prompts dict
chore: remove lingering comment
Update mteb/models/model_implementations/saga_models.py
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>
update meta
change parameters and memory usage
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>
Co-authored-by: Roman Solomatin <36135455+Samoed@users.noreply.github.com> (d7c521c)
fix: Handle quantization in sentence transformers as an experiment
fix: Handle quantization in sentence transformers as an experiment (#4367)
handle sbert quants
move inside model encode
update type of prompt type
fix tests (e45cbaa)
`` DeprecationWarning: The model 'mteb/baseline-random-encoder' has been renamed to 'mteb/baseline-random-encoder'. To prevent this warning use the ne…
This gave the following incorrect warning:
DeprecationWarning: The model 'mteb/baseline-random-encoder' has been renamed to 'mteb/baseline-random-encoder'. To prevent this warning use the new name.
model = mteb.get_model_meta("mteb/baseline-random-encoder") ([`b65730d`](https://github.com/embeddings-benchmark/mteb/commit/b65730d833b3e321be759b243f62090066faad45))
## Unknown
* model: add BidirLM text embedding family (270M, 0.6B, 1B, 1.7B) (#4374)
* model: add BidirLM text embedding family (270M, 0.6B, 1B, 1.7B)
* Apply suggestions from code review
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>
* run lint
---------
Co-authored-by: Isaac Chung <chungisaac1217@gmail.com>
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com> ([`e8a4069`](https://github.com/embeddings-benchmark/mteb/commit/e8a40693b00005cb0362bd6c5798d61287a196f3))
* [MVEB] PE-AV Model, Kinetics400 Dataset, RavdessAV Dataset (#4199)
* fix: Reclassify SIBFLEURS as AudioClassification instead of AudioMultilabelClassification
* fix: Move SIBFLEURS descriptive stats to AudioClassification
* Adding video modality
* Add Kinetics-400 dataset
* Add pe_av model
* fix typo
* fix collator bug
* Edit selecting column in classification abstask
* Properly handle frames in PE_AV
* add self kwarg to method
* Add audio collator
* fix type error
* fix audio_video embeds object handling
* Add Ravdess_av clustering
* fix task metadata
* start video integration
* start video integration
* upd task structure
* upd video input type
* combine video and audio to dict
* fix task side
* fix pe_av model
* lower writer batch size
* fix col labels
* lint
* add pe_av model metadata
* fix datasets metadata
* remove accidently commited files
* remove nested list structure from datasets
* edit collator to handle one video item
* multimodal collator + fix comments
* lint
* metadata update
* using forward pass to get embeds
* replace forward pass + add audio to msrvtt
* fix category metadata
* edit get embeddings
* add n_embedding_parameters
* change input col name to list
* lint + type check
* add classvar
* add str to classvar
* Change list to sequence
* lint + type check error
* edit dataloader and msrvtt handling of input column
* move seqeuence out of type checking
* fix random baseline
* add collator to random baseline
* restore previous dict structure + make audio optional
* clean structure
* lint
* safety check
* decrease writer batch size
* match msrvtt format
* type check fix
* refactor: keep video and audio as separate dataset columns
* fix: handle single-string input_column correctly in _prepare_dataset
* review fixes
* lint
* type hins fix
* address review: simplify input_column_name, remove VideoInputItem, fix collator output
- Revert input_column_name from Mapping[str, str] to str | Sequence[str]
- Remove VideoInputItem wrapper, pass frames tensor directly
- Make VideoCollator return BatchedInput (consistent with AudioCollator)
- MultimodalCollator uses static methods instead of chaining collators
* fix: update clustering_evaluator to use Sequence instead of Mapping
* fix: handle Sequence input_column_name in second create_dataloader call
* fix: skip statistics and text cleaning for multi-column video tasks
* fix: pass explicit None for TypedDict fields in multi-column statistics
* address Kenneth review: rename collators, update docs, simplify annotations
- Rename VideoCollator -> FramesCollator, MultimodalCollator -> VideoCollator
- Update VideoInput docstring to clarify frames-only, audio in AudioInput
- Update input_column_name docs in classification/clustering base classes
- Use ClassVar[Sequence[str]] for video task input_column_name
- Extract isinstance check to top of zeroshot evaluator __call__
- Improve task_pipelines.py skip comment for multi-column tasks
- Add TODO for MSR-VTT dataset reupload
* docs: link to encoder I/O types for default column names in input_column_name
* fix: raise NotImplementedError for multi-column task cleaning
* refactor: use tuples for input_column_name to avoid ClassVar
* refactor: move Sequence handling into create_dataloader, simplify callers
---------
Co-authored-by: Roman Solomatin <36135455+Samoed@users.noreply.github.com> ([`5d3c845`](https://github.com/embeddings-benchmark/mteb/commit/5d3c8453db615a1a4ae4d53033711af21c1d502e))
* dataset: add BrowseComp-Plus (#4226)
* dataset: Add BrowseComp-Plus
* fix linting errors
* fixing bibtext formatting
* Split BrowseCompPlusRetrieval into gold_only and gold_and_evidence subsets
* fix: remove qa as a valid tag for metadata files
* simplify data loading by reuploading the data
---------
Co-authored-by: Kenneth <kennethenevoldsen@gmail.com> ([`e722b76`](https://github.com/embeddings-benchmark/mteb/commit/e722b7640ed1abee68c3df5023a186b36a15325f))
fix: add language extraction from HF model cards to get_model_meta
fix: add language extraction from HF model cards to get_model_meta (#4278)
feat: add language extraction from HF model cards to get_model_meta
Extract language codes from HuggingFace model card metadata and convert them to MTEB's internal ISOLanguageScript format (e.g., "eng-Latn").
Closes #3694
The language codes (ful, som, yor, etc.) only appear in JSON files which are already excluded by [tool.typos.files] extend-exclude.
Address KennethEnevoldsen's review comments:
Fixes PLW0603 lint error (global statement discouraged). The private function names are kept per KennethEnevoldsen's review.
Co-authored-by: Kenneth Enevoldsen <kennethcenevoldsen@gmail.com> (0fe9fe2)
colpali_engine models query processing (#4361)fix bug (2ea3fe7)
Update codefuse_models.py (#4363) (1db4399)
model: add webAI's ColVec1 Models (#4358)
Added webai_models and ViDoRe v3 runner (to be removed) ...
Separation between .cache dir and output dir for runpod eval ...
Removed run_vdrv3, modified model implementation, TODO: Need to add commit hash after updating model cards ...
Added latest commit hash, ready to submit PR ...
Apply suggestion from @KennethEnevoldsen
Co-authored-by: Kenneth Enevoldsen <kenevoldsen@pm.me>
Co-authored-by: Kenneth Enevoldsen <kenevoldsen@pm.me>
Co-authored-by: Kenneth Enevoldsen <kenevoldsen@pm.me>
Model meta removal, applied when loading the model?
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>
Co-authored-by: Kenneth Enevoldsen <kenevoldsen@pm.me>
Co-authored-by: Kenneth Enevoldsen <kenevoldsen@pm.me>
Co-authored-by: Kenneth Enevoldsen <kenevoldsen@pm.me>
Added changes based on pr suggestions ...
Fix format for linter check ...
Update mteb/models/model_implementations/webai_models.py
Update mteb/models/model_implementations/webai_models.py
Update mteb/models/model_implementations/webai_models.py
fix format
Co-authored-by: Kenneth Enevoldsen <kenevoldsen@pm.me>
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>
Co-authored-by: Roman Solomatin <36135455+Samoed@users.noreply.github.com> (d146a2a)
2627fdd)fix: reupload SDSKoPubVDRT2IRetrieval to mteb org
Missed it in #4348 (3cfcf6e)
dataset: Add SDS KoPub-VDR (#4348)
dataset: Add SDS KoPub-VDR
Update mteb/tasks/retrieval/kor/sds_kopub_vdr_t2it_retrieval.py
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>
update dataset path and revision
Apply suggestion from @KennethEnevoldsen
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>
Co-authored-by: Roman Solomatin <36135455+Samoed@users.noreply.github.com>
Co-authored-by: Kenneth Enevoldsen <kenevoldsen@pm.me> (00da722)
skip leaderboard tests if gradio is not installed (0d13dd0)
fix: Reclassify SIBFLEURS as AudioClassification instead of AudioMultilabelClassification
fix: Reclassify SIBFLEURS as AudioClassification instead of AudioMultilabelClassification (#4353)
fix: Reclassify SIBFLEURS as AudioClassification instead of AudioMultilabelClassification
fix: Move SIBFLEURS descriptive stats to AudioClassification (ea83696)
fix: Change main score from accuracy to LRAP
For MAEB MultiLabelClassification tasks.
I suspect we also need to change the submitted results to match this change (f0f540d)
activate PLR rules (#4317)
activate PLR rules
fix tests
add noqa
add more noqa
remove most of noqa
remove commented rules
fix remove breaking changes
update max args
fix typing
fix lint (fe1f74c)
fix: Add training_datasets annotation to audio model implementations
training_datasets annotation to audio model implementations (#4345)Add training_datasets to audio model implementations
Populate training_datasets field for all audio models using closest
matching MTEB task names where available (e.g. AudioSetMini, FleursA2TRetrieval,
CommonVoiceMini17A2TRetrieval). Datasets without MTEB equivalents are noted
as comments within the set. (fff86d6)
ci: remove inline comments from Makefile for Windows compatibility
fix: remove inline comments from Makefile for Windows compatibility
Inline # comments in recipe lines are passed to the shell, and
Windows cmd.exe does not treat # as a comment character. Remove
these comments to fix Makefile usage on Windows.
Closes #2300 (3254563)
fix: Fixed a bug where benchmarks could not be recorded. (#4325)
fix - Fixed a bug where benchmarks could not be recorded.
feat - add event_logger unit tests (487802d)
fix: Add support for the lower bound of dependencies (#4343)
test lowest deps
custom cache
upd name
switch to pip
add venv creation
change resolution type
bump scipy
bump numpy
add transformers to dependencies
bump torch
lower transformers version
bump pydantic
bump pydantic
try to fix
format
fix branch history
add comment to action
fix comment
try to fix
add bitextparser
Co-authored-by: Kenneth Enevoldsen <kennethcenevoldsen@gmail.com> (c247f7e)
dataset: Add MuPLeR (#4324)
dataset: Add MuPLeR
Update mteb/tasks/retrieval/multilingual/mupler_retrieval.py
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>
Fix bibtex formatting
Removed unnecessary load_data function, added eupl-1.2 license
Apply suggestion from @KennethEnevoldsen
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>
Co-authored-by: Kenneth Enevoldsen <kenevoldsen@pm.me> (334e908)
fix: Showing loading screen while changing benchmark
fix: Showing loading screen while changing benchmark (#4304)
Showing loading screen while changing benchmark
fix lock
Co-authored-by: Kenneth Enevoldsen <kennethcenevoldsen@gmail.com> (9e3936d)
model: add Octen INT8 models and update HuggingFace org to Octen (#4347)
feat: add Octen INT8 models and update HuggingFace org to Octen
feat: add output_dtypes
fix: add instruction_template_8b_int8 with '- ' doc prefix for 8B INT8 model
fix: remove list wrapper from output_dtypes (30aedd8)
docs: update links in README for task and usage sections
docs: update links in README for task and usage sections (#4342)
docs: update links in README for task and usage sections
upd links
Co-authored-by: Roman Solomatin <36135455+Samoed@users.noreply.github.com> (e32aa74)
fix: Not possible to get benchmarks on leaderboard without leaderboard deps (#4330)
move menuentry
fix typing
fix imports
make menu private (c36bc41)
Add more prompts to codefuse_models.py (#4336) (`b1c36f9`)
fix import (17e1c4d)
Add more prompts to codefuse_models.py (#4336) (b1c36f9)
remove -query suffix from harrier prompts_dict keys (#4331)
fix: remove -query suffix from harrier prompts_dict keys
The mteb prompts_dict lookup uses task_metadata.name (e.g. 'News21InstructionRetrieval'), not the suffixed form. Keys like 'News21InstructionRetrieval-query' were never matched. Remove the '-query' suffix so prompts are correctly resolved.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> (1d0dac4)
apply changes from copilot review
fix: Make ModelMeta hashable (#4315)
Make ModelMeta hashable
apply changes from copilot review
convert embed_dim to hashable
apply changes from review
removed all in-place mutation in ModelMeta
fix tests
fiix typechecking
add typehinting comment on correct line
change equality function
remove froze=true and add setattr (7363cae)
model: Add harrier-oss-v1 model implementations (270m, 0.6b, 27b) (#4326)
Add harrier-oss-v1 model implementations (270m, 0.6b, 27b)
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
lint
Simplify instruction_template from @Samoed
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Co-authored-by: Kenneth Enevoldsen <kennethcenevoldsen@gmail.com>
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com> (39261ea)
fix: activate flake8-bandit rules
fix: activate flake8-bandit rules (#4319)
activate flake8-bandit rules
upd rules
fix rules (588fb1e)
fix: Conan-embedding-v1 HF link 🤗
fix: Conan-embedding-v1 HF link (9311eb7)
docs: Remove unused icon for models
remove unused icon for models (07141c6)
docs: Simplify task types in documentation (#4212)
WIP: suggestion for tasktype simplification
As discussed in: https://github.com/embeddings-benchmark/mteb/pull/4203
This is mostly intended to be useful when structering the documentation.
This is just a suggestion. We could also do much simpler simplifications e.g. removing modality and multilinguality of types.
update tasks
resolve indentation warning
add no-wrap for merge
update index site
Update mteb/abstasks/task_metadata.py
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>
Co-authored-by: Isaac Chung <chungisaac1217@gmail.com>
fixed headings
format
fix removed MockMultilingualImageMultilabelClassificationTask
fix spelling lint errors
spelling lint
add fixes from comments
re-order and fix links
added description for tasks pages and removed autogenerated segments
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>
Co-authored-by: Isaac Chung <chungisaac1217@gmail.com> (5738e39)
public_training_code for F2LLM models (#4313)discussed here:
https://github.com/embeddings-benchmark/mteb/issues/3237#issuecomment-4148299969 (bb7aaf1)
tests: activate t20 rules (#4318)
activate t20 rules
add log too
remove log
Update pyproject.toml
Co-authored-by: Kenneth Enevoldsen <kennethcenevoldsen@gmail.com>
Co-authored-by: Kenneth Enevoldsen <kennethcenevoldsen@gmail.com> (9726d2b)
fix: Ensure that experiments_kwargs are passed for get_model_meta
fix: Ensure that experiments_kwargs are passed for get_model_meta (#4308)
fix: Ensure that experiments_kwargs are passed for get_model_meta
fixes #4307
remove test files from tests
fix based on comment
avoid manipulating modelmeta
make sure model meta is only assigned in load_model
re-add loader kwargs (72083fe)
ci: update actions for node v24
ci: update actions for node v24 (#4303)
update actions
change to release version
change back to commit (7720fdf)
Add Thai to leaderboard language selector
Add MTEB(tha, v1) to the Language-specific section of the benchmark selector so it appears in the HuggingFace leaderboard UI.
The benchmark definition was merged in #4213.
Co-authored-by: anusoft <anu@anusoft.com>
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com> (42ffe62)
fix: Ensure that experiments_kwargs are passed for get_model_meta
feat: Evaluate Compression Methods on Retrieval Tasks
feat: Evaluate Compression Methods on Retrieval Tasks (#3950)
Add quantization support
Refactor quantization support into wrapper class
Remove quantization from CLI and update compression wrapper
Change quantization level to string enum
Change enum type to HelpfulStrEnum
Use torch.dtypes for compression levels
Fix linter
Always quantize queries
Fix linter
Define quantization levels as enum
Rename variables
Rename variables and add documentation
Update docs for adding models and what's new section
update docs
better handle of kwargs
fix typing
Raise error on invalid clipping margins
Fix lint
Co-authored-by: Roman Solomatin <36135455+Samoed@users.noreply.github.com> (931683e)
model: Add qihoo360/Zhinao-ChineseModernBert-Embedding (#4266)
add: sbert zhinao_modernbert.
add: sbert zhinao_modernbert.
update: zhinao-chinesemodernbert training datasets.
update: zhinao-chinesemodernbert info.
update: zhinao-chinesemodernbert info.
lint
update: zhinao-chinesemodernbert contacts.
lint
Co-authored-by: caiheng3 <caiheng3@360.cn>
Co-authored-by: Roman Solomatin <36135455+Samoed@users.noreply.github.com> (5ecdd29)
fix: Show models with 0 active params on performance per model size plot
fix: Show models with 0 active params on performance per model size plot (#4289)
Show models with 0 active params on performance per model size plot
Update it to show same Active Parameters count on hover
Fix double showing of Active Parameters in Hover Data
updated lock
improved figure sizes
fix sizes of circles
fix ruff
Co-authored-by: Kenneth Enevoldsen <kennethcenevoldsen@gmail.com>
Co-authored-by: Roman Solomatin <36135455+Samoed@users.noreply.github.com> (1a9ccd4)
fix: update dataset card with descriptive category and paper title reference
fix: update dataset card template with descriptive category and paper title reference
Closes #2807 (ef5edd4)
Pin uv docker (#4301)
pin uv docker
try (5648336)
deprecated_evaluator.py: defer import mteb to method body
fix: skip encode for single-label cluster sets in legacy clustering
For cluster sets with only one unique label (e.g. cluster set #29 in ArxivClusteringP2P with only 'quant-ph'), skip the expensive encode and clustering steps. The v_measure is always 1.0 for these sets, so we add it directly to preserve backward-compatible scores while saving compute.
Fixes #1866 (47caa30)
fix: resolve circular imports and add CI enforcement (#4277)
fix: resolve circular imports and add CI enforcement
Move top-level import mteb statements that created circular dependency
chains through mteb/__init__.py into function bodies or replace with
direct submodule imports. This eliminates 38 circular import chains.
Files fixed:
import mteb to method bodyimport mteb to function bodiesimport mteb with direct import from
mteb.models.get_model_metaimport mteb to method bodyimport mteb to method bodiesAdd scripts/check_circular_imports.py that detects runtime circular
imports via AST analysis and integrate it into both make lint and
make lint-check targets so CI prevents future regressions.
Closes #3937
Move imports from mteb.get_tasks, mteb.models.get_model_meta, and
mteb.benchmarks.get_benchmark to top-level in files where it doesn't
cause circular dependencies (abs_encoder, cache, deprecated_evaluator).
Keep inline imports only in task_metadata.py and task_result.py where
circular dependency chains make top-level imports impossible.
Replace import mteb with direct submodule imports throughout.
Remove scripts/check_circular_imports.py and its Makefile references per reviewer consensus. Move get_task import in abs_encoder.py to function-level to break the circular import chain: mteb.abstasks → abstask → mteb.models → abs_encoder → get_tasks → mteb.abstasks
The import was changed from mteb.get_model_metas to a direct import in mteb.cache, so the mock needs to patch mteb.cache.get_model_metas.
Co-authored-by: Kenneth Enevoldsen <kenevoldsen@pm.me>
Apply suggestion from @KennethEnevoldsen
fix patch
fix: patch get_model_metas where it's looked up, not where it's defined
unittest.mock.patch must target the namespace where the name is used.
Since cache.py does from mteb.models.get_model_meta import get_model_metas,
the correct patch target is mteb.cache.get_model_metas, not
mteb.models.get_model_meta.get_model_metas.
Co-authored-by: Kenneth Enevoldsen <kenevoldsen@pm.me>
Co-authored-by: Roman Solomatin <36135455+Samoed@users.noreply.github.com> (7795c9b)
fix: resolve API models in performance-over-time plot (#4276)
fix: resolve API models in performance-over-time plot
API models (voyage, openai, cohere, etc.) were being filtered out of the performance-over-time plot because their reference URLs are blog posts or docs pages, not HuggingFace URLs. The existing lookup only worked for HuggingFace URLs where org/model could be extracted from the URL path.
Add fallback model meta resolution: first try HuggingFace URL extraction (existing behavior), then match by (reference_url, display_name), then by display_name alone. Also build lookup dicts once per plot instead of per-row for better performance.
Fixes #4141
Instead of reverse-engineering model metadata from markdown cells in the plot function, add Release Date as a column in all summary table variants. This makes the performance-over-time plot work for all models (including API models) without needing URL-based fallback resolution.
Pass raw summary data through gr.State for plot functions while
dropping Release Date from the displayed gr.DataFrame table. (cca4a52)
fix: Deprecate sentence_transformers_loader in favor of SentenceTransformerEncoderWrapper (#4279)
deprecate: sentence_transformers_loader in favor of SentenceTransformerEncoderWrapper
sentence_transformers_loader was just a thin wrapper around
SentenceTransformerEncoderWrapper. This deprecates the function
(with a DeprecationWarning) and updates all internal usages to use
SentenceTransformerEncoderWrapper directly as the model loader.
The function remains exported for backward compatibility.
Closes #3739
f9abce8)fix: Refactor queries and document dataloader to allow multiple modalities
fix: Refactor queries and document dataloader to allow multiple modalities (#4232)
Refactor queries and document dataloader to allow multiple modalities
update passing collate
update bm25 and bb25 models based on refactor
remmove _create_text_queries_dataloader function after updating bm25 and bb25
remove _create_dataloader_for_retrieval_corpus, _create_text_dataloader_for_queries, _create_dataloader_for_queries_conversation functions
refactor create_dataloader function
remove _create_image_dataloader, _create_audio_dataloader, and _create_video_dataloader under refactoring
Remove _create_queries_dataloader and _create_document_dataloader
Remove _create_dataloader_from_texts
Added _create_dataloader_from_texts again
small refactor
fix check for failing tests
apply changes from review (ef1f598)
dataset: added new STS task for Ukrainian: Sed small (#4297)
added new STS task for Ukrainian: Sed small
updated commit id for the dataset
added descr stats (6a798bc)
Fix formatting in citation entry in README (#4294) (b09e6dc)
Pin uv version (#4296)
pin uv version (1fbcff7)
fix: add contributed_by field to TaskMetadata
fix: add contributed_by field to TaskMetadata (#4275)
feat: add contributed_by field to TaskMetadata (#3920)
Add optional contributed_by field to specify who provided datasets, especially useful for private datasets where source info is harder to find. The field is surfaced in dataset card descriptions and the leaderboard task info table.
3f4c00b)fix: ensure that display_on_leaderboard actually reflect whether the benchmark is displayed
fix: ensure that display_on_leaderboard actually reflect whether the benchmark is displayed (#4288)
fix: ensure that display_on_leaderboard actually reflect whether the benchmark is displayed
I believe the previous attribute was a leftover from an earlier version of the leaderboard
fix typing (d4daab0)
fix: Add modality filtering to get_model_metas (#4262)
Add modality filtering to get_model_metas
Add tests for modality filtering and update typing
Add tests for modality filtering
Fix lint formatting (newline + spacing)
Fix modality tests to use string values
Fix lint (final newline)
Fix lint + finalize modality filtering
Add test for unfiltered models and fix lint
Fix final newline
lint
Co-authored-by: David Schechter <davidschechter@davids-air-2.lan>
Co-authored-by: David Schechter <davidschechter@Davids-MacBook-Air-2.local>
Co-authored-by: Roman Solomatin <36135455+Samoed@users.noreply.github.com> (3813b2d)
fix: fix tasktype aggregation (#4283)
fix: remove broken logo
fix: vidore leaderboard
ff91ca2)fix: Revert gradio version bump
fix: don't fetch from hub when calling get_model_meta, but do when calling get_model
fix: don't fetch from hub when calling get_model_meta, but do when calling get_model (#4284)
Fix behaviour when getting metadata for non-existing models
Fix naming of cross-encoder/ms-marco-TinyBERT-L2-v2 (1c90fcb)
Added potion-base-32m and potion-retrieval-32m (f3e4dfb)
Fix display of vidore benchmarks on the leaderboard (#4282)
fix: remove broken logo
fix: vidore leaderboard
3149207)fix: make sure that the leaderboard build as intended
fix: make sure that the leaderboard build as intended (#4269)
fix: make sure that the leaderboard build as intended
includes the fix from #4268 along with a few additional minor fixes. Notably gradio was to allow for pandas v3 and update makefile to ensure that the leaderboard is run with the correct set of dependencies
Co-authored-by: Munot Ayush Sunil <munotayush6@kgpian.iitkgp.ac.in> (fd49226)
fix: InstructSentenceTransformers embed dim
fix: InstructSentenceTransformers embed dim (#4264)
fix: InstructSentenceTransformers embed dim
introduced in #4170
Co-authored-by: Roman Solomatin <36135455+Samoed@users.noreply.github.com> (b517688)
fix: Remove Memory Usage column and Add active Paramters column (#3979)
Update Leaderboard - Remove Memory Usage column and Add active Parameters column
Rename columns for better visibility
fix active parameter formatting on LB
Add exact embedding parameter value for KALM Model (109ef0b)
feat: Add Matryoshka support when loading a model
feat: Add Matryoshka support when loading a model (#4170)
add Matryoshka support
raise error
upd embed_dim in leaderboard
fix tests
fix typcheck
add tests
upd check
add docs (44e5947)
Merge branch 'main' of https://github.com/embeddings-benchmark/mteb (9a2dbd8)
benchmark: Add Thai benchmark MTEB(tha, v1) (#4213)
Add Thai benchmark: MTEB(tha, v1)
Add a Thai language benchmark with 28 tasks spanning 6 task types:
Results for 13 models already merged: embeddings-benchmark/results#428
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>
Address review feedback from @KennethEnevoldsen:
Removed (12 tasks):
Kept (15 tasks):
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Co-authored-by: anusoft <anu@anusoft.com>
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com> (03dcfc8)
model: Add NanoVDR-S-Multi with custom AbsEncoder for asymmetric VDR (#4242)
Add model meta for nanovdr/NanoVDR-S-Multi
add n_embedding_parameters
Implement NanoVDRWrapper as custom AbsEncoder with asymmetric routing
Query encoding uses the lightweight NanoVDR-S-Multi student (69M, text-only). Document encoding uses the frozen Qwen3-VL-Embedding-2B teacher (2B, VLM). Teacher is lazy-loaded only when document encoding is needed.
Fix ruff lint and formatting
Default to student encoder for non-retrieval tasks
Add error for unsupported image-query tasks
Remove trust_remote_code=True (model class built locally by MTEB)
Apply suggestion from @Samoed
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>
Co-authored-by: Roman Solomatin <36135455+Samoed@users.noreply.github.com>
Co-authored-by: Kenneth Enevoldsen <kenevoldsen@pm.me>
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com> (2e1e513)
docs: fix formatting in the documentation (#4258) (`7ff0dea`)
7ff0dea)fix: Added Embedding Parameter Value for Image and Audio Models (#4167)
Added Embedding Parameter Value for Image and Audio Models
Remove models from _MISSING_N_EMBEDDING_MODELS
Fix tests
Remove comment from CLAP models
Added embedding parameter value=0 for only image/audio models
update tests and modalities (14f33b5)
Fix ColQwen3.5-4.5B - ColPaliEngineWrapper.encode() for multimodal datasets (#4245)
fix: improve input encoding logic for ColPali models to handle text and image features correctly
fix: refactor encoding logic in ColPali and ColQwen models to improve handling of text and image features
fix: refactor ColQwen3.5 wrapper to enhance input handling and support for image-text embeddings
fix: linting
Update mteb/models/model_implementations/colqwen_models.py
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>
fix: image processing logic in ColQwen3.5 wrapper to match how
Update mteb/models/model_implementations/colqwen_models.py
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>
process_queries adds the query prefix and augmentation tokens that the model was trained with. process_texts is plain tokenization without these, leading to ~0.02-0.03 lower nDCG scores.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com> (8d80cdd)
fd53427)fix: remove unused calculate_probs method from model wrappers
fix: remove unused calculate_probs method from model wrappers
Closes #4201 (c28635b)
feat: detect frameworks from HuggingFace tags instead of hardcoding PyTorch
Expand _get_frameworks_from_hf_tags to detect pytorch, tf, jax, and openvino from HuggingFace model tags instead of assuming PyTorch by default. Add JAX and OpenVINO to the FRAMEWORKS literal type.
Closes #4104 (1ea8628)
fix colqwen3.5 (#4243) (c6b9389)
Model: Add athrael-soju/ColQwen3.5-v3 (#4241)
feat: add ColQwen3.5 model wrapper and metadata
Update mteb/models/model_implementations/colqwen_models.py
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com> (f0f10c2)
dataset: Add MMDocIR (#4230)
dataset: Add MMDocIR
change dataset path and revision (4f1688f)
docs: remove mieb and mmteb contribution docs
I don't think we maintain these anymore. I think they are fine to remove (823236c)
fix docs paths (0d07d33)
6787a17)fix: metadata getting computed for existing MTEB model (#4231)
Fix behaviour while getting metadata of existing MTEB model
Added basic metadata in overwrite
Updated CrossEncoderWrapper with same changes (973a5a1)
New model revision (f913ed8)
Fix zeroentropy/zembed-1 metadata (revision, release_date, max_tokens)
The metadata added in #4202 had incorrect values for three fields:
3cd67fd)Add Zeroentropy models (#4228)
Add Zeroentropy models
correct metadata
Correct loader_kwargs for rerankers (791a185)
model: nvidia/llama-nemotron-embed-vl-1b-v2 for ViDoRe (#4192)
Adds nvidia/llama-nemotron-embed-vl-1b-v2 model
Update mteb/models/model_implementations/nvidia_nemotron_vl_models.py
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>
Fixing tests and linting issues
Update mteb/models/model_implementations/nvidia_nemotron_vl_models.py
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>
Nemotron Embed VL 1B: Setting the number of tiles an image can be split
Fixing lint issue
Update mteb/models/model_implementations/nvidia_nemotron_vl_models.py
Disabling image modality by default
Update mteb/models/model_implementations/nvidia_nemotron_vl_models.py
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com> (2a8c2d3)
fix: Error in siglip output conversion
fix: Error in siglip output conversion (#4205)
fix: Error in siglib output conversion
add mean pool to siglip
format
Apply suggestions from code review
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>
add missing depencies
added fix for siglip dependencies
format
fix dependencies
added image normalization
This should happen here: https://github.com/embeddings-benchmark/mteb/blob/ce7590dcc9c620450ca192a3ec101a62631e6b55/mteb/_create_dataloaders.py#L291-L292
Not sure why it is needed
relax protobuf dependency
lint
update pyproject.toml dependencies
Co-authored-by: Your Name <you@example.com>
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com> (ec20d1e)
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com> Co-authored-by: Roman Solomatin <36135455+Samoed@users.noreply.github.com> (`5530493`)
fix: Add ViDoRe(v3.1) (#4220)
fix: Add ViDoRe(v3.1)
Apply suggestion from @Samoed
add to init
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>
Co-authored-by: Roman Solomatin <36135455+Samoed@users.noreply.github.com> (5530493)
fixed overview for models and benchmarks
docs: migrate to zensical (#4203)
migrate to zeniscal
added breadcrumbs
added navigation icons
minor docs fix
fix annotations
change to links
fixed overview for models and benchmarks
try to use zensical
add copy paste button for models
add copy-paste button to tasks and benchmarks as well
remove plugins
get back mieb and mmteb
rename api back
add tasks to overview
reorder overview page
update lock file
Co-authored-by: Kenneth Enevoldsen <kennethcenevoldsen@gmail.com> (a484cfd)
Now diplay main score in task results. As well as the task_res.main_score property.
Also added "..." to indicate that there are more attributed than what is being shown.
res = mteb.evaluate(model, task)
res
res[0]
# currently displays:
# ModelResult(model_name=mteb/baseline-random-encoder, model_revision=1, task_results=[...](#1))
# TaskResult(task_name=LccSentimentClassification, scores=...)
# with PR:
# ModelResult(model_name=mteb/baseline-random-encoder, model_revision=1, task_results=[...](#1), ...)
# TaskResult(task_name=LccSentimentClassification, main_score=0.32, scores=...)
``` ([`7c831b0`](https://github.com/embeddings-benchmark/mteb/commit/7c831b068b0e8341485b745e05803239a5d16c2f))
## Unknown
* model: Qwen3-VL-Embedding (#4198)
* add qwen3-vl-embedding implementation
* lint and test
* lint
* handle image+text mode
* address review comments
* address comments
* fix resolve dependency
---------
Co-authored-by: Kenneth Enevoldsen <kennethcenevoldsen@gmail.com> ([`0bb0917`](https://github.com/embeddings-benchmark/mteb/commit/0bb0917c596247b1fa336a73cbbc71b3a1ac01f9))
* Added model zeroentropy/zembed-1 (#4202)
* Added model zeroentropy/zembed-1
- [y] I have filled out the ModelMeta object to the extent possible
- [y] I have ensured that my model can be loaded using
- [y] `mteb.get_model(model_name, revision)` and
- [y] `mteb.get_model_meta(model_name, revision)`
- [y] I have tested the implementation works on a representative set of tasks.
- [y] The model is public, i.e., is available either as an API or the weights are publicly available to download
* Apply suggestion from @Samoed
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>
* lint
---------
Co-authored-by: Ryan Wang <ryanwang@DN0a249162.SUNet>
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>
Co-authored-by: Roman Solomatin <36135455+Samoed@users.noreply.github.com> ([`69421f9`](https://github.com/embeddings-benchmark/mteb/commit/69421f95fafb55e7f183b89544a31fed4e56fa18))
* Add Reason-ModernColBERT (#4218)
* Add Reason-ModernColBERT
* Fix variable name + add size
* Fix variable name + add size ([`b877424`](https://github.com/embeddings-benchmark/mteb/commit/b8774246a712cc5c9878a8abc022d1e155158708))
Your coding agent can read these notes before it upgrades. Set up the MCP server →