NewYour coding agent can read the release notes before it upgrades.Set up the MCP server →
PyPI · #4752 most downloaded on PyPI
Massive Text Embedding Benchmark
Last release today
04 Oct 2026
Ships on a steady schedule
a new release about every 9 days
Nearly every release is documented
notes for 58 of the last 60 stable releases
Nothing withdrawn
no release was ever pulled
4 years old
814 releases · first in 2022
One column per quarter.
Nothing published for this version
Nothing published for this version
chore: update the citation cache [skip ci]
87f5bcd)NeuCLIR2023Retrieval loaded mteb/NeuCLIR2022Retrieval at the same
revision as NeuCLIR2022Retrieval, so it ran on the 2022 queries and
qrels. Point it to mteb/NeuCLIR2023Retrieval, whose queries and qrels
match the original mteb/neuclir-2023.
Regenerate its descriptive statistics and list it under
KNOWN_ISSUES["zero_relevant_docs"]: 4 queries only have score-0
judgments, as in the original qrels. (9572cf4)
test: check final scores in model-task integrations (#5321)
test: check final model-task scores
test: isolate model-task score cases
test: account for media codec score baselines
test: check multimodal pair classification scores
test: account for pair classification codecs
test: check scores in library integrations
test: account for dataset score environments
test: structure score baselines by model
test: remove redundant model parametrization
simplify modelinfo
remove comment
update after merge
add future
print actual and expected scores
fix test
upd prescision
Co-authored-by: Roman Solomatin <36135455+Samoed@users.noreply.github.com> (33a8aa9)
upd (5a29ea9)
model: add ColNanoVDR late-interaction query towers (ColVec1.1-4b/8b) (#5518)
model: add ColNanoVDR late-interaction query towers (ColVec1.1-4b/8b)
Two asymmetric late-interaction retrievers: a 150M text-only multi-vector
student encodes queries, the frozen ColPali-style teacher it was distilled
from encodes page images, and scoring is MaxSim.
The student needs sentence-transformers>=6 for MultiVectorEncoder, so this
adds a colnanovdr requirement group; heavy imports stay inside functions.
Also corrects training_datasets for nanovdr/NanoVDR-S-Multi, which was
trained on the same data: TAT-DQA was missing and TabFQuAD listed instead.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
n_embedding_parameters is the student's input embedding matrix
(50368 x 768); n_parameters is the exact count of the packaged model
rather than a rounded figure.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Documents now go through the registered ColVec1.1 wrapper, so the teacher's
pinned revision, processor settings and handling of text-only and image
documents are exactly those of its own entry. The teacher is built in
init.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com> (2852c40)
model: add nvidia/llama-nemotron-rerank-vl-1b-v2 (#5565) (c9f9329)
Leaderboard integation with experiments (#4900)
recreate experiments
update after frontend testing
parse model meta from experiments
fix test
partly read meta
Address PR #4900 review feedback
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
make working natively
simplify experiments handling
refactor
simplify comments
fix typing
optimize loading
simplify comments
fix upload
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com> (60b2a1e)
Custom task grouping for leaderboard (#5174)
Add CustomGrouping: multi-dimension custom task aggregations
Generalizes the existing TASK_TYPES aggregation pattern so a Benchmark
can declare one or more named custom grouping dimensions directly in
its aggregations sequence (mixed with plain BenchmarkAggregation
flags, no new field, no new enum member). Each CustomGrouping produces
its own dynamic per-group mean columns in the leaderboard summary,
namespaced 'dimension::label' to avoid collisions between dimensions.
Migrates LMEB to use this for a 'Memory Type' breakdown (episodic /
dialogue / semantic / procedural) instead of registering four
duplicate Benchmark objects, resolving the reviewer objection on
PR #4914 for issue #4898.
404 tests pass; ruff check/format clean.
Groups the 20 tasks into Short (12) / Long (8) via the same
CustomGrouping mechanism used for LMEB's Memory Type breakdown --
the 8 domains with a Long variant land in both groups, the 4
code/math domains (Leetcode, Aops, TheoremQA*) only have Short.
Adds test_bright_document_length_grouping_covers_all_tasks, asserting
full task coverage and that group membership matches the Long/Short
suffix on each task name.
The language sidebar filter isn't gated on Benchmark.language_view --
it reads tasksMeta[].languages directly and triggers a server refetch
for any benchmark with more than one task language. scores_by_custom_group
was previously left frozen at unfiltered values there (LMEB/BRIGHT(v1.1)
are both eng-only today, which is why this never surfaced, not because
of language_view as the old comment claimed).
start refactor
Merge duplicated bucket-and-average logic behind a shared _bucket_means
_recompute_lenient_means and _recompute_lenient_custom_groups both
independently bucketed scores_by_task by a task->key mapping and
averaged each bucket. Extracted the shared primitive (_bucket_means);
both callers now just supply their task_to_key mapping(s) and combine
the per-bucket results into their own return shape (task-type recompute
also derives the two scalar means; custom-group recompute returns one
bucket-dict per dimension).
No behavior change -- 409 tests still pass.
refactor
fix
remove comments
Move lenient-recompute functions out of mteb.api into _benchmark_metrics
_bucket_means/_recompute_lenient_means/_recompute_lenient_custom_groups
are pure functions over plain dicts, not tied to FastAPI/pydantic --
moved to mteb.benchmarks._benchmark_metrics (mteb.api.aggregators now
imports them) alongside the other aggregation helpers that already
live there.
Also extracted _bucket_task_result_scores, the TaskResult-based analog
of _bucket_means, and refactored _compute_task_types and
_compute_custom_group_means to share it instead of each re-implementing
the same bucket-by-key + null-tracking loop.
Tests moved from tests/test_api/ (now removed) into
test_benchmark_score.py alongside the other _benchmark_metrics.py
coverage. No behavior change -- 407 tests pass.
simplify comments
Apply suggestions from code review
Co-authored-by: Copilot Autofix powered by AI <175728472+Copilot@users.noreply.github.com>
move import
lint
remove unnecessary aggregations
Send custom-group task membership so the frontend can recompute under sidebar filters
CustomGroupSchema previously only sent {label, description} — the
frontend had no way to know which tasks belong to a group, so the
task-type/domain/modality sidebar filters (client-side, no server
round-trip) left scoresByCustomGroup frozen at stale values while
scoresByTaskType correctly recomputed alongside it.
Adds tasks: list[str] to CustomGroupSchema, populated from
CustomGroup.tasks in both BenchmarkSchema.from_benchmark (static
declaration) and aggregators.py's data-driven summary construction
(joined back via the same declared_by_dim lookup already used for
descriptions).
414 tests pass; verified live against LMEB.
Bypassing pre-commit: the typos hook fails on a pre-existing, unrelated
acronym false-positive (fgmcaps_retrieval.py's 'retrievAl' in FIGMA),
not on anything in this change.
add persubset
simplify
remove comments
fix test
update skip rule
Apply suggestion from @KennethEnevoldsen
Co-authored-by: Kenneth Enevoldsen <kennethcenevoldsen@gmail.com>
don't compute leniently
remove vs
Co-authored-by: Copilot Autofix powered by AI <175728472+Copilot@users.noreply.github.com>
Co-authored-by: Kenneth Enevoldsen <kennethcenevoldsen@gmail.com> (6cb5891)
fix: count single-colour images and flag unanswerable queries
fix: count single-colour images and flag unanswerable queries (#5347)
feat: count single-colour images and flag unanswerable queries
ImageStatistics gains constant_images -- images that are one colour
everywhere. Such an image carries no visual signal; it is usually a failed
fetch silently replaced by a blank frame, which makes it missing data rather
than noise.
Auditing the image tasks available offline (2.2M images) found that what a
constant image costs depends entirely on which side it sits on:
So the check is scoped by evaluation semantics rather than by column name.
A query or labelled sample is itself the evaluated unit, so a constant one
is always a defect. A document is only unretrievable when it is the whole
of some query's gold set, which qrels decide -- reported separately as
RelevantDocsStatistics.queries_with_all_gold_constant. Corpus columns are
no longer skipped; they are judged by the rule that fits them.
Counting is close to free: calculate_image_statistics already walks every
image for width and height, and the corpus flags are computed once and
shared between the image statistics and the qrels intersection.
Stats generated before these fields exist simply omit them and are skipped,
so both activate per task as descriptive statistics are regenerated. No
KNOWN_ISSUES entries are added.
Flagging any single-colour image treated template/placeholder art (a
solid red or brand-colour background) the same as a failed image fetch
silently replaced by a blank frame. Only pure black/white reliably
indicates the latter, so narrow detection to that and rename the
helpers/fields/messages so they no longer imply otherwise.
Two fixes were needed for the narrowed check to be correct rather than
just stricter: palette-mode images store indices, not colours, so an
index of 0 isn't black until resolved through the palette; and alpha
must be constant but excluded from the colour check itself, since an
ordinary opaque black/white pixel has alpha=255, not 0.
Qrels-aware semantics are unchanged: a query is only flagged when its
entire gold set is black/white, and corpus images no qrel references
stay unflagged.
fix: simplify black and white image checks following review
fix: convert to RGB for the black/white check and trim its test
Drop the per-mode handling (palette, CMYK, HSV, wide-integer white points)
in favour of a single RGB conversion, reduce the test to one small case,
and revert two formatting-only hunks.
Only the sign of the score matters to the count, and the mixed 1/2 pair
read as if it meant something.
A black or white image can be valid content, so both checks are
diagnostic heuristics rather than correctness invariants. Route them
through _WARNING_CHECK_KINDS, drop the failure language from the
docstrings, and name the gold-set statistic as the stronger signal for
manual review. The unit test now names its corpus ids instead of using
list indices.
The overall split of a multilingual retrieval task prefixes corpus and
qrel ids with split and subset, but the black/white doc id set was built
from the raw concatenated corpus ids, so
queries_with_all_gold_black_or_white was always 0 at the overall level.
Regenerate descriptive stats for VisualNewsI2TRetrieval (14 black/white
query images) and WITT2IRetrieval (21 black/white corpus images, 21
queries whose only gold document is black/white), matching the offline
audit.
Regenerating its stats added num_documents and num_queries, which lets
the duplicate checks run on this task for the first time. The duplicates
are a property of WIT and are already registered for WITI2TRetrieval.
Co-authored-by: Roman Solomatin <36135455+Samoed@users.noreply.github.com> (78aa9a6)
chore: update the citation cache [skip ci]
53d0e4c)push the citation cache with the RELEASE token (e4a85b2)
feat: add sentence transformers multi vector models (#5338)
implement multivector models
separate wrappers
Restore multi-vector encoder test and fix mypy type mismatches post-merge
upd docs
fix: dtype mismatch in rerank_top_ranked_documents reranking path
Restore the torch.as_tensor() wrap around candidate_embeddings that
existed in the pre-refactor SearchEncoderWrapper._rerank_documents.
Without it, _convert_to_tensor in similarity_functions.py force-casts
the (already numpy) candidate embeddings to float32 while an
already-tensor query embedding keeps its original dtype (e.g. float64),
causing a dtype-mismatch RuntimeError in cos_sim's matmul during
reranking-based retrieval/reranking tasks.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SHVK39hJCn1fTqsEgtkMwZ
keep conversion
don't raise SparseEncoder/MultiVectorEncoder ImportError for unrelated models
remove comments
remove comments
remove comments
bump version in what's new
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com> (5be29e1)
reupload (58ae94b)
fix: bge-small-en and bge-small-en-v1.5 embed_dim is 384, not 512
fix: bge-small-en and bge-small-en-v1.5 embed_dim is 384, not 512 (#5548) (033dfed)
fix: read task modalities in result filtering (#5550)
TaskResult had no modalities attribute, so the getattr in ModelResult._filter_tasks and ModelResult.modalities always returned []. Filtering by modality dropped every task and .modalities always fell back to ['text']. (ad793aa)
Point the four tasks at their re-uploaded mteb/ datasets and drop the
custom load_data, the same change as DS1000Retrieval (#5542). The four
tasks shared that loader.
fix: repair NanoVDR document encoding path
fix: repair NanoVDR document encoding path (#5552)
fix: repair NanoVDR document encoding path
nanovdr/NanoVDR-S-Multi fails at document-encoding time with an
ImportError, and has done since beee210 (#4699). Two bugs, both past
_load_teacher():
_load_teacher imports _build_qwen3_vl_for_embedding_class from
qwen3_vl_embedding_models, but #4699 deleted that factory when it
moved Qwen3VLEmbeddingWrapper from AbsEncoder to
InstructSentenceTransformerModel. Recovered it verbatim from
beee2102^ and vendored it here: it was private, and this wrapper is
its only remaining consumer, so re-adding it to that module would
partly undo #4699. Only additions are the type annotations current
lint requires; Cache/Qwen3VLConfig stay under TYPE_CHECKING so
the module remains importable without torch (#5463).
_encode_queries passed convert_to_numpy=False while leaving
convert_to_tensor at its default, which sentence-transformers
documents as returning list[Tensor]. mteb._convert_to_tensor then
raises "only one element tensors can be converted to Python scalars".
Now convert_to_tensor=True, plus .cpu() to match
_encode_documents -- cos_sim does no device harmonisation, so a
CUDA query matrix against a CPU corpus matrix would fail.
The import is lazy and only runs when encoding documents, so
mteb.get_model() and query encoding both worked and the breakage went
unnoticed for ~3.5 months.
Verified end to end on VidoreTabfquadRetrieval: the teacher loads, the
corpus encodes, and the run completes past the similarity step that
previously raised.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The vendored _build_qwen3_vl_for_embedding_class turns out to be
redundant. Qwen3VLModel is already the LM-head-less trunk and its
forward returns Qwen3VLModelOutputWithPast.last_hidden_state, which is
all the pooling in _encode_documents needs. Its
_checkpoint_conversion_mapping = {} override was dead weight too: that
attribute is not set on Qwen3VLPreTrainedModel, Qwen3VLModel, or
Qwen3VLForConditionalGeneration in current transformers.
The one thing the wrapper did provide was resetting rope_deltas per
forward -- Qwen3VLModel.forward reads self.rope_deltas but never
resets it -- so that moves to the call site.
Verified against the previous commit:
from_pretrained loads 625/625 weights with no init warnings, so theQwen3VLModel directlylast_hidden_state (max absMockAny2AnyRetrievalT2I gives identical values for all 149 metricsCo-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com> (75b4800)
Point DS1000Retrieval at the re-uploaded mteb/DS1000Retrieval dataset
and drop the custom load_data, same as #5495.
fix bugs in docs (#5535)
fix bugs in docs
merged other task types
fix links related to get_tasks (d98f640)
Revert "dataset: add MM-BRIGHT retrieval tasks" (#5531)
Revert "dataset: add MM-BRIGHT retrieval tasks (#5168)"
fix: don't place pair classification thresholds between tied scores
Add voyageai/* -> mongodb/* entries to _MODEL_RENAMES so existing callers of get_model_meta keep working with a deprecation warning.
fix: fix jina-v3 and arctic-embed dependencies (#5475)
fix dependency
update
simplify xformers version (d6a934c)
Model/voyage mongodb rebranding (#5510)
model: move Voyage AI models to the mongodb/ namespace
Voyage AI is now part of MongoDB. Rename every voyageai/<model> ModelMeta
name (and matching adapted_from / superseded_by references) to
mongodb/<model>, so newly generated results land under the new
organisation. HF reference URLs keep the existing huggingface.co/voyageai
paths because the Hub organisation has not moved.
Add voyageai/* -> mongodb/* entries to _MODEL_RENAMES so existing callers
of get_model_meta keep working with a deprecation warning.
mongodb/voyage-4-nano is only the registry name; the checkpoint still lives
at huggingface.co/voyageai/voyage-4-nano. Wrap the SentenceTransformer loader
so the HF id stays pinned to the voyageai organisation while the model is
registered (and its results stored) under the mongodb/ namespace.
Verified with mteb run -m mongodb/voyage-4-nano -t AILAStatutes. (58a5f7e)
fix: remaining changes before making import mteb torch-free and mteb-core compatible
fix: remaining changes before making import mteb torch-free and mteb-core compatible (#5502)
fix: remaining changes before making import mteb torch-free and mteb-core compatible
changes from review
fix: clean up stale comment and dead code from torch-free import refactor
Update a comment in test_ensure_no_torch_at_import.py that claimed
import mteb still pulls in torch, contradicted by
test_import_mteb_does_not_import_torch further down in the same file.
Also remove the now-unused logger in _set_seed.py left over after the
tensorflow debug-log call was dropped.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Co-authored-by: Isaac Chung <chungisaac1217@gmail.com>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com> (e07c2e3)
dataset: Add chinaopen multilingual video (#5393)
task: add ChinaOpen multilingual video retrieval (t2v, v2t)
Adds the first multilingual video retrieval tasks in mteb, built from
ChinaOpen-1k (ACM MM 2023): 1,092 Bilibili videos whose Chinese captions
are written by human annotators and whose English captions are
translations of them. Both language subsets describe the same videos, so
zho-Hans vs eng-Latn is a controlled comparison rather than a difference
in visual content.
Queries are deduplicated by caption text and every video carrying a
caption is marked relevant, so the few captions shared by more than one
video are multi-positive instead of being scored as misses. Uploader
video titles are not used, only the human-written manual captions.
The datasets are published in the standard retrieval layout with one
config triple per language, so no custom load_data is needed. The
construction script is included.
Both language subsets of the ChinaOpen tasks describe the same 1,092
videos, which is the point of the controlled comparison, so the corpus
is intentionally shared and the duplicate check counts it twice. This
is the same situation as XM3600 and XFlickr30k-Co, which are listed
under duplicate_image for the same reason.
Rebuild the ChinaOpen dataset as one HF repo (mteb/ChinaOpen1k) with a
shared videos config plus per-language <lang>-texts/<lang>-links
configs, instead of two full repos that each duplicated the video bytes
per direction. ChinaOpenT2VRetrieval and ChinaOpenV2TRetrieval now share
a _load_chinaopen(task, direction) loader that swaps which side of each
caption<->video link is the query, keeping the original dedup and
multi-positive qrels logic intact.
Verified end-to-end: rebuilt from the raw ChinaOpen-1k source, pushed to
mteb/ChinaOpen1k, and reran both tasks with the random-encoder baseline,
reproducing the original PR's reference ndcg@10 scores.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Recompute descriptive_stats JSON for ChinaOpenT2VRetrieval and
ChinaOpenV2TRetrieval against the new mteb/ChinaOpen1k dataset via
calculate_descriptive_statistics(overwrite_results=True). Values match
the original PR's stats exactly (num_samples=4360, num_queries=2176/2184,
num_documents=2184/2176), confirming the restructuring didn't change the
underlying data. Also includes a ruff-format pass on create_data.py.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Co-authored-by: Isaac Chung <chungisaac1217@gmail.com>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com> (52531fc)
fix: Add tests to validate citations and update citations
fix: Add tests to validate citations and update citations (#5478)
fix: Add tests to vaidate citations and update citations
A little less than 100 cases, where found. I manually went though these.
If a case was found missing (even if just by surface features then i replaced it with a google scholar reference, ACL, etc.. If you came up a second time I wrote an exception.
Please take care to review this in detail.
fix duplicates
fix citations
clean up
lint
format
add cit for tateoba
remove none citations
re-add swebench cit
add timeout
fail on unverified
add from comments
add missing doi
added missing doi
remove non-paper citations from models
update uv.lock
more fixes
finished remaining
format
Modify commit message to skip CI for citation cache
Add [skip ci] to citation cache update commit message.
Co-authored-by: Copilot Autofix powered by AI <175728472+Copilot@users.noreply.github.com>
ci: split citation checks and restore monthly refresh
ci: stop pinning uv to 0.10.12 in citation_check.yml
Reproduced the 45-minute hang from PR #5478's citation-check-pr run:
uv sync --group test --frozen with uv 0.10.12 deadlocks right after
creating the virtualenv, with zero output, matching the CI log exactly.
uv 0.11.22 fixed a resolver deadlock ("more deadlock-resistant concurrent
hashmap"); test.yml doesn't pin a version and gets 0.12.17, which installs
in seconds. Drop the pin so citation checks track the same unpinned,
patched uv version as test.yml.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
The doshi2026tableir @misc entry has no DOI or arXiv id, so refaudit
can't confirm it and a title-only lookup false-matched an unrelated
Crossref record, failing the citation check. Keep only the
herzig-etal-2021-open citation, which the dataset is adapted from and
verifies cleanly.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Same root cause as the reference-models hang: uv sync without --frozen
triggers a full universal re-lock (recomputing metadata across every
platform/Python version for all 80 extras), which stalled in CI for the
extract-and-run job. Add install-for-model-load-test to run uv sync --frozen, separate from running the test itself, matching the
install-for-tests/test split used elsewhere.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Co-authored-by: Isaac Chung <chungisaac1217@gmail.com>
Co-authored-by: Copilot Autofix powered by AI <175728472+Copilot@users.noreply.github.com>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com> (da37120)
fix: keep the benchmark when filtering BenchmarkResults by model
fix(cde): pass the prompt text, not its name, to the model
fix(cde): pass the prompt text, not its name, to the model (#5503)
fix(cde): resolve the prompt text, not its name, before encoding
CDEWrapper.encode passed the result of AbsEncoder.get_prompt straight
into SentenceTransformer.encode(prompt=...), which takes literal prompt
text. get_prompt returns the key of model_prompts, so cde-small-v1/v2
prepended "query" instead of "search_query: ".
Use _resolve_prompt(), the same helper SentenceTransformerEncoderWrapper
and SparseEncoderWrapper already call, which also replaces the hand-rolled
logging block.
test(cde): check the prompt text that reaches the model
Delete tests/test_models/test_cde_prompts.py
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com> (9971d0c)
fix: migrate to bitextparser v2
fix: migrate to bitextparser v2 (#5500)
Initial plan
chore: update bibtexparser v2 lockfile
Co-authored-by: isaac-chung <48971969+isaac-chung@users.noreply.github.com>
Co-authored-by: isaac-chung <48971969+isaac-chung@users.noreply.github.com>
Co-authored-by: copilot-swe-agent[bot] <198982749+Copilot@users.noreply.github.com>
Co-authored-by: Roman Solomatin <36135455+Samoed@users.noreply.github.com>
Co-authored-by: isaac-chung <48971969+isaac-chung@users.noreply.github.com> (0aa1af1)
fix: make model implementations importable without torch
fix: make model implementations importable without torch (#5463)
fix: make model implementations importable without torch
add gme, mixedbread and semantic router models changes
add mps in get_device
updated docs
lintter
changes from review
updated docs (88058a3)
fix: warm the benchmark-schema cache key the route actually uses (#5494)
fix: warm the benchmark-schema cache key the route actually uses
_prewarm_list_schemas submitted _benchmark_schemas_bytes with no arguments,
but /v1/benchmarks always calls it positionally. functools.lru_cache keys on
the argument tuple, so the warmed () entry is never read and the first request
rebuilds the list itself.
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com> (21f8148)
fix: pass num_proc to every dataloader created during evaluation (#5489)
fix: pass num_proc to every dataloader created during evaluation
simplify test
Co-authored-by: Roman Solomatin <36135455+Samoed@users.noreply.github.com> (322d09f)
fix: apply scripts filter in TaskResult.get_score (#5490)
fix: apply scripts filter in TaskResult.get_score
Apply suggestion from @Samoed
lint
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>
Co-authored-by: Roman Solomatin <36135455+Samoed@users.noreply.github.com> (9de9d1c)
fix: accept ISO 639-3 codes in get_model_metas(languages=...) (#5484)
fix: accept ISO 639-3 codes in get_model_metas(languages=...)
simplify
simplify
Co-authored-by: Roman Solomatin <36135455+Samoed@users.noreply.github.com> (8ecdd97)
fix: accept language-script and programming-language codes in filter_tasks (#5483)
fix: accept language-script and programming-language codes in filter_tasks
simplify check
simplify check
Co-authored-by: Roman Solomatin <36135455+Samoed@users.noreply.github.com> (67bd1b8)
fix: allow validate_and_filter in load_results without tasks (#5480) (af2708f)
fix: do not mutate eval_splits in calculate_descriptive_statistics (#5481) (c42a4e6)
fix: pass the selected benchmarks in create-model-results (#5477) (f9b5cf2)
fix: fix citations (#5473)
fix citations (0d9d487)
The lower-bound check was nested inside 'if upper is not None', so a
(lower, None) range filtered nothing. Hoist the guard so either bound
activates the filter; upper-only behaviour is unchanged. (629aa89)
Addresses issues mentioned in #5471 (closing only once we have added tests) (952bb7f)
np.memmap does not infer a -1 dimension, so NumpyCache.load() failed on every
existing cache directory (OverflowError on numpy<2.2, ValueError on numpy>=2.2)
and CachedEmbeddingWrapper could never reuse embeddings from an earlier run.
Derive the row count from the file size instead. (00407be)
model: add albertobarnabo/bge-m3-italian (#5488)
model: add albertobarnabo/bge-m3-italian
model: drop HF-only citation for bge-m3-italian (ee57f4e)
dataset: add MM-BRIGHT retrieval tasks (#5168)
dataset: add MM-BRIGHT reranking tasks
dataset: defer Pillow import for MM-BRIGHT
reupload
Co-authored-by: Roman Solomatin <36135455+Samoed@users.noreply.github.com> (5d3212e)
Add NQ-Tables retrieval (#5461)
Add unregistered NQ-Tables retrieval engineering prototype
Remove NQ-Tables implementation notes from prototype PR
Remove dangling qrel explanatory comment
Add test-only NQ-Tables descriptive statistics
Remove NQ-Tables validation artifacts from task PR
Register NQ-Tables as a standard retrieval task
Promote NQ-Tables from prototype status
Address NQ-Tables review feedback
format citation
Co-authored-by: Roman Solomatin <36135455+Samoed@users.noreply.github.com> (a0ee007)
dataset: add JuriFindITRetrieval (#5464)
dataset: add JuriFindITRetrieval
dataset: build a single test split for JuriFindITRetrieval (29b21ba)
model: restore bidirectional-attention fallback for PyPI colpali_engine (#5469)
fix(evie): restore bidirectional-attention fallback for PyPI colpali_engine
ColQwen3_5.enable_bidirectional_attention ships with the EVIE release of
colpali_engine but is absent from the published PyPI wheels that the evie
extra resolves to, so EvieWrapper.init raised AttributeError on a real
load (mteb#5451). Mock tasks never call from_pretrained, so CI missed it.
Restore the guard with an equivalent local implementation. Both the config
and the module flag have to be flipped: config.is_causal is what
create_causal_mask reads (and what sdpa honours on padded batches), while
Qwen3_5Attention.is_causal is what flash_attention_2 reads. Flipping only
the module flag leaves sdpa silently causal.
Verified bit-exact against the released implementation on padded batches
under both sdpa and flash_attention_2.
review: link the reference implementation, drop redundant comment
lint
Co-authored-by: Roman Solomatin <36135455+Samoed@users.noreply.github.com> (c9f5262)
Add model: Singaraj/sante-embed (#5466)
Add model: Singaraj/sante-embed
Delete mteb_mock_run_results.md
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com> (8765064)
Add NGA-KR/ja-embed-sts ModelMeta (46003a1)
dataset: add NayanaIR monolingual visual document retrieval (t2i, 8 Indic languages) (#5458)
dataset: add NayanaIR monolingual visual document retrieval (t2i, 8 Indic languages)
Update mteb/tasks/retrieval/multilingual/nayanair_monobench_retrieval.py
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>
Co-authored-by: gitgod-debug <186448411+gitgod-debug@users.noreply.github.com>
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com> (fc92859)
Create aurola_omni_models (#5453)
Create aurola_omni_models.py
Co-Authored-By: nehal309 <20415842+nehal309@users.noreply.github.com>
Co-Authored-By: nehal309 <20415842+nehal309@users.noreply.github.com>
Co-Authored-By: nehal309 <20415842+nehal309@users.noreply.github.com>
Co-authored-by: nehal309 <20415842+nehal309@users.noreply.github.com> (3ab2623)
feat: Functionality to clean task and remove low-quality samples
feat: Functionality to clean task and remove low-quality samples (#5107)
Functionality to clean task and remove low-quality samples
Move the filters to mteb.quality and support every modality, only added dedupliation filter
remove the code related to other filters
update docs based on review and based on only deduplication filter
changes from reviews, added all filters to single file, make all functions private, added modality check and threshold for min examples per label, update docs, and support for pair input tasks
update docs, change function names, clarified key naming, move hashes to separate file
changes from review
run lintter
changes from review, removed task_modified and _check_unusable_data
added whats_new section
correct link
update example in docs, added fix for shallow copy and made docstring more clearer for ValueError and KeyError
update what_new.md
changes from review
added skip in tests
changes from copilot review
lintter
fixed CI failure
remove _get_symmetric_sides from abstask and kep in _filters.py
change name from quality to data_cleaning
minor changes (de206dc)
fix: declare ModelMeta load dtypes with OutputDType instead of torch.dtype
fix: declare ModelMeta load dtypes with OutputDType instead of torch.dtype (#5441)
fix: declare ModelMeta load dtypes with OutputDType instead of torch.dtype
address copilot comments
address remaining copilot review comments
Co-authored-by: Isaac Chung <chungisaac1217@gmail.com> (c4f968c)
dataset: add OmniWikiRetrieval (1443790)
refactor: move Nano retrieval tasks to the v2 dataset format
All 13 Nano* tasks carried an identical v1 load_data() override that read the
zeta-alpha-ai copies into self.corpus/self.queries/self.relevant_docs. The
re-uploads under the mteb org are in the v2 layout, so the override can go and
the default AbsTaskRetrieval loader handles them.
Verified per task that the v2 dataset loads identical content to the original:
corpus ids and text, query ids and text, and qrels all match for 13/13. The v1
override hardcoded every relevance score to 1; the uploaded qrels also carry
score 1 throughout, so scores are unchanged.
Refs #3424
Co-authored-by: Claude Opus 5 <noreply@anthropic.com> (f18a069)
Add tencent/EVIE-8B and tencent/EVIE-4.5B (#5450)
Add tencent/EVIE-8B and tencent/EVIE-4.5B
Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
evie pins transformers>=5.13.1, which is incompatible with visrag-ret and
other extras (transformers<4.53). Add evie to [tool.uv] conflicts so uv can
resolve, and regenerate the lockfile.
Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
Delete mteb_mock_run_results.md
fix typing
Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>
Co-authored-by: Roman Solomatin <36135455+Samoed@users.noreply.github.com> (a25a026)
dataset add ColDeRReranking benchmark (#5409)
feat(tasks): add ColDeRReranking benchmark (#2709)
test: add duplicate_text exemption for ColDeRReranking in task quality tests
fix(reranking): clarify ColDeR benchmark semantics
fix: normalize ColDeR citation formatting
fix(reranking): remove unused ColDeR logger
Apply suggestion from @Samoed
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com> (5427a1f)
model: Add litillabs/litil-embed-0.6b (#5456)
Add litillabs/litil-embed-0.6b
Remove mock run report (40a77d2)
model: Add models multi-modal-embed [MOEB] (#5011)
[MOEB] Add models multi-modal-embed
adding n_embedding parameter
fix large model
fix lint
robustness auto model or model
lock update + remove numpy cast + shared base for large and small + remove _COMMON
remove librosa, cherry pick custom code from large model, and remove patches
remove labels
simplifying implementation and pinning revisions
fix lint
update number of parameters (ef8b40f)
task: add COCO Modality Equivalence retrieval tasks (#5384)
feat: add COCO Modality Equivalence retrieval tasks (issue #5358)
fix: bibtex formatting and ruff format in coco modality equivalence
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Fill dataset revision and fix build script for COCO modality equivalence
Fix ruff formatting in create_data.py
Add COCO modality equivalance task, analysis script and gap results
Add directional asymmetry analysis script and results (issue #5360)
Add sampling budget analysis script and results (issue #5362)
feat(analysis): PCA latent dimensions analysis for MTEB issue #5367
Fix COCO modality-equivalence retrieval tasks and add descriptive stats
Co-authored-by: Rakshitha Ireddi <rakshithaireddi@Rakshithas-MacBook-Pro.local>
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
Co-authored-by: Isaac Chung <chungisaac1217@gmail.com> (1dfd69a)
Create Ops MM models (2B, 7B) (#5332)
Create ops_mm_models
Co-Authored-By: Deep Shah <21212684+deep9539@users.noreply.github.com>
Update ops_mm_models.py
lint
Co-authored-by: Deep Shah <21212684+deep9539@users.noreply.github.com>
Co-authored-by: Roman Solomatin <36135455+Samoed@users.noreply.github.com> (4e24e0c)
Add WeMM embedding model. (#5333)
Add WEMM models
Update pyproject.toml
Use InstructSentenceTransformerModel
fix video and better support instruction prompt
Update wemm_models.py
Update wemm_models.py
reformat and fix batch padding
fix padding
fix padding
Update wemm to use sentence_transformer
Add init method
Reformat to pass linter
Add comment to explain the permute
address comments
Co-authored-by: Roman Solomatin <36135455+Samoed@users.noreply.github.com> (951d86b)
Add RzenEmbed model (#5339)
first commit
handle empty task instruction
resolve naming discrepancy
fix things
add tensor support
dimension handling
fix error
visual embeddings to language model dimension
Update rzen_embed_model.py
Fix system prompt, image + video coprocessing.
Co-Authored-By: nehal309 <20415842+nehal309@users.noreply.github.com>
Co-Authored-By: nehal309 <20415842+nehal309@users.noreply.github.com>
Co-Authored-By: nehal309 <20415842+nehal309@users.noreply.github.com>
Co-Authored-By: nehal309 <20415842+nehal309@users.noreply.github.com>
Co-Authored-By: nehal309 <20415842+nehal309@users.noreply.github.com>
Add change to make rzen embed compatible to >= 4.57.0
add rzen training data and embedding_params
lint
Update rzen_embed_model.py
Co-authored-by: nehal309 <20415842+nehal309@users.noreply.github.com>
Co-authored-by: Roman Solomatin <36135455+Samoed@users.noreply.github.com> (74ee71c)
dataset: add ToolRetrieval benchmark task (#5410)
feat(retrieval): add ToolRetrieval benchmark task (#3628)
fix(retrieval): reproduce published ToolRet results
ToolRetrieval scored ~11 NDCG@10 above the paper because it differed from the
reference implementation in three ways, each making retrieval easier:
eval_retrieval(category="all") doesweb, whose tasks rangePool the corpus, expose the 35 retrieval tasks as subsets so MTEB's
cross-subset mean matches the paper's aggregation, and add
ToolRetrievalInstruction for the w/ inst. setting (Table 5) alongside
ToolRetrieval (w/o inst., Table 4).
all-MiniLM-L6-v2 w/ inst., NDCG@10, paper Table 5 in parentheses:
web 13.23 (12.77), code 32.38 (31.59), customized 33.29 (32.24). Mean absolute
delta drops from 10.91 to 0.76. Reproduce with scripts/reproduce_toolret.py,
which also raises max_seq_length to 512 as the reference implementation does;
this model's card ships 256 and web swings ~5 NDCG@10 on that alone.
The example script now carries the reference implementation's per-model
handling, without which the published numbers do not come out: SentenceTransformer
vs fp16 AutoModel dispatch, per-family pooling, the L2-normalization skip for
contriever and gtr-t5, per-family prompt templates, min(max_position_embeddings,
2048), and word-level truncation. It takes --model/--all/--settings and prints a
PASS/FAIL against the published values, defaulting to Table 5.
Over 9 baselines the paper's findings replicate:
Exact per-cell agreement is not available: the reference print_results() computes
a size-weighted (micro) mean while the published tables match an unweighted
(macro) one, so the paper's numbers were not produced by the released code.
Two bugs in the reproduction script, both found by running it over every
baseline rather than one:
Coverage is now 10 of Table 5's 26 rows -- every single-vector dense retriever
under 7B that runs in this harness. gtr-t5-large is included despite scoring
below gtr-t5-base, inverting the paper's ordering for that pair; it is reported
as measured rather than dropped, and the docstring records both the headline
figures and the figures excluding it.
fix(scripts): use native MTEB ToolRetrieval evaluation
chore: remove Modal evaluation helper
reupload
remove unnecessary script
Co-authored-by: Roman Solomatin <36135455+Samoed@users.noreply.github.com> (a768963)
Add ModelMeta for NGA-KR/ko-embed-cls (#5446)
Add ModelMeta for NGA-KR/ko-embed-cls
Fix ModelMeta for NGA-KR/ko-embed-cls (lint, model_type, prompts)
Add n_embedding_parameters
Co-authored-by: Isaac Chung <isaac.chung@foam.io> (cb65aad)
task: add BioVITA multimodal retrieval (#5153)
task: add BioVITA multimodal retrieval
fix: use fixed-format datasets for BioVITA
fix: simplify BioVITA task metadata and loading
fix: classify BioVITA tasks as reranking
fix: address BioVITA reranking review
style: format BioVITA data script
fix: return BioVITA audio fallbacks
fix: add BioVITA descriptive statistics
test: update BioVITA reranking quality allowlist
docs: clarify BioVITA tie handling
refactor: remove _BioVITAReranking base class, add task_specific_scores to each class explicitly
Also fix BibTeX formatting in webvid_covr files to satisfy pre-commit hook.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Co-authored-by: Michelle Yang <myang333@users.noreply.github.com>
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>
Co-authored-by: Isaac Chung <chungisaac1217@gmail.com>
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com> (6ee088e)
Add ModelMeta for NGA-KR/ko-embed-nli (#5447)
Add ModelMeta for NGA-KR/ko-embed-nli
Add n_embedding_parameters (8b85f40)
add Nemotron model citations and public training references (#5445)
fix: add Nemotron citations and Embed-VL training code link
fix: link Nemotron 3 public training datasets
fix: link the full Nemotron 3 training data sections (4fe4737)
add OpenMDW license metadata for corresponding models (#5444)
fix: update OpenMDW license metadata for Nemotron models (5826242)
fix: recognize OpenMDW-1.1 as an open license (1183517)
Add ModelMeta for NGA-KR/ko-embed-v0 (#5440)
Add ModelMeta for NGA-KR/ko-embed-v0
lint
Co-authored-by: Roman Solomatin <36135455+Samoed@users.noreply.github.com> (a8bce9d)
docs: annotate kwargs with Any in adding_a_dataset.md
docs: annotate **kwargs with Any in adding_a_dataset.md (#5395)
docs: annotate **kwargs with Any in adding_a_dataset.md examples
ci: skip tests for docs-only PRs
Revert "ci: skip tests for docs-only PRs"
This reverts commit abe747b. (ed47a25)
Fine-tuned from LaBSE on the same Mauritian Creole corpus as
Singaraj/morisien-embed, which was fine-tuned from multilingual-e5-base.
The first model is now marked superseded_by this one. On a retrieval pool
built from FLORES+ with the released xSIM++ augmentation it scores below
untrained LaBSE, while this one is the first configuration measured that
beats it: error rate 0.2932 against 0.3343, McNemar exact p = 0.00083. (48c1f79)
Add model: sshalimov04/ru-reranker-edge-150m (#5401)
Add ModelMeta for sshalimov04/ru-reranker-edge-150m
Add ModelMeta for sshalimov04/ru-reranker-edge-150m
Add ModelMeta for sshalimov04/ru-reranker-edge-150m
Add ModelMeta for sshalimov04/ru-reranker-edge-150m
fix lint
Co-authored-by: Roman Solomatin <36135455+Samoed@users.noreply.github.com> (79a9dad)
dataset: add FineGrainOCR image-text clustering (#5341)
Explore FineGrainOCR cross-modal clustering
Build and benchmark FineGrainOCR clustering subset
Move FineGrainOCR build scripts off task branch
dataset: add FineGrainOCR image-text clustering
dataset: use full FineGrainOCR validation split
dataset: simplify FineGrainOCR task evaluation (5050aab)
dataset: add Spoken Wikipedia speech-text retrieval (a2t, t2a, 5 languages) (#5368)
dataset: add Spoken Wikipedia speech-text retrieval (5 languages)
Co-authored-by: Claude Opus 5 <noreply@anthropic.com> (5681a90)
dataset: add Multi30k multilingual image-text retrieval (t2i, i2t) (#5322)
dataset: add Multi30k multilingual image-text retrieval (t2i, i2t)
Adds Multi30kT2IRetrieval and Multi30kI2TRetrieval covering English,
Czech, German and French. Neither direction was previously in mteb.
Multi30k stores one row per Flickr30k image with parallel human
translations in four columns, so the image side is identical across
language subsets. The loader builds it once per split and varies only
the caption side.
Part of the multilingual coverage tracked in #4842.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Two CI failures:
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Addresses review feedback: the loader now writes
task.dataset[lang][split] = RetrievalSplitData(...) instead of returning
separate corpus/queries/relevant_docs dicts, and load_data on both classes
delegates to the shared helper.
Also drops the modality columns and the None filler columns on both sides.
They only existed to satisfy the old format; with RetrievalSplitData the
modality follows from the columns present, so the image side is id + image
and the text side is id + text.
Behaviour is unchanged - clip-vit-base-patch32 on en gives 0.7433 (t2i) and
0.7598 (i2t), identical to the scores in the PR description.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Addresses review feedback asking for the source and a transformation script.
The task previously reshaped a third-party upload at load time. The reshaped
form is now published in mteb's standard retrieval format, so the task file is
metadata only and uses the default loader, and the transformation is
reproducible from scripts/data/multi30k_retrieval/create_data.py rather than
implicit in load_data.
Source is romrawinjp/multi30k at revision 110e827, MIT, unchanged. The image
side is identical across the four language subsets, so it is stored once and
every language config points at the same file rather than duplicating it.
The image table is written through the datasets API so the parquet carries the
Image feature in its schema metadata; writing the struct directly produces a
plain {bytes, path} column that does not decode on load.
Behaviour is unchanged: clip-vit-base-patch32 on en gives 0.7433 (t2i) and
0.7598 (i2t), identical to before, and the descriptive statistics are
byte-identical.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
fix: correct license to cc-by-sa-4.0 and add W17-4718 citation for French extension
fix: reformat BibTeX citations to pass citation formatting test
fix: sort BibTeX fields alphabetically to pass citation formatting test
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
Co-authored-by: Isaac Chung <chungisaac1217@gmail.com> (20c5a7b)
model: add ImageBind model (Track 4, MOEB) (#5015)
feat: add ImageBind model wrapper (Track 4, MOEB)
Adds ImageBindWrapper and ModelMeta for nielsr/imagebind-huge — Meta's
joint multimodal embedding model supporting text, image, and audio in a
single 1024-dim embedding space.
imagebind package (optional dep group)imagebind optional dep group to pyproject.tomlCo-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
T alias to transforms (lowercase module as uppercase)Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
fix: ruff ICN001 - use from torchvision import transforms
fix: sort imports in _load_images (isort I001)
fix: remove temp WAV files, add source comment for image transforms
fix: ruff format - wrap long get_clip_timepoints call
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Co-authored-by: Rakshitha Ireddi <rakshithaireddi@Rakshithas-MacBook-Pro.local>
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
Co-authored-by: Isaac Chung <chungisaac1217@gmail.com> (61129c3)
fix: isolate embedding caches by prompt type
fix: isolate embedding caches by prompt type (#5386)
fix: isolate embedding caches by prompt type
test: verify prompt-specific cache directories
test: generalize prompt cache isolation coverage
Update docs/get_started/advanced_usage/cache_embeddings.md
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>
Co-authored-by: Roman Solomatin <36135455+Samoed@users.noreply.github.com> (3eb1cf9)
fix(openai): add %s placeholder so retry logging does not raise TypeError
Both OpenAI embedding retry handlers passed the caught exception as a
positional argument to logger.info() while the format string had no %
placeholder, so LogRecord.getMessage() raised TypeError: not all arguments converted during string formatting.
The CLI installs RichHandler (build_cli.py) whose emit() calls
self.format(record) without the try/except that stdlib StreamHandler
uses, so the TypeError propagates out of logger.info(), out of the
except block, and neither time.sleep() nor the retry ever runs. With
the default --verbosity 2 the mteb logger sits at INFO, so a single
transient 429 aborts the run and masks the real API error.
Adding the placeholder restores the intended message and matches the
existing style in google_text_embedding.py. The # noqa: PLE1205
suppressions are removed because the rule no longer fires and RUF100
(unused-noqa) is enabled.
Co-authored-by: Claude Opus 5 <noreply@anthropic.com> (0a05983)
fix: repair broken Vultron revision pins
fix: repair broken Vultron revision pins (#5378)
fix: repair broken Vultron revision pins
fix: propagate ColQwen revision to all loaders
test: narrow ColQwen revision assertions
fix: update Vultron revisions
remove adapter_kwargs
lint
Co-authored-by: Roman Solomatin <36135455+Samoed@users.noreply.github.com> (c9da0cb)
dataset: add TAU Urban Acoustic Scenes 2022 clustering (a2a) (#5371)
dataset: add TAU Urban Acoustic Scenes 2022 clustering (a2a)
Closes #5019. mteb already scores this dataset with labels through
TAUAcousticScenes2022Mobile; this adds the unsupervised counterpart the issue
asks for, on the same stratified subsample at the same revision so the two are
directly comparable.
CLAP reaches v_measure 0.1680 against a 0.0049 random floor. That it does well
here while sitting at chance on speech retrieval is consistent with its training
on audio captions of environmental sound.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
reference pointed at the dataset record, so the task had no paper behind it.
Mesaros, Heittola and Virtanen describe this recording collection and the
device-generalization setup in the DCASE 2020 workshop proceedings, which is
also what soundata cites for this dataset family. The Zenodo release stays in
the bibtex as the source of the actual files.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com> (178c79f)
Adds VaaniA2TRetrieval and VaaniT2ARetrieval over 45 Indian languages, from
Vaani's transcribed release.
That release ships an official test split, so unlike the main Vaani corpus the
evaluation set is held out at source rather than sampled from training data.
45 of the 64 languages with a test split are kept; the rest fall under 25 usable
utterances. 4,001 clips.
Scripts were determined from the transcripts rather than assumed. Two are not
what the language name suggests: Chakma is romanised in this release despite
having its own script, and Tulu is written in Kannada.
Transcripts carry annotation markup - <noise>, <pause> and similar event tags
plus {...} braces marking code-switched English. Tags are removed as annotation
artefacts; braces are unwrapped so the code-switched word survives.
Byte-identical audio and repeated transcript text are dropped per language so
an identical query is not marked relevant to only one of the clips it matches.
main_score is hit_rate_at_5, matching every existing audio-text retrieval task
in mteb.
Co-authored-by: Claude Opus 5 <noreply@anthropic.com> (dca9d97)
ci: defer datasets import to avoid slow cold-start when reading model metadata
ci: defer datasets import to avoid slow cold-start when reading model metadata (#5385)
fix: defer datasets import to avoid slow cold-start when reading model metadata
fix: support Python <3.11 in import hygiene test (enum.StrEnum fallback)
fix: stub HelpfulStrEnum as str in test, works on Python 3.10+ (819d04d)
Enables ANN401 (flake8-annotations: any-type) (#5381)
Enables ANN401 (flake8-annotations: any-type)
changes from copilot review
changes from review
lintter
minor change (5403187)
[MOEB] Add EMID pair classification (#5291)
[MOEB] Add EMID pair classification and EMID a2i + i2a retreival datasets
add stats
fix types
reduce size and dedup
adapt eval for multi-model pair classif
remove retrieval updates
move data to mteb org
fix metadata
remove monkey patch
fix mock pair classif
revert changes specific to pair classif eval (fc1d87d)
fix: update stale tokenizer revisions for Transformers v5
fix: update stale tokenizer revisions for Transformers v5 (#5377)
fix: update stale tokenizer revisions
test: remove brittle tokenizer fingerprints (dd3cd65)
Enables ANN205 (missing-return-type-static-method) Ruff Rule (#5380) (2e54d27)
Enables ANN (flake8-simplify) Ruff Rule (#5372)
Enables SIM (flake8-simplify) Ruff Rule
enable ANN in mteb/models folder
enable ANN in mteb/tasks folder (56fa292)
model: improve MIRACL query prompt for voyage-4-large (#5374)
Evolutionary prompt search over the 18-language macro. nDCG@10 0.65193 -> 0.67583,
better on all 18 languages, validated on 6 languages held out of the search.
Co-authored-by: fzowl <zoltan@voyageai.com> (d684be9)
model: add BioVITA multimodal encoder (#5346)
model: add BioVITA multimodal encoder
fix: address BioVITA model review comments
Delete mteb_mock_run_results.md
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com> (ae93300)
dataset: add Omnilingual ASR speech-text retrieval (50 new languages)
Co-authored-by: Claude Opus 5 <noreply@anthropic.com> (f2abe47)
dataset: add GLAMI-1M retrieval (#5331)
dataset: add GLAMI-1M multimodal classification
dataset: address GLAMI-1M review feedback
dataset: keep GLAMI names and descriptions separate
dataset: add GLAMI-1M image-to-text retrieval
Rename glami_1m_t2i_retrieval.py to glami_1m_retrieval.py
Fix GLAMI retrieval import after rename
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com> (e217c5a)
dataset: add CAMEO multilingual speech emotion classification (5 languages)
Co-authored-by: Claude Opus 5 <noreply@anthropic.com> (e253d04)
task: add Lombard GRID audiovisual retrieval tasks (#5337)
task: add Lombard GRID i2va retrieval
task: add Lombard GRID audiovisual retrieval tasks (4306a1b)
dataset: add Afri-MCQA multilingual speech-image retrieval (a2i, i2a) (#5356)
dataset: add Afri-MCQA multilingual speech-image retrieval (16 African languages)
Co-authored-by: Claude Opus 5 <noreply@anthropic.com> (fbda84a)
[MOEB] Add XModBench any-to-any retrieval tasks (#5294)
Add XModBench any-to-any retrieval tasks
Signed-off-by: jupyterjazz <saba.sturua@jina.ai>
Signed-off-by: jupyterjazz <saba.sturua@jina.ai>
Signed-off-by: jupyterjazz <saba.sturua@jina.ai>
Signed-off-by: jupyterjazz <saba.sturua@jina.ai>
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>
Signed-off-by: jupyterjazz <saba.sturua@jina.ai>
Signed-off-by: jupyterjazz <saba.sturua@jina.ai>
Signed-off-by: jupyterjazz <saba.sturua@jina.ai>
Signed-off-by: jupyterjazz <saba.sturua@jina.ai>
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com> (b287699)
feat(models): add kazalbrur/bangla-embed-e5-small-banglish
Bengali/Banglish sentence encoder (118M) distilled from BGE-M3 on top of
intfloat/multilingual-e5-small, with cross-script (romanized Bengali)
retrieval support. Uses the E5 query:/passage: prompt convention and a
1024-dim dense projection head.
Co-authored-by: Kazal Chandra Barman <kazal.chandra@technonext.com> (3334caa)
feat: register nlpai-lab/KURE-v2 Korean ColBERT model
Add the ModelMeta for nlpai-lab/KURE-v2, a PyLate ColBERT (late-interaction,
multi-vector) model adapted from skt/A.X-Encoder-base, using the
MultiVectorModel loader. Training data is lightonai/embeddings-fine-tuning
plus a subsample of lightonai/embeddings-pre-training-curated.
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> (8dc44a3)
Adds a2t/t2a, v2t/t2v and va2t/t2va over AVCaps, plus descriptive statistics.
AVCaps captions each clip three ways - audio only, visuals only, and both
together - so the audio-only, video-only and combined directions can be scored
independently on identical clips rather than inferred from one mixed caption set.
Audio is demuxed to a separate mono 16 kHz stream and each task exposes only the
media its captions describe, so an audio-caption task cannot be solved off the
video track. Official test split only: 200 clips.
Repeated caption text is dropped (725/788/808 -> 715/780/802). Annotators
sometimes write the same short sentence for different clips, which would leave
an identical query marked relevant to only one of the clips it describes.
Co-authored-by: Claude Opus 5 <noreply@anthropic.com> (a1268f9)
Add CrisisMMD image-text classification tasks (#5340)
Add CrisisMMD image-text classification tasks
Add CrisisMMD zero-shot classification tasks
Revert "Add CrisisMMD zero-shot classification tasks"
This reverts commit 6123c90.
Inline CrisisMMD classification settings (12aab77)
dataset: add MVL-SIB sentence-to-image retrieval (#5334) (8487c1a)
fix: Make it such that torch is not a required import of mteb.types (#5329)
make mteb.types torch-free
changes from review
rename file to tests/test_ensure_no_torch_at_import.py (5f1167a)
fix: add recording-level FleursT2ARetrieval.v2 / FleursA2TRetrieval.v2
fix: add recording-level FleursT2ARetrieval.v2 / FleursA2TRetrieval.v2 (#5307)
fix: add recording-level FleursT2ARetrieval.v2 / FleursA2TRetrieval.v2
FLEURS id identifies a sentence, not a recording: each sentence is read by
up to six different native speakers who all share one id. Both Fleurs
retrieval tasks key queries, corpus and qrels on that id, so all recordings of
a sentence collapse into a single id. They are not dropped -- every row is
still encoded and scored -- but for T2A they merge under one result key, so
aggregation over speakers happens implicitly and depends on row order rather
than being declared as multi-positive qrels; for A2T they are never evaluated
as independent queries. Across the 102 subsets 77,809 recordings carry only
33,018 distinct ids.
The FLEURS paper defines both directions explicitly (Conneau et al., SLT 2022,
Sec. 4): speech-to-text retrieves "the correct text segment", and text-to-speech
scores "retrieving any of the speakers who speaks the correct textual query".
That is exactly hit_rate_at_k (success@k), which both tasks already use as
main_score -- the metric was right, only the qrels were not.
v2 matches those semantics:
Recording ids are {sentence_id}-{rank}, ranked by the globally unique FLEURS
audio filename, so an id is invariant to row order. In ln_cd two distinct
sentence ids carry byte-identical text; those documents are indistinguishable,
so both are marked positive rather than merging the ids.
v1 is left intact so published results stay valid; it only gains a
superseded_by pointer, following the BSARDRetrieval -> BSARDRetrieval.v2
convention for a one-to-many qrel fix. get_tasks(tasks=[...]) short-circuits
before exclude_superseded, so MAEB(beta) still resolves v1 by name.
Descriptive stats for the new tasks reuse v1's audio statistics, justified by
verifying the audio column is byte-identical between the v1 and v2
constructions; the text and qrel statistics are recomputed. Truncated audio in
the upstream data is a separate data-quality issue and is deliberately not
filtered here.
Fixes #5270
Delete tests/test_tasks/test_fleurs_v2.py
lint
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>
Co-authored-by: Roman Solomatin <36135455+Samoed@users.noreply.github.com> (ef85903)
dataset: add WIT image-to-text retrieval (0d7a65c)
dataset: add XM3600 image-to-text multilingual retrieval (#5325)
dataset: add XM3600 image-to-text retrieval
dataset: address XM3600 I2T review feedback (05a344c)
model: Pass kwargs to bm25s (#5268)
pass kwargs to bm25s (9e96b20)
dataset: add XFlickr30k-Co image-to-text retrieval (#5324)
dataset: add XFlickr30k-Co image-to-text retrieval
fix: address XFlickr30k-Co review feedback (2fa7131)
Enable A (flake8-builtins) ruff rule (#5328)
enable A ruff rule (d40e259)
Adds DROIDIT2VRetrieval, composed image+text -> video retrieval over the
DROID in-the-wild Franka manipulation dataset. 1,500 successful episodes
with unique language instructions (5-60 s at 15 fps), evenly sampled
across the collection. The query pairs the initial scene image from the
exterior_1 camera with the instruction; the corpus holds exterior_2
videos of the same episodes, so queries and documents never share a
viewpoint and exact frame matching cannot solve the task. Relevance is
instance-level 1:1.
Reference runs (ndcg@10): mteb/baseline-random-encoder 0.0028;
jinaai/jina-embeddings-v5-omni-nano 0.1691.
Claude-Session: https://claude.ai/code/session_01Vue9R33gcLojTcxp8ZDkW3
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com> (86079bf)
dataset: add BridgeData V2 v2v robot manipulation retrieval (MOEB) (#5259)
dataset: add BridgeData V2 v2v robot manipulation retrieval (MOEB)
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Vue9R33gcLojTcxp8ZDkW3
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Vue9R33gcLojTcxp8ZDkW3
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Vue9R33gcLojTcxp8ZDkW3
Co-authored-by: Claude Fable 5 <noreply@anthropic.com> (f1a3c51)
fix: mock_tasks not supported by get_instruction
fix: mock_tasks not supported by get_instruction (#5298)
fix mock_tasks not supported by get_instruction
fix typecheck
fixing to get prompt from base class
added prompts for zero-shot classification and Image-Text Pair classification in
remove code from get_tasks (b36fc25)
fix image dataloader (#5309)
fix image loader
rename
fix typing (d384f76)
Update SugarCrepe task description to reflect word order (#5314)
SugarCrepe: note that 572/7511 pairs differ only in word order
Counted on mteb/SUGARCREPE_fmt: 572 of 7511 pairs have a caption and
negative_caption that are the same multiset of words, all from the
swap_obj and swap_att subsets. An order-insensitive text representation
scores chance on those by construction, capping text_acc at 0.962. (1c9514d)
ci: limit PR dataset checks to added task files
ci: limit PR dataset checks to added task files (#5293)
ci: limit PR dataset checks to added task files
ci: narrow dataset fallback fix
Co-authored-by: nightcityblade <nightcityblade@gmail.com>
Co-authored-by: Isaac Chung <chungisaac1217@gmail.com> (dff378f)
fix: Remove legacy and add superseeded by
Since we have the "newer version" below I think it is better to remove legacy. If we keep legacy then it might be nice to add a rule to check that superseeded by is specified.
For jpn it seems to be legacy, but I can't see any newer version?
For farsi I added superseeded by. (10d8943)
Add DNA-VL-STEER-2B (#5313)
model: Add DNA-VL-STEER-2B
DNA-VL-STEER-2B is a language-bias-calibrated variant of
Qwen/Qwen3-VL-Embedding-2B. Identical architecture and loader, so the meta
inherits the base entry's capacity fields; it differs in declaring all 36
calibrated languages.
Set adapted_from instead of describing the base model in a comment (293e9ad)
dataset: add Spanish Wikinews clustering tasks (#5308)
Add Spanish Wikinews clustering tasks
Add Spanish Wikinews descriptive statistics
chore: allow Spanish prompt token in typos
fix: format Spanish Wikinews citations
Co-authored-by: clemente Ranokau <clemente@ranokau.com> (5b35ec8)
Add amgix/static-retrieval-multilingual-69m-v1 to misc_models (a92ff3e)
Add dense webvid retrieval (#5215)
First commit dense_webvid_retrieval
Fix revision and pass test
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com> (775411c)
Create speech edit acoustic AT2A dataset (#5285)
Create speech edit acoustic AT2A dataset
Co-Authored-By: nehal309 <20415842+nehal309@users.noreply.github.com>
Co-Authored-By: nehal309 <20415842+nehal309@users.noreply.github.com>
Co-Authored-By: nehal309 <20415842+nehal309@users.noreply.github.com>
Co-Authored-By: nehal309 <20415842+nehal309@users.noreply.github.com>
update bibtex to pass test
some nits and license fix
Add SpeechEditAcousticRetrieval to duplicate_text
uv ruff fix
Update data_prep.py
Update test_task_quality.py
Co-authored-by: nehal309 <20415842+nehal309@users.noreply.github.com> (be7d3a1)
Add ACM composed audio dataset (#5187)
ACM dataset
fix test_task_quality.py
Update test_task_quality.py
Update init.py
Apply suggestion from @Samoed
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com> (8e9f956)
Add narcolepticchicken/octen-law-8b-v1 (#5305)
Add Octen Law 8B v1 model
Disclose model training datasets
style: format Octen model registration
Use ScoringFunction enum for Octen Law
Remove mock run results
Update Octen Law namespace to Litil Labs (9d2de3a)
Add NSynth instrument family clustering task (#5111)
task: add NSynth instrument family clustering
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> (d7adced)
dataset: add VCDB core video and audio-video retrieval (#5269)
dataset: add VCDB core video retrieval
dataset: add VCDB core audio retrieval
dataset: document VCDB duplicate audio
dataset: mark VCDB license unspecified
dataset: replace VCDB audio task with audio-video (b398a41)
dataset: add ManiSkill i2v/v2i robot manipulation retrieval (MOEB) (#5257)
dataset: add ManiSkill i2v/v2i robot manipulation retrieval (MOEB)
Adds ManiSkillI2VRetrieval and ManiSkillV2IRetrieval, image<->video
retrieval over ManiSkill3 motion-planning demonstrations (8 tabletop
tasks, 150 successful episodes each, replayed via environment states and
rendered at 256x256 from two viewpoints: base sensor camera for
goal-state images, human render camera for videos). Per task, episodes
split into disjoint query (10) and corpus (140) pools; relevance is
task-level and multi-positive with the query's source episode held out.
Instance-level 1:1 designs over the near-duplicate corpus measured at
chance level for current models and were rejected. Adds the v2i task
category and Robotics domain (same additions as the LIBERO PR; merges
cleanly in either order).
Reference runs (ndcg@10): mteb/baseline-random-encoder 0.1236 (i2v) /
0.1129 (v2i); jinaai/jina-embeddings-v5-omni-nano 0.4193 (i2v) /
0.4766 (v2i).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Vue9R33gcLojTcxp8ZDkW3
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Vue9R33gcLojTcxp8ZDkW3
The official motion-planning demo files reuse episode seeds, so the
first-150 slice contained 25 duplicate corpus episodes (caught by
test_dataset_quality). Select the first 150 unique-seed successful
episodes instead, rebuild both datasets, refresh revisions, descriptive
stats and reference results.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Vue9R33gcLojTcxp8ZDkW3
Co-authored-by: Claude Fable 5 <noreply@anthropic.com> (739d8ef)
dataset: add LIBERO i2v/v2i robot manipulation retrieval (MOEB) (#5255)
dataset: add LIBERO i2v/v2i robot manipulation retrieval (MOEB)
Adds LIBEROI2VRetrieval and LIBEROV2IRetrieval, image<->video retrieval
over the LIBERO robot manipulation benchmark (40 tasks, 1,693 episodes,
256x256 @ 10 fps). Queries are goal-state images (final frames) of
held-out episodes; relevance is task-level and multi-positive with the
query's source episode excluded from the corpus, so exact frame matching
cannot solve the task. Adds the v2i task category (previously empty
direction) and a Robotics task domain.
Reference runs (ndcg@10): mteb/baseline-random-encoder 0.0297 (i2v) /
0.0261 (v2i); jinaai/jina-embeddings-v5-omni-nano 0.3537 (i2v) /
0.3312 (v2i).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Vue9R33gcLojTcxp8ZDkW3
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Vue9R33gcLojTcxp8ZDkW3
Co-authored-by: Claude Fable 5 <noreply@anthropic.com> (718d295)
Add Clotho moment dataset (AT2A) (#5287)
Add clotho-moment retrieval task
update version
Co-Authored-By: nehal309 <20415842+nehal309@users.noreply.github.com>
Co-Authored-By: nehal309 <20415842+nehal309@users.noreply.github.com>
Co-Authored-By: nehal309 <20415842+nehal309@users.noreply.github.com>
Move descriptive stage to right directory.
Fix formatting in Clotho moment retrieval file
Update test_task_quality.py
Update test_task_quality.py
Co-authored-by: nehal309 <20415842+nehal309@users.noreply.github.com> (bb05784)
dataset: add EDIR (#5247)
dataset: add EDIR
update bibtex format and mark known duplicates
add script for processing data
lint format
add prompt
update query prompt and delete document prompt (2440f68)
Update GreenLeaf Law Embed Tiny: add 35+ languages (#5304)
Update GreenLeaf languages: add 35+ supported languages
Model trained on multilingual legal corpus covering 35+ languages
including English, German, French, Spanish, Chinese, Japanese,
Korean, Arabic, Hindi, and others.
Co-authored-by: Surya-saka <surya@judicialmind.ai> (626de72)
[MOEB] Add UniME-V2-LLaVA-OneVision-8B model (#5296)
Add UniME V2 model
Signed-off-by: jupyterjazz <saba.sturua@jina.ai>
Signed-off-by: jupyterjazz <saba.sturua@jina.ai>
Signed-off-by: jupyterjazz <saba.sturua@jina.ai>
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com> (9ef064e)
ci: fix dataset_loading workflow hang by replacing uv sync with uv pip install
ci: fix dataset_loading workflow hang by replacing uv sync with uv pip install (#5290)
ci: isolate uv cache to diagnose dataset_loading hang
ci: trigger dataset_loading workflow on changes to itself
ci: add venv creation diagnostics and restore pre-created venv workaround
ci: install Python before diagnostic step
ci: add trace logging and --no-install-project isolation test
ci: scope RUST_LOG to uv= to reduce trace volume
ci: replace uv sync with pip install to fix hang
uv sync hangs indefinitely because it must parse uv.lock (54 MB, 19k lines)
before doing anything else. Rust serde deserialization of this file likely
exhausts the runner's available memory, causing the process to stall.
The dataset-loading test only needs mteb, pytest, and pytest-rerunfailures,
so we bypass the lockfile entirely with a direct pip install.
ci: use uv pip install instead of uv sync to avoid lockfile parsing
ci: add --system flag to uv pip install
ci: create venv via python -m venv then uv pip install
ci: add bibtexparser dependency
ci: add gitpython dependency (8df869d)
ci: use minimal deps and fallback venv for dataset_loading workflow (#5289) (f5f504b)
ci: use Python 3.10 in dataset_loading workflow (#5286)
ci: use Python 3.10 in dataset_loading workflow to share uv cache with lint
ci: add 30m timeout and verbose output to dataset_loading install step (526cd73)
ci: use --frozen flag in dataset_loading and lint install steps (#5277)
ci: use --frozen flag in dataset_loading and lint install steps
Bypasses dependency resolution in CI where the lockfile should always
be used as-is, speeding up the install step.
ea5bbe6)fix: Enable PT ruff rule (#5238)
enable PT ruff rule
PT018 is not enforced and PT006 default apply (d088296)
update invalid links setting (3bf8702)
Fix jinav4 expriments (#5095)
use device during compute
fix experiment kwarg pass
fix dense multimodal
change model type to list (1459881)
model: Add gve models (#4975)
model: add GVE video embedding models (3B, 7B)
Adds Alibaba-NLP/GVE-{3B,7B}, general video embedders built on
Qwen2.5-VL that support text, image, video, and composed queries.
The HF repos ship a custom Qwen25VLForEmbedding class, but it is a
plain subclass of Qwen2_5_VLForConditionalGeneration with no extra
weights (and its remote code is incompatible with transformers
>= 4.56), so the native class is loaded instead and embeddings are
read from output_hidden_states, skipping vocab logits via
logits_to_keep=1. Pooling follows the model card: L2-normalized
last-token hidden state with left padding. Video frame sampling
reuses FramesCollator (fps=1, max 8 frames, per the model card).
Verified on MPS: text, image, and video inputs each produce
normalized 2048-dim embeddings for GVE-3B.
Closes #3770
transformers 5.0 video processors read the pixel budget from
size["longest_edge"] and ignore the legacy max_pixels attribute,
so videos were processed at full resolution and overflowed
max_length, truncating vision tokens. Set both fields.
The 8-frame demo settings from the model card under-sample videos for
retrieval; MSRVTT R@1 came in well below the paper. fps=2 capped at 32
frames matches other mteb video wrappers while keeping video tokens
(~800) within the 1200 max_length budget.
32-frame videos tokenize past the demo's 1200 max_length, truncating
vision tokens which the processor rejects. 4096 leaves headroom; text
batches are unaffected since padding is to longest-in-batch.
The released GVE checkpoints pool the <|endoftext|> token appended
after the assistant turn (the convention of the team's GME codebase),
not the bare generation prompt shown in the model card demo. Without
it, retrieval quality drops sharply and task instructions actively
hurt. Verified on the authors' own UVRB MSRVTT split (1,000 JSFusion
pairs, 8 uniform frames, 200 tokens/frame): R@1 improves from 0.312
to 0.440 vs the paper's reported 0.431, and instruction sensitivity
collapses to noise (all placements 0.424-0.440).
lint: explicit strict=True in encode batch zip (a233646)
[MOEB] model: add ViCLIP video-language model (L-14, B-16) (#5013)
model: add ViCLIP video-language model (L-14, B-16)
Adds ViCLIPWrapper and ModelMeta for OpenGVLab/ViCLIP-L-14-hf and
OpenGVLab/ViCLIP-B-16-hf. Model from InternVid (ICLR 2024,
arXiv:2307.06942). Covers T2V and V2T retrieval tasks in MOEB.
Closes #5012, part of #4842.
style: fix ruff formatting in viclip_models.py
fix lint
fix: add mean/std source comment, remove defensive tensor checks
Co-authored-by: Rakshitha Ireddi <rakshithaireddi@Rakshithas-MacBook-Pro.local>
Co-authored-by: Roman Solomatin <36135455+Samoed@users.noreply.github.com> (0045490)
task: add ABOI2VRetrieval image-to-video retrieval (i2v) (#5284)
Add ABOI2VRetrieval: natively cross-modal image-to-video retrieval
Amazon Berkeley Objects ships, for the same product, a 360-degree turntable
"spin" photographed in a rig and separate catalog photographs shot at a
different time under different lighting. A query is therefore never a frame of
its own positive video and no crop, re-encode or temporal-neighbour
relationship links the two, so frame leakage is structurally impossible rather
than filtered out after the fact.
corpus 2857 videos, one per product, encoded from that product's spin by
selecting azimuth % 3 == 0. That single rule yields exactly 24 frames at
15-degree steps for all 8209 ABO spins (the 8116 dense ones store
azimuths 0..71, the 93 sparse ones store exactly {0,3,...,69}), so the
corpus is homogeneous without dropping any sequence. h264 / 384px long
side / 12 fps / 2.0 s, yuv420p limited-range bt709. Every output was
ffprobe-validated to decode to 24 frames.
queries 2857 catalog photographs, one per product, from the "context" bucket:
the product shot in a room or scene. Images sharing a perceptual hash
with another product (boilerplate, size charts), failing a zero-shot
category gate (swatches, macro crops, dimension diagrams), within a
pHash radius of the product's own spin, or shot on a white studio sweep
are all excluded.
qrels 1:1, score 1. One product per spin sequence, so no two queries share a
relevant document.
Queries are selected by a semantic category rather than by distance from the
answer, so every query stays an answerable depiction of the item; difficulty
comes from corpus size and the domain gap between a styled room photo and a
turntable render. Product types are restricted to five volumetric home-goods
categories; flat goods (RUG, WALL_ART) are excluded because a turntable
rotation of a flat object is close to degenerate.
ABO is CC BY 4.0, which permits redistribution of adapted material with
attribution. Note the bucket still carries a stale LICENSE-CC-BY-NC-4.0.txt
from 2021 and the AWS open-data registry entry was never updated after the
2023 relicense; the current README, all four subdirectory READMEs and the
project page all state CC BY 4.0.
i2v only: "v2i" is not in the TaskCategory literal.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The initial push produced only the auto-generated card, which has no license
or attribution. CC BY 4.0 Section 3(a) requires the creator credit, license
notice, warranty disclaimer, source link and an indication that the material
was modified to be present where the material is shared, so the card now
carries all of it plus the note about the bucket's stale NC license file.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> (429b5a6)
Add Qwen3 voice models (#5240)
Add Qwen3 voice models
Co-Authored-By: nehal309 <20415842+nehal309@users.noreply.github.com>
Co-authored-by: nehal309 <20415842+nehal309@users.noreply.github.com> (33345dc)
dataset: add MovingFashion bidirectional image-video retrieval (#5272)
dataset: add MovingFashion video-to-image retrieval
dataset: add reverse MovingFashion retrieval (d1fb15f)
Add judicialmind/greenleaf-law-embed-tiny (#5274)
Add judicialmind/greenleaf-law-embed-tiny model implementation
Add official mteb mock-run results (28/28 passed)
Update revision hash to match scrubbed model repo (bff06c94)
Remove custom GreenLeafEmbedWrapper, use SentenceTransformerEncoderWrapper directly with trust_remote_code kwarg
Remove mteb_mock_run_results.md
Add public_training_data link to judicialmind/legal-training-dataset
Update mteb/models/model_implementations/greenleaf_models.py
Co-authored-by: Surya-saka <surya@judicialmind.ai>
Co-authored-by: Surya-saka <sakasurya@users.noreply.huggingface.co>
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com> (d1665ae)
task: Add Breakfast Clustering and Breakfast Pair Classification datasets (#5275)
[MOEB] Add Breakfast Clustering and Breakfast Pair Classification datasets
remove var _BIBTEX
add descriptive stats
move data to mteb (2c1ceab)
dataset: Add EMID A2I + I2A retrieval datasets (#5279)
[MOEB] Add EMID pair classification and EMID a2i + i2a retreival datasets
add stats
fix types
reduce size and dedup
remove pc files
move data to mteb org (43ed975)
Add Webvid covr dataset (#5216)
Add WebVid dataset and retrieval task
WebVid provides Video1 + Edit = Video2 examples. Rather than using Video1, they have proposed to use middle frame of the video as Image, which makes the task Image + Test -> Video retrieval task.
Co-authored-by: Nehal Kathrotia <nehalkathrotia@gmail.com>
Exclude corrupted video
Add descriptive stats
fix revision and arxiv url
fix reference and main_score
Add WebVid dataset and retrieval task
WebVid provides Video1 + Edit = Video2 examples. Rather than using Video1, they have proposed to use middle frame of the video as Image, which makes the task Image + Test -> Video retrieval task.
Co-authored-by: Nehal Kathrotia <nehalkathrotia@gmail.com>
Exclude corrupted video
Add descriptive stats
fix revision and arxiv url
fix reference and main_score
add WebVidCoVRIT2VRetrieval to known duplicate exception
Co-authored-by: Nehal Kathrotia <nehalkathrotia@gmail.com>
Co-authored-by: Deep Shah <shahdeep@google.com> (aeb24c4)
task: add EVVE event video retrieval (v2v) (#5197)
dataset: add EVVE retrieval
dataset: align EVVE with review conventions
dataset: finalize EVVE review alignment
fix: download EVVE construction metadata
fix: pin updated EVVE dataset revision (0c0f6e5)
Remove AfriMTEB task dataset_transform() (#4905)
Fix AfriMTEB task schema issues
Remove unrelated benchmark changes
Remove unnecessary AfriXNLI dataset transform
Remove unrelated benchmark changes
Remove commented out code in AfriHate and KinNews classification
Remove dataset_transform from AfriHate and KinNews classification
Co-authored-by: Isaac Chung <chungisaac1217@gmail.com> (0fb6484)
dataset: Add multilingual MMarco retrieval task (#5068)
feat: add multilingual MMarco retrieval task
added revision MMarcoRetrievalMultilingual
Trigger CI after updating PR description
added MMarcoRetrievalMultilingual and stats
added MMarcoRetrievalMultilingual to mteb/tasks/retrieval/multilingual/init.py
changed task description
Marked MMarcoRetrieval as superseded by the new MMarcoRetrievalMultilingual task
MMarcoRetrievalMultilingual task contributed_by=None
removed contributed_by and is_beta
Add MMarcoRetrievalMultilingual to KNOWN_ISSUES in test_task_quality
Add MMarcoRetrievalMultilingual to KNOWN_ISSUES short_text
Co-authored-by: Isaac Chung <chungisaac1217@gmail.com> (ad874e1)
dataset: Add SciRepEval classification tasks (DRSM, biomimicry, FoS, MeSH) (#5026)
Add SciRepEval DRSM classification task
Add SciRepEvalDRSMClassification, a 5-way single-label text
classification task from the SciRepEval benchmark (Singh et al., EMNLP
2023). Given a biomedical paper's title and abstract, the task predicts
its Disease Research State Model category.
The allenai/scirepeval drsm config exposes only a single evaluation
split, so dataset_transform builds a stratified train/test split for the
classification evaluator and subsamples the test split to 2048 examples.
Includes TaskMetadata (revision-pinned), registration in the eng
classification init, and precomputed descriptive statistics.
Part of #591.
Extend the SciRepEval coverage beyond DRSM with three more
classification tasks from the allenai/scirepeval benchmark:
For FoS and MeSH only the "evaluation" split is loaded (the upstream
train splits are hundreds of thousands to millions of rows). Each task
ships pinned metadata, registration, and precomputed descriptive
statistics.
Reference runs (accuracy, CPU): random-encoder vs multilingual-e5-small
Part of #591.
Set license to odc-by for the four SciRepEval tasks (aggregate benchmark
is released under ODC-BY per the SciRepEval repo). Set biomimicry
annotations_creators to human-annotated (labels are manually annotated
gold tags from the PeTaL database). Rewrite the four task descriptions in
a what-it-tests / how / attributes style.
Switch single-label SciRepEval tasks to cross-validation
Keep evaluation split name instead of renaming to train
Per Samoed's review comment, the cross-validation SciRepEval tasks
(DRSM, biomimicry, MeSH descriptors) no longer rename their single
split to train. eval_splits and train_split are both set to
evaluation, and the descriptive stats files are updated to match.
Co-authored-by: Claude <noreply@anthropic.com> (14b295c)
benchmark: Add Slovak tasks and SK-MTEB benchmark (MTEB(slk, v1)) (#4788)
Add Slovak tasks and SK-MTEB benchmark (MTEB(slk, v1))
Adds new Slovak-language tasks across 7 task types and registers the full SK-MTEB benchmark. Tasks cover retrieval, STS, pair classification, classification, reranking, clustering, and bitext mining for Slovak.
fix: Update Slovak citations and improve dataset references
Update mteb/benchmarks/benchmarks/benchmarks.py
Co-authored-by: Kenneth Enevoldsen <kenevoldsen@pm.me>
Co-authored-by: Kenneth Enevoldsen <kenevoldsen@pm.me>
fix: Update citation for SkMTEB benchmark to reflect new publication details
refactor: Simplify Multi-EuP Slovak classification tasks by removing custom dataset loader and mixin
fix: Update SkMTEB tasks after resolving conflicts and update reference with arXiv link
Fix SlovakRTE label polarity
fix: Update dataset_transform signature
Enable trust_remote_code for gte-multilingual-base
Alibaba-NLP/gte-multilingual-base ships its architecture as custom
code in the HF repo (Alibaba-NLP/new-impl); without this it fails to
load. Mirrors the existing pattern in arctic_models.py.
fix: Improve dataset descriptions for Slovak NLI, RTE, SkQuadReranking, SMESumRetrieval, and STS
fix: Update method signatures across Slovak tasks to standardize dataset_transform and load_data definitions
refactor: Point SlovakSumURLClustering, SlovakSTS, SMESumRetrieval at pre-built mteb datasets
fix: Exempt SlovakPharmacyDrMaxReranking and OpusSlovakEnglishBitextMining from dataset quality checks
Both fail new duplicate/short-text checks due to genuine, negligible source-data
characteristics rather than bugs: OPUS-100 naturally repeats short common
subtitle/legal-document phrases (24/2000 test pairs), and DrMax's real search-query
log contains a couple of 1-character queries (2/4676).
Co-authored-by: Kenneth Enevoldsen <kenevoldsen@pm.me> (c8a519a)
model: add Qwen3-VL-Reranker (2B, 8B) (#5057)
model: add Qwen3-VL-Reranker (2B, 8B)
fill languages from model card (33 languages)
trigger CI re-run
reranker: add use_instructions flag instead of mutating kwargs
Co-authored-by: Hubert Lu <hubielu@email.com> (233478f)
fix: repair Memotion retrieval dataset
model: Add LanguageBind video and audio model wrapper (#4557)
model: add LanguageBind video, audio, and image wrappers with compatibility patches
feat: HF mirror auto-download for LanguageBind + extras group for reproducibility
wip: phase 1 review fixes for GPU testing
refactor: trim LanguageBind compat shims to the two that are required
Tested each of the five compatibility patches individually on GPU against the
pinned transformers 4.45.2 / torchvision 0.20.1 / torchaudio 2.5.1:
Added inline comments documenting why the two retained shims are required.
Move the _patch_compatibility() call into _ensure_languagebind_source() so the
shims only run when a LanguageBind model is actually loaded, rather than
globally at module import time.
Instead of monkey-patching transformers/torchvision/torchaudio at runtime from
the MTEB wrapper, the compatibility fixes now live in the vendored LanguageBind
source itself (pinned HF revision d2d0f6f):
The wrapper now only downloads and imports the pinned source revision, with no
runtime patching of external libraries.
Verified clean import under torch 2.11 / torchvision 0.26 / transformers 4.45.2.
Add lower bounds to the languagebind extra deps for reproducibility:
einops, decord, opencv-python-headless, pytorchvideo, peft.
refactor to use python package and added temp tests
fix: functional_tensor shim for modern torchvision + pin transformers<4.46 + use get_model in tests
Delete tests/test_models/test_language_bind.py
fix languagebind video transform and add missing audio/video deps
restore PLW0717 ruff ignore rule
Co-authored-by: Kenneth Enevoldsen <kennethcenevoldsen@gmail.com>
Co-authored-by: Isaac Chung <chungisaac1217@gmail.com> (aa46f13)
add missing qrel ID statistics (#5226)
feat: add missing qrel ID statistics
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Use exact num_missing_query_ids / num_missing_corpus_ids statistics
instead of inferring dangling qrels from
unique_relevant_docs > num_documents. Keep the old bound as a fallback
for statistics generated before these fields existed.
Regenerate OVENIT2ITRetrieval statistics and allowlist the four tasks
flagged by the new check. See PR description for details.
Co-authored-by: Xu Liu <lxer@Xus-MacBook-Pro.local>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> (d5a4595)
dataset: add IntPhys2 video zero-shot task (#5225)
dataset: add IntPhys2 video zero-shot task
Move IntPhys2 constants into task class
dataset: use MTEB IntPhys2 artifact
dataset: add IntPhys2 conversion script (436d67f)
dataset: add Beehive States audio classification task (#5262)
dataset: add Beehive States audio classification task
fix: remove unnecessary writer batch setting
Update mteb/abstasks/classification.py
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com> (768fac5)
Claude-Session: https://claude.ai/code/session_01Vue9R33gcLojTcxp8ZDkW3
Co-authored-by: Claude Fable 5 <noreply@anthropic.com> (4e4b5b1)
add giga models (#5252)
add giga models
Update mteb/models/model_implementations/giga_models.py
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>
added instructions, changes model loader to InstructSentenceTransformerModel
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>
Co-authored-by: Roman Solomatin <36135455+Samoed@users.noreply.github.com> (4ca789c)
Add model: OpenGVLab/InternVideo2-CLIP-1B-224p-f8 (#5078)
feat: add InternVideo2-CLIP-1B-224p-f8 (WIP)
chore: pin InternVideo2 CLIP revision
fix: disable LoRA wrapper so InternVL-C text weights load
feat: add InternVideo2-CLIP-1B-224p-f8
refactor: drop dead peft remap, fix install hint, add n_embedding_parameters
refactor: drop dead peft remap, fix install hint, add n_embedding_parameters
chore: declare internvideo2 as a conflicting extra, regenerate uv.lock
refactor: load InternVideo2-CLIP via SentenceTransformers
style: fix InternVideo2 lint
Update mteb/models/model_implementations/internvideo2_models.py
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>
Co-authored-by: Hubert Lu <hubielu@email.com>
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com> (2902127)
make sparse models on all tasks
feat: Add sparse encoder (#5087)
add sparse encoder
make sparse models on all tasks
improve tests
improve max sim
fix typing
simplify comments
Update mteb/models/sentence_transformer_wrapper.py
Co-authored-by: Kenneth Enevoldsen <kennethcenevoldsen@gmail.com>
Co-authored-by: Kenneth Enevoldsen <kennethcenevoldsen@gmail.com>
Co-authored-by: Copilot Autofix powered by AI <175728472+Copilot@users.noreply.github.com>
Co-authored-by: Kenneth Enevoldsen <kennethcenevoldsen@gmail.com>
Co-authored-by: Copilot Autofix powered by AI <175728472+Copilot@users.noreply.github.com> (44f4bed)
dataset: add MMEdit audio editing retrieval (#5246) (dc9ec0f)
Add UEmbed implementation. (#5200)
Add UEmbed implementation.
Signed-off-by: sighingsnow <songtingyu220@gmail.com>
Signed-off-by: sighingsnow <songtingyu220@gmail.com>
Signed-off-by: sighingsnow <songtingyu220@gmail.com> (82edbff)
fix: add GME dependencies and support image-only batches
dataset: add Mars-VL-Pairs bidirectional retrieval (#5148)
dataset: add Mars-VL-Pairs retrieval
fix: address Mars-VL-Pairs review feedback
fix: require TLS verification for Mars recovery
fix: use standard metrics for full-gallery MRR
fix top k calculation
fix test
remove k values from mock
fix index
remove files
Co-authored-by: Roman Solomatin <36135455+Samoed@users.noreply.github.com> (14b5040)
fix: skip R_cap division when a query has no relevant docs
fix: skip R_cap division when a query has no relevant docs (#5253)
fix: skip R_cap division when a query has no relevant docs
test: cover recall_cap with no relevant documents (ea20326)
add total parameters to performance size plot (#5245)
feat: add total parameters to performance size plot
remove test
Co-authored-by: Roman Solomatin <36135455+Samoed@users.noreply.github.com> (fe08f24)
model: add OmniRetriever-7B (#5164)
model: add OmniRetriever-7B
docs: add OmniRetriever mock-run results
fix: address OmniRetriever review feedback
Delete mteb_mock_run_results.md
format
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>
Co-authored-by: Roman Solomatin <36135455+Samoed@users.noreply.github.com> (c6801b8)
Addresses #5230 (in code, we still need the results)
intfloat/e5-mistral-7b-instruct uses last-token pooling — the embedding is the hidden
state of the EOS </s>. Under transformers>=5 that </s> is no longer added, so we
pool the last content token instead and the embeddings are wrong.
The model repo contradicts itself:
tokenizer_config.json sets "add_eos_token": truetokenizer.json's post-processor adds only <s>transformers<5 resolved that in favour of the config and rewrote the post-processor at load
time; transformers>=5 takes tokenizer.json as-is. Nothing warns.
transformers 4.57.6 post_processor: <s> + A + </s>
transformers 5.15.0 post_processor: <s> + A
Passing add_eos_token=True on its own is not a complete fix: on transformers>=5 it
rebuilds the post-processor from add_bos_token, which defaults to False for this repo, so
you get the </s> back but lose the <s>. Both flags are needed.
bBSARDNLRetrieval (nDCG@10), intfloat/e5-mistral-7b-instruct, batch size 8, same GPU.
The before / add_eos_token only / after rows are mteb 2.19.3 with the two tokenizer
kwargs applied through a loader_kwargs override, which is exactly what this patch writes
into the file:
| mteb | transformers | tokenizer | score |
|---|---|---|---|
| 2.18.1 | 4.57.6 | <s> … </s> |
0.37365 |
| 2.18.4 | 4.57.6 | <s> … </s> |
0.37365 |
| 2.18.7 | 4.57.6 | <s> … </s> |
0.37365 |
| 2.19.3 | 4.57.6 | <s> … </s> |
0.37365 |
| before | 5.15.0 | <s> … |
0.06504 |
add_eos_token only |
5.15.0 | … </s> |
0.34506 |
| after | 5.15.0 | <s> … </s> |
0.37360 |
For reference the bBSARD paper reports 0.3770 for this model.
@tomaarsen I will just make you aware of this. I am not sure if it is something that we want to address upstream (if nothing else good to know that it exists).
Re the GritLM → Sentence Transformers migration (#4085) — the GritLM loader scores 0.05508
on main today, it truncates at max_length=512 and it prefixes documents with "Instruct: \nQuery: ". Fixing these give similar scores.
I scanned to see if there were other cases of this, mostly it found cases of misspecified dependencies or missing. I have fixed those.
The only other case was Salesforce/SFR-Embedding-Code-2B_R, but that doesn't work with never versions of ST and it is unclear
if BeastyZ/e5-R-mistral-7b is intended to run with mean pooling (ships no modules.json or 1_Pooling, so ST falls back to mean pooling)
Co-authored-by: Kenneth <kennethenevoldsen@gmail.com> (a20055d)
model: add Gemini Embedding 2 (#5220)
model: add Gemini Embedding 2
truncate audio inputs to duration limit
use query and body fields for text-only retrieval
remove tests
Co-authored-by: Roman Solomatin <36135455+Samoed@users.noreply.github.com> (de40885)
fix: break p-MRR rank ties by doc id to match the other retrieval metrics
fix: Enable B905 and B028 Ruff rules
fix: Enable B905 and B028 Ruff rules (#5205)
Enable B905 and B028 Ruff rules
made strict=False for BUCC (510fff6)
model: add Amazon Nova Multimodal Embeddings (text, image, audio, video) (#5206)
model: add Amazon Nova Multimodal Embeddings (text, image)
Bedrock synchronous InvokeModel wrapper for
amazon.nova-2-multimodal-embeddings-v1:0.
Uses embeddingPurpose for asymmetric retrieval (GENERIC_RETRIEVAL for
queries, GENERIC_INDEX for documents) and maps Classification and
Clustering tasks to their dedicated purposes.
Audio and video are deferred to a follow-up: mteb decodes both before
they reach the encoder, so sending them would require re-encoding
decoded media back into a container.
Closes #5204
style: satisfy ruff format and no-self-use
refactor: move Nova wrapper into bedrock_models and share helpers
Per review, NovaMultimodalEmbeddingsModel now lives alongside BedrockModel
instead of in its own file.
Extracts three pieces both classes use:
Nova stays a separate class rather than a third provider branch on
BedrockModel: BedrockModel.encode is text-only by construction, and Nova's
request body and response shape differ from Titan's.
Addresses review:
Audio arrives from mteb as a float array plus sampling rate, so it can be
encoded to WAV losslessly via the stdlib wave module. Verified against
Bedrock: text, image and audio all return embeddings at the requested dim.
Video remains unsupported: mteb decodes video to a frame tensor with no
frame rate and no audio track before the encoder sees it, so reconstructing
a container would require choosing an arbitrary fps.
Video arrives as a torchcodec VideoDecoder, so frames are re-encoded to mp4
at the source frame rate (metadata.average_fps) and sent inline as base64.
Pass-through of the source container is not an option: Nova validates the
container against the declared format and MSVD ships AVI, which is not in
the accepted enum.
Uses AUDIO_VIDEO_COMBINED, which returns a single vector per item.
AUDIO_VIDEO_SEPARATE would return one vector per stream and break the
one-embedding-per-item contract.
Two guards on the decode: num_frames from the container header can overshoot
what actually decodes, so trailing frames are dropped until the read
succeeds (same approach as FramesCollator), and frames are cropped to even
dimensions since h264 requires them. Segments are capped at 30s, Nova's
limit.
MSVDT2VRetrieval nDCG@10 = 0.846 over 660 videos.
Co-authored-by: Hubert Lu <hubielu@email.com> (cd6cd8f)
model: add VideoMAE video encoders (#5094)
model: add VideoMAE and TimeSformer video encoders
model: add VideoMAE video encoders
Split TimeSformer into a separate file and PR per review.
Load the checkpoint state dict directly instead of via a lazy closure.
model: pin VideoMAE to transformers v4
model: return the encoder output for VideoMAE
Drop fc_norm per review; it belongs to the classification head. That also
removes _load_checkpoint_tensors, unused now that the transformers v4 pin
handles the attention biases.
fix: pass frames as HWC lists for the transformers v4 image processor
Update mteb/models/model_implementations/videomae_models.py
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>
fix: restore logger import
model: use AutoVideoProcessor and drop the transformers v4 pin
Restore q_bias/v_bias in the wrapper instead, which v5 drops on load.
Same HMDB51Clustering score as the pinned path and roughly twice as fast.
AutoVideoProcessor only resolves videomae from transformers 5.0.0 onward.
Co-authored-by: Hubert Lu <hubielu@email.com>
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com> (32072a9)
[MOEB research]fix: strip empty caption segments in Clotho retrieval tasks (#5062)
fix: strip empty caption segments in Clotho retrieval tasks
ClothoT2ARetrieval and ClothoA2TRetrieval build queries/corpus by
splitting captions on ".", but never stripped or filtered the
resulting segments. Captions ending in a period produced a trailing
empty-string segment that became a real query/document with its own
qrel entry.
Add .strip() + skip-if-empty to both load_data() implementations and
regenerate the committed descriptive_stats JSON to reflect the cleaned
data (num_queries/num_documents: 5585 -> 4680 on the affected side,
905 empty segments removed, zero non-empty duplicates).
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Per review feedback: existing model results were already scored under
ClothoT2ARetrieval/ClothoA2TRetrieval as originally defined, so we
should not silently mutate those tasks in place. Instead, following
MTEB's established versioning convention (see STS22 -> STS22.v2,
PoemSentimentClassification.v2), this reverts the in-place fix and
adds ClothoT2ARetrievalV2/ClothoA2TRetrievalV2 (name suffix ".v2")
with the corrected load_data().
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
This reverts commit 940ed4b.
Per review feedback: existing model results were already scored under
ClothoT2ARetrieval/ClothoA2TRetrieval as originally defined, so we
should not silently mutate those tasks in place. Instead, following
MTEB's established versioning convention (see STS22 -> STS22.v2,
PoemSentimentClassification.v2, DBPediaHardNegatives.v2), this adds
ClothoT2ARetrievalV2/ClothoA2TRetrievalV2 (name suffix ".v2") with
the corrected load_data(), leaving the original tasks untouched.
Not included in this commit: updating mteb/benchmarks/benchmarks.py
to point MAEB(beta) at the .v2 task (left as a separate decision).
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
docs: shorten ClothoT2ARetrieval.v2 description per review suggestion
Co-authored-by: Kenneth Enevoldsen <kenevoldsen@pm.me>
Mirrors the ClothoT2ARetrieval.v2 description shortened per
KennethEnevoldsen's review suggestion (applied directly on GitHub in
4f7e166). Also fixes the pre-existing "datasetst" typo in both v2
descriptions.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
ClothoT2ARetrieval.v2 and ClothoA2TRetrieval.v2 now load from
lxercode/clotho_t2a_v2 and lxercode/clotho_a2t_v2 (pinned revisions)
instead of running a custom load_data() against mteb/Clotho. These
repos contain the same empty-query-fixed data the custom loaders
already produced, materialized and uploaded ahead of time, so the
tasks no longer need a load_data() override.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Co-authored-by: Xu Liu <lxer@Xus-MacBook-Pro.local>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
Co-authored-by: Kenneth Enevoldsen <kenevoldsen@pm.me> (dc1ac0a)
model: add UNITE (friedrichor/Unite-Base-Qwen2-VL-2B) (#5196)
model: add UNITE (friedrichor/Unite-Base-Qwen2-VL-2B)
model: fix UNITE metadata (memory, revision, training data)
model: lint fixes for UNITE wrapper
model: load UNITE via base Qwen2VL class, subclassing breaks key remapping on transformers 5.x
model: cap UNITE video frames at 360*420 to match reference inference
model: fix UNITE video path, tensor-aware frame resize and fps sampling
style: ruff format unite_models.py
refactor: drop unused num_frames param, fps mode only
Update mteb/models/model_implementations/unite_models.py
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>
model: move video cap into processor, batch encode
Update mteb/models/model_implementations/unite_models.py
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>
docs: note why UNITE loads via the base class
model: use UNITE subclass with transformers v4
model: use 32 frames for UNITE video evaluation
Update mteb/models/model_implementations/unite_models.py
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>
Co-authored-by: Hubert Lu <hubielu@email.com>
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com> (a7c7086)
dataset: add IncompeBench (#5194)
dataset: add IncompeBench
add IncompeBenchLenientRetrieval and rename the strict task (fa7ce6b)
[MOEB] Add VELOCITI video pair classification task (#4980) (#4982)
[MOEB] Add VELOCITI video pair classification task (#4980)
fix: annotate mutable class attributes with ClassVar for lint
fix: rebuild VELOCITI-PC to avoid video duplication, rewrite description
Addresses review on PR #4982:
Addresses review on PR #4982: instead of downloading a raw videos.zip
and manually extracting it, videos now live in a second HF dataset
config (video_id, video) with 864 unique rows using a proper Video()
feature, joined against the (video_id, text, label) table at load time
via concatenate_datasets(axis=1) -- an Arrow-level column join, no
decode/re-encode round trip. Removes zipfile/tempfile handling
entirely; simplifies both the wrapper code and standalone usage of the
dataset outside mteb, per Samoed's review comment.
CI's test_dataset_quality caught 3,469 duplicated (video_id, text) pairs
in the reshaped dataset -- VELOCITI's source data reuses captions across
different test categories/events for the same video. Deduplicated
17,584 -> 11,669 rows, and dropped the single pair that had conflicting
labels across source rows (unsafe to keep either). Labels are no longer
perfectly balanced (7,463 label=0 / 4,206 label=1) as a result -- a real
property of the deduplicated data, not a bug.
CI's persistent HF cache (workflow key 'Linux-hf', never invalidated
across commits) was silently serving a pre-dedup Arrow snapshot of the
small (video_id, text, label) table with zero network calls, producing
stale duplicate-pair failures on the correct, verified-clean data.
Confirmed via job log: 'Cache restored from key: Linux-hf' followed by
no download activity for VELOCITI-PC at all. force_redownload makes
this cheap 401KB table immune to that class of staleness going forward.
The previous force_redownload fix addressed stale-cache reuse but not
the real bug: CI runs the suite with pytest-xdist (-n auto), and
datasets' Arrow-cache build for this table isn't safe when a concurrent
test worker touches the same shared HF cache dir at the same time --
that's a race, not staleness, so force_redownload could never have
fixed it (confirmed: it didn't -- CI still failed identically on that
commit).
Switched to hf_hub_download + pandas for this small (401KB) table
instead of load_dataset, since a single atomic revision-pinned file
fetch has no cache-build step for another worker to race. Verified
correct under 5 concurrent processes hitting the same cache
directory simultaneously, in addition to a normal single-process load.
test_dataset_quality reads task.metadata.descriptive_stats, a committed
JSON file under mteb/descriptive_stats/ -- it never calls load_data().
That file was generated once, before any of the zip->parquet rebuild,
dedup, or cache fixes, and was never regenerated afterward, so it kept
reporting the original 17,584-row/11,670-unique-pair pre-dedup numbers
regardless of what the dataset actually contained. Every prior fix
(force_redownload, then hf_hub_download) was correct for load_data()
but irrelevant to this specific failure.
Regenerated via task.calculate_descriptive_statistics(overwrite_results=True)
against the live, correct dataset: num_samples/unique_pairs now 11669/11669
(zero duplicates), unique_videos still 864 as expected.
The hf_hub_download switch was based on an unconfirmed race-condition
theory (CI runs pytest -n auto, multiple workers sharing one HF cache
dir). It was never actually proven necessary: the test that was failing
(test_dataset_quality) doesn't call load_data() at all, it reads a
committed descriptive-stats JSON, which was the real fix. Reverting to
plain load_dataset() per review -- simpler, and there's no evidence the
extra complexity was ever solving anything.
Verified end-to-end on a fully cleared cache: 11,669 rows, zero
duplicate pairs, correct columns. (b7f1504)
fix: add missing boto3 extras group
The four bedrock/* models declare extra_requirements_groups=["boto3"],
but no boto3 group exists in [project.optional-dependencies].
_validate_extras_groups raises ValueError, so mteb.get_model() fails for
all four Bedrock models.
Co-authored-by: Hubert Lu <hubielu@email.com> (0f2d655)
model: add VideoPrism video encoders (base, large, LVT base, LVT large) (#5193)
model: add VideoPrism video encoders (base, large, LVT base, LVT large)
docs: reference the transformers VideoPrism page and explain the pinned revisions
model: add fps/max_frames args and simplify the frame cast for VideoPrism
Update mteb/models/model_implementations/videoprism_models.py
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>
Co-authored-by: Hubert Lu <hubielu@email.com>
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com> (072909a)
fix: set explicit query prefix for ColPali wrappers
fix: Add support for more metrics in ZeroShot classification
fix: Add support for more metrics in ZeroShot classification (#5179)
Add more zeroshot metrics
remove
add test for metrics
fix typing
fix tests
don't convert to tensor/array explicitly (f414a06)
The task's bibtex_citation listed only the upstream MorisienMT data paper, so
users of the task had no citation path to the work that built the MTEB task
itself. Adds the Zenodo record for morisien-embed alongside it; the data
citation is unchanged. (35cca9a)
feat: add OpenAIEndpointWrapper for HTTP-based vLLM servers
feat: add OpenAIEndpointWrapper for HTTP-based vLLM servers (#4834)
feat: add VllmEndpointWrapper for HTTP-based vLLM servers
Add a new wrapper to enable MTEB benchmarking against remote vLLM
servers via HTTP API, complementing the existing VllmEncoderWrapper
which requires local in-process instantiation.
The existing VllmEncoderWrapper only supports local vLLM instantiation,
which has several limitations:
New VllmEndpointWrapper class that:
Communicates with vLLM servers via OpenAI-compatible /v1/embeddings API
Works with any vLLM backend (CPU or GPU)
Supports remote endpoints, authentication, and SSL configuration
Auto-detects max_length from model metadata
Includes retry logic and response validation
Enables server reuse across multiple benchmark runs
Benchmarking production vLLM deployments
Testing remote vLLM CPU servers
Evaluating vLLM endpoints without local instantiation overhead
CI/CD integration with existing vLLM infrastructure
mteb/models/vllm_endpoint_wrapper.py: New HTTP-based wrapper
mteb/models/init.py: Export VllmEndpointWrapper
tests/test_models/test_vllm_endpoint_wrapper.py: Basic tests
docs/get_started/advanced_usage/vllm_wrapper.md: Documentation
Validated against production vLLM CPU servers running:
Co-Authored-By: Claude Sonnet 4.5 <noreply@anthropic.com>
Signed-off-by: Maryam Tahhan <mtahhan@redhat.com>
Co-Authored-By: Claude Sonnet 4.5 <noreply@anthropic.com>
Signed-off-by: Maryam Tahhan <mtahhan@redhat.com>
Changes per maintainer review:
All 8 tests pass. Tested successfully with live vLLM server running ibm-granite/granite-embedding-english-r2.
Co-Authored-By: Claude Sonnet 4.5 <noreply@anthropic.com>
Signed-off-by: Maryam Tahhan <mtahhan@redhat.com>
Refactored OpenAI-compatible API wrappers to eliminate code duplication and
add reranking support:
OpenAIBaseWrapper (shared HTTP logic)
├── OpenAIAPIWrapper (embeddings via /v1/embeddings)
└── OpenAIRerankWrapper (reranking via /v1/rerank)
Addresses reviewer feedback for reranking API support while maintaining clean,
maintainable architecture.
Co-Authored-By: Claude Sonnet 4.5 <noreply@anthropic.com>
Signed-off-by: Maryam Tahhan <mtahhan@redhat.com>
Co-Authored-By: Claude Sonnet 4.5 <noreply@anthropic.com>
Signed-off-by: Maryam Tahhan <mtahhan@redhat.com>
Signed-off-by: Maryam Tahhan <mtahhan@redhat.com>
Changes based on review from @Samoed:
Co-Authored-By: Claude Sonnet 4.5 <noreply@anthropic.com>
Signed-off-by: Maryam Tahhan <mtahhan@redhat.com>
Co-Authored-By: Claude Sonnet 4.5 <noreply@anthropic.com>
Signed-off-by: Maryam Tahhan <mtahhan@redhat.com>
Co-Authored-By: Claude Sonnet 4.5 <noreply@anthropic.com>
Signed-off-by: Maryam Tahhan <mtahhan@redhat.com>
Co-Authored-By: Claude Sonnet 4.5 <noreply@anthropic.com>
Signed-off-by: Maryam Tahhan <mtahhan@redhat.com>
This makes it clear that OpenAIAPIEncodeWrapper is for encoding (embeddings)
and adds "API" to OpenAIAPIRerankWrapper for consistency. Both classes wrap
OpenAI-compatible HTTP APIs.
Updated all references in:
Co-Authored-By: Claude Sonnet 4.5 <noreply@anthropic.com>
Signed-off-by: Maryam Tahhan <mtahhan@redhat.com>
Address PR review feedback:
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Signed-off-by: Maryam Tahhan <mtahhan@redhat.com>
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Signed-off-by: Maryam Tahhan <mtahhan@redhat.com>
add whats new
add multimodality support
fix typing
run tasks
Signed-off-by: Maryam Tahhan <mtahhan@redhat.com>
Co-authored-by: Claude Sonnet 4.5 <noreply@anthropic.com>
Co-authored-by: Kenneth Enevoldsen <kennethcenevoldsen@gmail.com>
Co-authored-by: Roman Solomatin <36135455+Samoed@users.noreply.github.com> (dc7176b)
data: add FGMCaps (#5191)
data: add FGMCaps
lint format (2270401)
Remove dead dataloader code (#5192)
remove dead dataloader code (3347cd1)
Add SSW60 dataset (#5154)
Add SSW60 first commit
Modify the license
Add descriptive stats
Remove not required Exception handling
ran lint
pin revision
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com> (6e38137)
ci: Revert copilot dataset pr review
fix: break MRR ties by doc id to match pytrec_eval
fix: break MRR ties by doc id to match pytrec_eval (#5136)
fix: break MRR ties by doc id to match pytrec_eval
test: cover MRR tie-break order-independence and pytrec_eval parity (2922866)
fix: leaderboard partial split aggregation
fix: leaderboard partial split aggregation (#5102)
null per-task scores
extend the split-completeness guard to Mean (Subset)
Fix score computation
simplify comments
check every aggregation in the Mean partial-coverage test
add skip mark
add skip mark
add skip mark (56a33ab)
…can load it, instead of relying on the deprecated Model2VecModel. All 186 tasks re-run against the new revision. (`ed285c7`)
Fix Copilot review comment command
Updated the comment command to use '@copilot' instead of '@github-copilot'. (04403ce)
ci: add Copilot dataset PR review via label trigger (#5163)
add Copilot review instructions for dataset PRs
ci: trigger Copilot review via label instead of applyTo
ci: clarify results table and add random encoder run command
ci: remove unverifiable dataset runs check, add gap and size checks (728e9cb)
fix: require datasets>=4.0.0 for video extra (#5173)
fix: require datasets>=4.0.0 for video extra
regenerate uv.lock (1d587bf)
fix: combine subsets in run_settings.jsonl (#5099)
fix combine subsets in run_settings.jsonl
copilot suggestion
matched implementation with existing run_settings
fixed typecheck (c6b7e1b)
Enable some ruff rules (#5108) (c4f9d80)
model: add PS3 (nvidia/PS3-1.5K/4K-SigLIP and SigLIP2) (#5077)
Co-authored-by: Hubert Lu <hubielu@email.com> (df925fd)
Add Flickr dataset I2A and A2I (#5137)
Add Flickr dataset I2A and A2I
Co-Authored-By: Deep Shah <21212684+deep9539@users.noreply.github.com>
refactor and directory update
format lint files
Co-authored-by: Deep Shah <21212684+deep9539@users.noreply.github.com> (0e07d11)
dataset: add CoVR-R (#5116)
dataset: add CoVR-R
reupload dataset and update descriptive stats (cce148b)
dataset: add REAL-MM-RAG (#5106)
dataset: add REAL-MM-RAG
test: allow duplicate images in REAL-MM-RAG
Update mteb/tasks/retrieval/eng/real_mm_rag_retrieval.py
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com> (4b1fbb0)
task: add Stanford I2V image-to-video+audio retrieval (i2va) (#5151)
dataset: add Stanford I2V retrieval
dataset: add Stanford I2V audio
dataset: align Stanford I2V license metadata (8e170e9)
model: add Dasheng audio encoders (base, 0.6B, 1.2B) (#5118)
wip: Dasheng audio wrapper
model: add Dasheng audio encoders (base, 0.6B, 1.2B)
deps: declare einops for Dasheng
revert unrelated uv.lock changes
meta: add citation and training datasets for Dasheng
Co-authored-by: Hubert Lu <hubielu@email.com> (cf631a3)
model: add Cosmos-Embed1 (nvidia/Cosmos-Embed1-224p/336p/448p) (#5133)
model: add Cosmos-Embed1 (nvidia/Cosmos-Embed1-224p/336p/448p)
model: add Cosmos-Embed1 (nvidia/Cosmos-Embed1-224p/336p/448p)
fix: GPU-path bugs in Cosmos-Embed1 (int device, dtype cast, mixed-resolution clips)
review: access projections directly instead of a getattr helper
Co-authored-by: Hubert Lu <hubielu@email.com> (7525053)
model: add NVIDIA RADIO family (RADIO-B/L/H) (#5061)
model: add NVIDIA RADIO family (RADIO-B, RADIO-L, RADIO-H)
fix: cite RADIOv2.5 paper alongside AM-RADIO
Apply suggestions from code review
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>
fix: lint after applying review suggestions
chore: refresh mock-run results
refactor: drop redundant modality guards, remove mock-run results
Update mteb/models/model_implementations/radio_models.py
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com> (d35c86e)
dataset: add SpokenCOCO (#5121) (f2dfbd3)
task: add CaReBench video retrieval (#5112)
task: add CaReBench video retrieval
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> (e040e04)
Add Audio flamingo 3 (#5080)
Add audio flamingo 3 model
Fix type errors
Fixing type error and minor refactoring.
Co-authored-by: Nehal Kathrotia <nehalkathrotia@gmail.com>
fix the cast error using chat template Solves expected scalar type Float but found BFloat16
Co-authored-by: Nehal Kathrotia <nehalkathrotia@gmail.com>.
Co-authored-by: Nehal Kathrotia <nehalkathrotia@gmail.com>
embeddings was bfloat16 (produced under torch.autocast(dtype=torch.bfloat16)), and NumPy has no bfloat16 dtype, so the final torch.cat(...).numpy() call raised TypeError: Got unsupported ScalarType BFloat16. Casting to float32 before moving off-GPU fixes it — only the small pooled embedding tensor is upcast, not the full model.
Co-authored-by: Nehal Kathrotia <nehalkathrotia@gmail.com>
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>
Minor nits
Change init func signature to capture torch_dtype and device_map
Add n embedding params
fix lint
fix lint
fix import order
Apply suggestions from code review
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>
Co-authored-by: Nehal Kathrotia <nehalkathrotia@gmail.com>
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com> (09e0c88)
dataset: add MMLongBench-Doc (#5138)
dataset: add MMLongBench-Doc
test: register MMLongBenchDocRetrieval duplicate images (7bd35b3)
model: add malteos/most-embed-de (German retrieval, Nemotron-3-Embed-1B fine-tune) (#5149)
Adds a ModelMeta for malteos/most-embed-de, a 1.1B German retrieval embedding model fine-tuned from nvidia/Nemotron-3-Embed-1B-BF16.
The base model is already implemented, so this reuses SentenceTransformerEncoderWrapper and the existing nemotron-3-embed extras group. The 4096 model_max_length cap mirrors the base model's entry so results stay comparable to the baseline.
Verified with mteb mock-run (28/28 text tasks passing).
Claude-Session: https://claude.ai/code/session_01LaEay2SrYBRYnojLT8mJi4 (5e36d33)
Add KiteFishAI/Nano-Em1-0.6B-v2.1 (#5150)
Add KiteFishAI/Nano-Em1-0.6B-v2.1
Add new training datasets to kite_fish_models.py
Refactor training datasets to use dictionary format (8df42c1)
dataset: add Greatest Hits audio-video material retrieval (a2v, v2a) (#5003)
task: add Greatest Hits audio-video material retrieval (a2v, v2a)
Impact-sound <-> video retrieval from Greatest Hits / Visually Indicated Sounds (CC-BY-4.0). Match an impact across modalities by material (17 materials). 992 impacts. LCO-Omni: a2v 15.2 / v2a 17.4 vs random ~10.5 nDCG@10 (map@10 3x random). Includes descriptive statistics.
The audio side was stored as file-path references (via Dataset.to_parquet
with an Audio() column), so it only decoded on the original build machine.
Rebuilt with the audio bytes embedded directly in the parquet (matching the
video side); refreshed to 1000 clips over 17 materials. Updated dataset
revisions and regenerated descriptive statistics. (e73b747)
Remove gradio lb tests and ci (#5103)
remove gradio lb tests and ci
remove duplicated import (409a355)
task: add ADVANCE audio-image retrieval (a2i, i2a) (#5119)
task: add ADVANCE audio-image retrieval (a2i, i2a)
Adds ADVANCEA2IRetrieval and ADVANCEI2ARetrieval, filling the audio<->image gap under MOEB Track 1 (#4842). 5,075 geotagged locations pairing FreeSound field recordings with co-located Google Earth aerial imagery across 13 land-cover classes, license CC-BY-4.0 (verified via the dataset's Zenodo record, the primary source repo). Reshaped from blanchon/ADVANCE into BEIR-style corpus/queries/qrels layout, pushed to yaswanth169/ADVANCE-A2I and yaswanth169/ADVANCE-I2A. Note: ~6.7GB at original resolution.
Closes #5006
Per Samoed's review on #5119 -- the script that reshapes blanchon/ADVANCE
into BEIR-style corpus/queries/qrels layout was only linked in a PR
comment; now committed under scripts/ as part of the diff itself. (a3feb86)
model: add ColQwen-Omni (vidore/colqwen-omni-v0.1) (#5115)
model: add ColQwen-Omni (vidore/colqwen-omni-v0.1)
fix: add n_embedding_parameters to ColQwen-Omni meta
fix: set n_embedding_parameters for ColQwen-Omni
revert unrelated uv.lock changes
fix: set do_sample_frames=False once at init for ColQwen-Omni
Co-authored-by: Hubert Lu <hubielu@email.com> (3c1f36d)
add bright pro benchmark (#5129) (44e33bf)
model: add Singaraj/morisien-embed (Mauritian Creole, mfe) (#5125) (1276f64)
benchmark: add BRIGHT-Pro retrieval (7 StackExchange domains) (#4929)
task: add Bright-Pro retrieval (7 StackExchange domains)
Adds seven per-domain retrieval tasks built on yale-nlp/Bright-Pro:
biology, earth_science, economics, psychology, robotics, stackoverflow,
sustainable_living. Each task loads the documents and examples HF
configs and exposes binary qrels from gold_ids for standard nDCG@10
evaluation. The dataset's reasoning-aspect annotations (aspects config)
are not consumed by the standard retrieval task and remain available on
the Hub for users who want aspect-aware evaluation.
Closes #4623
Register a top-level BRIGHT_PRO Benchmark grouping the 7 per-domain
BrightPro* retrieval tasks so users can run mteb.get_benchmark("BRIGHT-Pro")
in one call. Mirrors how BRIGHT/BRIGHT(v1.1) are registered.
Address review comment from @Samoed on #4651.
Update each BrightPro{Domain}Retrieval task's query prompt from the BRIGHT-v1.1-style 'Represent this {domain} post for searching relevant passages: ' to the BRIGHT-Pro paper's 'Given a {domain} post, retrieve relevant passages that help answer the post'.
Matches the prompt body used by BRIGHT-Pro's reference evaluation harness across all instruction-tuned retrievers (qwen3-embed, reasonir, gte-Qwen2, gritlm, etc.) so MTEB users running these models on BrightPro tasks see the paper-protocol numbers.
Generated via task.calculate_descriptive_statistics() as required by tests/test_tasks/test_metadata.py; fixes the failing test CI jobs.
Mirror the existing per-task BRIGHT entries with the 7 BrightPro tasks, verbatim-matching the task metadata prompts so scores reproduce out of the box (requested in PR #4651 review).
reupload
task: use natural domain names in BrightPro prompts
The raw subset slugs (earth_science, sustainable_living, stackoverflow) leaked into the query prompts because BRIGHT's config template substitutes the subset key directly (Given a {task} post, ...), and BRIGHT-Pro inherited that mechanism. Replace them with the StackExchange site names they refer to, and fix the article agreement (a -> an) for Earth Science / Economics.
Keeps mteb/models/model_implementations/reasonir_model.py in sync.
The leaderboard renders the first lines of a description, so state the retrieval quality being measured before the details. Move the benchmark attribution to the end and rephrase it as 'Was developed as part of', since a task can belong to several benchmarks.
Also set the benchmark display_name to BRIGHT-Pro.
This reverts commit b021cef5.
@Samoed is right that rewriting the prompts shifts the scores, which would leave the leaderboard numbers not matching the ones published with the benchmark. Keeping the prompts as the paper ran them takes priority over the cosmetics of the subset slugs, so this goes back to the original wording.
The slug leakage (a "sustainable_living post", "a earth_science post") is real but has to be weighed against reproducibility; see the PR thread for the measured effect.
This reverts commit 4bc04e6c07a008ae71d547e9a23d7c972855b7b0.
Co-authored-by: Roman Solomatin <36135455+Samoed@users.noreply.github.com> (f0cb9fd)
dataset: add MorisienMTBitextMining (Mauritian Creole, mfe) (#5110)
dataset: add MorisienMTBitextMining (Mauritian Creole, mfe)
dataset: set MorisienMTBitextMining license to MIT
The upstream MorisienMT dataset has been relicensed to MIT by its author, so the task metadata and the repackaged dataset now declare MIT.
dataset: repoint MorisienMTBitextMining to mteb-hosted copy (a0cf754)
model: add erikkaum/lattice-retrieval (#5105)
feat: add lattice retrieval model
test: add lattice mock-run results
Delete mteb_mock_run_results.md
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com> (98f3b8f)
dataset: add LPMusicCapsMTT A2T and T2A retrieval tasks (#5071)
dataset: add LPMusicCapsMTT A2T and T2A retrieval tasks
dataset: add creation script for LPMusicCapsMTT
Co-authored-by: Hubert Lu <hubielu@email.com> (a57299e)
model: moca-embed/MoCa-Qwen25VL-3B (#3589) (#5076)
feat: add MoCa-Qwen25VL-3B (closes #3589)
fix: add n_embedding_parameters for MoCa
fix: address review on MoCa (default prompt, training datasets, transformers pin)
chore: regenerate uv.lock for the moca extra
Co-authored-by: Hubert Lu <hubielu@email.com> (09a4a30)
model: add cnmoro/static-nomic-384-pten-v2-st (static pt/en Model2Vec) (#5075)
model: add cnmoro/static-nomic-384-pten-v2
Static (Model2Vec/Tokenlearn) pt/en embedding model distilled from nomic-embed-text-v2-moe. Adds Model2VecStaticModelWrapper because the checkpoint uses vocabulary quantization, which sentence-transformers' StaticEmbedding cannot load.
Publishes a materialized (one row per token) export of the model so the standard
SentenceTransformerEncoderWrapper can load it, instead of relying on the deprecated
Model2VecModel. All 186 tasks re-run against the new revision. (ed285c7)
Declare Common Voice (and derived CommonLanguage) in fusion-embedding training_datasets (0204e96)
fix: EmotionAnalysisPlus handle empty validation split + load real train split
fix: EmotionAnalysisPlus handle empty validation split + load real train split (#4937)
fix(EmotionAnalysisPlus): handle empty validation split + load real train split
EmotionAnalysisPlus: drop split pin and dataset_transform per review
Co-authored-by: Nicolas Helmeyer <helmeyen@login-3.server.mila.quebec>
Co-authored-by: Roman Solomatin <36135455+Samoed@users.noreply.github.com> (10f655f)
Update quality check (#5090)
update quality check
simplify comments
Update tests/test_tasks/test_task_quality.py
Co-authored-by: Kenneth Enevoldsen <kennethcenevoldsen@gmail.com>
Co-authored-by: Kenneth Enevoldsen <kennethcenevoldsen@gmail.com> (90da6ee)
Apply suggestions from code review
ci: Add pr length check (#5088)
add pr check
remove edited
fix tests
Apply suggestions from code review
Co-authored-by: Kenneth Enevoldsen <kenevoldsen@pm.me>
don't close pr
Label PRs with long descriptions
Add a label to PRs with long descriptions to manage them effectively.
Co-authored-by: Kenneth Enevoldsen <kennethcenevoldsen@gmail.com>
Co-authored-by: Kenneth Enevoldsen <kenevoldsen@pm.me>
Co-authored-by: Isaac Chung <chungisaac1217@gmail.com>
Co-authored-by: Kenneth Enevoldsen <kennethcenevoldsen@gmail.com> (881ee1f)
model: Add missing requirements for nomic models (8e30887)
fix: mock zeroshot tasks must declare the text modality
Zeroshot classification scores inputs against candidate text prompts, so these tasks require a text tower. The audio, video and video+audio mocks omitted "text" from metadata.modalities, so get_compatible_mock_tasks selected them for audio-only and video-only models, which then fail with KeyError on the input column. The image mock and every real zeroshot task already declare text.
Co-authored-by: Hubert Lu <hubielu@email.com> (f7c910d)
c1c888e)fix: count symmetric STS pairs once
fix: count symmetric STS pairs once (#4958)
test: detect symmetric STS pair leakage
refactor: reuse STS unique pair statistics
refactor: centralize symmetric pair counting (f49193c)
model: Split jinav4 models (#5082)
split jina
improve model_meta
keep device map (e05fa29)
model: add TIPSv2 (google/tipsv2-b14/l14/so400m14/g14) (#5067)
model: add TIPSv2 family (b14, l14, so400m14, g14)
remove mock run results file
Co-authored-by: Hubert Lu <hubielu@email.com> (7eea8b1)
docs: add multiple choice retrieval task example
Add a dedicated example to the 'Adding a Task' docs showing how to implement a multiple choice retrieval task. It explains that such tasks compare each query against a fixed candidate set via the top_ranked split and use accuracy as the main score, and demonstrates building the top_ranked split from relevant_docs in dataset_transform, following the BLINK task pattern.
Closes #4408 (3748f63)
.n_active_parameters so that active parameters correct when using the API (#5066)Active parameters is incorrect when using the API Fixes #5064
@Samoed anything I am missing here? (ee5689e)
Fix/afrie5 loader kwargs + changing AfriE5 revision (#5059)
fix: pass dtype via model_kwargs and drop invalid normalized kwarg for AfriE5-Large-instruct
fix: bump AfriE5-Large-instruct revision to include the modules.json,sentence_xlm-roberta_config.json, and 1_Pooling/config.json file additions
Co-authored-by: Nicolas Helmeyer <helmeyen@login-3.server.mila.quebec>
Co-authored-by: Nicolas Helmeyer <helmeyen@login-1.server.mila.quebec> (d0bad1e)
Add Hanno-Labs/dinghy-law-8b-v1 (legal embedding model) (#5058) (4ce470f)
Move data to mteb HF repo (mmvu, covers80, flare, miao, seavl, shs100k) (#5056) (0456b96)
benchmark: Expand MTEB(kor, v*) with retrieval/STS/NLI/clustering tasks (#4870)
Add 7 Korean retrieval tasks to MTEB(kor, v1)
Extend MTEB(kor, v1) retrieval coverage with LawIRKo, SQuADKorV1Retrieval, AutoRAGRetrieval, PublicHealthQA, BelebeleRetrieval, MultiLongDocRetrieval, and MrTidyRetrieval (kor subset), matching the KURE Korean retrieval set.
Extend MTEB(kor, v1) with STS17 (ko-ko), KLUE-NLI and PawsXPairClassification (NLI / pair classification), and SIB200ClusteringS2S, KlueMrcDomainClustering, KlueYnatMrcCategoryClustering (clustering). Kor-NLI is omitted (no mteb task).
Register dragonkue/BGE-m3-ko, dragonkue/multilingual-e5-small-ko, dragonkue/snowflake-arctic-embed-l-v2.0-ko, exp-models/dragonkue-KoEn-E5-Tiny, jhgan/ko-sroberta-multitask, nlpai-lab/KURE-v1, nlpai-lab/KoE5, telepix/PIXIE-Rune-v1.5, upskyy/bge-m3-korean so their Korean results render on the leaderboard. Metadata fetched from the HF Hub; loaders mirror each base family.
Add 'kor' to the bm25 language map with character-level tokenization (matching the existing jpn/zho handling for no-space scripts), so the bm25s reference baseline can index Korean retrieval tasks in MTEB(kor, v1).
(n_embedding_parameters is already set on all 9 entries.)
The prior commit only captured the file rename; the benchmarks.py / init.py / _leaderboard_menu.py edits were not staged. This commit adds them:
Sources: each model's card at the pinned revision (+ KURE GitHub repo).
BGE-m3-ko: training_datasets set() (author-confirmed, no mteb data in fine-tune)
Move pixie_rune_v1_5 to pixie_models.py (review request)
Relocate the telepix/PIXIE-Rune-v1.5 ModelMeta from korean_models.py to pixie_models.py alongside PIXIE-Rune-v1.0, reusing that file's PIXIE_RUNE_V1_PROMPTS (identical query-prefix scheme). Entry values unchanged.
8194 is max_position_embeddings (incl. XLM-R's 2 offset slots), not the usable input length: these models' tokenizer_config model_max_length and sentence-transformers max_seq_length are both 8192, which is what ModelMeta.max_tokens is documented to mean. Fixes the three Korean bge-m3 fine-tunes (per review suggestion) and the same error inherited in BAAI/bge-m3 itself, manu/bge-m3-custom-fr, GreenNode VN x2, AITeamVN/Vietnamese_Embedding, deepvk/USER-bge-m3, jina-embeddings-v3.
KorNLI is now available in mteb; include it in the v2 pair-classification set (v2: 19 -> 20 tasks). Merged main to pick up the task definition.
Per review, the 9 Korean community ModelMetas now live in #4921 so that embeddings-benchmark/results#581 can merge independently; this PR stays a benchmark-definition change (MTEB(kor, v2), bm25s Korean tokenization, max_tokens fixes).
Address review on #4870:
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Co-authored-by: dgyu <dgyu@sionic.ai>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> (db49990)
model: Add dragonkue/multilingual-e5-small-ko-v2 and sionic-ai/comsat-e… (#5054)
Register dragonkue/multilingual-e5-small-ko-v2 and sionic-ai/comsat-embed-ko-8b-preview
Two Korean community embedding models evaluated on MTEB(kor, v2):
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> (53a9c40)
Add perplexity-ai/pplx-embed-v1-late-0.6b model meta (#4813)
Add perplexity-ai/pplx-embed-v1-late-0.6b model meta
Closes #4691. Adds the ModelMeta entry for perplexity-ai/pplx-embed-v1-late-0.6b, a PyLate/ColBERT late-interaction (MaxSim) embedding model with 128-dim token-level vectors, continued-trained from perplexity-ai/pplx-embed-v1-0.6b. Metadata sourced from the HF model card/API and config files (revision, MIT license, 128-dim projection, 512 document length, ~596M params).
Add embedding parameter count for pplx late embed (7c2712d)
dataset: Repoint dead dataset to mteb's supported for AfriSentiLangClassification (#4912)
Fixed AfriSentiLangClassification: repointed dead HausaNLP dataset to mteb/afri_senti_lang mirror, removed redundant tweet->text preprocessing step
remove processing
Co-authored-by: Nicolas Helmeyer <helmeyen@login-2.server.mila.quebec>
Co-authored-by: Roman Solomatin <36135455+Samoed@users.noreply.github.com> (74044b9)
fix: Add MockTask for testing new model implementations
fix: Add MockTask for testing new model implementations (#4814)
Add MockTask for testing new model implementations
fix import in test file
Added under mteb package and remove from tests
restore unnecessary changes
fix lintter
remove from init.py
move mocktasks under mteb/tests_mock
fix imports
added mteb.tests_mock to not typecheck in ptproject.toml
rename folder and move files
remove mock_tasks.py and add separate file under mteb/tests/mock_task folder
remove unnecessary files and simplify
fix test and make removed legacy MockTask exposing to public
remove TYPE_CHECKING import
remove unncessary rules added to lintter for mteb/tests
rename folder from mteb/tests to mteb/mocks
remove tests/mock_tasks.py and update imports in tests
remove mteb/mocks/task_grid.py and add all things to mteb/mocks/init.py
removed utils.py and add create_mock_samples.py
change MockFastClusteringTask to LegacyMockFastClusteringTask and move multilingual_eval_langs to individual files
added mock-run CLI command for Mocktasks verification
fix MockRetrieval task import in test_hybrid_search.py
minor markdown formatting fix
update table format
add reason column
fix typecheck error
update formatting of terminal output as per review
changes from review
Separate core functionality to separate script and simplify CLI command, other docs changes frm reviws
rename function to match CLI and other changes from review
added MockRunResults and simplified other things
small fix
fix import
lintter
minor fixed in mock_task
minor error fix
minor
lintter (f5337a7)
bdce12d)ci: test reference models on all benchmarks
ci: test reference models on all benchmarks (#5030)
test reference models on all benchmarks
skip coderag
use benchmark/task combo
format citations (136b399)
fix: compute relevance scores in float32 to avoid low-precision ties (HPS) (#4933)
fix: compute relevance scores in float32 to avoid low-precision ties (HPS)
The similarity primitives in mteb.similarity_functions converted non-tensor
inputs to float32 but left existing tensors at their incoming dtype. When an
encoder returns float16/bfloat16 embeddings (e.g. models run in low precision),
cosine/dot scoring was therefore performed in low precision. Reduced-mantissa
formats coarsely bucket values in a narrow range, collapsing distinct relevance
scores into spurious ties; ranking then depends on arbitrary tie-breaking,
inflating retrieval-metric variance and bias.
Upcast float16/bfloat16 embeddings to float32 before scoring (High-Precision Scoring). The upcast is value-preserving — every float16/bfloat16 value is exactly representable in float32 — and a no-op for embeddings that are already float32, so leaderboard results for full-precision models are unchanged.
Reference: "Reliable Evaluation Protocol for Low-Precision Retrieval" (Yang et al., 2026), https://aclanthology.org/2026.acl-short.33
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Refactor tensor conversion to handle low-precision types more efficiently.
Removed low-precision upcasting for tensors.
Upcast sub-float32 floats to float32 in tensor conversion.
Co-authored-by: Kisu Yang <lab4@vaiv.kr>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Co-authored-by: Roman Solomatin <36135455+Samoed@users.noreply.github.com> (4b4f5de)
model: Add BGE Small Structural Separator model implementation (#4928)
Add BGE Small Structural Separator model implementation
Load structural separator from full Hugging Face checkpoint
Address structural separator review feedback
Address remaining structural separator review feedback
Use standard encoder protocol (1019eed)
fix: upcast reranker logits to float32 before softmax (HPS)
fix: upcast reranker logits to float32 before softmax (HPS) (#4934)
fix: upcast reranker logits to float32 before softmax (HPS)
The Qwen3 and monoT5-family cross-encoder rerankers apply log_softmax
directly to logits produced by a model running in bfloat16/float16. The
reduced mantissa of these formats coarsely buckets the (0, 1) probability
range, so distinct query-document relevance scores collapse into spurious
ties. Retrieval/reranking metrics then depend on arbitrary tie-breaking,
which inflates their variance and bias (e.g. up to ~38%p MRR@10 range and
+9.08%p bias for Qwen3-Reranker under bfloat16).
Apply High-Precision Scoring (HPS): upcast the logits to float32 immediately before the softmax, leaving the forward pass in low precision. The upcast is value-preserving, adds negligible cost, and restores near-full-precision ranking stability.
Reference: "Reliable Evaluation Protocol for Low-Precision Retrieval" (Yang et al., 2026), https://aclanthology.org/2026.acl-short.33
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>
Co-authored-by: Kisu Yang <lab4@vaiv.kr>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com> (ce4fb38)
[MOEB] Add dataset SHS100K (#5045)
[MOEB] Add dataset SHS100K
fix first results issue
fix spelling lint issue (6d97b70)
model: Add mDenseOn/mLateOn definitions (#5048) (0b75fb8)
[MOEB] Add several FLARE-based tasks (#4991)
[MOEB] Add several FLARE-based tasks
formatting
fix description
add the descriptive stats
fix URL
fix citation
Co-authored-by: Roman Solomatin <36135455+Samoed@users.noreply.github.com> (aadb54b)
[MOEB] Add MIAO-based datasets (#5001)
[MOEB] Add MIAO-based datasets
add descriptive stats
edit create data
update description
pin dataset version (db53995)
model: Add Bekko embedding models (#5043)
feat(models): add Bekko embedding models
docs(models): add Bekko paper citation
fix(models): use Bekko public release date
docs(models): clarify Bekko language metadata
feat(models): register Bekko core languages (ccd0914)
fix: avoid stacking variable-length tensors in JinaV4Wrapper.encode (#5050) (`8e2ac63`)
8e2ac63)[MVEB] Add STARBench video-centric QA task (#4746)
[MVEB] Add STARBench video-centric QA task
[MVEB] Split STARBench into per-config tasks (feasibility, interaction, prediction, sequence)
fix: run ruff format on star_bench.py
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Co-authored-by: Yashwanth Devavarapu <yashwanthdevavarapu@Yashwanths-MacBook-Pro.local>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com> (a1e96c8)
model: add RTriever-4B and DIVER-Retriever-4B (#5039)
model: add RTriever-4B and DIVER-Retriever-4B
Two reasoning-intensive retrievers, both Qwen3-Embedding-4B fine-tunes, so both reuse q3e_instruct_loader: last-token pooling, cosine similarity, and an "Instruct: ...\nQuery:" prefix on queries only. That matches the template in each checkpoint's config_sentence_transformers.json verbatim.
DIVER-Retriever-4B declares the BRIGHT corpora in training_datasets. Its model card lists reasonir/reasonir-data among its training sets, and the hard-query half of that dataset stores positive documents by id and reconstructs them from xlangai/BRIGHT, so those corpora are part of its training data transitively.
RTriever-4B is trained on synthetic data with no MTEB task overlap.
Adds Diver-Retriever-4B, -1.7B and -0.6B alongside the -1020 release. All four share the same interface and reuse q3e_instruct_loader; the shared training-data annotation is factored into DIVER_TRAINING_DATA.
The 1.7B checkpoint is built on the Qwen3-1.7B base LM rather than a Qwen3-Embedding checkpoint. Its 1_Pooling/config.json carries the 4B model's word_embedding_dimension (2560) while the checkpoint's hidden size is 2048, so ModelMeta records the real value.
Per review: borrowing another family's loader function couples these models to the Qwen3 module, so a change there would silently reach RTriever and DIVER. Each file now carries its own instruction_template and passes InstructSentenceTransformerModel with explicit loader_kwargs.
The template is byte-identical to what q3e_instruct_loader produced
("Instruct: {instruction}\nQuery:" on queries, empty on documents), so scores
are unchanged. (604a73c)
update GZTAN reference (d3f982e)
Add BirdCLEF Species Audio Clustering task (closes #5017) (#5023)
Add BirdCLEF Species Audio Clustering task (closes #5017)
Part of MOEB: Massive Omni Embedding Benchmark (tracking issue #4842)
fix: update BirdCLEF bibtex citation to Cañas et al. 2025
fix: correct bibtex field order for BirdCLEF citation
Co-authored-by: Rakshitha Ireddi <rakshithaireddi@Rakshithas-MacBook-Pro.local> (23b79d1)
[MOEB] Add MMVU dataset (#5038)
[MOEB] Add MMVU dataset
fix lowest dep issue and updating reference
simplify mmvu with push_to_hub, adding create data for reference
Apply suggestion from @Samoed
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com> (b7e6584)
[MOEB] Add SEA-VL: Multicultural VL Dataset for Southeast Asia (#5040)
[MOEB] Add SEA-VL: Multicultural VL Dataset for Southeast Asia
fix task and add stats
remove shs import
simplify sea_vl by pre-computing dataset (9c30471)
model: Change attention for codefuse models (#5032)
change attention for codefuse (2c772e6)
Registers the German-focused mxbai embedding model (mixedbread-ai/deepset-mxbai-embed-de-large-v1, 487M params, XLM-RoBERTa backbone, 1024-dim) in the model implementations registry using the generic SentenceTransformerEncoderWrapper reference loader.
This model already has 3.3M downloads on the Hub but was missing from the MTEB registry, which blocks the pending results PR (embeddings-benchmark/results#647) that reports its scores on the full MTEB(deu, v1) benchmark.
Co-authored-by: Simon Dittrich <ai@cloudsoziologe.de> (a4046b9)
model: Add jina-reranker-v3.5 metadata (#5041)
model: add jina-reranker-v3.5 metadata
fix: support jina reranker v3.5 API
remove tests
Co-authored-by: Roman Solomatin <36135455+Samoed@users.noreply.github.com> (e123184)
Dump vLLM to v0.26.0 (#5022)
init
Signed-off-by: wang.yuqi <yuqi.wang@daocloud.io>
refine
refine
Signed-off-by: wang.yuqi <yuqi.wang@daocloud.io>
Signed-off-by: wang.yuqi <yuqi.wang@daocloud.io>
Signed-off-by: wang.yuqi <yuqi.wang@daocloud.io> (fbdd40e)
model: Add LightOn-rerank models (3 pointwise + 3 listwise) (#4961)
add LightOn-rerank models (3 pointwise + 3 listwise)
address review: drop PW wrapper, move padding_side to tokenizer_config, add modality checks
restore uv.lock from upstream main
predict now accepts batch_size correctly (f38c969)
update ruff to 0.16 (`df38f9b`)
update ruff to 0.16 (df38f9b)
fix(bm25): fall back to language-agnostic tokenization for unknown languages
Co-authored-by: Nicolas Helmeyer <helmeyen@login-4.server.mila.quebec> (787bf7c)
add mteb-pt to homepage (8847790)
model: add webAI ColVec1.1 4B and 8B models (#5010)
model: add webAI ColVec1.1 models
model: declare ColVec1.1 transformers requirement
fix(model): address initial ColVec1.1 review feedback
fix(model): make SDPA the ColVec1.1 default (6e72309)
model: Update fusion-embedding-2 revision to v0.3-preview (1720d8b1) (#5028) (46a2c21)
Add ModelMeta: minetta/nemotron-3-embed-8b-legal (#5027)
Add ModelMeta: minetta/nemotron-3-embed-8b-legal
license as URL (openmdw-1.1 not in Licenses literal)
Fill n_embedding_parameters; set public_training_code/data explicitly
revert public_training_* to None (schema expects str|None)
Co-authored-by: banyaneth <banyaneth@users.noreply.github.com> (6089992)
Add KiteFishAI/Nano-Em1-0.6B-v2 (#4993)
Add KiteFishAI/Nano-Em1-0.6B-v2
Update model revision hash in qwen3_models.py
Update Nano-Em1-0.6B-v2 model metadata
Updated model metadata for Nano-Em1-0.6B-v2 including release date, parameter counts, and memory usage.
Update qwen3_models.py
Refactor training datasets format in qwen3_models.py
Update qwen3_models.py
Add ScoringFunction import to qwen3_models.py
format & move
format & move
Co-authored-by: Roman Solomatin <36135455+Samoed@users.noreply.github.com> (c67d000)
Part of MOEB: Massive Omni Embedding Benchmark (tracking issue #4842)
Co-authored-by: Rakshitha Ireddi <rakshithaireddi@Rakshithas-MacBook-Pro.local> (609b43f)
task: add Song Describer text-music retrieval (t2a, a2t) (#4988)
task: add Song Describer text-music retrieval (t2a, a2t)
Human-written music captions from the Song Describer Dataset (CC-BY-SA-4.0, MTG-Jamendo audio) as text<->music retrieval. 746 captions over 547 tracks. CLAP t2a hit_rate@5 6.3 vs random 1.3. Complements MusicCaps with human captions and a published retrieval benchmark. Includes descriptive statistics.
Rebuild from the full Zenodo SDD release (was the 547-track valid subset), matching the paper's corpus. laion/larger_clap_general reproduces the paper's Table 5 CLAP retrieval almost exactly (T2A R@1/5/10 = 4.79/17.72/29.48 vs paper 4.42/17.02/26.01). Updated revisions, descriptions, and descriptive statistics.
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com> (d206dae)
update links to code for Nemotron models (18005a3)
[MOEB] Add AESDD dataset (#4978)
[MOEB] Add AESDD dataset
remove corrupt audiofile
simplifying implementation with AESDD fixed (8751596)
task: add CASTELLA audio moment retrieval (t2a) (#4984)
task: add CASTELLA audio moment retrieval (t2a)
First temporal-localization retrieval task for the audio benchmark: captions retrieve the 10-second window of a long recording containing the described moment, built on CASTELLA (arXiv 2511.15131), the DCASE 2026 Task 6 evaluation set. 12,046 windows from 566 recordings, 1,347 queries, graded by >=50 percent overlap with annotated moments.
chore: add descriptive statistics for CASTELLAAMRRetrieval
task: switch CASTELLA-AMR to full-recording retrieval
Replace the 10s-window corpus with the 566 complete recordings (60-300s), one gold recording per caption. Matches the paper's audio length and simplifies the task; recomputed descriptive statistics.
style: ruff format castella_amr (885f740)
model: add SigLIP2 family (15 checkpoints) (#4973)
model: add SigLIP2 family (15 checkpoints)
Adds ModelMeta entries for the google/siglip2-* fixed-resolution checkpoints (base/large/so400m/giant-opt). They reuse the existing SiglipModelWrapper since these checkpoints load as SiglipModel. NaFlex variants are excluded as they need different processor handling.
Closes #2301
SigLIP2 ships a fast tokenizer so sentencepiece/protobuf are not
required; the image extra is added automatically. (62c5293)
[MOEB]: Add Covers80 dataset (#4986)
[MOEB]: Add Covers80 dataset
update description
formatting
add descriptive stats
small improvement script (be0f623)
Add Hanno-Labs/dinghy-law-4b-v1 (legal embedding model) (#4992)
Add Hanno-Labs/dinghy-law-4b-v1 (legal embedding model)
Apply suggestions from code review
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com> (41a5310)
fix: type mismatch between some whisper models and mel features (#4990) (`acbf4ca`)
acbf4ca)[MOEB] Register whisper-large-v3-turbo (#4989) (f095887)
Cite the Fusion Embedding arXiv paper in both ModelMetas (#4987) (0ae18d6)
task: add VSC2022 video-to-video copy-detection retrieval (v2v) (#4985)
task: add VSC2022 video-to-video copy-detection retrieval (v2v)
chore: add descriptive statistics for VSC2022Retrieval
lint
Co-authored-by: Roman Solomatin <36135455+Samoed@users.noreply.github.com> (80f6e3b)
task: add MomentSeeker composed video retrieval (it2v, vt2v) (#4967)
task: add MomentSeeker composed video retrieval (it2v, vt2v)
Adapts MomentSeeker long-video moment retrieval (CC-BY-NC-SA-4.0) into corpus retrieval: 30-second non-overlapping 360p chunks of the source videos form the corpus, with relevance from overlap against annotated answer intervals. Two directions: text+image queries and text+video queries, 400 each. Self-contained: adds the it2v TaskCategory code. Closes #4944.
chore: add descriptive statistics
style: apply ruff format to moment_seeker.py
task: use map_at_5 as main_score for MomentSeeker (matches the paper) (9082aa7)
Add new version for LCO-3B (#4971)
Add new version for LCO-3B model. This version release on May 2025 and has shown better performance compared to it's predecessors. (4f1a28c)
Update fusion-embedding-2 revision to v0.2 (9451b840f0d1)
The v0.2 release fine-tunes on the expanded AudioCaps 2.0 pool; non-audio
paths are unchanged (bitwise-verified on the released artifact). Weights and
card at the new revision; the v0.1 pin remains valid in history. (7967939)
Three related bugs surfaced while running PairClassification tasks with jina-embeddings-v4:
get_text_embeddings/get_image_embeddings called task_type.startswith() without a None guard. get_prompt_name legitimately returns None when a task isn't in the model's prompt map, which the surrounding code already handles elsewhere (jina_task_name = model_prompts.get(task_type) if task_type else None) -- these two call sites just missed the same guard.
jina_embeddings_v4's model_prompts was missing a PairClassification entry that every other jina registration in this file already has (v3 maps it to "classification"; v4 has no classification adapter, so text-matching is the correct equivalent -- the same adapter v4 already uses for STS). Without it, PairClassification tasks silently fell back to the retrieval adapter instead of text-matching.
JinaV4Wrapper.encode() returned raw GPU tensors (or lists of GPU tensors)
without converting to numpy, unlike JinaWrapper (v3)'s equivalent method
which already does this. Downstream evaluators that call np.asarray() on
the result (e.g. PairClassificationEvaluator) crashed on GPU tensors. (5f517f5)
task: add VimSketch query-by-vocal-imitation retrieval (a2a) (#4963)
task: add VimSketch query-by-vocal-imitation retrieval (a2a)
Adds the first audio-to-audio retrieval task: 2,168 vocal-imitation queries (max 4 per reference class) against 544 reference sounds from the VimSketch dataset (Vocal Imitation Set + VocalSketch, CC-BY-4.0, Zenodo 2596911). Ground truth is the instance-level imitation-to- reference mapping encoded in the dataset. Closes #2242.
Clusters the 2,168 vocal imitations by imitated reference class (542 classes), derived from the retrieval qrels of the same dataset. Data lives in a clustering config of the existing dukesun99/VimSketch repo.
chore: add descriptive statistics
Update citation
Update citation
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com> (9d15297)
task: add CLD composed audio retrieval (at2a) (#4965)
task: add CLD composed audio retrieval (at2a)
Adds the first composed audio-text-to-audio retrieval task, built from the CLD subset of ADIFF (ICLR 2025, MIT annotations): a source Clotho v2.1 evaluation clip plus a language description of how the target differs, retrieving the target among 1,045 clips. 2,000 queries sampled with a fixed seed from the 5,225 evaluation pairs. Data repackaged in standard corpus/queries/qrels layout. Related to #4943 (AudioDiffCaps itself cannot be rebuilt publicly: its repo ships no s2 JAMS).
update metadata
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>
chore: add descriptive statistics
task: add CLDA2TRetrieval, audio-to-text difference retrieval (a2t)
chore: add descriptive statistics for CLDA2TRetrieval
format
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>
Co-authored-by: Roman Solomatin <36135455+Samoed@users.noreply.github.com> (3e4ec52)
task: add SPEECH-COCO spoken-caption image retrieval (a2i, i2a) (#4964)
task: add SPEECH-COCO spoken-caption image retrieval (a2i, i2a)
Adds audio-image cross-modal retrieval in both directions using SPEECH-COCO (arXiv 1707.08435, CC-BY-4.0): spoken captions synthesized with eight TTS voices over MS-COCO images. Data is a deterministic downsample of the mteb/SpeechCoco validation split originally processed in PR #3070 (2,048 images, 1,000 queries per direction), repackaged in standard corpus/queries/qrels layout. Requires the a2i/i2a TaskCategory extension. Closes #2298.
chore: add descriptive statistics (e8f0b21)
model: add DINOv3 ViT family (6 checkpoints) (#4976)
Adds ModelMeta entries for the facebook/dinov3-vit* web-image (LVD-1689M) checkpoints, reusing the existing DINOModel wrapper. ConvNeXt and satellite (SAT-493M) variants are left out for now: ConvNeXt outputs spatial feature maps that the CLS-pooling wrapper does not handle. Repos are gated behind the DINOv3 license.
Closes #3031 (e490c78)
dataset: add cleaned STSBenchmark v2 (#4876)
dataset: add cleaned STSBenchmark v2
reupload
Co-authored-by: Roman Solomatin <36135455+Samoed@users.noreply.github.com> (7f64db9)
Add voiceclap-large-v2 (#4972) (b1d6882)
task: add SoundingEarth audio-image retrieval (a2i, i2a) (#4966)
task: add SoundingEarth audio-image retrieval (a2i, i2a)
Adds cross-modal retrieval between co-located field recordings and aerial images from SoundingEarth (CC-BY-4.0 on Zenodo): 2,048 locations sampled with a fixed seed, 1,000 queries per direction, instance-level ground truth. Audio is sourced per recording from Internet Archive via the authors' pipeline. Self-contained: adds the a2i/i2a TaskCategory codes. Closes #4942.
81958bb)fix: correctly isolate splits when aggregating scores in aggregated tasks
fix: correctly isolate splits when aggregating scores in aggregated tasks (#4897)
fix: correctly isolate splits when aggregating scores in aggregated tasks
test: add regression test for isolated split aggregation in AbsTaskAggregate
chore: format regression test with ruff
test: refactor regression test to use existing MockAggregatedTask
test: fix state leakage in MockAggregatedTask by using model_copy
chore: format test with ruff
test: specify results for both inner tasks of MockAggregatedTask to properly simulate aggregation
feat: validate task results match requested tasks in AbsTaskAggregate
chore: fix import sorting with ruff check
fix: set main_score to None when missing task results instead of raising ValueError
test: add dev split assertion to missing task test
minor fix
Co-authored-by: Kenneth Enevoldsen <kennethcenevoldsen@gmail.com> (023b0b1)
docs: add icon to hybrid model page
The hybrid model page was the only page under advanced usage missing
frontmatter, so it rendered without an icon in the docs navigation.
Add a title and the lucide/merge icon to match the other pages. (0c71e16)
fix: ModelMeta Validation (#4904)
fix ModelMeta Validation
fix loader signature in other methods
revert changes in model_meta and add fix in result_cache
add as a method in modelmeta
added test for model_validate_json_resolved (439298f)
model: Add fusion-embedding-2 (EximiusLabs/fusion-embedding-2-2b-preview) (#4952)
model: Add fusion-embedding-2 (EximiusLabs/fusion-embedding-2-2b-preview)
Second generation of the fusion-embedding family: modality-gated deep adapters on the same frozen base (60.6M trained parameters). The existing wrapper serves both generations; the repository's remote code attaches and gates the adapters internally. Training corpus is identical to generation 1.
Both fusion-embedding-1 and fusion-embedding-2 now reference the same
@software entry (the family report covers both generations), fixing the
duplicate-citation check that flagged the near-identical per-generation
titles. (8d5edc2)
Cap Seq Length for Some Models (#4873)
fix(e5-omni): cap text encode length at 512 tokens (paper config)
fix(tevatron-omniembed): cap text encode length at 512 tokens
fix(lco-embedding): truncate text to 300 tokens matching author eval
Apply suggestions from code review
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com> (cd18288)
leaderboard: add MTEB(por, v1) to Language-specific section (#4938)
leaderboard: add MTEB(por, v1) to Language-specific section
run make lint (5594cc9)
fix: load gte-Qwen2-7B-instruct in bf16 with trust_remote_code
fix: load gte-Qwen2-7B-instruct in bf16 with trust_remote_code (#4935)
fix: load gte-Qwen2-7B-instruct in bf16 with trust_remote_code
Three issues with the stock loader for Alibaba-NLP/gte-Qwen2-7B-instruct:
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com> (f4a5870)
model: add StaRSE model (#4936)
added starse model meta
added starse model meta (ae87966)
Set training_datasets for comsat-embed-ja models (empty set + inheritance) (#4930)
Set training_datasets=set() for comsat-embed-ja models
Their own fine-tuning data contains no mteb datasets, so use an empty set
(instead of None) — this lets the base models' training data be inherited
via adapted_from (Qwen/Qwen3-Embedding-8B, cl-nagoya/ruri-v3-310m), giving a
meaningful zero-shot percentage on the leaderboard instead of N/A. (3cdda31)
model: Add fusion-embedding-1 (EximiusLabs/fusion-embedding-1-2b-preview) (#4909)
model: Add fusion-embedding-1 (EximiusLabs/fusion-embedding-1-2b-preview)
Use commit SHA for ModelMeta revision
Load via AutoModel trust_remote_code; add fusion-embedding extras group
Point code links at the renamed family repo (Eximius-Labs/fusion-embedding)
Add image support to the wrapper
get_image_embeddings routes through the frozen base model's own image path (exposed by the repo's remote code); images are embedded one at a time, as in other merged wrappers. modalities now lists image. Fused image+text inputs raise NotImplementedError.
9ea1d63)Add Hanno-Labs/dinghy-law-0.6b-v1 (legal embedding model) (#4926)
Add Hanno-Labs/dinghy-law-0.6b-v1 (legal embedding model)
Address review: use SentenceTransformerEncoderWrapper + model_prompts (verified byte-identical to q3e loader, 65.85 -> 65.85 on all 8 MTEB(Law) tasks); trim verbose comments to model card
Use InstructSentenceTransformerModel with a custom instruction_template (per @Samoed); same mechanism q3e wraps, so 65.85 unchanged; bare per-task instructions via prompts_dict
Co-authored-by: Stephen Solka <stephen@standd.io> (a954de2)
InjongoIntent: dropped eng config (#4913)
InjongoIntent: dropped eng config (as mteb mirror was missing eng/test split)
style: apply ruff format
Co-authored-by: Nicolas Helmeyer <helmeyen@login-4.server.mila.quebec>
Co-authored-by: Nicolas Helmeyer <helmeyen@login-1.server.mila.quebec> (651d53f)
Fixup logging message for multimodal sbert models (#4927)
fix modalities
add log message
fix cross-encoder
fixup logging message (30ebfee)
fix: Use only supported modalities for multimodal sentence_transformers
fix: Use only supported modalities for multimodal sentence_transformers (#4923)
fix modalities
add log message
fix cross-encoder (22092d6)
model: Add LFM2.5 and LFM2 models (#4835)
Added LFM2.5 and LFM2 models
correct embed_dim for colbert models
changes from review (9d3e05c)
model: Nemotron 3 Embed models (#4922)
Nemotron 3 Embed models (ceffb8b)
model: Add VoiceNet voiceclap-large and voiceclap-small (#4864) (aa68e9c)
Register Korean community embedding models (split from #4870) (#4921)
Register Korean community embedding models (split out of #4870)
Adds ModelMeta for 9 Korean community models so their results can appear on the leaderboard: dragonkue/BGE-m3-ko, dragonkue/multilingual-e5-small-ko, dragonkue/snowflake-arctic-embed-l-v2.0-ko, exp-models/dragonkue-KoEn-E5-Tiny, jhgan/ko-sroberta-multitask, nlpai-lab/KURE-v1, nlpai-lab/KoE5, telepix/PIXIE-Rune-v1.5 (placed in pixie_models.py), upskyy/bge-m3-korean.
Split out of #4870 per review, so that embeddings-benchmark/results#581
(which carries these models' results) can merge independently of the
benchmark-definition changes. (7723dcc)
fix: Change spanish dataset repo to mteb
upd spanish dataset repo (f2bf512)
models: added models for mteb-pt evaluation (#4894)
models: added models for mteb-pt evaluation
fix: ran make lint
updated mteb-pt citation
removed #todo tag
Update mteb/models/model_implementations/aegis_models.py
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>
fix: unit tests citation
moved models to portuguese_models.py
ran make lint
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com> (08f42e7)
model: Add ModelMeta for minishlab/potion-code-16M-v2
Adds the CoIR-focused static code embedding model, distilled from
nomic-ai/CodeRankEmbed and trained on CornStack via Tokenlearn and
contrastive fine-tuning. (193e3f6)
support porcessing images and text together (0ab45ef)
dataset: Added NANOBEIR Extended Benchmark (#4837)
Added NANOBEIR Extended Benchmark
fix tasks naming
fix citation formatting and tests
minor changes
update dataset path after uploading to mteb hf and remove custom formating
added description for each dataset and revert reference link
change names (702c9fa)
Add LingoIITGN/qwen-indic-v1 to sentence_transformers_models.py (#4884)
Add LingoIITGN/qwen-indic-v1 to sentence_transformers_models.py
Update mteb/models/model_implementations/sentence_transformers_models.py
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>
Update ModelMeta with from_hub hardware values
format
linting issues fix
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>
Co-authored-by: Roman Solomatin <36135455+Samoed@users.noreply.github.com> (500eff0)
feat: Add openness score to model metadata and documentation
feat: Add openness score to model metadata and documentation (#4877)
Add openness score to model metadata and documentation
add model card
refactor how openness is displayed in docs to make it more condensed
add whats new documentation
added script for generatin new structure
added a few visualizations to illustrate the state of openness
Co-authored-by: Kenneth Enevoldsen <kennethcenevoldsen@gmail.com> (77c6aa2)
add kalm-reranker-v1 (#4882)
add kalm-reranker-v1
reformated kalm-reranker-v1
reformated kalm-reranker-v1
fix (e84b288)
chore: fix Pydantic 2.11 deprecation warning
Replace instance-level access of model_fields with class-level access to clear pytest warnings and future-proof for Pydantic V3. (6f5d73d)
always rebuild image (28bc435)
feat: auto-install optional model dependencies via MTEB_AUTO_INSTALL_EXTRAS (#4875)
feat: auto-install optional model dependencies via MTEB_AUTO_INSTALL_EXTRAS
add whats new
Co-authored-by: Kenneth Enevoldsen <kennethcenevoldsen@gmail.com> (00fb3c6)
Japanese dense embedder fine-tuned from cl-nagoya/ruri-v3-310m (ModernBERT,
315M params, 768-dim, max 8192, cc-by-nc-4.0). Loads as a plain
sentence-transformers model; its Japanese query/document prefixes ship in
config_sentence_transformers.json and are picked up automatically. (786b31c)
fix: add missing arXiv ID and update publication year for mFollowIR citations (09efbdc)
fix por benchmark (#4844) (9d21470)
Add ModelMeta for sionic-ai/comsat-embed-ja-8b-preview (#4885)
Japanese-specialized dense embedder fine-tuned from Qwen/Qwen3-Embedding-8B
(8B params, 4096-dim, last-token pooling + L2 norm, cc-by-nc-4.0). Loads as a
plain sentence-transformers model; its instruct query prompt ships in
config_sentence_transformers.json and is picked up automatically. (8aebdf8)
This is technically my first contribution to mteb. I found a small typo in the README.md, so I thought I d fix it. (1fdb489)
fix(JinaV5OmniWrapper): text-task prefix + channels-last video frames
fix(JinaV5OmniWrapper): text-task prefix + channels-last video frames (#4881)
fix(JinaV5OmniWrapper): use Query/Document prefix for all text tasks
The text-matching, classification and clustering adapters were trained WITH the "Query: "/"Document: " prefix, not without it. Running them without any prefix causes dramatic score drops (e.g. STS12: 0.85→0.38, Banking77: 0.90→0.30, SprintDuplicateQuestions: 0.96→0.18).
New logic: audio/video non-retrieval tasks skip the prefix (unchanged); text and image tasks always use the prefix regardless of adapter type.
Update unit tests to reflect the corrected expected behavior.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
torchcodec's VideoDecoder yields (T, C, H, W) uint8 frame batches, but the model's HF remote code detects video only for channels-last (T, H, W, 3|4) arrays. The mismatched tensor fell through to the text fallback and was embedded as str(array) — every video collapsed to near-identical garbage embeddings (e.g. Shot2Story20KAT2VRetrieval nDCG 0.0009).
Verified on GPU: same video as channels-first tensor vs channels-last array gives cosine 0.14 between embeddings; the channels-first embedding matches the embedding of the array's string repr.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
Co-authored-by: Roman Solomatin <36135455+Samoed@users.noreply.github.com> (598a7eb)
dataset: KorNLI (#4880) (e1fbfe8)
Fix task crashes in EBind (short audio) (#4872)
fix(ebind): pad sub-window audio clips so short samples don't crash the batch (3fdcb0d)
model: Add Byrne-Embed model implementation (#4847)
Add Byrne-Embed model implementation (Quazim0t0/Byrne-Embed)
Byrne-Embed: load via trust_remote_code AutoModel (drop snapshot_download)
Update mteb/models/model_implementations/byrne_models.py
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>
Co-authored-by: SeanceTable <apoetyouknow@gmail.com>
Co-authored-by: Kenneth Enevoldsen <kenevoldsen@pm.me>
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com> (8b8f169)
Your coding agent can read these notes before it upgrades. Set up the MCP server →