NewYour coding agent can read the release notes before it upgrades.Set up the MCP server →
PyPI · #2804 most downloaded on PyPI
Massive Text Embedding Benchmark
Last release today
01 Oct 2026
Ships fairly regularly
a new release about every 1 weeks
Nearly every release is documented
notes for 60 of the last 60 stable releases
Nothing withdrawn
no release was ever pulled
4 years old
811 releases · first in 2022
One column per quarter.
fix: Download cached results zip from cached-data branch
fix: Download cached results zip from cached-data branch (#3795)
Optimize leaderboard startup by downloading cached results from cached-data branch
🤖 Generated with Claude Code
Co-Authored-By: Claude Sonnet 4 <noreply@anthropic.com>
make lint
Fix leaderboard stability test with enhanced debugging
🤖 Generated with Claude Code
Co-Authored-By: Claude Sonnet 4 <noreply@anthropic.com>
Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com>
🤖 Generated with Claude Code
Co-Authored-By: Claude Sonnet 4 <noreply@anthropic.com>
The cached results file has grown to ~92.7MB, exceeding the previous 50MB limit. This change increases the limit to 500MB to accommodate current and future file sizes.
🤖 Generated with Claude Code
Co-Authored-By: Claude Sonnet 4 <noreply@anthropic.com>
GitHub Actions were failing because cachetools was not installed during CI test runs. The leaderboard extra was already defined with cachetools>=5.2.0 but wasn't included in the install-for-tests target used by CI.
🤖 Generated with Claude Code
Co-Authored-By: Claude Sonnet 4 <noreply@anthropic.com>
Addresses PR comment feedback indicating the log flushing optimization was unnecessary at this stage. Removes:
Leaderboard functionality remains unchanged and tests pass.
🤖 Generated with Claude Code
Co-Authored-By: Claude Sonnet 4 <noreply@anthropic.com>
remove _validate_benchmark_json
Refactor leaderboard caching to use ResultCache and consolidate tests
Move download_cached_results_from_branch to ResultCache class and reduce TestDownloadCachedResultsFromBranch from 23 to 13 test cases while maintaining full coverage.
🤖 Generated with Claude Code
Co-Authored-By: Claude Sonnet 4 <noreply@anthropic.com>
Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com>
lint and remove unreachable code
Move shared test fixtures to parent conftest.py
All 25 tests now passing (23 in test_result_cache.py, 2 in test_integration.py)
🤖 Generated with Claude Code
Co-Authored-By: Claude Sonnet 4.5 <noreply@anthropic.com>
make method private
Fix content type validation test to match implementation behavior
The test_content_type_handling test was expecting warnings for unexpected content types, but the actual implementation raises exceptions. Updated test to use pytest.raises() for proper exception validation.
🤖 Generated with Claude Code
Co-Authored-By: Claude Sonnet 4 <noreply@anthropic.com>
update cache based on review comments
type check
Remove unused leaderboard_test_config fixture
fix: remove unused mock_invalid_json fixture
rm AGENTS/,d
reduce number of excepts in app.py
Co-authored-by: Claude Sonnet 4 <noreply@anthropic.com>
Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com> (3303087)
add model: bflhc/Octen-Embedding-4B (#3816) (6dcbf9f)
Add filter for model type (#3799)
Add filter for model type
fix literal issue
fix
remove white space
remove logic in filter_tasks
remove info in leaderboard
add tests
update tests
add default in model types
fix model filter
Co-authored-by: Roman Solomatin <36135455+Samoed@users.noreply.github.com> (6f4627e)
This completes the integration of the new leaderboard CLI command into the project's build system and removes deprecated direct module execution.
feat: Add leaderboard CLI command (#3802)
feat: add leaderboard CLI command with cache-path option
test: add comprehensive tests for leaderboard CLI command
try to fix install
fix: lazy-load leaderboard to avoid requiring deps for CLI
Update mteb/cli/build_cli.py
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>
make lint
remove AGENTS.md
move import to top of file
log the default cache path
Improve leaderboard tests to verify actual cache paths
Address PR feedback by modifying leaderboard tests to verify the actual cache paths passed to get_leaderboard_app instead of mocking ResultCache.
This provides better test coverage by validating that the cache objects passed to the leaderboard app have the correct paths, as suggested in PR comment: https://github.com/embeddings-benchmark/mteb/pull/3802#discussion_r2650719614
🤖 Generated with Claude Code
Co-Authored-By: Claude Sonnet 4 <noreply@anthropic.com>
Address PR feedback by combining test_leaderboard_custom_cache_path and test_leaderboard_default_cache into a single parametrized test.
This addresses PR comment: https://github.com/embeddings-benchmark/mteb/pull/3802#discussion_r2650721879 "Can be combined with the following test using a parametrize argument"
🤖 Generated with Claude Code
Co-Authored-By: Claude Sonnet 4 <noreply@anthropic.com>
Address PR feedback by updating the project to use the new leaderboard CLI:
python -m mteb leaderboard
instead of python -m mteb.leaderboard.appif __name__ == "__main__": block from mteb/leaderboard/app.py
as this functionality is now handled by the CLI commandThis completes the integration of the new leaderboard CLI command into the project's build system and removes deprecated direct module execution.
🤖 Generated with Claude Code
Co-Authored-By: Claude Sonnet 4 <noreply@anthropic.com>
feat: add theme and head parameters to leaderboard CLI
fix: suppress leaderboard warnings on CLI launch
test: update leaderboard tests for theme and head params
Revert "Update make run-leaderboard to use new CLI and remove app.py main block"
This reverts commit d4df501a4c5111b80899624620398c2e72a308a2.
Co-authored-by: Kenneth Enevoldsen <kennethcenevoldsen@gmail.com>
docs: update leaderboard CLI usage
update docs to show defaults
fix: apply ruff formatting
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>
Co-authored-by: Claude Sonnet 4 <noreply@anthropic.com>
Co-authored-by: Kenneth Enevoldsen <kennethcenevoldsen@gmail.com> (1a64ed6)
docs: add benchmark filtering examples
docs: add benchmark filtering examples (#3805)
docs: add benchmark filtering examples
Apply suggestion from @Samoed
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>
docs: remove custom benchmarks subsection
docs: expand filtering section with content tabs
docs: fix code block indentation in content tabs
build: include docs deps in dev group
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com> (b1aae79)
fix: repo exists check (#3813)
fix repo exists check
add test (6ebc5fd)
Update the API of Bytedance/Seed1.6-embedding-1215 (#3814)
update reference website of Seed1.6-embedding-1215
update Bytedance/Seed1.6-embedding-1215 model (ab2d494)
update generate_model_card with get_benchmark_result() (#3796)
update generate_model_card with get_benchmark_result()
add support for list of benchmarks
split parameters
fix type
generate card
add tests
add tests
add tabulate to test dependencies
correct tests
Co-authored-by: Roman Solomatin <36135455+Samoed@users.noreply.github.com> (9e867f5)
Add function for creating mock images (#3803)
create function for creating mock tasks
add annotations (48f137e)
Add benchmark aliases (#3767)
add benchmark aliases
split to aliases
move aliases
create aliases in separate function
simplify a bit
add test
Apply suggestions from code review
Co-authored-by: Kenneth Enevoldsen <kennethcenevoldsen@gmail.com>
add default value
add MTEB alias
Co-authored-by: Kenneth Enevoldsen <kennethcenevoldsen@gmail.com> (480f1b9)
Update mteb/results/benchmark_results.py
fix: add typecheck (#3550)
add pytyped
start typing
finish evaluators
add more types
Update mteb/results/benchmark_results.py
Co-authored-by: Kenneth Enevoldsen <kennethcenevoldsen@gmail.com>
apply comments
continue typechecking
fix typehint
typechecking
fix tests
fix type errors again
fix cache
add more types
fix method
roll back pyproject
activate PGH
install more types
almost finish
fix search wrappers
add ci
fix tests
fix 3.10 types
rollback overload
fixes after merge
change to iterable
add fixes
remove summarization scores hint
simplify deprecated_evaluator
simplify model conversion
add comment for typechecking
remove casts
remove duplicated function
Co-authored-by: Kenneth Enevoldsen <kennethcenevoldsen@gmail.com> (a99557d)
save kwargs passed to get_model in model_meta (#3785)
save kwargs passed to get_model in model_meta
add save_kwargs to load_model
removed copy of meta
Update mteb/models/model_meta.py
try to run with kwargs
try to move kwargs
add tests
change model in tests
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>
Co-authored-by: Roman Solomatin <36135455+Samoed@users.noreply.github.com> (3ec1f63)
fix: Added warnings.warn when logging warnings
fix: Added warnings.warn when logging warnings (#3753)
Added warnings.warn when logging warnings
address comments
Added depreciation warning
made better
address comments
address comments
address comments
Co-authored-by: Roman Solomatin <36135455+Samoed@users.noreply.github.com>
Co-authored-by: Kenneth Enevoldsen <kennethcenevoldsen@gmail.com> (10e6bc5)
Replaced deprecated .apply(keep_best) pattern
docs: update MIEB contributing guide for MTEB v2 AbsTask structure (#3787)
docs: update MIEB contributing guide for MTEB v2 AbsTask structure
Update docs/mieb/readme.md
Update docs/mieb/readme.md (fb53f57)
fix: Add model_type in model_meta for all models (#3751)
Add model_type in model_meta for all models
added literal for model_type
update jina embedding model type
Added model_type to from_cross_encoder() method
update test
change location in model_meta to pass test
update late_interaction model and fix test
update late_interaction for colnomic models
update test
Update mteb/models/model_meta.py
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>
fix naming
remove is_cross_encoder field and convert it into property
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com> (522eecc)
Optimize validate filter scores only (#3792)
feat: add detailed timing logs to leaderboard initialization
Add comprehensive timing information to track performance of each step in the leaderboard building process:
Each step logs start and completion times with elapsed duration to help identify performance bottlenecks during leaderboard initialization.
Implemented 3 high-impact optimizations to reduce benchmark processing time:
Cache get_model_metas() calls using @functools.lru_cache
Replace pandas groupby().apply() with vectorized operations
Cache version string parsing with @functools.lru_cache
Performance improvements:
This significantly improves leaderboard startup time by reducing the benchmark processing bottleneck.
perf: optimize validate_and_filter_scores filtering logic
Update mteb/results/task_result.py
Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com>
Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com> (d3bc4cc)
Add leaderboard timing logs and join_revisions() speedups (#3790)
feat: add detailed timing logs to leaderboard initialization
Add comprehensive timing information to track performance of each step in the leaderboard building process:
Each step logs start and completion times with elapsed duration to help identify performance bottlenecks during leaderboard initialization.
Implemented 3 high-impact optimizations to reduce benchmark processing time:
Cache get_model_metas() calls using @functools.lru_cache
Replace pandas groupby().apply() with vectorized operations
Cache version string parsing with @functools.lru_cache
Performance improvements:
This significantly improves leaderboard startup time by reducing the benchmark processing bottleneck.
Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com>
This resolves the issue where tasks with different original revisions that mapped to the same cleaned value would be grouped together non-deterministically.
refactor: use default lru_cache maxsize for _get_cached_model_metas
refactor: remove optimization markers from comments
Apply suggestion from @isaac-chung
Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com> (57a4b0c)
fix: legacy clustering processing
fix clustering processing (eee248a)
feat: Added get_benchmark_result() to BenchmarkResults to obtain a benchmark table
feat: Added get_benchmark_result() to BenchmarkResults to obtain a benchmark table (#3771)
Update BenchmarkResults to output results of benchmark
added score column and correct TYPE_CHECKING
address comments
address comments
fix import
fix tests
fix tests
change BenchmarkResults to Pydantic dataclass
change benchmark to pydantic dataclass
fix tests
fix model
fix
lint
remove future
fix after review
add test
reapply comments from review
remove mock benchmark
add documentation
added actual results
Update docs/usage/loading_results.md
Co-authored-by: Kenneth Enevoldsen <kenevoldsen@pm.me>
Co-authored-by: Roman Solomatin <36135455+Samoed@users.noreply.github.com>
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>
Co-authored-by: Kenneth Enevoldsen <kenevoldsen@pm.me> (92fa705)
Renaming model forward_passages to forward_images (0a0e398)
dataset: Add Turkish Constitutional Court violation classification task (#3777)
Add Turkish Constitutional Court violation classification task
Format Turkish task files with ruff
Update mteb/tasks/classification/tur/turkish_constitutional_court.py
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>
Fix dataset revision to pinned commit hash
Remove results directory from task PR
task dataset path updated
load data deleted
ruff formatted
bibtex fix
descriptive stats added
descriptive stats file name fix
Fix dataset duplicates/overlap and bibtex
Fix dataset duplicates, remove validation, and clean BibTeX
Format TurkishConstitutionalCourtViolation task with ruff
upd statistics
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>
Co-authored-by: Roman Solomatin <36135455+Samoed@users.noreply.github.com> (d3ce3d5)
model: add c2llm & fix f2llm (#3782)
add c2llm & fix f2llm
Update c2llm via sentence_transformer_wrapper
add c2llm languages
format
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com> (9b00964)
dataset: add SQuADKorV1Retrieval task for Korean (#3779)
dataset: add SQuADKorV1Retrieval task for Korean
Add Korean SQuAD v1.0 retrieval task based on the yjoonjang/squad_kor_v1 dataset.
This dataset provides:
Statistics:
fd339dd)update reference website of Seed1.6-embedding-1215 (f5182fb)
fix pre commit (#3775) (1a7b4cf)
fix_mod_embedding_oom (#3774) (02bdcd8)
model: added Bytedance/Seed1.6-embedding-1215 (#3760)
add model: Bytedance/Seed1.6-embedding-1215
make lint
fix typo
remove get_text_embedding() and get_image_embedding(), only reserve get_fused
move "max_tokens" and "available_embed_dims" into model implementations
fix typo
Apply suggestions from code review
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com> (0d73c3d)
Added revision parameter for results repo (#3764)
Added revision parameter for results repo
fix clone_cmd
Added depth parameter in clone command (e8c02b1)
fix: Added missing model citations
fix: Added missing model citations (#3762)
Added model citations
Apply suggestions from code review
Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com>
fix syntax and remove citation of 1 model
Delete scripts/model_citations_report.csv
add missing "@" in citations
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>
Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com> (4f7016b)
fix bflhc/MoD-Embedding loader (#3759)
Rename prompt_dicts to prompts_dict in mod_models.py
move to meta (1113e56)
fix: instruction selection for InstructSentenceTransformerModel
fix: instruction selection for InstructSentenceTransformerModel (#3758)
fix doc
fix get_instruction (b0f677c)
fix prompt_dict to prompt_dicts (#3756) (e9e7cd2)
model: mod_models.py (#3749)
Add new model: mod_models.py
Update mteb/models/model_implementations/mod_models.py
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>
The instruction_template function already adds "Instruct:" and "Query:" prefixes automatically, so these should not be duplicated in the PREDEFINED_PROMPTS dictionary values.
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com> (ad4a8b6)
remove transformers version (`80fef47`)
fix: Add optional "per language table"
fix: Add optional "per language table" (#3617)
fix: external links in hf space
try to fix model name
feat: Add per-language table creation
refactor: Simplify per-language table creation by removing unnecessary comments and print statements
feat: Enhance per-language table functionality
feat: Enhance per-language table functionality with support flag and styling improvements
feat: redo
feat: refacto language filtering support
feat: update per-language table to support 'all' option and improve styling for wide tables
feat: simplify per-language table styling by rounding values for wide tables
Update mteb/benchmarks/benchmarks/benchmarks.py
Co-authored-by: Kenneth Enevoldsen <kenevoldsen@pm.me>
feat: enhance language view condition to include 'all' option in per-language table
fix
refactor: update button display options in per-language table styling
fix: check emptiness before further analysis
fix: set column width for task and language tables
Co-authored-by: Roman Solomatin <36135455+Samoed@users.noreply.github.com>
Co-authored-by: Kenneth Enevoldsen <kenevoldsen@pm.me> (46612af)
add PawanEmbd-68M model metadata (#3703)
add PawanEmbd-68M model metadata
add adapted_from & lint file (05e8b79)
Adds baseline model for MTEB(Scandinavian) (XLM-R models are also relevant elsewhere)
Adds baseline model for MTEB(Scandinavian) (XLM-R models are also relevant elsewhere)
closes #3679
closes #3678
closes #3677 (fd395fd)
fix: Add hamming score to multilabel classification
fix: Add hamming score to multilabel classification (#3700)
test: add mixed performance case for hamming_score
add hamming score to multilabel classification metrics
🤖 Generated with Claude Code
Co-Authored-By: Claude Sonnet 4 <noreply@anthropic.com>
move to eval folder
remove try except
Co-authored-by: Claude Sonnet 4 <noreply@anthropic.com> (bc2e24e)
dataset: Replace JMTEB(v2) with MTEB(jpn,v1) for general-purpose Japanese evaluation (#3699)
Replace JMTEB(v2) with MTEB(jpn,v1) for general-purpose Japanese evaluation
fix get_benchmark as Japanese benchmark updated
Replace MTEB(jpn, v1) with JMTEB(v2)
Revert "Replace MTEB(jpn, v1) with JMTEB(v2)"
This reverts commit b49f67b4409becbc81e65318b5f79ea195e40304.
keep MTEB(jpn, v1) as legacy
revert class name
revert change (11a2e9e)
dataset: RuSciBenchBitextMining small dataset cleanup (#3690)
RuSciBenchBitextMining small dataset cleanup
RuSciBenchBitextMining new version
Add RuSciBenchBitextMining.v2 stats (cb4b3ed)
model: add NbAiLab/nb-bert-large and NbAiLab/nb-bert-base (#3688)
Initial plan
Add NbAiLab/nb-bert-large model to MTEB
Co-authored-by: KennethEnevoldsen <23721977+KennethEnevoldsen@users.noreply.github.com>
Apply suggestions from code review
add models
fix naming and linting
Co-authored-by: copilot-swe-agent[bot] <198982749+Copilot@users.noreply.github.com>
Co-authored-by: KennethEnevoldsen <23721977+KennethEnevoldsen@users.noreply.github.com>
Co-authored-by: Kenneth Enevoldsen <kennethcenevoldsen@gmail.com> (97834cb)
Add dataset task filter (#3685)
init task filter
fix typehint
split filters and pipelines
change filters logging (3afd6f8)
fix: Don't pass embed dim to openai/text-embedding-ada-002
openai/text-embedding-ada-002 (#3689)It does not support embed dim.
closes #3687
tested imp. with the below code to ensure that it works with both old and new models
import mteb
mdl = mteb.get_model("openai/text-embedding-ada-002")
task = mteb.get_task("STSBenchmark", eval_splits=["test"])
mteb.evaluate(mdl, task, cache=None)
mdl = mteb.get_model("openai/text-embedding-3-small")
mteb.evaluate(mdl, task, cache=None)
Co-authored-by: Kenneth <kennethenevoldsen@gmail.com> (25af7b9)
Merge branch 'main' of https://github.com/embeddings-benchmark/mteb (`c04effb`)
Merge branch 'main' of https://github.com/embeddings-benchmark/mteb (c04effb)
Fix: fix Performance per Model Size display issues (#3691)
fix - filtering Mean(Task)=0 (86e9235)
model: Kowshik24/bangla-sentence-transformer-ft-matryoshka-paraph… (#3661)
Add Model: Kowshik24/bangla-sentence-transformer-ft-matryoshka-paraphrase-multilingual-mpnet-base-v2
Updated the memory usage and public training code
fix: correct public_training_code syntax and initialize training_datasets as an empty set
fix: update import path for ModelMeta in kowshik24_models.py (a39c011)
model: add two Japanese models sarashina-embedding-{v1, v2}-1b (#3683)
Add sarashina embedding models
fix public training data info
fix prompt setting for v2 model
Use InstructionSentenceTransformerModel wrapper for instruction model
Update mteb/models/model_implementations/sarashina_embedding_models.py
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com> (db74fc5)
fix: Change git clone results depth to 1
change git clone depth (874ddf2)
resolve search wrapper warning (46fb44c)
model: add ruri-series Japanese embedding models (#3684)
Add ruri models
fix typo and add trust_remote_code when necessary (5de8d98)
dataset: Add hebrew v3 (#3607)
add hebrew v3
add to init
add v4
v4 -> v3 (3982f33)
dataset: Add new benchmark JMTEB(v2) (#3660)
Add benchmark: JMTEB(v2)
fix bibtex format
Fix bibtex format, description and contacts for JMTEB v2
Fix bib of JMTEB
Fix dataset version (d14923c)
fix: get_model now correctly assumed SentenceTransformer
get_model now correctly assumed SentenceTransformer (#3673)fix: get_model now correctly assumed SentenceTransformer if
unkown
fixes #3670 (21cf638)
ci: remove unnecessary items on disk
ci: remove unnecessary items on disk (#3664)
remove unnecessary items on disk
don't remove on windows (2c51e3a)
ci: Update broken links in pull request template (#3656) (095803c)
Updated image sources in README to use raw links. This ensures that the readme has images in pypi (8e2929d)
663bb87)Removed unnecessary docker system prune command from the workflow.
Close https://github.com/embeddings-benchmark/mteb/issues/3669 (6d92cbc)
fix: Reduce the number of decimals for the number of parameters (#3668) (397e0dd)
fix: cohere import error (#3665)
fix cohere (745ad84)
fix: Remove "Unknown" for int on leaderboard causing them to be unsortable (#3653)
fix: Remove "Unknown" for int on leaderboard causing them to be unsortable
Fixes #3579
instruction_template as a kwargs (#3654)I assume it is intended to use instruct_wrapper
fix some formatting
fixed
update
removed incorrect text
fix: GoogleTextEmbeddingModel were given multiple model_name (#3658)
fix: GoogleTextEmbeddingModel were given multiple model_name
fix based on comments
model: Added dfm-sentence-encoder models (#3655)
fix: Linq-Embed-Mistral loader does not take instruction_template as a kwargs
I assume it is intended to use instruct_wrapper
94979ca)Model: add e5-nl models (#3646)
e5-nl models
e5-nl moved to a new file (f3e5763)
fix: fix display for task information and improve UI for benchmark filtering
fix: fix display for task information and improve UI for benchmark filtering (#3629)
fix: bump gradio to v6
fixes #3601
more fixes
refactor themes (might need some more refactors)
feat - issue #3569
feat - issue #3569
feat - issue #3569
feat - issue #3616
feat - CheckboxGroup
feat - ruff check
bump gradio
fix - task options bug
fix - fix task filter condition
Co-authored-by: Kenneth Enevoldsen <kennethcenevoldsen@gmail.com> (a882295)
dataset: ruMTEB v1.1 (#3631)
ruMTEB_v1
ruMTEB_v1
ruMTEB_v1
ruMTEB_v1
ruMTEB_v1.1 description and contacts
Update benchmarks.py
edited MTEB(rus, v1) display name (c228707)
fix: colpali_training_set & updated JinaVDR and ViDoRe tasks annotation
fix: colpali_training_set & updated JinaVDR and ViDoRe tasks annotation (#3636)
add adapted annotation
fix training set annotation
update nemotriever datasets (1ce74c2)
fix: add flag to run public only tasks (#3563)
run public only tasks
pass to evaluate
add tests
update cli
add task error
fix metadata name
fix tests
Update mteb/evaluate.py
Co-authored-by: Kenneth Enevoldsen <kennethcenevoldsen@gmail.com>
remove from cli
add tests
fix exception
fix renaming
raise error if public_only False
rollback co2
Apply suggestions from code review
Co-authored-by: Kenneth Enevoldsen <kennethcenevoldsen@gmail.com>
Co-authored-by: Kenneth Enevoldsen <kennethcenevoldsen@gmail.com>
remove test
fix test
Co-authored-by: Kenneth Enevoldsen <kennethcenevoldsen@gmail.com> (f905b68)
model: Add IEITYuan/Yuan-embedding-2.0-en model (#3630)
add
Update mteb/models/model_implementations/yuan_models_en.py
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>
Co-authored-by: Kenneth Enevoldsen <kenevoldsen@pm.me>
Co-authored-by: Kenneth Enevoldsen <kenevoldsen@pm.me>
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>
Co-authored-by: Kenneth Enevoldsen <kenevoldsen@pm.me> (d31f1cd)
model: Add tomoro-colqwen3-embed embedding models (#3627)
feat(colqwen3): add wrapper and model metadata
feat(colqwen3): update ColQwen3Wrapper to use bfloat16 and enhance similarity scoring
fix(colqwen): require transformers>=4.57 and refresh metadata, set revision
refactor(colqwen): reorder wrappers and metadata definitions for clarity
chore(colqwen): set release date for tomoro colqwen3 8b
chore(colqwen): remove unused methods and fix lint errors
feat(colqwen3): add fused image-text encoding path
refactor(colqwen): unify encode method with get_fused_embeddings
chore(colqwen): update encoding progress message
chore(colqwen): update model revisions for colqwen models
docs(colqwen): update train data annotation
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>
Co-authored-by: tankm <kyemin.tan@tomoro.ai>
Co-authored-by: Huang Xin <hxssg1124@gmail.com>
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com> (71ac96c)
model: euler legal embedding (#3640)
add model implementation
modify the correct model path
Update mteb/models/model_implementations/euler_models.py
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>
Co-authored-by: Roman Solomatin <36135455+Samoed@users.noreply.github.com> (a1fbdb9)
model: jina-reranker-v3 (#3645)
add: jina-reranker-v3
Update mteb/models/model_implementations/jina_models.py
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>
fix: remove useless check
fix: use only one class which inherits from CrossEncoderWrapper
Update mteb/models/model_implementations/jina_models.py
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>
fix: remove unused import
Update mteb/models/model_implementations/jina_models.py
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>
fix: ruff format
add: cross_encoder label
add: datasets
Update mteb/models/model_implementations/jina_models.py
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com> (1ecd892)
dataset: MultiLongDocReranking (#3642)
Add MultiLongDocReranking dataset
Add descriptive stats for MultiLongDocReranking
Fix information
reformat
fix hash
Update multi_long_doc_reranking.py
lint
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>
Co-authored-by: Roman Solomatin <36135455+Samoed@users.noreply.github.com> (4f7774d)
Remove deprecated torch_dtype (`7e2fa98`)
ci: Add HF_TOKEN to dataset loading and merge CI (#3622) (4ffef40)
ci: update action versions (#3623)
update action versions (bcf4e82)
docs: Update "speeding up"-section to include bumping version (#3634)
Update "speeding up"-section to include bumping version
Also corrected grammar and punctuation for clarity.
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com> (392186f)
feat: add search encoder backend (#3492)
add search backend
make faiss optional
fix import
use faiss in reranking
add support for multiple similarities
remove check
update index check
rename and move files
add missing files
fix import
rename to add documents
make index backend optional
remove streaming backend
fix test
add doc
add memory
add API to docs (4ed7ef4)
fix model memory (73168c6)
model: Add eager-embed embedding model (#3602)
Add eagerembed model
Address CR comments to make code cleaner
Refactor code to remove unnecessary dataloader. Use prompt_type
Update model revision. Move tokenizer config to file
Add support for unified encoding
Fix vidore2 retrieval language filter
Remove unused methods. Fix vidore2 filtering
Remove deprecated torch_dtype (7e2fa98)
# 2.2.2 (2025-11-25) ## Fix * fix: vidore loading (#3618) fix vidore loading (`ca8e7c4`)
fix: Avoiding stating warning if what is logged is not a warning
Stating warning here might lead devs to think it should be considered a warning. It is just a info message.
Also changes to that "..." appears at the end (b0d6c7b)
feat: make STS and PairClassification asymmetric
feat: make STS and PairClassification asymmetric (#3568)
make STS and PairClassification asymmetric
update logging
make terra v2
fix
fix tests (5010468)
model: Add IEITYuan/Yuan-embedding-2.0-zh model
fix: Correcting the cohere lstrip bug in cohere
fix: Correcting the cohere lstrip bug in cohere (#3610)
Correcting the cohere lstrip bug
Update mteb/models/model_implementations/cohere_models.py
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com> (398b31b)
add note for BUCC tasks about using train split; (3af54eb)
add missing citation for Vietnamese retrieval datasets (#3608)
add missing citation
Update mteb/tasks/retrieval/vie/green_node_table_markdown_retrieval.py
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com> (027bc17)
tests: Add tests for dataset quality (#3603)
tests: Add test to prevent low-quality tasks
This tests ensures a minimum quality of task for future submissions
Currently it tests for:
I suspect we can likely add many more tests to this in future PRs
collect all errors before raising
Update tests/test_tasks/test_task_quality.py
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com> (9b898ea)
fix: improve messages for running missing splits
It now looks like:
INFO:mteb.evaluate:Found existing results for MassiveIntentClassification, only running missing splits (subset): {'validation': ['hy', 'jv', 'fr', 'hu', 'tl', 'az', 'zh-CN', 'ms', 'ar', 'pt', 'ja', 'tr', 'hi', 'lv', 'sw', 'nl', 'en', 'ko', 'mn', 'zh-TW', 'kn', 'am', 'he', 'my', 'sq', 'vi', 'fi', 'ru', 'cy', 'it', 'pl', 'el', 'de', 'te', 'af', 'ro', 'sl', 'fa', 'ur', 'ml', 'is', 'bn', 'es', 'km', 'ka', 'th', 'ta', 'id'], 'test': ['hy', 'jv', 'fr', 'hu', 'tl', 'az', 'zh-CN', 'ms', 'ar', 'pt', 'ja', 'tr', 'hi', 'lv', 'sw', 'nl', 'en', 'ko', 'mn', 'zh-TW', 'kn', 'am', 'he', 'my', 'sq', 'vi', 'fi', 'ru', 'cy', 'it', 'pl', 'el', 'de', 'te', 'af', 'ro', 'sl', 'fa', 'ur', 'ml', 'is', 'bn', 'es', 'km', 'ka', 'th', 'ta', 'id']}
``` ([`5d7b78b`](https://github.com/embeddings-benchmark/mteb/commit/5d7b78bd84437f14f01149d86e362b1016d4e88f))
Found a couple of issue where the cache is not hit
fix: issues on cache hits (#3558)
fix: issues on cache hits
Found a couple of issue where the cache is not hit
overwrite_strategy == "only-missing" and overwrite_strategy == OverwriteStrategy.ONLY_MISSING would never be met as it is never both (so we always rerun all splits)remote vs remote/results), which means that it is never hit.add test
format
fix: Overwrite / ignore existing results if not mergeable
add fest from review
fix match
Co-authored-by: Kenneth <kennethenevoldsen@gmail.com> (9d7d4df)
update gradio (09021df)
fix: typo for attn_implementation kwargs in jasper models (#3592)
add prompt dict for Jasper_Token_Compression_600M
use sdpa for Jasper_Token_Compression_600M
remove JasperTokenCompressionLoader
fix typo for attn_implementation (0d33bd3)
add prompt dict for Jasper_Token_Compression_600M (#3587)
add prompt dict for Jasper_Token_Compression_600M
use sdpa for Jasper_Token_Compression_600M
remove JasperTokenCompressionLoader (5bca292)
tests: Added test for ensuring training datasets can be computed (#3566)
fix: Fix adapted from points to the models itself
We should probably add a test to prevent this in the future.
Added test
Update tests/test_models/test_model_meta.py (4636b24)
utilize max_seq_length (`f7b481e`)
max_seq_length (#3588)utilize max_seq_length (f7b481e)
add training code and citation for Jasper_Token_Compression_600M (#3584) (76be959)
model: Add spartan8806/atles-champion-embedding model (#3575)
Add ATLES Champion Embedding model wrapper
Move ATLES Champion embedding wrapper into mteb/models and export it
Update spartan8806_atles_champion.py with current date and time
Update the file init.py with new content
Update spartan8806_atles_champion.py with the latest content
fix: Use sentence_transformers_loader per MTEB guide
fix: Move model to model_implementations and clean up per @Samoed feedback
fix: Remove old files and revert init.py per @Samoed feedback
fix: Remove old model file from wrong location
fix: Remove stray root file
fix: Restore proper line breaks in model file
fix(lint): correct module docstring in model file
feat: Add training dataset info to ModelMeta
fix: Update revision to commit hash and add training datasets
fix: Restore file with proper newlines (encoding fix)
fix: Update training_datasets format per @Samoed feedback
Update mteb/models/init.py
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>
Update mteb/models/model_implementations/spartan8806_atles_champion.py
fix: Add required metadata fields (memory_usage_mb, open_weights, public_training_code, public_training_data)
fix: Correct release_date to November (2025-11-16)
fix: Update release_date to 2025-11-15 (actual training date)
fix: Add adapted_from and convert public_training_code/data to URLs
Update mteb/models/model_implementations/spartan8806_atles_champion.py
Co-authored-by: spartan8806 <spartan8806@users.noreply.github.com>
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com> (723fd98)
fix: benchmark references links
fix: benchmark references links (#3560)
fix: external links in hf space
try to fix model name
Co-authored-by: Roman Solomatin <36135455+Samoed@users.noreply.github.com> (07f1e6e)
fix: Set default input_type for VoyageMultiModalModelWrapper
fix: Set default input_type for VoyageMultiModalModelWrapper (#3567)
fix: Set default input_type for VoyageMultiModalModelWrapper
Apply suggestions from code review
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com> (711e7cb)
fix: Fix adapted from points to the models itself
We should probably add a test to prevent this in the future. (08b8ec7)
fix: MTEB-NL switches to v2 datasets with new prompts
fix: resolve hash randomization in retrieval task ID generation
This commit fixes non-deterministic query ID assignment in three retrieval tasks caused by Python hash randomization when using enumerate(set()).
Affected tasks:
0c4f099)fix: Added leaderboard Vidore V3
fix: Added leaderboard Vidore V3 (#3542)
feat:initial leaderboard proposal
feat: update summary table for ViDoRe V3 to reflect Document Understanding tasks
refactor: update leaderboard references
fix: update VISUAL_DOCUMENT_RETRIEVAL to use VidoreBenchmark
fix: update JinaVisualDocumentBenchmark summary table creation method
fix: add VisualDocumentRetrieval to previous benchmark names
fix: remove JinaVisualDocumentBenchmark
Co-authored-by: Antoine Edy <antoine.edy@illuin.tech> (ab390ce)
model: Added emillykkes scandi models (#3521)
model: Added EmbeddingGemma-Scandi-300m
from sentence_transformers import SentenceTransformer
# test that is loads with sentence transformers
model = SentenceTransformer("emillykkejensen/EmbeddingGemma-Scandi-300m")
# OSError: Can't load the model for 'emillykkejensen/EmbeddingGemma-Scandi-300m'. If you were trying to load it from 'https://huggingface.co/models', make sure you don't have a local directory with the same name. Otherwise, make sure 'emillykkejensen/EmbeddingGemma-Scandi-300m' is the correct path to a directory containing a file named pytorch_model.bin, tf_model.h5, model.ckpt or flax_model.msgpack.
Can't get the model to load
added other models
fix remaining metada issue (1ad433f)
ci: fix false positive check in typos (#3540) (`b2b9599`)
b2b9599)make single line descriptions (fe83e27)
fix: Pass encode kwargs in all dataloaders (#3548)
pass all encode kwargs in dataloaders
fix tests
fix tests (eaec6cb)
Model : Tarka Embedding 350M V1 (#3549)
Add : Tarka Embedding 350M V1
removing custom wrapper
removing device map and minor changes
minor fix (27c10e9)
Add concurrency to tests (#3543)
add concurrency to tests (57e179d)
dataset: Benchmark/VidoreV3 (#3514)
vidore_3_tasks
update init
add benchmark
fix oopsies + linting
add private tasks + apply some reco
update private datasets path
refactor: update Vidore3 retrieval classes and paths for improved organization
update descriptions for public datasets
fix loading error
sort imports with updated names
update: Vidore3 retrieval references and citations
use batched image processing
don't process if
fix: update dataset revisions for Vidore3 retrieval classes + remove custom load_data methods
add private datasets
fix private test
feat: added descriptive statistics for public ViDoRe V3 datasets
format benhcmark citation
feat: enhance ViDoRe V3 benchmark with detailed description and update task domains
Update mteb/benchmarks/benchmarks/benchmarks.py
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>
feat: better leaderboard
fix:update description and sample creation
Update mteb/leaderboard/benchmark_selector.py
Co-authored-by: Antoine Edy <antoine.edy@illuin.tech>
Co-authored-by: Roman Solomatin <36135455+Samoed@users.noreply.github.com>
Co-authored-by: antoineedy <antoineedy@outlook.fr>
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>
Co-authored-by: Kenneth Enevoldsen <kenevoldsen@pm.me> (1a02edb)
d98a008)prompts for vabb_clustering fixed
fix: MTEB-NL prompts (#3516)
adding prompts to MTEB-NL
prompts for vabb_clustering fixed
arguana-nl fixed
update retrieval/nld/init
descriptive stats for v2
prompts in Dutch for MTEB-NL
descriptions added to the Dutch prompts for MTEB-NL
Update mteb/tasks/retrieval/nld/argu_ana_nl_retrieval.py
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com> (8f3f806)
fix: Add support for python 3.14 (#3450) (`632a83a`)
632a83a)model: add kalm_models.py ModelMeta (#3519)
model: add kalm_models.py ModelMeta
fix: model revision
fix: recover ESCIReranking for train data
fix: recover the original auto-generated kalm_training_data
fix: restore docs logo files
Co-authored-by: xinshuohu <xinshuohu@tencent.com> (28f9c54)
model: add EvoQwen2.5-VL-Retriever (18ff6f3)
fix: materialize corpus id to speed up evaluation
reupload winogrande (`fe43f73`)
fix: aggregated task evaluation
fix aggregated task evaluation (5eae04c)
fix: remove set_float32_matmul_precision
set_float32_matmul_precision (#3509)remove set_float32_matmul_precision (ce07dfd)
fix ReasonIR instruction (#3506) (8189108)
Update links in leaderboard description (#3503)
Update links in leaderboard description
Resolve 404's for leaderboard contribution docs
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com> (976fadf)
dataset: Add MTEB-NL to the leaderboard (#3489)
adding WebFAQRetrieval to MTEB-NL
adding MTEB-NL to the leaderboard
MTEB-NL renamed to Dutch in the leaderboard (da9feef)
add benchmark to minor releases (`b649e6f`)
ci: New release workflow (#3448)
add sematic release
add main release
fix: update action
merge semantic with release
run release after tests passed
add benchmark to minor releases (b649e6f)
Update links in readme (b0b0e7d)
docs: fix broken links (#3483)
docs: fix broken links
Update docs/usage/defining_the_model.md
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com> (dfd516a)
fix: Rollback to semantic release (#3502)
rollback to semantic release
update pyproject.toml (1325328)
fix: simplify release (#3494)
fix release CI
skip if None
simplify workflow
remove ifs
make action condition as and instead or (4484112)
fix: add prompts to hardnegative tasks (#3469)
add prompts to hardnegative tasks
fix superseded_by
add stats
add more description (7b7bdd0)
fix: verify languages during filtering (#3472)
verify lang code
add support for language script code
fix
add check for language names
update language check
remove problematic tasks (799b869)
fix: release CI (#3493)
fix release CI
skip if None (21223ed)
fix: top_k document selection in two stage reranking (#3486)
fix topk
update test for correct top_k selection (16ae6ff)
fix: task metadata was not passed in Jina implementation (#3485)
Update jina_models.py
Add missing keyword argument
Add missing keyword arguments (ea1bac1)
fix: qrels selection (31c8329)
Add spell checker (#3476)
add spell checker
remove arg (9e683fe)
Correcting the VoyageAI multimodal code (#3491) (cf81dd1)
Remove skip for tasks (#3475) (fcd3b71)
Descriptive stats, MIRACLVisionRetrieval (#3473) (5b92f73)
Update text_segments.py (#3474) (181490f)
Almost final descriptive stats (#3463)
fix tasks
final statistics
remove persiantexttone, was renamed to SynPerTextToneClassification
remove duplicated tasks statistics
fix category
run test on all tasks
remove image/text: None
try to fix none in batch
add vdr statistics (0ead029)
fix miracl loading (#3466) (179702e)
docs: Ignore overview docs (#3456) (`5e903b1`)
5e903b1)feat: Add mteb nl (#3464)
classification tasks for MTEB-NL
clustering tasks for MTEB-NL
ml classification tasks for MTEB-NL
pair classification tasks for MTEB-NL
sts tasks for MTEB-NL
cherry-picked commits from MTEB-NL for tasks
descriptive stats for MTEB-NL
Update mteb/tasks/classification/nld/DutchColaClassification.py
Co-authored-by: Kenneth Enevoldsen <kenevoldsen@pm.me>
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>
imports fixed for MTEB-NL
SICKNLPairClassification the HF location fixed
DutchColaClassification the HF location fixed
DutchGovernmentBiasClassification the HF location fixed
XLWICNLPairClassification the HF location fixed
BBSARDNLRetrieval the HF location fixed
LegalQANLRetrieval the HF location fixed
cls and mlcls deduplication for MTEB-NL
SICK-NL deduplication for MTEB-NL
pep 8 to MTEB-NL (2)
bBSARDNLRetrieval deduplicated
imports fixed for pair classification MTEB-NL
Co-authored-by: Kenneth Enevoldsen <kenevoldsen@pm.me>
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com> (4f9f157)
__init__ (#3458)update root init (f8c07ff)
fix license (6fbe482)
fix: vdr category (#3465) (`4b384bb`)
4b384bb)Add more statistics (#3462)
add statistics
add PatchCamelyonZeroShot stats (93d638d)
model: Kalm model (#3461)
readd (5c111ea)
Add stats from fzoll (#3460)
Correcting the get_tasks filtering issue
Correcting the get_tasks filtering issue
Descriptive stats, last part
Descriptive stats, last part
Co-authored-by: fzoll <fodizoltan@gmail.com> (4ef51f2)
To avoid the text focus (`0a6fe95`)
To avoid the text focus (0a6fe95)
Don't shorten embedding size (be20185)
fix: Roll back setting OMP_NUM_THREADS for clustering (#3444)
test: disable flaky test
Added issue to readd #3441
This rolls back #3400 as it didn't work (as shown in most recent CI on main) (38e7bc7)
6eab159)fix: speedup retrieval computation
fix: speedup retrieval computation (#3454)
speedup retrieval computation
lint (01f3a19)
Automatically generated by python-semantic-release
ci: Updating docs ci (#3445)
feat!: Updating to v2
2.0.0
Automatically generated by python-semantic-release
Co-authored-by: semantic-release <semantic-release> (3af1aa0)
add citations to models (#3435)
fix: add citations to models
Update mteb/models/model_implementations/salesforce_models.py
Co-authored-by: Yongbin Choi <whybe.choi@gmail.com> (a04d78b)
Fix: Cache invalidation (#3393)
feat - issue #3381
feat - issue #3381
feat - add task_select.change (9b08e8c)
Merge statics (#3452)
Recalculating desc. stats, part 6 (#3451)
Correcting the get_tasks filtering issue
Correcting the get_tasks filtering issue
Descriptive stats, part 6
merge statisrics
Co-authored-by: fzoll <5575946+fzoll@users.noreply.github.com> (5e6542e)
48b11a7)Add english code retriever model by @fedor28 in https://github.com/embeddings-benchmark/mteb/pull/3302
docs/adding_a_benchmark.md by @whybe-choi in https://github.com/embeddings-benchmark/mteb/pull/3344Full Changelog: https://github.com/embeddings-benchmark/mteb/compare/1.39.7...2.0.1
fix: Change language for task SlovakMovieReviewSentimentClassification (#3296) (`0a902a3`)
fix: add prompt for MIRACLRetrievalHardNegatives
fix: add prompt for MIRACLRetrievalHardNegatives (#3266)
add prompt for MIRACLRetrievalHardNegatives
add MIRACLRetrievalHardNegatives.v2
Update mteb/tasks/Retrieval/multilingual/MIRACLRetrieval.py
Co-authored-by: Kenneth Enevoldsen <kenevoldsen@pm.me>
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>
Co-authored-by: Kenneth Enevoldsen <kenevoldsen@pm.me> (9b6f320)
fix: Add retry and token counting in Cohere models
fix: Add retry and token counting in Cohere models (#3253)
Retry and token counting in Cohere models
Retry and token counting in Cohere models
Retry and token counting in Cohere models
Co-authored-by: Roman Solomatin <36135455+Samoed@users.noreply.github.com> (e81c94f)
fix: resolve flash-attention dependency issue
fix: resolve flash-attention dependency issue (#3265)
fix: Only pin model name and rank
Currently we pin 3 columns, this makes it hard or impossible to view on phones. The 3rd column is also no longer garuanteed as RTEB leaderboard does not use the zero-shot column
This has been tested and works.
fixed Resolve flash-attention dependency issues
Fixes #3240 (1e29385)
fix: Only pin model name and rank
Currently we pin 3 columns, this makes it hard or impossible to view on phones. The 3rd column is also no longer garuanteed as RTEB leaderboard does not use the zero-shot column (58a81a9)
fix: Move zero-shot percentage calculation to the end of summary (#3231)
Refactor: Move zero-shot percentage calculation to the end of summary table creation which only apply to RTEB table.
Update RTEB benchmark name from "RTEB(beta)" to "RTEB" for consistency in display.
feat - RTEB(beta)
feat - remove Zero-shot
Co-authored-by: ethan <smiletoye@gmail.com> (65829bd)
model: Add ReasonIR (#3221)
model: Add ReasonIR
Update mteb/models/reasonir_model.py
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>
Co-authored-by: Niklas <n.muennighoff@gmail.com>
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>
Co-authored-by: Niklas <n.muennighoff@gmail.com> (f2504bd)
fix bm25 on small datasets (#3261) (237d8dc)
Update tasks & benchmarks tables (c8ae52c)
Added Japanese to Retrieval (#3252)
feat - add Japanese
feat - use mteb.get_benchmark
fix - 3.9 test error
Revert "fix - 3.9 test error"
This reverts commit 6bfee53cff48304cc22d8248aa275dcc9e385475.
fix - 3.9 test error (53b1c29)
Fix AbsTaskTextRegression task (#3257)
Fix AbsTaskTextRegression (08b98cd)
Update tasks & benchmarks tables (89bec7d)
Aggregate by subset for HUMEv1 (#3255)
aggregate by subset for HUMEv1 (36901eb)
docs: Update adding benchmark documentation
docs: Update adding benchmark documentation (#3229)
update adding_a_benchmark.md documentation
fix numbers (50aa4ac)
fix: Further specified macro-language code for Norwegian (#3228)
fix: Further specified macro-language code for Norwegian
"nor" is a macro-language code that covers bokmål and nynorsk (both norwegian), but this means that these datasets will be missed if using "nob" or "nno". Specifying it like this should allow this.
a2f7488)Update tasks & benchmarks tables (810ae28)
Remove 'HUME(v1)' from leaderboard benchmark (#3236)
Remove 'HUME(v1)' from leaderboard benchmark
lint (e419b54)
Update tasks & benchmarks tables (9a606a0)
dataset: add human tasks and benchmark (#3214)
Human Subsets Tasks
Fixed Multilingual Classification Subset
linting
fix citations format
make lint
fix tests
remove human folder
fix relative imports
add adapted_from for all human subsets
fix pydantic errors
add benchmark object
make benchmark discoverable
bibtex test
Apply suggestion
Co-authored-by: Kenneth Enevoldsen <kennethcenevoldsen@gmail.com>
Co-authored-by: Kenneth Enevoldsen <kennethcenevoldsen@gmail.com>
rename & reupload
upd tests
upd tests again
add model
add benchmark to leaderboard
change branch of leaderboard
remove branch of load data
fix model meta path
make mteb importable
update repo
Update mteb/benchmarks/benchmarks/benchmarks.py
Update mteb/leaderboard/benchmark_selector.py
Update mteb/load_results/load_results.py
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>
Co-authored-by: Adnan El Assadi <aassadi22@ku.edu.tr>
Co-authored-by: Isaac Chung <chungisaac1217@gmail.com>
Co-authored-by: Kenneth Enevoldsen <kennethcenevoldsen@gmail.com>
Co-authored-by: AdnanElAssadi56 <115242814+AdnanElAssadi56@users.noreply.github.com> (48a01fc)
fix: Add submission references for RTEB
fix: Add submission references for RTEB (#3233)
fix: Add rteb submission references and improve descriptions.
Added evaluation request
added field for tasks (600c290)
feat: Officially include RTEB in the leaderboard
feat: Officially include RTEB in the leaderboard (#3222)
feat - adjust Rteb's Benchmark
feat - add blank
fix menu names
Update mteb/leaderboard/benchmark_selector.py
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>
moving around tasks
fix: Update RTEB summary columns (#3226)
fix(models): ensure prompt_type is passed to format_instruction (#3216)
1.38.58
Automatically generated by python-semantic-release
Adding Cohere's output_dimension and embedding_type parameter (#3204)
Adding Cohere's output_dimension and embedding_type parameter Cohere's embed-v4 binary and int8
Correcting due to comments
dataset: add swedish cpc patent classifications to mteb (#3072)
feat: add swedish cpc patent classifications to mteb
fix: formatting and init imports
fix: update mteb task according to feedback
fix: perform citation and code formatting
fix: add train and test split for both datasets
fix: AttributeError in ColPaliEngineWrapper similarity method (#3177)
fix: delete kwargs for similarity score in ColPaliEngineWrapper for method behavior
chore: fix colpali_models similarity handle device
Update tasks & benchmarks tables
1.38.59
Automatically generated by python-semantic-release
fix: prevent EOS token truncation (#3218)
fix(models): prevent EOS token truncation for BMRetriever
refactor(models): refactor tokenizer setup in InstructSentenceTransformerWrapper
fix(models): correct eos token handling in BMRetrieverWrapper
1.38.60
Automatically generated by python-semantic-release
Update giga embeddings (#3210)
update giga embeddings
update giga embeddings
3b-september-2025
fixed
lint
Update mteb/models/ru_sentence_models.py
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>
change revision due to flash-attn dependency
change apply_instruction_to_passages
Co-authored-by: Kolodin Egor <eikolodin@sberbank.ru> Co-authored-by: Roman Solomatin <36135455+Samoed@users.noreply.github.com> Co-authored-by: Roman Solomatin <samoed.roman@gmail.com> Co-authored-by: Неизвестный Пользователь722497 <dolegosmirnov@sberbank.ru>
fix: Refactor split create_tables into static Benchmark methods (#3126)
feat - Split create_tables into static Benchmark methods
feat - format
Update mteb/leaderboard/table.py
Co-authored-by: Kenneth Enevoldsen <kenevoldsen@pm.me>
feat - remove search query;take benchmark result as input;addressing the circular import,
feat - format
Update mteb/benchmarks/benchmark.py
Co-authored-by: Kenneth Enevoldsen <kenevoldsen@pm.me>
Co-authored-by: Kenneth Enevoldsen <kenevoldsen@pm.me>
feat - use to_dataframe;clean table.py;move creat_table
feat - fix circular import
feat - clean-up
feat - format
Co-authored-by: Kenneth Enevoldsen <kenevoldsen@pm.me>
Automatically generated by python-semantic-release
Adding another voyageai model
Update tasks & benchmarks tables
feat - filter_by_privacy
feat - add new fields for rteb part
feat - getattr
feat - adjust privacy filter logic
feat - enhance summary table column renaming and add 'is_public' field mapping
fix: remove unused 'is_public' attribute from TaskResult
Co-authored-by: Yongbin Choi <whybe.choi@gmail.com> Co-authored-by: semantic-release <semantic-release> Co-authored-by: fzoll <5575946+fzoll@users.noreply.github.com> Co-authored-by: Atheer <atheer2104@protonmail.com> Co-authored-by: Yong woo Song <ywsong.dev@kakao.com> Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com> Co-authored-by: Egor <31567312+ekolodin@users.noreply.github.com> Co-authored-by: Kolodin Egor <eikolodin@sberbank.ru> Co-authored-by: Roman Solomatin <36135455+Samoed@users.noreply.github.com> Co-authored-by: Roman Solomatin <samoed.roman@gmail.com> Co-authored-by: Неизвестный Пользователь722497 <dolegosmirnov@sberbank.ru> Co-authored-by: Kenneth Enevoldsen <kenevoldsen@pm.me> Co-authored-by: smile <smile@pinai.io> Co-authored-by: ethan <smiletoye@gmail.com>
removed show_rteb args
avoid defining function where we can just use the metadata
minor fixes
minor fixes
fix: Correct logic for filtering public tasks in ModelResult class (#3230)
Co-authored-by: ethan <smiletoye@gmail.com>
Co-authored-by: q275343119 <275343119@qq.com>
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>
Co-authored-by: 笑尿伊人 <44760272+q275343119@users.noreply.github.com>
Co-authored-by: Yongbin Choi <whybe.choi@gmail.com>
Co-authored-by: fzoll <5575946+fzoll@users.noreply.github.com>
Co-authored-by: Atheer <atheer2104@protonmail.com>
Co-authored-by: Yong woo Song <ywsong.dev@kakao.com>
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
Co-authored-by: Egor <31567312+ekolodin@users.noreply.github.com>
Co-authored-by: Kolodin Egor <eikolodin@sberbank.ru>
Co-authored-by: Roman Solomatin <36135455+Samoed@users.noreply.github.com>
Co-authored-by: Неизвестный Пользователь722497 <dolegosmirnov@sberbank.ru>
Co-authored-by: smile <smile@pinai.io>
Co-authored-by: ethan <smiletoye@gmail.com> (11f9c1d)
Update tasks & benchmarks tables (867105f)
Update tasks & benchmarks tables (65f29e6)
dataset: Add Software Issue Localization Datasets (#3178)
add software issue localization datasets
add software issue localization datasets
update and add multilingual datasets
fix citation format issues
Update mteb/tasks/Reranking/eng/SWEbenchVerifiedReranking.py
fix linting issues
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com> (e56e7c4)
model: Update Youtu embedding model (#3227)
add youtu models
add a blank line
fix the optional dependencies and lint the code
remove unused dependencies and reformat
revise prompt_type
update youtu_models
Co-authored-by: springxchen <springxchen@tencent.com> (0000ae2)
model: New qzmodel (#3211)
Update qzhou_models.py
Update qzhou_models.py
reformat script code
Update configuration
According to our new decision, the model name has been changed to "QZhou-Embedding-Zh".
Fix variable naming issues. (e299345)
Update tasks & benchmarks tables (7f5990a)
Extending the RTEB benchmark (#3223)
Adding another voyageai model (4f58684)
fix: Refactor split create_tables into static Benchmark methods
fix: Refactor split create_tables into static Benchmark methods (#3126)
feat - Split create_tables into static Benchmark methods
feat - format
Update mteb/leaderboard/table.py
Co-authored-by: Kenneth Enevoldsen <kenevoldsen@pm.me>
feat - remove search query;take benchmark result as input;addressing the circular import,
feat - format
Update mteb/benchmarks/benchmark.py
Co-authored-by: Kenneth Enevoldsen <kenevoldsen@pm.me>
Co-authored-by: Kenneth Enevoldsen <kenevoldsen@pm.me>
feat - use to_dataframe;clean table.py;move creat_table
feat - fix circular import
feat - clean-up
feat - format
Co-authored-by: Kenneth Enevoldsen <kenevoldsen@pm.me> (cb03bd4)
Update giga embeddings (#3210)
update giga embeddings
update giga embeddings
3b-september-2025
fixed
lint
Update mteb/models/ru_sentence_models.py
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>
change revision due to flash-attn dependency
change apply_instruction_to_passages
Co-authored-by: Kolodin Egor <eikolodin@sberbank.ru>
Co-authored-by: Roman Solomatin <36135455+Samoed@users.noreply.github.com>
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>
Co-authored-by: Неизвестный Пользователь722497 <dolegosmirnov@sberbank.ru> (15f9909)
fix: prevent EOS token truncation
fix: prevent EOS token truncation (#3218)
fix(models): prevent EOS token truncation for BMRetriever
refactor(models): refactor tokenizer setup in InstructSentenceTransformerWrapper
fix(models): correct eos token handling in BMRetrieverWrapper (f58ac2b)
Your coding agent can read these notes before it upgrades. Set up the MCP server →