NewYour coding agent can read the release notes before it upgrades.Set up the MCP server →
PyPI · #2804 most downloaded on PyPI
Massive Text Embedding Benchmark
Last release today
21 Sep 2026
Ships fairly regularly
a new release about every 1 weeks
Nearly every release is documented
notes for 60 of the last 60 stable releases
Nothing withdrawn
no release was ever pulled
4 years old
805 releases · first in 2022
One column per quarter.
add to docs to say what this script does
ci: Dataset check on new PR (#3103)
add dataset check on new PR
add extract datasets
run as module
update startswith
update workflow name
add GitPython
export var
same shell session
address review comments
add to docs to say what this script does
add docs (6e8eba1)
fix: add voyage quantization models (#3092)
Adding quantization support
Update mteb/models/voyage_models.py
Co-authored-by: Kenneth Enevoldsen <kenevoldsen@pm.me>
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>
Simplifying the quantization/output_dtype
Update mteb/model_meta.py
Co-authored-by: Kenneth Enevoldsen <kenevoldsen@pm.me>
Co-authored-by: Kenneth Enevoldsen <kenevoldsen@pm.me>
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com> (9c7804c)
model: add Youtu-Embedding-V1 (#3115)
add youtu models
add a blank line
fix the optional dependencies and lint the code
remove unused dependencies and reformat
revise prompt_type
Co-authored-by: springxchen <springxchen@tencent.com> (652ff2b)
test out GH models with welcoming new comers (73a35e0)
- Added an include_private parameter to the get_tasks() function that defaults to False
fix: Allow closed datasets (#3059)
Added an include_private parameter to the get_tasks() function that defaults to False
This ensures that by default, tests only run on public datasets
Tests can explicitly set include_private=True when needed to test private datasets
Added is_public: bool | None = None field to TaskMetadata
The field is optional and defaults to None (treated as public)
Updated the is_filled() method to exclude is_public from required fields
Added documentation
Added an include_private parameter to the get_tasks() function that defaults to False
This ensures that by default, tests only run on public datasets
Tests can explicitly set include_private=True when needed to test private datasets
Added is_public: bool | None = None field to TaskMetadata
The field is optional and defaults to None (treated as public)
Updated the is_filled() method to exclude is_public from required fields
Added documentation
Correcting due to comments
Update mteb/abstasks/TaskMetadata.py
Co-authored-by: Kenneth Enevoldsen <kenevoldsen@pm.me>
Co-authored-by: Kenneth Enevoldsen <kenevoldsen@pm.me>
Removing the not used filter_tasks_by_privacy function
Correcting due to comments
Correcting due to comments
Correcting due to comments
Removing the test case
Rename the include_private parameter to exclude_private
Rename the include_private parameter to exclude_private
Add private tasks tests
Add private tasks tests
Update tests/test_tasks/test_private_tasks.py
Co-authored-by: Kenneth Enevoldsen <kenevoldsen@pm.me>
Add private tasks tests
Add private tasks tests
Add private tasks tests
Co-authored-by: Kenneth Enevoldsen <kenevoldsen@pm.me> (5844cc7)
model: Add ModelMeta for OrdalieTech/Solon-embeddings-mini-beta-1.1 (#3090)
Add ModelMeta for OrdalieTech/Solon-embeddings-mini-beta-1.1
Add training_datasets (common_corpus, fineweb, wiki_fr, private LLM-synth)
Format with ruff + add loader per review
Apply ruff format/fixes
Update mteb/models/ordalietech_solon_embeddings_mini_beta_1_1.py
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>
Register OrdalieTech/Solon-embeddings-mini-beta-1.1 in overview (ModelMeta + loader)
Update mteb/models/ordalietech_solon_embeddings_mini_beta_1_1.py
Co-authored-by: Kenneth Enevoldsen <kenevoldsen@pm.me>
fix import
Add memory_usage_mb=808.0 and required fields to ModelMeta
Fix 210 milions of parameters
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>
Co-authored-by: Isaac Chung <chungisaac1217@gmail.com>
Co-authored-by: Kenneth Enevoldsen <kenevoldsen@pm.me> (4774b74)
fix: Add @classmethod for @field_validators in TaskMetadata (#3100) (`4012517`)
fix: Updating the default batch size calculation in the voyage models (#3091) (`5851c7a`)
5851c7a)Combine Plots and Tables into a Single (#3047)
feat - Combine Plots and Tables into a Single Tab #3009
feat - Resize the plot to make it more readable
feat - Remove the (radar chart)
feat - Add a comment stating that it only shows the Top 5 models in the table.
feat - adjust layout
Update mteb/leaderboard/app.py
format
Co-authored-by: Kenneth Enevoldsen <kenevoldsen@pm.me>
Co-authored-by: Kenneth Enevoldsen <kennethcenevoldsen@gmail.com>
Co-authored-by: Isaac Chung <chungisaac1217@gmail.com> (9586697)
CI: Set upper limit for xdist version (#3098)
Commentout bibtex formatting
Remove -n auto
get back bibtex
try limiting versions
revert coverage
revert coverage
Co-authored-by: Isaac Chung <chungisaac1217@gmail.com> (17fa697)
fix: duplicate mteb multilingual variables
fix: duplicate mteb multilingual variables (#3080)
fix benchmark naming
format
lint (27be671)
fix: Improving validate_task_to_prompt_name logs and error messages (#3079)
Improving validate_task_to_prompt_name logs and error messages
linter fixes
Adding None prompts tests
Update test_benchmark_sentence_transformer
Update mteb/leaderboard/benchmark_selector.py
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com> (139fc73)
model: mdbr-leaf models (#3081)
added MDBR leaf models
fixed revision for mdbr-leaf-ir
added model prompts
updated training datasets
fixed linting
lotte task reference
Co-authored-by: Robin Vujanic <robin.vujanic@mongodb.com> (e4c2a95)
Update tasks & benchmarks tables (5bf303b)
Move dev to dependency groups (#3088)
add dependency groups (cd14ef6)
fix: run ruff check on all files during ci
fix: run ruff check on all files during ci (#3086)
fix: run ruff check on all files during ci
format (b46b633)
fix: Add beta version of RTEB related benchmarks
fix: Add beta version of RTEB related benchmarks (#3048)
Add RTEB related benchmarks
Add RTEB related benchmarks
Correcting the task names in the RTEB benchmarks
Update mteb/leaderboard/benchmark_selector.py
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>
Adding the CURE dataset to RTEB benchmarks
Use the right language subset
Fix broken finance icon URL in RTEB benchmarks
Replace broken libre-finance-dollar.svg with working libre-gui-price-tag.svg Validated all icon URLs and confirmed accessibility compliance
Add the rteb_benchmarks to the BENCHMARK_REGISTRY
Add the rteb_benchmarks to the BENCHMARK_REGISTRY
Add the rteb_benchmarks to the BENCHMARK_REGISTRY
Add the rteb_benchmarks to the BENCHMARK_REGISTRY
Add the rteb_benchmarks to the BENCHMARK_REGISTRY
Add the rteb_benchmarks to the BENCHMARK_REGISTRY
Add the rteb_benchmarks to the BENCHMARK_REGISTRY
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com> (1541318)
Fix reference link (d2c3570)
fix: Update revision for qzhou models (#3069) (`63a0c60`)
add bug label to bug issue template
ci: Add stale workflow (#3066)
add stale workflow
add permissions
add bug label to bug issue template
revert bug issue and only look at more info needed issues
more accurate name
override default (df719cc)
1f9641a)70724e7)fix: ensure that there are always relevant docs attached to query
fix: ensure that there are always relevant docs attached to query (#3058)
fix: ensure that there are always relevant docs attached to query
Here is brief test that it doesn't influence scores:
t1 = mteb.get_task("TwitterHjerneRetrieval")
meta = mteb.get_model_meta("minishlab/potion-base-2M")
eval = mteb.MTEB(tasks=[t1])
res = eval.run(model=meta.load_model())
# before fix:
res[0].get_score() # np.float64(0.02837)
res[0].scores
before_fix = {
"train": [
{
"ndcg_at_1": 0.02597,
"ndcg_at_3": 0.02213,
"ndcg_at_5": 0.0262,
"ndcg_at_10": 0.02837,
"ndcg_at_20": 0.04548,
"ndcg_at_100": 0.13527,
"ndcg_at_1000": 0.24507,
"map_at_1": 0.00866,
"map_at_3": 0.01317,
"map_at_5": 0.0149,
"map_at_10": 0.01562,
"map_at_20": 0.01898,
"map_at_100": 0.02968,
"map_at_1000": 0.03841,
"recall_at_1": 0.00866,
"recall_at_3": 0.02056,
"recall_at_5": 0.02922,
"recall_at_10": 0.03355,
"recall_at_20": 0.08268,
"recall_at_100": 0.43766,
"recall_at_1000": 1.0,
"precision_at_1": 0.02597,
"precision_at_3": 0.02165,
"precision_at_5": 0.01818,
"precision_at_10": 0.01039,
"precision_at_20": 0.01234,
"precision_at_100": 0.01481,
"precision_at_1000": 0.0034,
"mrr_at_1": 0.025974,
"mrr_at_3": 0.041126,
"mrr_at_5": 0.04632,
"mrr_at_10": 0.048485,
"mrr_at_20": 0.058356,
"mrr_at_100": 0.070186,
"mrr_at_1000": 0.071349,
"nauc_ndcg_at_1_max": 0.33969,
"nauc_ndcg_at_1_std": -0.202864,
"nauc_ndcg_at_1_diff1": -0.127,
"nauc_ndcg_at_3_max": 0.409376,
"nauc_ndcg_at_3_std": -0.039352,
"nauc_ndcg_at_3_diff1": -0.022816,
"nauc_ndcg_at_5_max": 0.250499,
"nauc_ndcg_at_5_std": -0.115263,
"nauc_ndcg_at_5_diff1": -0.057017,
"nauc_ndcg_at_10_max": 0.238696,
"nauc_ndcg_at_10_std": -0.138396,
"nauc_ndcg_at_10_diff1": -0.045287,
"nauc_ndcg_at_20_max": 0.154456,
"nauc_ndcg_at_20_std": -0.070635,
"nauc_ndcg_at_20_diff1": 0.074499,
"nauc_ndcg_at_100_max": -0.005753,
"nauc_ndcg_at_100_std": -0.074738,
"nauc_ndcg_at_100_diff1": -0.005851,
"nauc_ndcg_at_1000_max": 0.109439,
"nauc_ndcg_at_1000_std": -0.089797,
"nauc_ndcg_at_1000_diff1": -0.021634,
"nauc_map_at_1_max": 0.33969,
"nauc_map_at_1_std": -0.202864,
"nauc_map_at_1_diff1": -0.127,
"nauc_map_at_3_max": 0.385244,
"nauc_map_at_3_std": -0.080638,
"nauc_map_at_3_diff1": -0.060991,
"nauc_map_at_5_max": 0.294871,
"nauc_map_at_5_std": -0.119069,
"nauc_map_at_5_diff1": -0.06234,
"nauc_map_at_10_max": 0.285698,
"nauc_map_at_10_std": -0.132856,
"nauc_map_at_10_diff1": -0.055015,
"nauc_map_at_20_max": 0.236619,
"nauc_map_at_20_std": -0.100673,
"nauc_map_at_20_diff1": -0.002619,
"nauc_map_at_100_max": 0.15345,
"nauc_map_at_100_std": -0.138888,
"nauc_map_at_100_diff1": -0.02257,
"nauc_map_at_1000_max": 0.171402,
"nauc_map_at_1000_std": -0.134644,
"nauc_map_at_1000_diff1": -0.034477,
"nauc_recall_at_1_max": 0.33969,
"nauc_recall_at_1_std": -0.202864,
"nauc_recall_at_1_diff1": -0.127,
"nauc_recall_at_3_max": 0.375072,
"nauc_recall_at_3_std": -0.009643,
"nauc_recall_at_3_diff1": -0.089168,
"nauc_recall_at_5_max": 0.147691,
"nauc_recall_at_5_std": -0.128654,
"nauc_recall_at_5_diff1": -0.084259,
"nauc_recall_at_10_max": 0.141055,
"nauc_recall_at_10_std": -0.165932,
"nauc_recall_at_10_diff1": -0.060966,
"nauc_recall_at_20_max": 0.043863,
"nauc_recall_at_20_std": -0.028374,
"nauc_recall_at_20_diff1": 0.157575,
"nauc_recall_at_100_max": -0.157183,
"nauc_recall_at_100_std": -0.019437,
"nauc_recall_at_100_diff1": 0.013395,
# "nauc_recall_at_1000_max": nan,
# "nauc_recall_at_1000_std": nan,
# "nauc_recall_at_1000_diff1": nan,
"nauc_precision_at_1_max": 0.33969,
"nauc_precision_at_1_std": -0.202864,
"nauc_precision_at_1_diff1": -0.127,
"nauc_precision_at_3_max": 0.406318,
"nauc_precision_at_3_std": 0.007031,
"nauc_precision_at_3_diff1": -0.034709,
"nauc_precision_at_5_max": 0.178131,
"nauc_precision_at_5_std": -0.112493,
"nauc_precision_at_5_diff1": -0.045535,
"nauc_precision_at_10_max": 0.167897,
"nauc_precision_at_10_std": -0.150626,
"nauc_precision_at_10_diff1": -0.027811,
"nauc_precision_at_20_max": 0.081428,
"nauc_precision_at_20_std": -0.042304,
"nauc_precision_at_20_diff1": 0.17278,
"nauc_precision_at_100_max": -0.150619,
"nauc_precision_at_100_std": 0.016133,
"nauc_precision_at_100_diff1": -0.065571,
"nauc_precision_at_1000_max": -0.017244,
"nauc_precision_at_1000_std": 0.046614,
"nauc_precision_at_1000_diff1": -0.028258,
"nauc_mrr_at_1_max": 0.33969,
"nauc_mrr_at_1_std": -0.202864,
"nauc_mrr_at_1_diff1": -0.127,
"nauc_mrr_at_3_max": 0.409511,
"nauc_mrr_at_3_std": -0.064671,
"nauc_mrr_at_3_diff1": -0.01911,
"nauc_mrr_at_5_max": 0.319584,
"nauc_mrr_at_5_std": -0.103546,
"nauc_mrr_at_5_diff1": -0.025109,
"nauc_mrr_at_10_max": 0.309614,
"nauc_mrr_at_10_std": -0.117564,
"nauc_mrr_at_10_diff1": -0.019691,
"nauc_mrr_at_20_max": 0.262976,
"nauc_mrr_at_20_std": -0.092222,
"nauc_mrr_at_20_diff1": 0.024507,
"nauc_mrr_at_100_max": 0.256052,
"nauc_mrr_at_100_std": -0.094249,
"nauc_mrr_at_100_diff1": 0.012432,
"nauc_mrr_at_1000_max": 0.260112,
"nauc_mrr_at_1000_std": -0.098845,
"nauc_mrr_at_1000_diff1": 0.009697,
"main_score": 0.02837,
"hf_subset": "default",
"languages": ["dan-Latn"],
}
]
}
# with update:
res[0].get_score() # np.float64(0.02837)
res[0].scores
with_fix = {
"train": [
{
"ndcg_at_1": 0.02597,
"ndcg_at_3": 0.02213,
"ndcg_at_5": 0.0262,
"ndcg_at_10": 0.02837,
"ndcg_at_20": 0.04548,
"ndcg_at_100": 0.13527,
"ndcg_at_1000": 0.24507,
"map_at_1": 0.00866,
"map_at_3": 0.01317,
"map_at_5": 0.0149,
"map_at_10": 0.01562,
"map_at_20": 0.01898,
"map_at_100": 0.02968,
"map_at_1000": 0.03841,
"recall_at_1": 0.00866,
"recall_at_3": 0.02056,
"recall_at_5": 0.02922,
"recall_at_10": 0.03355,
"recall_at_20": 0.08268,
"recall_at_100": 0.43766,
"recall_at_1000": 1.0,
"precision_at_1": 0.02597,
"precision_at_3": 0.02165,
"precision_at_5": 0.01818,
"precision_at_10": 0.01039,
"precision_at_20": 0.01234,
"precision_at_100": 0.01481,
"precision_at_1000": 0.0034,
"mrr_at_1": 0.025974,
"mrr_at_3": 0.041126,
"mrr_at_5": 0.04632,
"mrr_at_10": 0.048485,
"mrr_at_20": 0.058356,
"mrr_at_100": 0.070186,
"mrr_at_1000": 0.071349,
"nauc_ndcg_at_1_max": 0.33969,
"nauc_ndcg_at_1_std": -0.202864,
"nauc_ndcg_at_1_diff1": -0.127,
"nauc_ndcg_at_3_max": 0.409376,
"nauc_ndcg_at_3_std": -0.039352,
"nauc_ndcg_at_3_diff1": -0.022816,
"nauc_ndcg_at_5_max": 0.250499,
"nauc_ndcg_at_5_std": -0.115263,
"nauc_ndcg_at_5_diff1": -0.057017,
"nauc_ndcg_at_10_max": 0.238696,
"nauc_ndcg_at_10_std": -0.138396,
"nauc_ndcg_at_10_diff1": -0.045287,
"nauc_ndcg_at_20_max": 0.154456,
"nauc_ndcg_at_20_std": -0.070635,
"nauc_ndcg_at_20_diff1": 0.074499,
"nauc_ndcg_at_100_max": -0.005753,
"nauc_ndcg_at_100_std": -0.074738,
"nauc_ndcg_at_100_diff1": -0.005851,
"nauc_ndcg_at_1000_max": 0.109439,
"nauc_ndcg_at_1000_std": -0.089797,
"nauc_ndcg_at_1000_diff1": -0.021634,
"nauc_map_at_1_max": 0.33969,
"nauc_map_at_1_std": -0.202864,
"nauc_map_at_1_diff1": -0.127,
"nauc_map_at_3_max": 0.385244,
"nauc_map_at_3_std": -0.080638,
"nauc_map_at_3_diff1": -0.060991,
"nauc_map_at_5_max": 0.294871,
"nauc_map_at_5_std": -0.119069,
"nauc_map_at_5_diff1": -0.06234,
"nauc_map_at_10_max": 0.285698,
"nauc_map_at_10_std": -0.132856,
"nauc_map_at_10_diff1": -0.055015,
"nauc_map_at_20_max": 0.236619,
"nauc_map_at_20_std": -0.100673,
"nauc_map_at_20_diff1": -0.002619,
"nauc_map_at_100_max": 0.15345,
"nauc_map_at_100_std": -0.138888,
"nauc_map_at_100_diff1": -0.02257,
"nauc_map_at_1000_max": 0.171402,
"nauc_map_at_1000_std": -0.134644,
"nauc_map_at_1000_diff1": -0.034477,
"nauc_recall_at_1_max": 0.33969,
"nauc_recall_at_1_std": -0.202864,
"nauc_recall_at_1_diff1": -0.127,
"nauc_recall_at_3_max": 0.375072,
"nauc_recall_at_3_std": -0.009643,
"nauc_recall_at_3_diff1": -0.089168,
"nauc_recall_at_5_max": 0.147691,
"nauc_recall_at_5_std": -0.128654,
"nauc_recall_at_5_diff1": -0.084259,
"nauc_recall_at_10_max": 0.141055,
"nauc_recall_at_10_std": -0.165932,
"nauc_recall_at_10_diff1": -0.060966,
"nauc_recall_at_20_max": 0.043863,
"nauc_recall_at_20_std": -0.028374,
"nauc_recall_at_20_diff1": 0.157575,
"nauc_recall_at_100_max": -0.157183,
"nauc_recall_at_100_std": -0.019437,
"nauc_recall_at_100_diff1": 0.013395,
# "nauc_recall_at_1000_max": nan,
# "nauc_recall_at_1000_std": nan,
# "nauc_recall_at_1000_diff1": nan,
"nauc_precision_at_1_max": 0.33969,
"nauc_precision_at_1_std": -0.202864,
"nauc_precision_at_1_diff1": -0.127,
"nauc_precision_at_3_max": 0.406318,
"nauc_precision_at_3_std": 0.007031,
"nauc_precision_at_3_diff1": -0.034709,
"nauc_precision_at_5_max": 0.178131,
"nauc_precision_at_5_std": -0.112493,
"nauc_precision_at_5_diff1": -0.045535,
"nauc_precision_at_10_max": 0.167897,
"nauc_precision_at_10_std": -0.150626,
"nauc_precision_at_10_diff1": -0.027811,
"nauc_precision_at_20_max": 0.081428,
"nauc_precision_at_20_std": -0.042304,
"nauc_precision_at_20_diff1": 0.17278,
"nauc_precision_at_100_max": -0.150619,
"nauc_precision_at_100_std": 0.016133,
"nauc_precision_at_100_diff1": -0.065571,
"nauc_precision_at_1000_max": -0.017244,
"nauc_precision_at_1000_std": 0.046614,
"nauc_precision_at_1000_diff1": -0.028258,
"nauc_mrr_at_1_max": 0.33969,
"nauc_mrr_at_1_std": -0.202864,
"nauc_mrr_at_1_diff1": -0.127,
"nauc_mrr_at_3_max": 0.409511,
"nauc_mrr_at_3_std": -0.064671,
"nauc_mrr_at_3_diff1": -0.01911,
"nauc_mrr_at_5_max": 0.319584,
"nauc_mrr_at_5_std": -0.103546,
"nauc_mrr_at_5_diff1": -0.025109,
"nauc_mrr_at_10_max": 0.309614,
"nauc_mrr_at_10_std": -0.117564,
"nauc_mrr_at_10_diff1": -0.019691,
"nauc_mrr_at_20_max": 0.262976,
"nauc_mrr_at_20_std": -0.092222,
"nauc_mrr_at_20_diff1": 0.024507,
"nauc_mrr_at_100_max": 0.256052,
"nauc_mrr_at_100_std": -0.094249,
"nauc_mrr_at_100_diff1": 0.012432,
"nauc_mrr_at_1000_max": 0.260112,
"nauc_mrr_at_1000_std": -0.098845,
"nauc_mrr_at_1000_diff1": 0.009697,
"main_score": 0.02837,
"hf_subset": "default",
"languages": ["dan-Latn"],
}
]
}
# check
with_fix == before_fix # True
* restructure
* format
* relax pytrec versions
* fix incorrect parsing ([`9c27f71`](https://github.com/embeddings-benchmark/mteb/commit/9c27f71e44612f190756d41f1fcffeb817b0f3e3))
## Unknown
* model: Add CoDi-Embedding-V1 (#3054)
* add codiemb-minicpm
* replace codiemb_minicpm with codi_model
* Update mteb/models/codi_model.py
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>
* Update mteb/models/codi_model.py
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>
* Update mteb/models/codi_model.py
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>
* update code
* update code
* reformat
---------
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com> ([`4994ea1`](https://github.com/embeddings-benchmark/mteb/commit/4994ea1758baf7d7bc183215280595fdb2b36c10))
* Update tasks & benchmarks tables ([`26468f8`](https://github.com/embeddings-benchmark/mteb/commit/26468f8db5633b06f328d163c43f31e195062a96))
* dataset: Add JinaVDR (#2942)
* feat: added jinavdr benchmark
* feat: added description for jinavdr
* feat: fixed licenses and added bibtex
* feat: made jinav4 compatible with vidore benchmark
* feat: corrected query numbers
* feat: removed print
* feat: added max pixel argument for jina models
* feat: score calculation on cpu
* feat: adjust jina model for new mteb code
* feat: code cleanup
* feat: corrected bibtex
* feat: make colpali run with jinavdr
* feat: fixed comments
* feat: better reference and fixed comments
* feat: added date for tasks
* feat: fixed missing metadata and bibtex
* feat: added descriptions per dataset ([`cf3e1bb`](https://github.com/embeddings-benchmark/mteb/commit/cf3e1bbe62b53abeec7a932373c52e7d852f97cd))
* Correcting the (new) DS1000 dataset's revision (#3063)
* Add DS1000 retrieval task
- Code retrieval task based on 1,000 data science programming problems
- Natural language queries matched to Python data science code
- Uses python-Code evaluation language for code-specific metrics
- Covers pandas, numpy, matplotlib, scikit-learn, and scipy libraries
* Add DS1000Retrieval to imports
* Add descriptive statistics for DS1000Retrieval
* Reformatting
* Reformatting
* Add DS1000Retrieval task implementation ([`8e1c354`](https://github.com/embeddings-benchmark/mteb/commit/8e1c3547b4b2b3b18241e4f6456fc48360c6a041))
* Update tasks & benchmarks tables ([`69099fe`](https://github.com/embeddings-benchmark/mteb/commit/69099fe216fe1ac4eaf8c5f68c0ba5b1dcd60cee))
* Add ChatDoctorRetrieval (#3045)
* Add ChatDoctorRetrieval
* Reformatting, correcting the revision
* Correct the dataset citation
* Correcting due to comments ([`e91cb8e`](https://github.com/embeddings-benchmark/mteb/commit/e91cb8ec97b24f8fa8547093ed573f4eea752989))
* Update tasks & benchmarks tables ([`d2fcbac`](https://github.com/embeddings-benchmark/mteb/commit/d2fcbac461f300e2f44b281f31e079e8e6ade348))
* dataset: Add ds1000 retrieval (#3038)
* Add DS1000 retrieval task
- Code retrieval task based on 1,000 data science programming problems
- Natural language queries matched to Python data science code
- Uses python-Code evaluation language for code-specific metrics
- Covers pandas, numpy, matplotlib, scikit-learn, and scipy libraries
* Add DS1000Retrieval to imports
* Add descriptive statistics for DS1000Retrieval
* Reformatting
* Reformatting ([`53f0986`](https://github.com/embeddings-benchmark/mteb/commit/53f09860ea6a68a1b5650b064cf86ce8c93333a2))
* Update tasks & benchmarks tables ([`e1ede42`](https://github.com/embeddings-benchmark/mteb/commit/e1ede4251b6d39ed86e59cbc9738863021193494))
* Add FreshStackRetrieval task (#3043)
* Add FreshStackRetrieval
* Reformatting, correcting the revision
* Dataset correction ([`a291a05`](https://github.com/embeddings-benchmark/mteb/commit/a291a055e8fc7814c2827a2becb1e32f7e4c7077))
* Update tasks & benchmarks tables ([`fe57390`](https://github.com/embeddings-benchmark/mteb/commit/fe573902ad7dba42a50b03186e813ec4ea854c61))
* Add FinanceBenchRetrieval task (#3044)
* Add FinanceBenchRetrieval
* Update mteb/tasks/Retrieval/eng/FinanceBenchRetrieval.py
---------
Co-authored-by: Isaac Chung <chungisaac1217@gmail.com> ([`4da11c6`](https://github.com/embeddings-benchmark/mteb/commit/4da11c6b8d23bef2ea85d89636c4e9224bd30832))
* Update tasks & benchmarks tables ([`fd8f89e`](https://github.com/embeddings-benchmark/mteb/commit/fd8f89ebc71a3fcb64deef94adbe00a3ec2b4bbd))
* Add finqa retrieval (#3042)
* Add FinQA retrieval task
- Financial numerical reasoning retrieval task based on FinQA dataset
- Numerical financial questions matched to relevant document data
- Covers earnings reports with tables and quantitative financial data
- Includes proper citations and descriptive statistics
* Add FinQARetrieval to imports
* Add descriptive statistics for FinQARetrieval
* Reformatting
* Reformatting
* Update mteb/tasks/Retrieval/eng/FinQARetrieval.py
---------
Co-authored-by: Isaac Chung <chungisaac1217@gmail.com> ([`7b57185`](https://github.com/embeddings-benchmark/mteb/commit/7b57185c0fadeabf3eddf534d00bfd3ca57a920a))
* Add hc3finance retrieval (#3041)
* Add HC3Finance retrieval task
- Financial retrieval task based on HC3 Finance dataset
- Financial questions matched to human and AI-generated content
- Covers financial explanations, analysis, and educational content
- Includes proper citations and descriptive statistics
* Add HC3FinanceRetrieval to imports
* Add descriptive statistics for HC3FinanceRetrieval
* Reformatting
* Reformatting, correcting the revision
* Update mteb/tasks/Retrieval/eng/HC3FinanceRetrieval.py
---------
Co-authored-by: Isaac Chung <chungisaac1217@gmail.com> ([`53d7d84`](https://github.com/embeddings-benchmark/mteb/commit/53d7d84b1e51111edbae63283f4630115156f3ab))
ci: Temporarily limit pytrec version to "pytrec-eval-terrier>=0.5.6, <0.5.8" to prevent errors
try to fix CI (6fa6efa)
fix: Add VN-MTEB benchmark and Leaderboard (#2995)
[ADD] 50 vietnamese dataset from vn-mteb
[UPDATE] task metadata
[UPDATE] import dependencies
[UPDATE] task metadata, bibtext citation
[UPDATE-TEST] test_model_meta
[UPDATE] sample_creation to machine-translated and LM verified
[ADD] sample creation machine-translated and LM verified
[ADD] VN-MTEB benchmark and leaderboard
[FIX] wrong benchmark name
[REMOVE] default fields metadata in Classfication tasks (0a6e855)
Update tasks & benchmarks tables (def1377)
fix MBPPRetrieval revision (#3055)
Update MBPPRetrieval.py
Co-authored-by: Roman Solomatin <36135455+Samoed@users.noreply.github.com> (ea801ec)
Update tasks & benchmarks tables (7da3cf9)
dataset: Added wikisql retrieval (#3039)
Add WikiSQL retrieval task
Add WikiSQLRetrieval to imports
Add descriptive statistics for WikiSQLRetrieval
Reformatting
Reformatting
Reformatting, correcting the revision (7b289f5)
Update tasks & benchmarks tables (1fff5ce)
dataset: Add mbpp retrieval (#3037)
Add MBPP retrieval task
Add MBPPRetrieval to imports
Add descriptive statistics for MBPPRetrieval
Reformatting
Reformatting (ac69263)
Fix 3 VN-MTEB Pair Classification tasks (#3053)
[ADD] 50 vietnamese dataset from vn-mteb
[UPDATE] task metadata
[UPDATE] import dependencies
[UPDATE] task metadata, bibtext citation
[UPDATE-TEST] test_model_meta
[UPDATE] sample_creation to machine-translated and LM verified
[ADD] sample creation machine-translated and LM verified
[ADD] Vietnamese Embedding models
[REMOVE] default fields metadata in Classfication tasks
[UPDATE] model to vi-vn language specific file
[FIX] lint
[FIX] model loader
[FIX] VN-MTEB 3 datasets PairClassification rename column (4e3fcd8)
ci: Updating rerun delays to prevent false positives errors (`e476dc3`)
ci: Updating rerun delays to prevent false positives errors (e476dc3)
ci: reduce parallel runs for when checking if a dataset exists (#3035)
The hope is that this will prevent many of the current errors (4aaf47e)
fix: Updated revision for jina-embeddings-v4 (#3046)
fix: jinav4 revision
Signed-off-by: admin <bo.wang@jina.ai>
Signed-off-by: admin <bo.wang@jina.ai>
Signed-off-by: admin <bo.wang@jina.ai>
Co-authored-by: admin <bo.wang@jina.ai> (c58b319)
model: add granite-embedding-english R2 models (#3050) (e08ec56)
model: Add GreenNode Vietnamese Embedding models (#2994)
[ADD] 50 vietnamese dataset from vn-mteb
[UPDATE] task metadata
[UPDATE] import dependencies
[UPDATE] task metadata, bibtext citation
[UPDATE-TEST] test_model_meta
[UPDATE] sample_creation to machine-translated and LM verified
[ADD] sample creation machine-translated and LM verified
[ADD] Vietnamese Embedding models
[REMOVE] default fields metadata in Classfication tasks
[UPDATE] model to vi-vn language specific file
[FIX] lint
[FIX] model loader (72f7b05)
Merge branch 'main' of https://github.com/embeddings-benchmark/mteb (d729d32)
Fix deprecated metadata_dict usage
The provided revisions doesn't seem to be present on: adrlau/navjordj-SNL_summarization_copy
Replacing with latest revision (5c65913)
Update tasks & benchmarks tables (a96f2e4)
dataset: Add HumanEvalRetrieval task (#3022)
Add HumanEvalRetrieval dataset
Fix TaskMetadata structure and remove descriptive_stats
Changed query_id/corpus_id to query-id/corpus-id to match actual dataset format.
Use self.metadata.dataset instead of self.metadata_dict for v2.0 compatibility.
d4e6223)model: Add granite-vision-embedding model (#3029)
Add files via upload
Address review comments
Address review comments
ruff format
Update mteb/models/granite_vision_embedding_models.py
lint error fix
Co-authored-by: Kenneth Enevoldsen <kenevoldsen@pm.me> (37d115a)
model: Add samilpwc_models meta (#3028)
model: Add samilpwc_models meta
Fix: Remove CONST
Fix: Reformat File
Update: model revision (96a7cc5)
fix: Add missing training sets for qzhou
fix: Add missing training sets for qzhou (#3023)
Supplement missing training sets
reformat code
Reorganize the data list format
update qzhou_model meta (20bc80c)
Update tasks & benchmarks tables (177997f)
Standardise task names and fix citation formatting (#3026)
fixes for name formatting (ea41e7a)
Add OpenAI models with 512 dimension (#3008)
Add OpenAI/text-embedding-3-small (512 dim) Add OpenAI/text-embedding-3-large (512 dim)
Correcting due to comments
Co-authored-by: fzowl <zoltan@voyageai.com> (d8b2910)
model: Add Cohere embed-v4.0 model support (#3006)
Add Cohere embed-v4.0 model support
🤖 Generated with Claude Code
Co-Authored-By: Claude <noreply@anthropic.com>
Update cohere_v.py and cohere_models.py to include the new embed-v4.0 model with proper configuration and integration.
🤖 Generated with Claude Code
Co-Authored-By: Claude <noreply@anthropic.com>
Co-authored-by: Claude <noreply@anthropic.com> (87eb27c)
Update tasks & benchmarks tables (4adf565)
dataset: Added 50 Vietnamese dataset from vn-mteb (#2964)
[ADD] 50 vietnamese dataset from vn-mteb
[UPDATE] task metadata
[UPDATE] import dependencies
[UPDATE] task metadata, bibtext citation
[UPDATE-TEST] test_model_meta
[UPDATE] sample_creation to machine-translated and LM verified
[ADD] sample creation machine-translated and LM verified
[REMOVE] default fields metadata in Classfication tasks (741b022)
lint: Correcting lint errors (#3004)
Adding Classification Evaluator test
Modifications due to the comments
Update tests/test_evaluators/test_ClassificationEvaluator.py
Co-authored-by: Kenneth Enevoldsen <kenevoldsen@pm.me>
Co-authored-by: Kenneth Enevoldsen <kenevoldsen@pm.me>
Modifications due to the comments
Modifications due to the comments
Correcting the lint errors
Co-authored-by: Kenneth Enevoldsen <kenevoldsen@pm.me> (01840ce)
model: BAAI/bge-m3-unsupervised Model (#3007)
Add BAAI/bge-m3-unsupervised Model (BAAI/bge_m3_retromae is commented out - the details are proper, but it fails during loading the model for me, so i commented out)
Remove the commented retromae model
Co-authored-by: fzowl <zoltan@voyageai.com> (042db73)
Add Voyage 3.5 model configuration
🤖 Generated with Claude Code
Co-authored-by: Claude <noreply@anthropic.com> (e5d386b)
qzhou-embedding model_meta & implementation (#2975)
qzhou-embedding model_meta & implementation
Update qzhou_models.py
Update qzhou_models.py
Processing todo items(Add default instruction)
correct bge datalist
correct 'public_training_data'
Update qzhou_models.py
Update qzhou_models.py
Update qzhou_models.py
Update qzhou_models.py
Update mteb/models/qzhou_models.py
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>
Co-authored-by: Kenneth Enevoldsen <kenevoldsen@pm.me>
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>
Co-authored-by: Kenneth Enevoldsen <kenevoldsen@pm.me> (6c1f1c6)
fix: Add new benchmark beRuSciBench along with AbsTaskTextRegression
fix: Add new benchmark beRuSciBench along with AbsTaskTextRegression (#2716)
Add RuSciBench
fix bitext mining lang
Add regression task
fix init
add missing files
Improve description
Add superseded_by
fix lint
Update regression task to match with v2
Add stratified_subsampling for regression task
Add boostrap for regression task
Rename task class, add model as evaluator argument
fix import
fix import 2
fixes
fix
Rename regression model protocol (36df9ca)
Update tasks & benchmarks tables (a86e2dd)
Update tasks & benchmarks tables (e4f30e9)
dataset: add BillSum datasets (#2943)
Added BillSum datasets
fixed billsumca
Updated BillSumCA description
Updated BillSumUS description
Update mteb/tasks/Retrieval/eng/BillSumCA.py
Co-authored-by: Kenneth Enevoldsen <kenevoldsen@pm.me>
Co-authored-by: Kenneth Enevoldsen <kenevoldsen@pm.me>
lint
lint
Co-authored-by: Kenneth Enevoldsen <kenevoldsen@pm.me>
Co-authored-by: Isaac Chung <chungisaac1217@gmail.com> (007d19f)
dataset: add GovReport dataset (#2953)
Added govreport task
Updated description (42dfe0d)
Update tasks & benchmarks tables (da46c8e)
dataset: Add BSARD v2, fixing the data loading issues of v1 (#2935)
BSARD loader fixed
BSARDv2 metadata fixed
Update mteb/tasks/Retrieval/fra/BSARDRetrieval.py
Co-authored-by: Kenneth Enevoldsen <kenevoldsen@pm.me> (8416541)
ci: bump semantic release (`4ef8571`)
4ef8571)docs: Update adding_a_dataset.md (#2947)
docs: Update adding_a_dataset.md
Update docs/adding_a_dataset.md (a78debf)
get_benchmark (#2939)The leaderboard would have (silent) errors where get_benchmark lead to a KeyError due to "selector_state" being passed as a default value. Setting DEFAULT_BENCMARK_NAME as the value solves this issue. (8496ec2)
fix: Only import SparseEncoder once sentence-transformer version have been checked (#2940)
fix: Only import SparseEncoder once sentence-transformer version have been checked
fixes #2936
Co-authored-by: Isaac Chung <chungisaac1217@gmail.com>
Co-authored-by: Isaac Chung <chungisaac1217@gmail.com> (79a43af)
5ed6c90)533ce59)fix: specify revision for opensearch
specify revision for opensearch (0ac0231)
mteb.get_model in adding_a_dataset.md (#2922)Update adding_a_dataset.md (c1922c8)
dataset: add BarExamQA dataset (#2916)
Add BareExamQA retrieval task
ran linter
updated details
updated details
fixed subtype name
fixed changes
ran linter again (1dcc6dc)
model: Add OpenSearch inf-free sparse encoding models (#2903)
add opensearch inf-free models
Co-authored-by: Isaac Chung <chungisaac1217@gmail.com> (5a868e3)
fix: change passage prompt to document
fix: change passage prompt to document (#2912)
change document to passage
fix prompt names
fix kwargs check
fix default prompt (a298fa9)
Update tasks & benchmarks tables (372fc4c)
dataset: Add JapaneseSentimentClassification (#2913)
Add JapaneseSentimentClassification (57438c2)
Update tasks & benchmarks tables (56c98ed)
Classification dataset cleaning (#2900)
Classification dataset cleaning
Update pull request number
Fix metadata test
fix formatting
add script for cleaning (aef1e33)
Evaluator tests (#2910)
Adding Classification Evaluator test
Modifications due to the comments
Update tests/test_evaluators/test_ClassificationEvaluator.py
Co-authored-by: Kenneth Enevoldsen <kenevoldsen@pm.me>
Co-authored-by: Kenneth Enevoldsen <kenevoldsen@pm.me>
Modifications due to the comments
Modifications due to the comments
Adding STSEvaluator and SummarizationEvaluator tests
Correcting due to the comments
Correcting due to the comments
Co-authored-by: Kenneth Enevoldsen <kenevoldsen@pm.me> (c7078af)
fix: update colpali engine models
fix: update colpali engine models (#2905)
adding vidore benchmarks
fix typo
clean vidore names + per lang eval
lint
vidore names
bibtex fix
fix revision
vidore v2 citation
update citation format and fix per-language mappings
lint: citations
typo citations
fix revisiions
lint
fix colnomic3b revision
fix colqwen2.5 revision + latest repo version
fix query agmentation tokens
colsmol revision (9864e2a)
Add Classification Evaluator unit test (#2838)
Adding Classification Evaluator test
Modifications due to the comments
Update tests/test_evaluators/test_ClassificationEvaluator.py
Co-authored-by: Kenneth Enevoldsen <kenevoldsen@pm.me>
Co-authored-by: Kenneth Enevoldsen <kenevoldsen@pm.me>
Modifications due to the comments
Modifications due to the comments
Co-authored-by: Kenneth Enevoldsen <kenevoldsen@pm.me> (4a47f90)
model: add kalm_models (kalm-emb-v2) ModelMeta (new PR) (#2889)
feat: add KaLM_Embedding_X_0605 in kalm_models
Update kalm_models.py for lint format
kalm-emb-v2
kalm-emb-v2
kalm-emb-v2
kalm-emb-v2
kalm-emb-v2
Co-authored-by: xinshuohu <xinshuohu@tencent.com>
Co-authored-by: Xinshuo Hu <yanshek.woo@gmail.com> (9ecac21)
model: add image support for jina embeddings v4 (#2893)
feat: unify text and image embeddings for all tasks
fix: uniform batch size
fix: update error message
fix: update code task
fix: update max length
fix: apply review suggestions (17be7e5)
fix datasets version (`00c95cf`)
fix datasets version (00c95cf)
Update tasks & benchmarks tables (5303fec)
dataset: Evalita dataset integration (#2859)
Added DadoEvalCoarseClassification
Removed unnecessary columns from DadoEvalCoarseClassification
Added EmitClassification task
added SardiStanceClassification task
Added GeoLingItClassification task
Added DisCoTexPairClassification tasks
Added EmitClassification, DadoEvalCoarseClassification, GeoLingItClassification, SardiStanceClassification inside the inits
changed import in DisCoTexPairClassification
removed GeoLingItClassification dataset
fixed citation formatting, missing metadata parameters and lint formatting
fixed metadata in XGlueWRPReranking
Added MKQARetrieval task
fixed type in XGlueWRPReranking
changed MKQARetrieval from cross-lingual to monolingual
formatted MKQARetrieval file
removed unused const
Co-authored-by: Mattia Sangermano <MattiaSangermano@users.noreply.huggingface.co> (ee17a6e)
model: add Hakim and TookaSBERTV2 models (#2826)
add tooka v2s
add mcinext models
update mcinext.py
Apply PR review suggestions
Update mteb/models/mcinext_models.py
Co-authored-by: mehran <mehan.sarmadi16@gmail.com>
Co-authored-by: Kenneth Enevoldsen <kenevoldsen@pm.me> (04dc6d4)
Update tasks & benchmarks tables (5be02c1)
Add and fix some Japanese datasets: ANLP datasets, JaCWIR, JQaRA (#2872)
Add JaCWIR and JQaRA for reranking
Fix ANLP Journal datasets
Add NLPJournalAbsArticleRetrieval and JaCWIRRetrieval
tackle test cases
Remove _evaluate_subset usage
Separate v1 and v2
Update info for NLP Journal datasets (70768b5)
Comment kalm model (#2877)
comment kalm model (a3ca95c)
model: add kalm_models ModelMeta (new PR) (#2853)
feat: add KaLM_Embedding_X_0605 in kalm_models
Update kalm_models.py for lint format
Co-authored-by: xinshuohu <xinshuohu@tencent.com> (b67bd04)
model: add listconranker modelmeta (#2874)
add listconranker modelmeta
fix bugs
use linter
lint
Co-authored-by: Roman Solomatin <36135455+Samoed@users.noreply.github.com> (5846f56)
fix tests to be compatible with SentenceTransformers v5 (#2875)
fix sbert v5
add comment (f346a37)
rename seed-1.6-embedding to seed1.6-embedding (#2870) (f27648b)
model: Adding nvidia/llama-nemoretriever-colembed models (#2861)
nvidia_llama_nemoretriever_colembed
correct 3b reference
lint fix
add training data and license for nvidia/llama_nemoretriever_colembed
lint
Co-authored-by: Isaac Chung <chungisaac1217@gmail.com> (4ff1413)
Bump gradio (a4388c2)
model: Adding Sailesh97/Hinvec (#2842)
Adding Hinvec Model's Meta data.
Adding hinvec_model.py
Update mteb/models/hinvec_models.py
Co-authored-by: Kenneth Enevoldsen <kennethcenevoldsen@gmail.com>
Co-authored-by: Kenneth Enevoldsen <kennethcenevoldsen@gmail.com> (e3286d5)
fix: prompt validation for tasks with -
fix: update training dataset info of Seed-1.6-embedding model
update seed1.6 model training data info (a8214e2)
docs: Fix some typos in docs/usage/usage.md
docs: Fix some typos in docs/usage/usage.md (#2835)
Update usage.md
Update usage.md
Update docs/usage/usage.md
Co-authored-by: Isaac Chung <chungisaac1217@gmail.com> (774a942)
fix: Update model selection for the leaderboard (#2855)
fix: Update model selection for the leaderboard
fixes #2834
This removed the lower bound selection, but generally I don't think people should care about the models being too small.
fix 1M --> 1B
format
rename model_size -> max_model_size (9a800d3)
model: add Seed-1.6-embedding model (#2841)
add Seed-1.6-embedding model
Update seed_1_6_embedding_models.py
update model meta info
support image encoder interface
error fix
fix: format seed_1_6_embedding_models.py with Ruff (8851bf0)
model: Add custom instructions for GigaEmbeddings (#2836)
add custom instructions
fixed
lint
fix last instruction
Co-authored-by: Kolodin Egor <eikolodin@sberbank.ru>
Co-authored-by: Roman Solomatin <36135455+Samoed@users.noreply.github.com> (d7ff1ab)
fix: Reuploaded previously unavailable SNL datasets
fix: Reuploaded previously unavailable SNL datasets (#2819)
fix: Reuploaded previously unavailable SNL datasets
closes #2477
removed exceptions from tests
temp fixes
added temporary fix
clean up commented out code
format (c790269)
Update tasks & benchmarks tables (74d17b2)
model: Added 3 HIT-TMG's KaLM-embedding models (#2478)
Added HIT-TMG_KaLM-embedding-multilingual-mini-instruct-v1 with instruct wrapper
Added KaLM_embedding_multilingual_mini_instruct_v1_5
Added model to overview.py
Fix Task Count Per Language Table in tasks.md
resolve conflicts
remove tasks.md
Modified get_instruction funcion
Added support for prompt dict in get_instruction
fix lang code
Address comments
Delete mteb/models/check_models.py
added prompts_dict support in InstructSentenceTransformerWrapper
corrected instruction format
corrected prompts format
added correct instruction format
fix implementation
remove if name main
add comment
Co-authored-by: Roman Solomatin <36135455+Samoed@users.noreply.github.com> (03e084b)
add description to issue template (#2817)
add description to template
fix typo (04c9511)
fix: Ensure bright uses the correct revision
fixes #2811 (56dc620)
fix: Adding client arg to init method of OpenAI models wrapper (#2803)
Adding OpenAI client arg to init method (e.g., for already initialized AzureOpenAI client)
To use OpenAI embedding models via Azure, the model wrapper needs to be initialized with a different client.
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>
Update mteb/models/openai_models.py
remove comment and format
Co-authored-by: Kenneth Enevoldsen <kennethcenevoldsen@gmail.com>
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com> (873ee76)
Add LGAI-Embedding
Add mteb/models/lgai_embedding_models.py
defined model metadata (3e291f3)
ci: fix config error for semantic release
discussed in: https://github.com/embeddings-benchmark/mteb/issues/2796 (3d8dd9e)
fix: Add adapted_from to Cmedqaretrieval (#2806)
fix: Add adapted_from to Cmedqaretrieval
Also snuck in a fix with form=None, which is no longer valid, but was still used in a few places.
fef1837)update training datasets
Co-authored-by: zhangzeqing <zhangzeqing@zhejianglab.com> (36a3c67)
Update tasks & benchmarks tables (5e6aa9d)
dataset: Add R2MED Benchmark (#2795)
Add files via upload
Add files via upload
Update benchmarks.py
Update init.py
Add files via upload
Update R2MEDRetrieval.py
Update run_mteb_r2med.py
Delete scripts/run_mteb_r2med.py
Update mteb/tasks/Retrieval/eng/R2MEDRetrieval.py
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>
Add files via upload
Delete mteb/descriptive_stats/Retrieval/R2MEDRetrieval.json
Add files via upload
Add files via upload
Add files via upload
Update R2MEDRetrieval.py
Add files via upload
Add files via upload
Add files via upload
Add files via upload
format citations
Update R2MEDRetrieval.py
Add files via upload
Add files via upload
Co-authored-by: Li Lei <34205771+ll0ruc@users.noreply.github.com>
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com> (b8e64e1)
model: add fangxq/XYZ-embedding (#2741)
add xyz model
add xyz model
add xyz model
update
update
update
update
update
update
update
lint
Co-authored-by: Roman Solomatin <36135455+Samoed@users.noreply.github.com>
Co-authored-by: Kenneth Enevoldsen <kennethcenevoldsen@gmail.com> (1c08974)
model: Add GeoGPT-Research-Project/GeoEmbedding (#2773)
add model: geogpt_models
update geogpt_models
use InstructSentenceTransformerWrapper
resolve pylint warning
format geogpt_models.py
Update mteb/models/geogpt_models.py
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>
Co-authored-by: zhangzeqing <zhangzeqing@zhejianglab.com>
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>
Co-authored-by: Kenneth Enevoldsen <kennethcenevoldsen@gmail.com> (8817670)
Update issue and pr templates (#2782)
Update issue templates
Update bug_report.md
test yaml template
add templates
update templates
add emojis
fix typo
Apply suggestions from code review
Co-authored-by: Kenneth Enevoldsen <kennethcenevoldsen@gmail.com>
update issue titles
update PR template
remove PR templates
Co-authored-by: Kenneth Enevoldsen <kennethcenevoldsen@gmail.com> (af7adbf)
bump ruff (#2784) (9e2e972)
model: Add Qwen3 Embedding model (#2769)
Init code
Remove extra config and lint code
use sentence transformer
add revisions
fix lint
Apply suggestions from code review
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>
fix lint
add framework
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com> (fe137d0)
fix: CachedEmbeddingWrapper issues in both documentation and code
Fixes #2772 (f7656d5)
fix: Update Caltech101 datasets to latest revision [v1]
fix: Update Caltech101 datasets to latest revision [v1] (#2778)
fix: Update Caltech101 datasets to latest revision [v2]
fixes: #2770 Fixes the issue, but only in v1
# tested using:
task: mteb.AbsTask = mteb.get_task("Caltech101ZeroShot")
task.load_data()
task.get_candidate_labels()
40f0841)ci: add new prefixes to releases
docs: Leaderboard simplifications
docs: Leaderboard simplifications (#2764)
docs: Leaderboard simplifications
Simplified sidebar, notably:
I also restructured the code so that nesting is easier.
Is it also possible to create a seperate section (see dummy screenshot)
refactor to reduce nesting
format (33fddfe)
fix: add xet support (#2603)
add xet version
add doc comment
change xet requirements
Update docs/usage/usage.md
Co-authored-by: Kenneth Enevoldsen <kennethcenevoldsen@gmail.com> (5ffcd63)
The type SIMILARITY is invalid. Correct one: SEMANTIC_SIMILARITY. See https://cloud.google.com/vertex-ai/generative-ai/docs/embeddings/task-types#supported_task_types (db43120)
ci: Delete cache in Model loading test only when model is loaded
ci: Delete cache in Model loading test only when model is loaded (#2761)
only delete cache when model loaded
testing it out (9827ec8)
fix: Add cadet-embed-base-v1 (#2727)
update
update overview.py for models
update
update (39a391d)
96706a8)docs: Updated description of FEVER
docs: Updated description of FEVER (#2745)
docs: Updated description of FEVER
Update the description to state that the corpus is the same as fever as we have have multiple questions on it
82f0bb9)fix: Update caltech101 (#2759)
docs: Updated description of FEVER
Update the description to state that the corpus is the same as fever as we have have multiple questions on it
Run both versions of one of the task using nomic-ai/nomic-embed-text-v1.5 and both scores match:
{
"dataset_revision": "851374102055782c84f89b1b4e9d128a6568847b",
"task_name": "Caltech101",
"mteb_version": "1.38.4",
"scores": {
"test": [
{
"accuracy": 0.897863,
{
"dataset_revision": "52439cf6d4f6ebf563d8cdc7f2c5371d9efd2686",
"task_name": "Caltech101",
"mteb_version": "1.38.4",
"scores": {
"test": [
{
"accuracy": 0.897929,
``` ([`1651f60`](https://github.com/embeddings-benchmark/mteb/commit/1651f60afeed767eba0fd0aae895108080301fee))
## Unknown
* Update Seed1.5 training data (#2749)
* update seed1.5 training data
* update seed1.5 training data ([`ccc5714`](https://github.com/embeddings-benchmark/mteb/commit/ccc5714938c690998a3c3e5807f55436e3e29af5))
* Update tasks & benchmarks tables ([`fe6729f`](https://github.com/embeddings-benchmark/mteb/commit/fe6729ff4ebbbe4373a984ab9dadbdf7df5a4f2c))
* Backfill task metadata for metadata for BigPatentClustering and AllegroReviews (#2755)
* big-patent
* allegro-reviews ([`e151fbd`](https://github.com/embeddings-benchmark/mteb/commit/e151fbd7920e5778a16c0cbeefcfb239362935e2))
fix: Correct embedding dimension for bge-m3
update metadata and add colsmol
fix: Add colpali models family (#2721)
add colpali models
add colpali as framework
add colpali as framework
update metadata and add colsmol
ix typos
account for revision
add training data info and lint
modify meta
correct colmodels meta and add colnomic 7b
fix typo in toml (colpali subdeps)
refine colmodel loading and metadata (6303839)
fix: Rename display name of VDR (#2734) (`b0988e2`)
fix: Promote Persian benchmark to v1
fix: Promote Persian benchmark to v1 (#2707)
Switch versioning from beta to v1 and add v1 to benchmark selector
Update Farsi benchmark display name, task IDs, and metadata
Add Hakim Model
fix hakim version
update
make lint
fix: Promote Persian benchmark to v1
Co-authored-by: mehran <mehan.sarmadi16@gmail.com>
Co-authored-by: Kenneth Enevoldsen <kennethcenevoldsen@gmail.com> (1098109)
274b28e)# 1.38.17 (2025-05-27) ## Fix * fix: IndicQARetrieval loader (#2729) * fix indic qa * add kwargs (`c3b66d9`)
clean vidore names + per lang eval
fix: Add vidore v2 benchmarks (#2713)
adding vidore benchmarks
fix typo
clean vidore names + per lang eval
lint
vidore names
bibtex fix
fix revision
vidore v2 citation
update citation format and fix per-language mappings
lint: citations
typo citations (175de94)
f3e706c)fix: Update Seed1.5-Embedding API
fix: Update Seed1.5-Embedding API (#2724)
update seed1.5-embedding api
update seed1.5-embedding api
update Seed1.5-Embedding API
update Seed1.5-Embedding resolve comments
update Seed1.5-Embedding lint
Update mteb/models/seed_models.py
Co-authored-by: Kenneth Enevoldsen <kennethcenevoldsen@gmail.com> (1da660e)
fix: Ara and ben classification dataset cleaning (#2632)
Improve classification datasets quality for ara and ben langs
add missing AJGT
fix format
change ajgt description
Fix numbers in description, add link to pull request
Add too short filter
Link in markdown format (4093099)
docs: fix number of tasks for eng, v2 in docs (#2720) (`7586624`)
fix: Integrate lightonai/GTE-ModernColBERT-v1
fix: Integrate lightonai/GTE-ModernColBERT-v1 (#2708)
fix: Integrate lightonai/GTE-ModernColBERT-v1
Fixes #2673
2b13659)fix: Rename gemini-embedding-exp-03-07 to gemini-embedding-001
fix: Rename gemini-embedding-exp-03-07 to gemini-embedding-001 (#2711)
Rename gemini-embedding-exp-03-07 to gemini-embedding-001
update referenfe link to the vertexAI API doc (0c0ad05)
docs: Updated the PR template and improved submission docs
docs: Updated the PR template and improved submission docs (#2704)
docs: Updated the PR template and improved submission docs
fixes #2568
Co-authored-by: Isaac Chung <chungisaac1217@gmail.com>
Co-authored-by: Isaac Chung <chungisaac1217@gmail.com> (835f6e6)
fix: Remove models from the leaderboard (#2705)
fix: Remove models from the leaderboard
I remove both models from the leaderboard by unlinking them from the import tree. I think this is the easiest way to add a model that not currently public.
78080cd)fix: Only install mteb into site packages
fix: Only install mteb into site packages (#2618)
Restrict installation directory
fix
namespace false
add star
add pont
fix import
fix import
add init files
fix setuptools find
fix image init
add missing templates
Co-authored-by: Roman Solomatin <36135455+Samoed@users.noreply.github.com> (1c803a1)
Fixes mistakes introduced in https://github.com/embeddings-benchmark/mteb/pull/2424
It seems like many of these requirements doesn't exist (voyageai>=1.0.0). @ayush1298 I am hoping you could clear up how this happened? (7222458)
Merge branch 'main' of https://github.com/embeddings-benchmark/mteb (272b20e)
Fix for Openai_Text-Embedding3-Small (#2702)
Fix for Openai_Text-Embedding3-Small
better syntax for readability (2c1fb62)
Fix for Openai_Text-Embedding3-Small (#2702)
Fix for Openai_Text-Embedding3-Small
better syntax for readability (94edb6d)
Correction in docs (#2688) (e97f606)
fix modality for OVENIT2TRetrieval
MTEB(Code, v1) languages (#2679)fix code languages (40ce571)
OVENIT2TRetrieval (#2678)fix modality (21506ed)
Leaderboard: UI simplifications for menus (#2672)
Leaderboard: UI simplifications for menus
Did a few things to improve the simplify the leaderboard UI.
Changes:
refactors:
fixed comment
fixes for sizes (debab47)
fix: Allow empty string for openai models
fix: Allow empty string for openai models (#2676)
fix for empty string input to openai/text-embedding-3-large
fix: Allow empty string in openai models
closes: #1650
fix based on review
Updated docstring
Co-authored-by: ayush1298 <munotayush6@kgpian.iitkgp.ac.in> (6f0b08d)
Update final version of Doubao-1.5-Embedding (Rename to Seed1.5-Embedding) (#2674)
update seed-embedding
update seed models
fix linting and tiktoken problem
fix tiktoken bug
fix lint
update name
Update mteb/models/seed_models.py
adopt suggestion
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>
update logging
update lint
update link
update revision
update Doubao-1.5-Embedding revision 3
rename Doubao-1.5-Embedding to Seed1.5-Embedding
Co-authored-by: zhangpeitian <zhangpeitian@bytedance.com>
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com> (5194c23)
fix: Update datasets wich can't be loaded with datasets>=3.0
datasets>=3.0 (#2661)fix: Update datasets wich can't be loaded with datasets>=3.0 (#1619)
reupload datasets
fix loader
remove commented code
lint
update pyproject dependencies (1ba6716)
rename model RELLE to CHAIN19 (#2671)
Add relle
defined model metadata for relle
Add mteb/models/relle_models.py
Update mteb/models/relle_models.py
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>
run after "make lint"
Add model into model_modules and lint check
rename model change model name
rename model change model name
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com> (bbbaa42)
fix: SIB200 machine translated > human translated
As correctly pointed out in:
https://huggingface.co/datasets/mteb/sib200/discussions/1 (ebdf0ca)
Add tests for leaderboard build (#2631)
Add tests for leaderboard build
add new action
remove build tests from other actions
fix tests
correct exclusion of test
added timeout constant (0b7f571)
fix: Update VisualSTS Aggregate task modalities
fix: Update VisualSTS Aggregate task modalities (#2597)
Update STS17MultilingualVisualSTS.py
fix STSBenchmarkMultilingualVisualSTS
Co-authored-by: Isaac Chung <chungisaac1217@gmail.com> (671dc04)
CI format citations (#2649)
ci format citations
add files
remove from lint CI
test lint
test lint
fix names (03b9a7f)
Remove typer dependency from citation script (#2629)
remove typer dependency from citation script (2f27544)
Update tasks & benchmarks tables (9deae69)
Revert "CI: fix infinitely committing issue (#2616)" (#2636)
This reverts commit 82dcb3dd03da0294d8a72ddca7a951680f33d67d. (6ee7e46)
remove irrelevant test (c8949be)
add Bilingual English-Danish parallel corpus from The Danish Medicines Agency (#2633)
add Bilingual English-Danish parallel corpus from The Danish Medicines Agency
bump dataset revision
format bibtex
format bibtex (c61ffb4)
Add Talemaader pair classification task (#2621)
Add talemaader pair classification task (a52ea2f)
fix: Removed missing dataset for MTEB(Multilingual) and bumped version
We should probably just have done this earlier to ensure that the multilingual benchamrk is runable. (f063638)
lint (2ecd7ad)
Merge branch 'main' of https://github.com/embeddings-benchmark/mteb (485941b)
Add ScandiSent dataset (#2620)
add scandisent dataset
add to init
typo (cb57999)
CI: fix infinitely committing issue (#2616)
fix token
try to trigger
add token
test ci
Update tasks & benchmarks tables
Update tasks & benchmarks tables
remove test lines
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com> (82dcb3d)
Update gradio version (#2558)
Update gradio version
Closes https://github.com/embeddings-benchmark/mteb/issues/2557
bump gradio (eabd9a5)
Update tasks & benchmarks tables (603aa5b)
CI: fix table (#2615) (4d09a1a)
Update tasks & benchmarks tables (20baefb)
Update tasks & benchmarks tables (69937da)
Update tasks & benchmarks tables (b2bfa6b)
Update tasks & benchmarks tables (94e7585)
Update tasks & benchmarks tables (73afd47)
Update tasks & benchmarks tables (b0c9e63)
Update tasks & benchmarks tables (60e1d2e)
Update tasks & benchmarks tables (979716d)
Update tasks & benchmarks tables (eb30080)
Update tasks & benchmarks tables (26cb06c)
Update tasks & benchmarks tables (e228e94)
Update tasks & benchmarks tables (1c1e179)
Update tasks & benchmarks tables (17120b2)
Update tasks & benchmarks tables (4844ab5)
Update tasks & benchmarks tables (43b364f)
Update tasks & benchmarks tables (d92e507)
Update tasks & benchmarks tables (e47b902)
Update tasks & benchmarks tables (df48ec9)
Update tasks & benchmarks tables (4584831)
Update tasks & benchmarks tables (b620a12)
Update tasks & benchmarks tables (44cec12)
Update tasks & benchmarks tables (480ba52)
Update tasks & benchmarks tables (4fd00cc)
Update tasks & benchmarks tables (11b5b33)
Update tasks & benchmarks tables (36e4172)
Update tasks & benchmarks tables (5b65218)
Update tasks & benchmarks tables (ee272f2)
Update tasks & benchmarks tables (309c51f)
Update tasks & benchmarks tables (b59392d)
Update tasks & benchmarks tables (4b97a83)
Update tasks & benchmarks tables (62a967b)
Update tasks & benchmarks tables (b720cfd)
Update tasks & benchmarks tables (5f4daf5)
Update tasks & benchmarks tables (5694e30)
Update tasks & benchmarks tables (e4935e2)
Update tasks & benchmarks tables (000d5bf)
Update tasks & benchmarks tables (ab42110)
Update tasks & benchmarks tables (bd7e85a)
Update tasks & benchmarks tables (ca5c3ad)
Update tasks & benchmarks tables (422fca2)
Update tasks & benchmarks tables (8e7d4f4)
Update tasks & benchmarks tables (e731eaa)
Update tasks & benchmarks tables (758be74)
Update tasks & benchmarks tables (258dd4e)
Update tasks & benchmarks tables (daa7807)
Update tasks & benchmarks tables (3e41806)
Update tasks & benchmarks tables (437b5e6)
Update tasks & benchmarks tables (2dd457c)
Update tasks & benchmarks tables (e2d43cf)
Update tasks & benchmarks tables (4e42192)
Update tasks & benchmarks tables (936dafb)
Update tasks & benchmarks tables (5f1b3d0)
Update tasks & benchmarks tables (e9b5706)
Update tasks & benchmarks tables (630c5bb)
Update tasks & benchmarks tables (d03650c)
Update tasks & benchmarks tables (1ae5750)
Update tasks & benchmarks tables (da5bf31)
Update tasks & benchmarks tables (49c33e2)
Update tasks & benchmarks tables (b78ea7d)
Update tasks & benchmarks tables (a2427d3)
Update tasks & benchmarks tables (fccc9b7)
Update tasks & benchmarks tables (4dcffb9)
Update tasks & benchmarks tables (17ae76e)
Update tasks & benchmarks tables (18faed2)
Update tasks & benchmarks tables (5d9332c)
Update tasks & benchmarks tables (5aafe93)
Update tasks & benchmarks tables (c32f3a9)
Update tasks & benchmarks tables (a2e14ae)
Update tasks & benchmarks tables (6ae644c)
Update tasks & benchmarks tables (6ed9b90)
Update tasks & benchmarks tables (0c24f8d)
Update tasks & benchmarks tables (23b999d)
Update tasks & benchmarks tables (350181f)
Update tasks & benchmarks tables (2168d9c)
Update tasks & benchmarks tables (f1f09f8)
Update tasks & benchmarks tables (6432ea8)
Update tasks & benchmarks tables (917263c)
Update tasks & benchmarks tables (952070e)
Update tasks & benchmarks tables (86069c7)
Update tasks & benchmarks tables (2c2ed55)
Update tasks & benchmarks tables (cd4670c)
Update tasks & benchmarks tables (cd83936)
Update tasks & benchmarks tables (54b863e)
Update tasks & benchmarks tables (b65f0ec)
Update tasks & benchmarks tables (3fd7bec)
Update tasks & benchmarks tables (9e5ce29)
Update tasks & benchmarks tables (2942557)
Update tasks & benchmarks tables (046ecf0)
Update tasks & benchmarks tables (90cd48a)
Update tasks & benchmarks tables (0eec584)
Update tasks & benchmarks tables (bd9bb89)
Update tasks & benchmarks tables (607eb6f)
Update tasks & benchmarks tables (f4d72bc)
Update tasks & benchmarks tables (f17902a)
Update tasks & benchmarks tables (296c1ee)
Update tasks & benchmarks tables (edbf218)
Update tasks & benchmarks tables (edb9c78)
Update tasks & benchmarks tables (72eea70)
Update tasks & benchmarks tables (c54e88f)
Update tasks & benchmarks tables (3703f11)
Update tasks & benchmarks tables (5b34e6a)
Update tasks & benchmarks tables (0665cd2)
Update tasks & benchmarks tables (8914793)
Update tasks & benchmarks tables (f9b747f)
Update tasks & benchmarks tables (61c611f)
Update tasks & benchmarks tables (ad232aa)
Update tasks & benchmarks tables (75db6fb)
Update tasks & benchmarks tables (37f86e2)
Update tasks & benchmarks tables (7060607)
Update tasks & benchmarks tables (114c273)
Update Doubao-1.5-Embedding revision (#2613)
update seed-embedding
update seed models
fix linting and tiktoken problem
fix tiktoken bug
fix lint
update name
Update mteb/models/seed_models.py
adopt suggestion
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>
update logging
update lint
update link
update revision
Co-authored-by: zhangpeitian <zhangpeitian@bytedance.com>
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com> (bcf532e)
Update tasks & benchmarks tables (3a7b723)
Update tasks & benchmarks tables (7bc22e2)
CI: update benchmark table (#2609)
update benchmark table
fix table (c020ebb)
Update Doubao-1.5-Embedding (#2611)
update seed-embedding
update seed models
fix linting and tiktoken problem
fix tiktoken bug
fix lint
update name
Update mteb/models/seed_models.py
adopt suggestion
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>
update logging
update lint
update link
Co-authored-by: zhangpeitian <zhangpeitian@bytedance.com>
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com> (7eba525)
add models from collection and revisions
fix: Add WebSSL models (#2604)
add 2 web SSL dino models
add models from collection and revisions
update memory_usage_mb and embed dim
use automodel instead (afb72ac)
fix mieb citation (#2606) (5a74754)
update Doubao-1.5-Embedding (#2575)
update seed-embedding
update seed models
fix linting and tiktoken problem
fix tiktoken bug
fix lint
update name
Update mteb/models/seed_models.py
adopt suggestion
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>
update logging
update lint
Co-authored-by: zhangpeitian <zhangpeitian@bytedance.com>
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com> (0bda363)
Add image only MIEB benchmark to LB left panel (#2596)
Update benchmarks.py
make lint
add to left side bar (039a965)
Add MIEB image only benchmark (#2590)
add vision only bench
add description
correct zs task modalities
specify tasks param (7b6d9d7)
fix codecarbon version (#2587) (ca10bac)
fix FlagEmbedding package name (#2588) (b1606ff)
Remove the comments from ImageEncoder (#2579) (`951bae3`)
fix: Add Encodechka benchmark (#2561)
add tasks
add benchmark
fix imports
update stsb split (0737e78)
Update tasks table (4f23d62)
Remove the comments from ImageEncoder (#2579) (951bae3)
move icon & name to benchmark dataclass (#2573) (fa5f034)
Update tasks table (e03333f)
Add missing annotations (#2498) (235906b)
Docs: Improve MIEB docs (#2569) (adfd92a)
Add ModelMeta for CodeSearch-ModernBERT-Crow-Plus (#2570)
Add files via upload
Update shuu_model.py
Update shuu_model.py
Update shuu_model.py
Update shuu_model.py
Update shuu_model.py
Update shuu_model.py
Update shuu_model.py
Update shuu_model.py
Update shuu_model.py
Update shuu_model.py
Update shuu_model.py
Update shuu_model.py
Update overview.py
Update shuu_model.py
Update shuu_model.py
Update shuu_model.py
Update mteb/models/shuu_model.py
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com> (713635a)
Update tasks table (e56aab5)
Backfill task metadata for metadata for GermanDPR and GermanQuAD (#2566)
Add metadata for GermanDPR and GermanQuAD
PR improvements (c0d3ca0)
Add relle (#2564)
Add relle
defined model metadata for relle
Add mteb/models/relle_models.py
Update mteb/models/relle_models.py
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>
run after "make lint"
Add model into model_modules and lint check
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com> (f11ac2a)
d475c7e)fix: jasper models embeddings having nan values (#2481) (`f7072d5`)
f7072d5)Fix leaderboard entry for BuiltBench (#2562)
Co-authored-by: Mehrzad Shahin-Moghadam <mehr@Mehrzads-MacBook-Pro.local> (4b755a3)
add USER2 (#2560)
add user2
add training code
update prompts (5ed6773)
Bumped gradio version to latest
feat: UI Overhaul (#2549)
Bumped gradio version to latest
Added new Gradio table functionality to leaderboard
Removed search bar
Changed color scheme in plot to match the table
Added new benchmark selector in sidebar
Changed not activated button type to secondary
Short-circuited callbacks that are based on language selection
Re-added column width calculation since it got messed up
Commented out gradient for per-task table as it slowed things down substantially
Styling and layout updates
Adjusted comments according to reviews
Converted all print statements to logger.debug
Removed pydantic version fix
Ran linting
Remove commented out code
Co-authored-by: Kenneth Enevoldsen <kennethcenevoldsen@gmail.com>
Moved English,v1 to Legacy section
Closed the benchmark sharing accordion by default
Adjusted markdown blocks according to suggestions
Ran linter
Co-authored-by: Kenneth Enevoldsen <kennethcenevoldsen@gmail.com> (0ab947b)
fix unintentional working of filters on leaderboard (#2535)
fix unintentional working of filters on leaderboard
address comments
make lint
address comments
rollback unnecessary changes (50d7e9e)
fix e5_R_mistral_7b (#2490)
fix e5_R_mistral_7b
change wrapper
address comments
Added kwargs for pad_token
correct lang format
address comments
add revision
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com> (4a6e539)
feat: Added dataframe utilities to BenchmarkResults
feat: Added dataframe utilities to BenchmarkResults (#2542)
fix: Added dataframe utilities to BenchmarkResults
get_results_table. I was considering renaming it to to_dataframe to align with tasks.to_dataframe. WDYT?Prerequisite for #2454:
@ayush1298 can I ask you to review this PR as well? I hope this give an idea of what I was hinting at. Sorry that it took a while. I wanted to make sure to get it right.
refactor to to_dataframe and combine common dependencies
ibid
fix revision joining after discussion with @x-tabdeveloping
remove strict=True for zip() as it is a >3.9 feature
updated mock cache (8fe5742)
fix me5 trainind data config to include xquad dataset (#2552)
fix: me5 trainind data config to include xquad dataset
Update mteb/models/e5_models.py
upddate: xquad key name
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com> (1f82b59)
Add xlm_roberta_ua_distilled (#2547)
defined model metadata for xlm_roberta_ua_distilled
Update mteb/models/ua_sentence_models.py
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>
included ua_sentence_models.py in overview.py
applied linting, added missing fields in ModelMeta
applied linting
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com> (3ff993d)
Add MIEB to README (75d3597)
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Your coding agent can read these notes before it upgrades. Set up the MCP server →