NewYour coding agent can read the release notes before it upgrades.Set up the MCP server →
PyPI · #3941 most downloaded on PyPI
A framework for evaluating language models
Last release 1 months ago
31 Aug 2026
Ships fairly regularly
a new release about every 2 months
Nearly every release is documented
notes for 17 of 18 stable releases
Nothing withdrawn
no release was ever pulled
5 years old
20 releases · first in 2021
One column per quarter.
Nothing published for this version
A fix-focused release. The main fixes are for few-shot leakage, a multiple-choice filter bug, and group stderr, alongside two new ONNX backends and ei
A fix-focused release. The main fixes are for few-shot leakage, a multiple-choice filter bug, and group stderr, alongside two new ONNX backends and eight new benchmark suites. Also updated most configs for datasets>=4, which accounts for much of the diff by volume.
Fixes that may shift previously reported numbers:
gen_prefix was resolved against the eval doc rather than the few-shot doc — splicing the evaluated question into every shot for RULER niah_single_1, humaneval_instruct, and humaneval_64_instruct by @adityasingh2400 in #3979MultiChoiceRegexFilter prefix-shadowing. Regex alternation is leftmost-wins, so a choice that prefixed a longer choice ("Guilty" vs "Guilty of Romance") matched inside it and scored correct answers as wrong — visible on BBH movie_recommendation by @iamsharduld in #3884weight_by_size: false. Groups reported the size-weighted pooled stderr even when the point estimate was an unweighted mean, giving error bars up to ~3x too narrow. Only unequal-sized subtasks change by @iamsharduld in #3882minerva_math answer normalization. sqrt shorthand no longer corrupts indexed roots (#4037), the thousands-separator strip no longer fuses digit tuples (0,1 → 01, #4039), an answer identical to the gold now scores correct (#4034), and the few-shot prompt LaTeX is corrected (#4045) by @feiiiiii5 and @nata2627. putnam_axiom shares these helpers and picks up the same fixes.onnxruntime — raw onnxruntime backend for Model Builder ONNX exports by @amd-sourjya in #3984onnxruntime-genai — cross-platform ONNX Runtime GenAI backend, with winml refactored on top of it by @thiagocrepaldi in #3960Install with pip install lm_eval[onnxruntime] or lm_eval[onnxruntime-genai]. For non-CPU execution providers install the matching wheel instead — onnxruntime-gpu / onnxruntime-rocm, or onnxruntime-genai-cuda / onnxruntime-genai-directml — these are mutually exclusive.
indicxnli_gu) by @bhaumik611 in #4056global_piqa_completions and global_piqa_prompted groups are now global_piqa_cloze and global_piqa_generation by @baberabb in #3816afrimmlu / afrimgsm / afrixnli task registration by @discobot in #3841, fixed the broken afrisenti / mafand prompt_2 group references by @DaoyuanLi2816 in #3847, and switched afrixnli prompt_1 doc_to_text to Jinja braces by @Solaris-star in #3944TypeError: unhashable type: 'dict' before inference when the tokenizer arrived as anything but a plain string (e.g. under local-chat-completions); the name is now resolved before the cached lookup by @nata2627 in #4048max_seq_lengths < 4096 by @vnayakde in #3372code_sim_score skips a blank leading line before extraction by @cameronshinn in #3921humaneval_random_span_infilling_light was registered under the wrong task name, causing a KeyError by @jaydeepborkar in #3768trasnlation typo corrected in translation task names by @DaoyuanLi2816 in #3824dataset_kwargs is now passed through (#3230) and init sets the task name (#3225) by @mprahlformat_span filter — normalizes labels only, leaving entity text untouched by @k-dickinson in #3887fewshot_config.split — a nested fewshot_config.split now takes precedence over the inherited top-level fewshot_split, as documented by @chuenchen309 in #3937think_end_token by @yaodong-shen in #3959auto:N batch size is parsed correctly in API models by @AbdullahRasheed45 in #3970, and batch_size="auto" no longer raises on the neuronx backend by @AbdullahRasheed45 in #3971max_length is detected from a nested text_config (Gemma3 multimodal) by @OrionArchitekton in #3916, max_cpu_memory is passed through to accelerate's max_memory by @Anai-Guo in #4016, and tokenizers that decode to an empty string fall back to tokenizer.eos_token by @ganeshr10 in #3657device argument by @Apeironics in #3803; sglang argument passing fixed by @wm901115nwpu in #3817--metadata accepts key=value pairs in addition to JSON by @baberabb in #4054--use_cache keys are hashable for multimodal image/byte requests by @feiiiiii5 in #4040delete_cache is a no-op when the directory is absent by @feiiiiii5 in #4035seed from a config file is normalized the same way the CLI does by @winklemad in #4003scripts/requests_caching.py entrypoint repaired — bad kwargs and a helper defined after its use by @Anai-Guo in #4059word_order=2) by @KrishVenky in #3780; TER metric direction corrected and the chrF docstring fixed by @borgr in #3993int64/int32, by @iamsharduld in #3885gen_prefix (RULER niah_single_1, humaneval_instruct, humaneval_64_instruct) and anywhere the sampler previously drew the eval document into its own shots. Prior numbers on those tasks may not be comparable.minerva_math and leaderboard_math_* now diverge. lm_eval/tasks/leaderboard/math/utils.py is deliberately frozen to reproduce Open LLM Leaderboard v2 scoring; the normalization fixes landed in minerva_math only.global_piqa task names changed — see Task Changes above.datasets<4 to keep script-based tasks working, you can unpin. Out-of-tree task YAMLs pointing at script-based datasets still need the same treatment.winml now sits on top of onnxruntime-genai and pulls in that extra.megatron_lm backend against Megatron-LM core_v0.18+: argument parsing moved out of initialize_megatron() by @oberpierre in #3991Note truncated.
Nothing published for this version
evaluate() accepts both shapes; load_task_or_group(...) and get_task_dict(...) are deprecated shims that return the old shape.
New release with four new model backends, tensor parallel support for transformers based models (hf), new benchmarks, a TaskManager refactor, and a long tail of task correctness fixes.
trt-llm) — NVIDIA TensorRT-LLM backend for optimized GPU inference by @Tracin in #3628megatron-lm) — Megatron-LM backend with TP/EP/DP support by @shangxiaokang in #3521 (with follow-up hardening in #3607)optimum-habana by @12010486 in #3550litellm) — Use LiteLLM as a unified API gateway for 100+ providers by @RheagalFire in #3721transformers models via tp_plan by @YangKai0616 in #3692TaskManager Refactor (#3549)TaskManager.load(...) returns a flat {tasks, groups} dict instead of the legacy nested {ConfigurableGroup: {name: Task}}. evaluate() accepts both shapes; load_task_or_group(...) and get_task_dict(...) are deprecated shims that return the old shape.Group class directly holds its child tasks; ConfigurableGroup is now a deprecated wrapper around it.include_path entries still override defaults.)SteeredHF renamed to SteeredModel — update imports if you're using the steering backend by @adrian-sauter in #3592>=0.18 as part of the data-parallel-with-Ray fixes by @baberabb in #3725enable_thinking is now disallowed for multiple_choice / loglikelihood tasks, and think_end_token is now required when enable_thinking=True. Configurations that combined these previously failed silently by @fxmarty-amd in #3675doc_to_text keeping a blank marker and dropping the question body by @Chessing234 in #3716doc_to_decontamination_query pointing at a nonexistent query field by @Chessing234 in #3718doc_to_decontamination_query pointing at nonexistent texte field by @Chessing234 in #3719dataset_path by @zhngstl in #3723!function imports to use absolute module paths by @Anai-Guo in #3731RephraseChecker.strip_changes greedy-regex bug by @Chessing234 in #3737CohereForAI org with CohereLabs by @juliafalcao in #3631--cache_requests always fails due to argparse type/choices conflict by @maxidl in #3588resolve_hf_chat_template by @DarkLight1337 in #3595CohereForAI organization to CohereLabs by @juliafalcao in #3631WatsonxLLM class mapping and errors by @Rafal-Chrzanowski-IBM in #3591enable_thinking with output_type: multiple_choice tasks / loglikelihood tasks; raise error in case think_end_token is not provided with enable_thinking=True by @fxmarty-amd in #3675gpu_memory_utilization=0.9 by @baberabb in #3732mmlu_pro by @baberabb in #3747Note truncated.
Minor release. Stay tuned for bigger changes next release.
Minor release. Stay tuned for bigger changes next release.
The following tasks have updated versions. Results from a previous task versions may not be directly comparable. See the linked PRs or individual task READMEs for changelogs.
afrobench_belebele (all variants): 2 → 3 in #3551
evalita_llm: 0.0 → 0.1 in #3551
include (all 90 language variants): 0.0 → 0.1 in #3551
mgsm_direct (all 11 language variants): 3.0 → 4.0 by @LakshyaChaudhry in #3574
headline_text field with headline by @Mr-Neutr0n in #3567eval() with ast.literal_eval in task configs for safer parsing by @baberabb in #3577hf_transfer import check by @baberabb in #3563modify_gen_kwargs call in vLLM VLMs by @hmellor in #3573gen_kwargs normalization inline to modify_gen_kwargs; fixed cached gen_kwargs mutation by @baberabb in #3582Full Changelog: v0.4.10...v0.4.11
Breaking Change: Lightweight Core with Optional Backends
The big change this release: the base package no longer installs model backends by default. We've also added new benchmarks and expanded multilingual support.
pip install lm_eval no longer installs the HuggingFace/torch stack by default. (#3428)
The core package no longer includes backends. Install them explicitly:
pip install lm_eval # core only, no model backends
pip install lm_eval[hf] # HuggingFace backend (transformers, torch, accelerate)
pip install lm_eval[vllm] # vLLM backend
pip install lm_eval[api] # API backends (OpenAI, Anthropic, etc.)Additional breaking change: Accessing model classes via attribute no longer works:
# This still works:
from lm_eval.models.huggingface import HFLM
# This now raises AttributeError:
import lm_eval.models
lm_eval.models.huggingface.HFLMThe CLI now uses explicit subcommands and supports YAML config files (#3440):
lm-eval run --model hf --tasks hellaswag # run evaluations
lm-eval run --config my_config.yaml # load args from YAML config
lm-eval ls tasks # list available tasks
lm-eval validate --tasks hellaswag,arc_easy # validate task configsBackward compatible when omitting run still works: lm-eval --model hf --tasks hellaswag
See lm-eval --help or the CLI documentation for details.
ContextSampler with new build_qa_turn helper (#3429)gen_kwargs with truncation_side support for vLLM (#3509)hf-mistral3) by @medhakimbedhief in #3487gen_prefix delimiter handling in multiple-choice tasks by @baberabb in #3508a=0 as valid answer index in build_qa_turn by @ezylopx5 in #3488fewshot_config not being applied to fewshot docs by @baberabb in #3461gsm8k_cot_llama target_delimiter issue by @baberabb in #3526max_length error by @baberabb in #3503vllm.transformers_utils.get_tokenizer import by @DarkLight1337 in #3482AutoModelForVision2Seq by @baberabb in #3522= sign in checkpoint names by @mrinaldi97 in #3517pretty_print_task for external custom configs by @safikhanSoofiyani in #3436Full Changelog: v0.4.9.2...v0.4.10
Resolved deprecated vllm.utils.get_open_port by @DarkLight1337 in #3398
This release continues our steady stream of community contributions with a batch of new benchmarks, expanded model support, and important fixes. A notable change: Python 3.10 is now the minimum required version.
A big wave of new evaluation tasks this release:
acc_norm_bytes metric by @baberabb in #3368Core Changes:
datasets library by @baberabb in #3316add_bos_token now defaults to None by @baberabb in #3347LOGLEVEL env var to LMEVAL_LOG_LEVEL to avoid conflicts by @fxmarty-amd in #3418Task Fixes:
error_type="ok" and display summary table by @fxmarty-amd in #3410, #3406crows_pairs dataset by @jannalulu in #3378add_bos_token not updating by @DarkLight1337 in #3206lambada_multilingual_stablelm by @jmichaelov, @HallerPatrick in #3294, #3222minerva_math by @baberabb in #3259Backend Fixes:
data_parallel_size>1 issue by @Dornavineeth in #3303vllm.utils.get_open_port by @DarkLight1337 in #3398additional_config parsing by @brian-dellabetta in #3393torch_dtype with dtype by @AbdulmalikDS in #3415tokenizer_info endpoint to avoid manual duplication by @m-misiura in #3185trust_remote_code: True from updated datasets by @Avelina9X in https://github.com/EleutherAI/lm-evaluation-harness/pull/3213add_bos_token not updated for Gemma tokenizer by @DarkLight1337 in https://github.com/EleutherAI/lm-evaluation-harness/pull/3206lambada_multilingual_stablelm by @HallerPatrick in https://github.com/EleutherAI/lm-evaluation-harness/pull/3222minerva_math by @baberabb in https://github.com/EleutherAI/lm-evaluation-harness/pull/3259acc_norm metric to BLiMP-NL by @jmichaelov in https://github.com/EleutherAI/lm-evaluation-harness/pull/3272acc_norm metric to ZhoBLiMP by @jmichaelov in https://github.com/EleutherAI/lm-evaluation-harness/pull/3271tokenizer_info endpoint to avoid manual duplication by @m-misiura in https://github.com/EleutherAI/lm-evaluation-harness/pull/3185humaneval_64_instruct task label in README to name in yaml file by @jmichaelov in https://github.com/EleutherAI/lm-evaluation-harness/pull/3344add_bos_token defaults to None by @baberabb in https://github.com/EleutherAI/lm-evaluation-harness/pull/3347vllm.utils.get_open_port by @DarkLight1337 in https://github.com/EleutherAI/lm-evaluation-harness/pull/3398error_type="ok" by @fxmarty-amd in https://github.com/EleutherAI/lm-evaluation-harness/pull/3410lambada_multilingual_stablelm by @jmichaelov in https://github.com/EleutherAI/lm-evaluation-harness/pull/3294LOGLEVEL to LMEVAL_LOG_LEVEL by @fxmarty-amd in https://github.com/EleutherAI/lm-evaluation-harness/pull/3418Full Changelog: https://github.com/EleutherAI/lm-evaluation-harness/compare/v0.4.9.1...v0.4.9.2
This release continues our steady stream of community contributions with a batch of new benchmarks, expanded model support, and important fixes. A notable change: Python 3.10 is now the minimum required version.
A big wave of new evaluation tasks this release:
acc_norm_bytes metric by @baberabb in #3368Core Changes:
datasets library by @baberabb in #3316add_bos_token now defaults to None by @baberabb in #3347LOGLEVEL env var to LMEVAL_LOG_LEVEL to avoid conflicts by @fxmarty-amd in #3418Task Fixes:
error_type="ok" and display summary table by @fxmarty-amd in #3410, #3406crows_pairs dataset by @jannalulu in #3378add_bos_token not updating by @DarkLight1337 in #3206lambada_multilingual_stablelm by @jmichaelov, @HallerPatrick in #3294, #3222minerva_math by @baberabb in #3259Backend Fixes:
data_parallel_size>1 issue by @Dornavineeth in #3303vllm.utils.get_open_port by @DarkLight1337 in #3398additional_config parsing by @brian-dellabetta in #3393torch_dtype with dtype by @AbdulmalikDS in #3415tokenizer_info endpoint to avoid manual duplication by @m-misiura in #3185trust_remote_code: True from updated datasets by @Avelina9X in #3213add_bos_token not updated for Gemma tokenizer by @DarkLight1337 in #3206lambada_multilingual_stablelm by @HallerPatrick in #3222minerva_math by @baberabb in #3259acc_norm metric to BLiMP-NL by @jmichaelov in #3272acc_norm metric to ZhoBLiMP by @jmichaelov in #3271tokenizer_info endpoint to avoid manual duplication by @m-misiura in #3185humaneval_64_instruct task label in README to name in yaml file by @jmichaelov in #3344add_bos_token defaults to None by @baberabb in #3347vllm.utils.get_open_port by @DarkLight1337 in #3398Note truncated.
This v0.4.9.1 release is a quick patch to bring in some new tasks and fixes. Looking aheas, we're gearing up for some bigger updates to tackle common
This v0.4.9.1 release is a quick patch to bring in some new tasks and fixes. Looking aheas, we're gearing up for some bigger updates to tackle common community pain points. We'll do our best to keep things from breaking, but we anticipate a few changes might not be fully backward-compatible. We're excited to share more soon!
think_end_token argument to strip intermediate reasoning from outputs for the hf, vllm, and sglang model backends. A related enable_thinking argument was also added for specific models that support it (e.g., Qwen).bbh_cot_fewshot prompts by @philipdoldo. (#3140)max_gen_toks to 2048 for HRM8K math benchmarks by @shing100. (#3124)DISABLE_MULTIPROC envar by @ankitgola005 and @neel04. (#3135, #3106)LMEVAL_HASHMM envar by @artemorloff in #2973include-path precedence handling to prioritize custom dir over default by @parkhs21 in #3068datasets < 4.0.0 temporarily to maintain compatibility with trust_remote_code by @baberabb. (#3172)fewshot_context by using kwargs by @kiersten-stokes in #3079TemplateError for chat_template by @baberabb in #3076subfolder from HF models by @younesbelkada in #3072LMEVAL_HASHMM envar by @artemorloff in #2973bbh_cot_fewshot: Removed repeated "Let''s think step by step." text from bbh cot prompts by @philipdoldo in #3140chat_template_args to vllm by @Avelina9X in #3164device from kwargs by @baberabb in #3181mmlu_continuation subgroup names to fit Readme and other variants by @lamalunderscore in #3137Full Changelog: v0.4.9...v0.4.9.1
Your coding agent can read these notes before it upgrades. Set up the MCP server →