NewYour coding agent can read the release notes before it upgrades.Set up the MCP server →
PyPI · #3941 most downloaded on PyPI
A framework for evaluating language models
Last release 1 months ago
31 Aug 2026
Ships fairly regularly
a new release about every 2 months
Nearly every release is documented
notes for 17 of 18 stable releases
Nothing withdrawn
no release was ever pulled
5 years old
20 releases · first in 2021
One column per quarter.
Breaking Changes & Important Updates
Enhanced Backend Support:
enable_thinking argument (#2947) and data parallel for V1 (#3011) by @anmarques and @baberabbMultimodal Capabilities:
Performance & Reliability:
quantization_config by @jerryzh168 in #2842yaml.CLoader for faster YAML loading by @giuliolovisotto in #2777Asian Languages:
European Languages:
African Languages:
Arabic Languages:
--examples argument for efficient multi-prompt evaluation by @felipemaiapolo and @mirianfsilva in #2520add_bos_token initialization by @baberabb in #2781softmax_dtype argument for HFLM by @Avelina9X in #2921cais/mmlu dataset source by @baberabb in #2918max_gen_toks to 2048 and max_length to 8192 for MMLU Pro tests by @dazipe in #2824We extend our heartfelt thanks to all contributors who made this release possible, including 43 first-time contributors who brought fresh perspectives and valuable improvements to the evaluation harness.
add_bos_token by @baberabb in #2781--examples Argument for Fine-Grained Task Evaluation in lm-evaluation-harness. This feature is the first step towards efficient multi-prompt evaluation with PromptEval [1,2] by @felipemaiapolo in #2520Note truncated.
Breaking Change : Python 3.8 support has been dropped as it reached end of life. Please upgrade to Python 3.9 or newer.
New Backend Support:
sparsify or sae_lens by @luciaquirke and @AMindToThinkBreaking Change: Python 3.8 support has been dropped as it reached end of life. Please upgrade to Python 3.9 or newer.
Added Support for gen_prefix in config, allowing you to append text after the <|assistant|> token (or at the end of non-chat prompts) - particularly effective for evaluating instruct models
Global Coverage:
Asian Languages:
European Languages:
Middle Eastern Languages:
We extend our thanks to all contributors who made this release possible and to our users for your continued support and feedback.
Thanks, the LM Eval Harness team (@baberabb and @lintangsutawika)
global_mmlu full version by @bzantium in #2636global_mmlu by @bzantium in #2652group from bigbench task configs by @baberabb in #2663assistant_prefill in docs to gen_prefix by @kiersten-stokes in #2683construct_requests kwargs in python tasks by @baberabb in #2700arithmetic: set target delimiter to empty string by @baberabb in #2701lighteval/MATH-Hard dataset with DigitalLearningGmbH/MATH-lighteval by @f4str in #2719Note truncated.
fix DeprecationWarning: invalid escape sequence '\s' for whitespace filter by @baberabb in #2560
This release includes several bug fixes, minor improvements to model handling, and task additions.
Python 3.8 support will be dropped in future releases as it has reached its end of life. Users are encouraged to upgrade to Python 3.9 or newer.
An important modification has been made to how delimiters are handled when applying chat templates in request construction, particularly affecting multiple-choice tasks. This change ensures better compatibility with chat models by respecting their native formatting conventions.
📝 For detailed documentation, please refer to docs/chat-template-readme.md
As well as several slight fixes or changes to existing tasks (as noted via the incrementing of versions).
Thanks, the LM Eval Harness team (@baberabb and @lintangsutawika)
metrics and filter to logged sample by @baberabb in #2517until by @baberabb in #2518loglikelihood_rolling across requests by @baberabb in #2559DeprecationWarning: invalid escape sequence '\s' for whitespace filter by @baberabb in #2560Full Changelog: v0.4.6...v0.4.7
We've now fully deprecated the use of group keys directly within a task's configuration file. The appropriate key to use is now solely tag for many ca…
We're excited to introduce prototype support for Vision Language Models (VLMs) in this release, using model types hf-multimodal and vllm-vlm. This allows for evaluation of models that can process text and image inputs and produce text outputs. Currently we have added support for the MMMU (mmmu_val) task and we welcome contributions and feedback from the community!
VLM models can be configured with several new arguments within --model_args to support their specific requirements:
max_images (int): Set the maximum number of images for each prompt.interleave (bool): Determines the positioning of image inputs. When True (default) images are interleaved with the text. When False all images are placed at the front of the text. This is model dependent.hf-multimodal specific args:
image_token_id (int) or image_string (str): Specifies a custom token or string for image placeholders. For example, Llava models expect an "<image>" string to indicate the location of images in the input, while Qwen2-VL models expect an "<|image_pad|>" sentinel string instead. This will be inferred based on model configuration files whenever possible, but we recommend confirming that an override is needed when testing a new model familyconvert_img_format (bool): Whether to convert the images to RGB format.lm_eval --model hf-multimodal --model_args pretrained=llava-hf/llava-1.5-7b-hf,attn_implementation=flash_attention_2,max_images=1,interleave=True,image_string=<image> --tasks mmmu_val --apply_chat_template
lm_eval --model vllm-vlm --model_args pretrained=llava-hf/llava-1.5-7b-hf,max_images=1,interleave=True --tasks mmmu_val --apply_chat_template
--apply_chat_template flag to ensure proper input formatting according to the model's expected chat template.max_images=1. Additionally, certain models expect image placeholders to be non-interleaved with the text, requiring interleave=False.We have currently most notably tested the implementation with the following models:
transformers from source)Several new tasks have been contributed to the library for this version!
New tasks as of v0.4.5 include:
As well as several slight fixes or changes to existing tasks (as noted via the incrementing of versions).
group versus tag splitWe've now fully deprecated the use of group keys directly within a task's configuration file. The appropriate key to use is now solely tag for many cases. See the v0.4.4 patchnotes for more info on migration, if you have a set of task YAMLs maintained outside the Eval Harness repository.
In HFLM, logic specific to handling inputs for Seq2seq (encoder-decoder models like T5) versus Causal (decoder-only autoregressive models, and the vast majority of current LMs) models previously hinged on a check for self.AUTO_MODEL_CLASS == transformers.AutoModelForCausalLM. Some users may want to use causal model behavior, but set self.AUTO_MODEL_CLASS to a different factory class, such as transformers.AutoModelForVision2Seq.
As a result, those users who subclass HFLM but do not call HFLM.__init__() may now also need to set the self.backend attribute to either "causal" or "seq2seq" during initialization themselves.
While this should not affect a large majority of users, for those who subclass HFLM in potentially advanced ways, see https://github.com/EleutherAI/lm-evaluation-harness/pull/2353 for the full set of changes.
We intend to further expand our multimodal support to a wider set of vision-language tasks, as well as a broader set of model types, and are actively seeking user feedback!
Thanks, the LM Eval Harness team (@baberabb @haileyschoelkopf @lintangsutawika)
evaluate by @baberabb in https://github.com/EleutherAI/lm-evaluation-harness/pull/2351eus_exams task configs by @baberabb in https://github.com/EleutherAI/lm-evaluation-harness/pull/2320self.backend from AUTO_MODEL_CLASS by @baberabb in https://github.com/EleutherAI/lm-evaluation-harness/pull/2353limit_mm_per_prompt by @baberabb in https://github.com/EleutherAI/lm-evaluation-harness/pull/2387Full Changelog: https://github.com/EleutherAI/lm-evaluation-harness/compare/v0.4.4...v0.4.5
…average score across multiple subtasks. See #Backwards Incompatibilities below for more information on changes and migration instructions.
This release includes the Open LLM Leaderboard 2 official task implementations! These can be run by using --tasks leaderboard. Thank you to the HF team (@clefourrier, @NathanHB , @KonradSzafer, @lozovskaya) for contributing these -- you can read more about their Open LLM Leaderboard 2 release here.
API support is overhauled! Now: support for concurrent requests, chat templates, tokenization, batching and improved customization. This makes API support both more generalizable to new providers and should dramatically speed up API model inference.
base_url to --model_args, for example, base_url=http://localhost:8000/v1/completions; concurrent requests are controlled with the num_concurrent argument; tokenization is controlled with tokenized_requests.--gen_kwargs as usual.local-completions using --apply_chat_template (either with or without tokenized_requests).
local-chat-completions (for e.g. with a OpenAI Chat API endpoint), but only the former supports loglikelihood tasks (e.g. multiple-choice). This is because ChatCompletion style APIs generally do not provide access to logits on prompt/input tokens, preventing easy measurement of multi-token continuations' log probabilities.lm_eval --model local-completions --model_args model=meta-llama/Meta-Llama-3.1-8B-Instruct,num_concurrent=10,tokenized_requests=True,tokenizer_backend=huggingface,max_length=4096 --apply_chat_template --batch_size 1 --tasks mmlulm_eval --model local-chat-completions --model_args model=meta-llama/Meta-Llama-3.1-8B-Instruct,num_concurrent=10 --apply_chat_template --tasks gsm8klocal-completions!We've reworked the Task Grouping system to make it clearer when and when not to report an aggregated average score across multiple subtasks. See #Backwards Incompatibilities below for more information on changes and migration instructions.
A combination of data-parallel and model-parallel (using HF's device_map functionality for "naive" pipeline parallel) inference using --model hf is now supported, thank you to @NathanHB and team!
Other new additions include a number of miscellaneous bugfixes and much more. Thank you to all contributors who helped out on this release!
A number of new tasks have been contributed to the library.
As a further discoverability improvement, lm_eval --tasks list now shows all tasks, tags, and groups in a prettier format, along with (if applicable) where to find the associated config file for a task or group! Thank you to @anthony-dipofi for working on this.
New tasks as of v0.4.4 include:
tags versus groups, and how to migratePreviously, we supported the ability to group a set of tasks together, generally for two purposes: 1) to have an easy-to-call shortcut for a set of tasks one might want to frequently run simultaneously, and 2) to allow for "parent" tasks like mmlu to aggregate and report a unified score across a set of component "subtasks".
There were two ways to add a task to a given group name: 1) to provide (a list of) values to the group field in a given subtask's config file:
# this is a *task* yaml file.
group: group_name1
task: my_task1
# rest of task config goes here...
or 2) to define a "group config file" and specify a group along with its constituent subtasks:
# this is a group's yaml file
group: group_name1
task:
- subtask_name1
- subtask_name2
# ...
These would both have the same effect of reporting an averaged metric for group_name1 when calling lm_eval --tasks group_name1. However, in use-case 1) (simply registering a shorthand for a list of tasks one is interested in), reporting an aggregate score can be undesirable or ill-defined.
We've now separated out these two use-cases ("shorthand" groupings and hierarchical subtask collections) into a tag and group property separately!
To register a shorthand (now called a tag), simply change the group field name within your task's config to be tag (group_alias keys will no longer be supported in task configs.):
# this is a *task* yaml file.
tag: tag_name1
task: my_task1
# rest of task config goes here...
Group config files may remain as is if aggregation is not desired. To opt-in to reporting aggregated scores across a group's subtasks, add the following to your group config file:
# this is a group's yaml file
group: group_name1
task:
- subtask_name1
- subtask_name2
# ...
### New! Needed to turn on aggregation ###
aggregate_metric_list:
- metric: acc # placeholder. Note that all subtasks in this group must report an `acc` metric key
- weight_by_size: True # whether one wishes to report *micro* or *macro*-averaged scores across subtasks. Defaults to `True`.
Please see our documentation here for more information. We apologize for any headaches this migration may create--however, we believe separating out these two functionalities will make it less likely for users to encounter confusion or errors related to mistaken undesired aggregation.
We're planning to make more planning documents public and standardize on (likely) 1 new PyPI release per month! Stay tuned.
Thanks, the LM Eval Harness team (@haileyschoelkopf @lintangsutawika @baberabb)
add_bos_token=True by @haileyschoelkopf in https://github.com/EleutherAI/lm-evaluation-harness/pull/2049trust_remote_code for Hellaswag by @haileyschoelkopf in https://github.com/EleutherAI/lm-evaluation-harness/pull/2029truthfulqa_gen by @haileyschoelkopf in https://github.com/EleutherAI/lm-evaluation-harness/pull/2090"recurrent_gemma" and other Gemma model types by @haileyschoelkopf in https://github.com/EleutherAI/lm-evaluation-harness/pull/2105lm_eval.caching module by @haileyschoelkopf in https://github.com/EleutherAI/lm-evaluation-harness/pull/2124scipy and skilearn dependencies by @nathan-weinberg in https://github.com/EleutherAI/lm-evaluation-harness/pull/2097revision kwarg dtype in edge-cases by @haileyschoelkopf in https://github.com/EleutherAI/lm-evaluation-harness/pull/2184simple_evaluate() LM TypeError by @haileyschoelkopf in https://github.com/EleutherAI/lm-evaluation-harness/pull/2258*ifeval tasks ( #2210 ) by @haileyschoelkopf in https://github.com/EleutherAI/lm-evaluation-harness/pull/2259loglikelihood_rolling caching ( #1821 ) by @haileyschoelkopf in https://github.com/EleutherAI/lm-evaluation-harness/pull/2187Full Changelog: https://github.com/EleutherAI/lm-evaluation-harness/compare/v0.4.3...v0.4.4
We're releasing a new version of LM Eval Harness for PyPI users at long last. We intend to release new PyPI versions more frequently in the future.
We're releasing a new version of LM Eval Harness for PyPI users at long last. We intend to release new PyPI versions more frequently in the future.
The big new feature is the often-requested Chat Templating, contributed by @KonradSzafer @clefourrier @NathanHB and also worked on by a number of other awesome contributors!
You can now run using a chat template with --apply_chat_template and a system prompt of your choosing using --system_instruction "my sysprompt here". The --fewshot_as_multiturn flag can control whether each few-shot example in context is a new conversational turn or not.
This feature is currently only supported for model types hf and vllm but we intend to gather feedback on improvements and also extend this to other relevant models such as APIs.
There's a lot more to check out, including:
Logging results to the HF Hub if desired using --hf_hub_log_args, by @KonradSzafer and team!
NeMo model support by @sergiopperez !
Anthropic Chat API support by @tryuman !
DeepSparse and SparseML model types by @mgoin !
Handling of delta-weights in HF models, by @KonradSzafer !
LoRA support for VLLM, by @bcicc !
Fixes to PEFT modules which add new tokens to the embedding layers, by @mapmeld !
Fixes to handling of BOS tokens in multiple-choice loglikelihood settings, by @djstrong !
The use of custom Sampler subclasses in tasks, by @LSinev !
The ability to specify "hardcoded" few-shot examples more cleanly, by @clefourrier !
Support for Ascend NPUs (--device npu) by @statelesshz, @zhabuye, @jiaqiw09 and others!
Logging of higher_is_better in results tables for clearer understanding of eval metrics by @zafstojano !
extra info logged about models, including info about tokenizers, chat templating, and more, by @artemorloff @djstrong and others!
Miscellaneous bug fixes! And many more great contributions we weren't able to list here.
We had a number of new tasks contributed. A listing of subfolders and a brief description of the tasks contained in them can now be found at lm_eval/tasks/README.md. Hopefully this will be a useful step to help users to locate the definitions of relevant tasks more easily, by first visiting this page and then locating the appropriate README.md within a given lm_eval/tasks subfolder, for further info on each task contained within a given folder. Thank you to @AnthonyDipofi @Harryalways317 @nairbv @sepiatone and others for working on this and giving feedback!
Without further ado, the tasks:
hendrycks_math task, the MATH task using the prompt and answer parsing from the original Hendrycks et al. MATH paper rather than Minerva's prompt and parsingThe save format for logged results has now changed.
output files will now be written to
{output_path}/{sanitized_model_name}/results_YYYY-MM-DDTHH-MM-SS.xxxxx.json if --output_path is set, and{output_path}/{sanitized_model_name}/samples_{task_name}_YYYY-MM-DDTHH-MM-SS.xxxxx.jsonl for each task's samples if --log_samples is set.e.g. outputs/gpt2/results_2024-06-28T00-00-00.00001.json and outputs/gpt2/samples_lambada_openai_2024-06-28T00-00-00.00001.jsonl.
See https://github.com/EleutherAI/lm-evaluation-harness/pull/1926 for utilities which may help to work with these new filenames.
In general, we'll be doing our best to keep up with the strong interest and large number of contributions we've seen coming in!
The official Open LLM Leaderboard 2 tasks will be landing soon in the Eval Harness main branch and subsequently in v0.4.4 on PyPI!
The fact that groups of tasks by-default attempt to report an aggregated score across constituent subtasks has been a sharp edge. We are finishing up some internal reworking to distinguish between groups of tasks that do report aggregate scores (think mmlu) versus tags which simply are a convenient shortcut to call a bunch of tasks one might want to run at once (think the pythia grouping which merely represents a collection of tasks one might want to gather results on each of all at once but where averaging doesn't make sense).
We'd also like to improve the API model support in the Eval Harness from its current state.
More to come!
Thank you to everyone who's contributed to or used the library!
Thanks, @haileyschoelkopf @lintangsutawika
neuralmagic models for sparseml and deepsparse by @mgoin in https://github.com/EleutherAI/lm-evaluation-harness/pull/1674--tasks list in README by @nairbv in https://github.com/EleutherAI/lm-evaluation-harness/pull/1726include by @lintangsutawika in https://github.com/EleutherAI/lm-evaluation-harness/pull/1749num_fewshot: 0 by @chujiezheng in https://github.com/EleutherAI/lm-evaluation-harness/pull/1769----hf_hub_log_args to --hf_hub_log_args by @MuhammadBinUsman03 in https://github.com/EleutherAI/lm-evaluation-harness/pull/1776--tasks list option in interface documentation by @sepiatone in https://github.com/EleutherAI/lm-evaluation-harness/pull/1792pretrained=gpt2 default by @haileyschoelkopf in https://github.com/EleutherAI/lm-evaluation-harness/pull/1775--hf_hub_log_args in interface documentation by @sepiatone in https://github.com/EleutherAI/lm-evaluation-harness/pull/1806docs by @zafstojano in https://github.com/EleutherAI/lm-evaluation-harness/pull/1863docs by @oneonlee in https://github.com/EleutherAI/lm-evaluation-harness/pull/1876batch_size=auto for HF Seq2Seq models (#1765) by @haileyschoelkopf in https://github.com/EleutherAI/lm-evaluation-harness/pull/1790lm_eval.logging -> lm_eval.loggers by @haileyschoelkopf in https://github.com/EleutherAI/lm-evaluation-harness/pull/1858higher_is_better tickers in output table by @zafstojano in https://github.com/EleutherAI/lm-evaluation-harness/pull/1893__main__.py by @sadra-barikbin in https://github.com/EleutherAI/lm-evaluation-harness/pull/1939docs/interface.md by @sadra-barikbin in https://github.com/EleutherAI/lm-evaluation-harness/pull/1955samples is newline delimited by @baberabb in https://github.com/EleutherAI/lm-evaluation-harness/pull/1930--gen_kwargs and VLLM (temperature not respected) by @haileyschoelkopf in https://github.com/EleutherAI/lm-evaluation-harness/pull/1800scripts.write_out error out when no splits match by @haileyschoelkopf in https://github.com/EleutherAI/lm-evaluation-harness/pull/1796fewshot_as_multiturn in results files by @haileyschoelkopf in https://github.com/EleutherAI/lm-evaluation-harness/pull/1995--trust_remote_code by @haileyschoelkopf in https://github.com/EleutherAI/lm-evaluation-harness/pull/1998LM dependency from build_all_requests by @baberabb in https://github.com/EleutherAI/lm-evaluation-harness/pull/2011trust_remote_code-related test failures by @haileyschoelkopf in https://github.com/EleutherAI/lm-evaluation-harness/pull/2024vllm by @baberabb in https://github.com/EleutherAI/lm-evaluation-harness/pull/2034exact_match HF Evaluate metric with install, don't call evaluate.load() on import by @haileyschoelkopf in https://github.com/EleutherAI/lm-evaluation-harness/pull/2045Full Changelog: https://github.com/EleutherAI/lm-evaluation-harness/compare/v0.4.2...v0.4.3
There were a few breaking changes to lm-eval's general API or logic we'd like to highlight:
We are releasing a new minor version of lm-eval for PyPI users! We've been very happy to see continued usage of the lm-evaluation-harness, including as a standard testbench to propel new architecture design (https://arxiv.org/abs/2402.18668), to ease new benchmark creation (https://arxiv.org/abs/2402.11548, https://arxiv.org/abs/2402.00786, https://arxiv.org/abs/2403.01469), enabling controlled experimentation on LLM evaluation (https://arxiv.org/abs/2402.01781), and more!
TemplateLM base class for lower-code new LM class implementations by @anjor--predict_only by @baberabbThere were a few breaking changes to lm-eval's general API or logic we'd like to highlight:
TaskManager APIpreviously, users had to call lm_eval.tasks.initialize_tasks() to register the library's default tasks, or lm_eval.tasks.include_path() to include a custom directory of task YAML configs.
Old usage:
import lm_eval
lm_eval.tasks.initialize_tasks()
# or:
lm_eval.tasks.include_path("/path/to/my/custom/tasks")
lm_eval.simple_evaluate(model=lm, tasks=["arc_easy"])
New intended usage:
import lm_eval
# optional--only need to instantiate separately if you want to pass custom path!
task_manager = TaskManager() # pass include_path="/path/to/my/custom/tasks" if desired
lm_eval.simple_evaluate(model=lm, tasks=["arc_easy"], task_manager=task_manager)
get_task_dict() now also optionally takes a TaskManager object, when wanting to load custom tasks.
This should allow for much faster library startup times due to lazily loading requested tasks or groups.
Previous versions of the library incorrectly reported erroneously large stderr scores for groups of tasks such as MMLU.
We've since updated the formula to correctly aggregate Standard Error scores for groups of tasks reporting accuracies aggregated via their mean across the dataset -- see #1390 #1427 for more information.
As always, please feel free to give us feedback or request new features! We're grateful for the community's support.
n-shot in table by @haileyschoelkopf in https://github.com/EleutherAI/lm-evaluation-harness/pull/1379write_out.py instructions in README by @haileyschoelkopf in https://github.com/EleutherAI/lm-evaluation-harness/pull/1371batch_size with auto defaults to 1 if No executable batch size found is raised by @pminervini in https://github.com/EleutherAI/lm-evaluation-harness/pull/1405evaluator.simple_evaluate signature by @Am1n3e in https://github.com/EleutherAI/lm-evaluation-harness/pull/1412True for HuggingFace datasets compatibility by @veekaybee in https://github.com/EleutherAI/lm-evaluation-harness/pull/1467True for HuggingFace datasets compatibility" by @haileyschoelkopf in https://github.com/EleutherAI/lm-evaluation-harness/pull/1474evaluater.evaluate by @baberabb in https://github.com/EleutherAI/lm-evaluation-harness/pull/1441WandbLogger to accept arbitrary kwargs by @baberabb in https://github.com/EleutherAI/lm-evaluation-harness/pull/1491--trust_remote_code by @haileyschoelkopf in https://github.com/EleutherAI/lm-evaluation-harness/pull/1517max_gen_toks generation kwarg default in code2_text. by @cosmo3769 in https://github.com/EleutherAI/lm-evaluation-harness/pull/1551max_gen_toks generation kwarg default in generative Bigbench by @haileyschoelkopf in https://github.com/EleutherAI/lm-evaluation-harness/pull/1546Full Changelog: https://github.com/EleutherAI/lm-evaluation-harness/compare/v0.4.1...v0.4.2
A new TaskManager object and the deprecation of lm_eval.tasks.initialize_tasks(), for achieving the easier registration of many tasks and configuratio…
This PR release contains all changes so far since the release of v0.4.0 , and is partially a test of our release automation, provided by @anjor .
At a high level, some of the changes include:
local-completions or local-chat-completions ( Thanks to @veekaybee @mgoin @anjor and others on this)!More frequent (minor) version releases may be done in the future, to make it easier for PyPI users!
We're very pleased by the uptick in interest in LM Evaluation Harness recently, and we hope to continue to improve the library as time goes on. We're grateful to everyone who's contributed, and are excited by how many new contributors this version brings! If you have feedback for us, or would like to help out developing the library, please let us know.
In the next version release, we hope to include
TaskManager object and the deprecation of lm_eval.tasks.initialize_tasks(), for achieving the easier registration of many tasks and configuration of new groups of taskslm_eval/api/samplers.py by @haileyschoelkopf in https://github.com/EleutherAI/lm-evaluation-harness/pull/1062main by @haileyschoelkopf in https://github.com/EleutherAI/lm-evaluation-harness/pull/1063evaluator.py by @lintangsutawika in https://github.com/EleutherAI/lm-evaluation-harness/pull/1104write_out by @baberabb in https://github.com/EleutherAI/lm-evaluation-harness/pull/1113evaluator.py" by @lintangsutawika in https://github.com/EleutherAI/lm-evaluation-harness/pull/1116qqp, mnli_mismatch: remove unlabled test sets by @baberabb in https://github.com/EleutherAI/lm-evaluation-harness/pull/1114hf modeling code by @haileyschoelkopf in https://github.com/EleutherAI/lm-evaluation-harness/pull/1096batch_size type by @xTayEx in https://github.com/EleutherAI/lm-evaluation-harness/pull/1128mps by @baberabb in https://github.com/EleutherAI/lm-evaluation-harness/pull/1133--gen_kwargs arg to None by @baberabb in https://github.com/EleutherAI/lm-evaluation-harness/pull/1145ruff by @baberabb in https://github.com/EleutherAI/lm-evaluation-harness/pull/1166mypy by @baberabb in https://github.com/EleutherAI/lm-evaluation-harness/pull/1193mamba_ssm) by @haileyschoelkopf in https://github.com/EleutherAI/lm-evaluation-harness/pull/1110utils.divide doc by @xTayEx in https://github.com/EleutherAI/lm-evaluation-harness/pull/1208process_docs() to fewshot_split by @haileyschoelkopf in https://github.com/EleutherAI/lm-evaluation-harness/pull/1276parallelize=True vs. accelerate launch distinction clearer in docs by @haileyschoelkopf in https://github.com/EleutherAI/lm-evaluation-harness/pull/1261nq_open / NaturalQs whitespacing by @haileyschoelkopf in https://github.com/EleutherAI/lm-evaluation-harness/pull/1289datasets dependency at 2.15 by @haileyschoelkopf in https://github.com/EleutherAI/lm-evaluation-harness/pull/1312datasets dependency to >=2.14 by @haileyschoelkopf in https://github.com/EleutherAI/lm-evaluation-harness/pull/1314local-completions support using OpenAI interface by @mgoin in https://github.com/EleutherAI/lm-evaluation-harness/pull/1277get_task_dict() in task registration / initialization by @haileyschoelkopf in https://github.com/EleutherAI/lm-evaluation-harness/pull/1331Filter docs not offset by doc_id by @baberabb in https://github.com/EleutherAI/lm-evaluation-harness/pull/1349lm_eval.tasks.initialize_tasks() to README by @haileyschoelkopf in https://github.com/EleutherAI/lm-evaluation-harness/pull/1330*args by @baberabb in https://github.com/EleutherAI/lm-evaluation-harness/pull/1369--gen_kwargs behavior by @haileyschoelkopf in https://github.com/EleutherAI/lm-evaluation-harness/pull/1329Full Changelog: https://github.com/EleutherAI/lm-evaluation-harness/compare/v0.4.0...v0.4.1
[Refactor] batch generation better for hf model ; deprecate hf-causal in new release by @haileyschoelkopf in https://github.com/EleutherAI/lm-evaluati…
triviaqa dataset link by @jon-tow in https://github.com/EleutherAI/lm-evaluation-harness/pull/364actions/setup-pythonin CI workflows by @jon-tow in https://github.com/EleutherAI/lm-evaluation-harness/pull/365triviaqa version by @jon-tow in https://github.com/EleutherAI/lm-evaluation-harness/pull/366lambada_openai multilingual data source by @jon-tow in https://github.com/EleutherAI/lm-evaluation-harness/pull/370hf-causal and hf-seq2seq model implementations by @haileyschoelkopf in https://github.com/EleutherAI/lm-evaluation-harness/pull/381huggingface.py model classes by @haileyschoelkopf in https://github.com/EleutherAI/lm-evaluation-harness/pull/427dtype command line flag to HFLM by @haileyschoelkopf in https://github.com/EleutherAI/lm-evaluation-harness/pull/523quantized and its default value for python < 3.11 compatibility by @passaglia in https://github.com/EleutherAI/lm-evaluation-harness/pull/532.to(device) by @haileyschoelkopf in https://github.com/EleutherAI/lm-evaluation-harness/pull/585--use_cache by @haileyschoelkopf in https://github.com/EleutherAI/lm-evaluation-harness/pull/619hf model ; deprecate hf-causal in new release by @haileyschoelkopf in https://github.com/EleutherAI/lm-evaluation-harness/pull/613device_map options for hf model type by @haileyschoelkopf in https://github.com/EleutherAI/lm-evaluation-harness/pull/625parallelize=True docs by @haileyschoelkopf in https://github.com/EleutherAI/lm-evaluation-harness/pull/638cuda:0 device assignment by @haileyschoelkopf in https://github.com/EleutherAI/lm-evaluation-harness/pull/647generation_kwargs by @lintangsutawika in https://github.com/EleutherAI/lm-evaluation-harness/pull/657winograd_schema by @lintangsutawika in https://github.com/EleutherAI/lm-evaluation-harness/pull/690num_fewshot by @lintangsutawika in https://github.com/EleutherAI/lm-evaluation-harness/pull/702big-refactor by @haileyschoelkopf in https://github.com/EleutherAI/lm-evaluation-harness/pull/686lm-eval by @lintangsutawika in https://github.com/EleutherAI/lm-evaluation-harness/pull/899greedy_until to generate_until by @lintangsutawika in https://github.com/EleutherAI/lm-evaluation-harness/pull/927generate_until rename by @haileyschoelkopf in https://github.com/EleutherAI/lm-evaluation-harness/pull/929gold_alias task YAML option by @haileyschoelkopf in https://github.com/EleutherAI/lm-evaluation-harness/pull/965Task subclasses by @haileyschoelkopf in https://github.com/EleutherAI/lm-evaluation-harness/pull/996examples/ folder by @haileyschoelkopf in https://github.com/EleutherAI/lm-evaluation-harness/pull/1018Full Changelog: https://github.com/EleutherAI/lm-evaluation-harness/compare/v0.3.0...v0.4.0
This release integrates HuggingFace datasets as the core dataset management interface, removing previous custom downloaders.
This release integrates HuggingFace datasets as the core dataset management interface, removing previous custom downloaders.
Task downloading to use HuggingFace.datasets by @jon-tow in https://github.com/EleutherAI/lm-evaluation-harness/pull/300TriviaQA by @jon-tow in https://github.com/EleutherAI/lm-evaluation-harness/pull/305SWAG by @jon-tow in https://github.com/EleutherAI/lm-evaluation-harness/pull/306numexpr by @jon-tow in https://github.com/EleutherAI/lm-evaluation-harness/pull/352TextSynth API by @jon-tow in https://github.com/EleutherAI/lm-evaluation-harness/pull/299LAMBADA dataset by @jon-tow in https://github.com/EleutherAI/lm-evaluation-harness/pull/357Full Changelog: https://github.com/EleutherAI/lm-evaluation-harness/compare/v0.2.0...v0.3.0
implemented description dict and deprecated provide_description
Major changes since 0.1.0:
--check_integrity flag to run integrity unit tests at eval time (#290)evaluate and simple_evaluate are now deprecated_CITATION attribute on task modules (#292)Nothing published for this version
Your coding agent can read these notes before it upgrades. Set up the MCP server →