NewYour coding agent can read the release notes before it upgrades.Set up the MCP server →
PyPI · #2338 most downloaded on PyPI
HuggingFace community-driven open-source library of evaluation
Last release 1 years ago
18 Sep 2025
Release timing varies
gaps range from 2 weeks to 10 months
Nearly every release is documented
notes for 14 of 14 stable releases
Nothing withdrawn
no release was ever pulled
4 years old
14 releases · first in 2022
One column per quarter.
Remove deprecated HfFolder by @Wauplin in #701
HfFolder by @Wauplin in #701
huggingface_hub>=1.0Full Changelog: v0.4.5...v0.4.6
Support datasets 4 by @lhoestq in #689
Support nltk>=3.9 to fix vulnerability by @albertvillanova in #629
Full Changelog: v0.4.3...v0.4.4
This release adds support for datasets>=3.0 by removing calls to deprecated code
This release adds support for datasets>=3.0 by removing calls to deprecated code
Full Changelog: https://github.com/huggingface/evaluate/compare/v0.4.2...v0.4.3
This release adds support for datasets>=3.0 by removing calls to deprecated code
Full Changelog: v0.4.2...v0.4.3
Update the documentation and citation of mauve by @krishnap25 in #416
Full Changelog: v0.4.1...v0.4.2
Add code example to docstrings by @stevhliu in #374
add method by @hazrulakmal in #424datasets import in Meteor metric by @mariosasko in #490Full Changelog: v0.4.0...v0.4.1
add trainer integration docs by @lvwerra in #325
scikit-learn install in spaces by @lvwerra in #345Evaluate usage for scikit-learn by @awinml in #368Full Changelog: v0.3.0...v0.4.0
add multilabel f1 eval usage by @fcakyon in #221
commit_hash to args by @lvwerra in #253handle_impossible_answer from the default PIPELINE_KWARGS in the question answering evaluator by @fxmarty in #272split and subset kwarg into other evaluators by @mathemakitten in #301HubEvaluationModuleFactory by @lvwerra in #314Full Changelog: v0.2.2...v0.3.0
Update CLI docs by @lvwerra in #218
Full Changelog: v0.2.1...v0.2.2
Add measurements to quality and style checks by @lvwerra in #203
Full Changelog: v0.2.0...v0.2.1
The evaluator has been extended to three new tasks:
evaluatorThe evaluator has been extended to three new tasks:
"image-classification""token-classification""question-answering"combineWith combine one can bundle several metrics into a single object that can be evaluated in one call and also used in combination with the evalutor.
evaluator tests by @lvwerra in https://github.com/huggingface/evaluate/pull/155input_texts to predictions in perplexity by @lvwerra in https://github.com/huggingface/evaluate/pull/157combine to compose multiple evaluations by @lvwerra in https://github.com/huggingface/evaluate/pull/150TextClassificationEvaluator test by @fxmarty in https://github.com/huggingface/evaluate/pull/172ImageClassificationEvaluator by @fxmarty in https://github.com/huggingface/evaluate/pull/173TokenClassificationEvaluator by @fxmarty in https://github.com/huggingface/evaluate/pull/167Full Changelog: https://github.com/huggingface/evaluate/compare/v0.1.2...v0.2.0
Fix trec sacrebleu by @lvwerra in https://github.com/huggingface/evaluate/pull/130
Full Changelog: https://github.com/huggingface/evaluate/compare/v0.1.1...v0.1.2
Fix broken links by @mishig25 in https://github.com/huggingface/evaluate/pull/92
pip install evaluate[evaluator] by @philschmid in https://github.com/huggingface/evaluate/pull/103evaluate dependency in spaces by @lvwerra in https://github.com/huggingface/evaluate/pull/88Full Changelog: https://github.com/huggingface/evaluate/compare/v0.1.0...v0.1.1
These are the release notes of the initial release of the Evaluate library.
These are the release notes of the initial release of the Evaluate library.
Goals of the Evaluate library:
evaluate.load(): The load() function is the main entry point into evaluate and allows to load evaluation modules from a local folder, the evaluate repository, or the Hugging Face Hub. It downloads, caches, and loads the evaluation modules and returns an evaluate.EvaluationModule.evaluate.save(): With save() a user can save evaluation results in a JSON file. In addition to the results from evaluate.EvaluationModule it can save additional parameters and automatically saves the timestamp, git commit hash, library version as well as Python path. One can either provide a directory for the results, in which case file names are automatically created, or an explicit file name for the result.evaluate.push_to_hub(): The push_to_hub function allows to push the results of a model evaluation to the model card on the Hugging Face Hub. The model, dataset, and metric are specified such that they can be linked on the hub.evaluate.EvaluationModule: The EvaluationModule class is the baseclass for all evaluation modules. There are three module types: metrics (to evaluate models), comparisons (to compare models), and measurements (to analyze datasets). The inputs can be either added with add (single input) and add_batch (batch of inputs) followed by a final compute call to compute the scores or all inputs can be passed to compute directly. Under the hood, Apache Arrow stores and loads the input data to compute the scores.evaluate.EvaluationModuleInfo: The EvaluationModule class is used to store attributes:
description: A short description of the evaluation module.citation: A BibTex string for citation when available.features: A Features object defining the input format. The inputs provided to add, add_batch, and compute are tested against these types and an error is thrown in case of a mismatch.inputs_description: This is equivalent to the modules docstring.homepage: The homepage of the module.license: The license of the module.codebase_urls: Link to the code behind the module.reference_urls: Additional reference URLs.evaluate.evaluator: The evaluator provides automated evaluation and only requires a model, dataset, metric, in contrast to the metrics in the EvaluationModule which require model predictions. It has three main components: a model wrapped in a pipeline, a dataset, and a metric, and it returns the computed evaluation scores. Besides the three main components, it may also require two mappings to align the columns in the dataset and the pipeline labels with the datasets labels. This is an experimental feature -- currently, only text classification is supported.evaluate-cli: The community can add custom metrics by adding the necessary module script to a Space on the Hugging Face Hub. The evaluate-cli is a tool that simplifies this process by creating the Space, populating a template, and pushing it to the Hub. It also provides instructions to customize the template and integrate custom logic.
@lvwerra , @sashavor , @NimaBoscarino , @ola13 , @osanseviero , @lhoestq , @lewtun , @douwekiela
Your coding agent can read these notes before it upgrades. Set up the MCP server →