NewYour coding agent can read the release notes before it upgrades.Set up the MCP server →
PyPI · #2234 most downloaded on PyPI
Hassle-free computation of shareable, comparable, and reproducible BLEU, chrF, and TER scores
Last release 8 months ago
12 Jan 2026
Release timing varies
gaps range from 2 weeks to 13 months
Nearly every release is documented
notes for 57 of the last 60 stable releases
Nothing withdrawn
no release was ever pulled
9 years old
73 releases · first in 2017
Dropped Python 3.8, added Python 3.13 (requires-python = ">=3.9")
spBLEU-1KspBLEU-1KUpdate CHANGELOG with note about v2.5.0
One column per quarter.
Update CHANGELOG with note about v2.5.0
Add WMT 2024 test sets.
Add WMT 2024 test sets.
Fix CI errors use pip install build instead of pip install .[build]
Fix CI errors
use pip install build instead of pip install .[build]
Allow printing of the domain (via --echo domain ) if available on the underlying test set.
Allow printing of the domain (via --echo domain) if available on the underlying test set.
Add project exports (closes #255 )
use refB by default for en-he and he-en
Version bump to 2.3.3
Version bump to 2.3.3 (#253)
Fixed:
Version bump to 2.3.2
Version bump to 2.3.2 (#246)
Fixed:
Added:
-tok spm is used (use explicit flores101 instead) (#238)Set lru_cache for SPM tokenizer
Set lru_cache for SPM tokenizer (#218)
Bugfix:
(#203) Added -tok flores101 and -tok flores200, a.k.a. spbleu. These are multilingual tokenizations that make use of the multilingual SPM models relea
Features:
-tok flores101 and -tok flores200, a.k.a. spbleu.
These are multilingual tokenizations that make use of the
multilingual SPM models released by Facebook and described in the
following papers:
--list SRC-TRG.
Thanks to Jaume Zaragoza (@ZJaume) for adding this feature.wmt22)--echo, e.g., sacrebleu -t wmt22 -l en-de --echo ?Bugfix: Standard usage was returning (and using) each reference twice.
Bugfix: Standard usage was returning (and using) each reference twice.
This release contains an inner reworking of the data representations, contributed by @BrightXiaoHan. This enables the following features:
This release contains an inner reworking of the data representations, contributed by @BrightXiaoHan. This enables the following features:
--echo (including origlang, docid, and genre, which are all available for most WMT corpora)We also added a Korean tokenizer (--tok ko-mecab), contributed by @NoUnique.
In addition, there are a number of bug fixes and minor fixes:
spm tokenizer, particularly for CJK languages..sacrebleu/wmt21/wmt_21.en-de.ref instead of .sacrebleu/wmt21/de-en.ref. The file extension corresponds to the field that gets passed to --echo.Features:
--echo now exposes document metadata where available (e.g., docid, genre, origlang)Under the hood:
--echo (e.g., "src")Many thanks to @BrightXiaoHan (https://github.com/BrightXiaoHan) for the bulk of the code contributions in this release.
Added -tok spm for multilingual SPM tokenization (#168) (thanks to Naman Goyal and James Cross at Facebook)
Features:
-tok spm for multilingual SPM tokenization (#168)
(thanks to Naman Goyal and James Cross at Facebook)Fixes:
This is a major release that introduces statistical significance testing for BLEU, chrF and TER. It should be noted that as of v2.0.0, the default out
This is a major release that introduces statistical significance testing for BLEU, chrF and TER. It should be noted that as of v2.0.0, the default output format of the CLI utility is json rather than the old single-line output. All tools should adapt to this change if they parse standard output.
Python < 3.6 support and migrate to f-strings.isinstance checks. If the user does not obey
to the expected annotations, exceptions will be raised. Robustness attempts lead to
confusions and obfuscated score errors in the past (fixes #121)colorama package.intl tokenizer: Use regex module. Speed goes from ~4 seconds to ~0.6 seconds
for a particular test set evaluation. (fixes #46)word_order argument. Added test cases against chrF++.py.
Exposed it through the CLI (--chrf-word-order) (fixes #124)--format/-f flag. The single-system output mode is now json by default.
If you want to keep the old text format persistently, you can export SACREBLEU_FORMAT=text into your
shell.tabulate package, the results are
nicely rendered into a plain text table, LaTeX, HTML or RST (cf. --format/-f argument).
The systems can be either given as a list of plain text files to -i/--input or
as a tab-separated single stream redirected into STDIN. In the former case,
the basenames of the files will be automatically used as system names.--confidence flag)
as well as paired bootstrap resampling (--paired-bs) and paired approximate
randomization tests (--paired-ar) when evaluating multiple systems (fixes #40 and fixes #78).Python < 3.6 support and migrate to f-strings.portalocker version pinning, add regex, tabulate, numpy dependencies.isinstance checks. If the user does not obey
to the expected annotations, exceptions will be raised. Robustness attempts lead to
confusions and obfuscated score errors in the past (#121)colorama package.intl tokenizer: Use regex module. Speed goes from ~4 seconds to ~0.6 seconds
for a particular test set evaluation. (#46)var if variable number of references is used.argparse.Namespace objects.Metric class is introduced to guide further
metric development. This class defines the methods that should be implemented
in the derived classes and offers boilerplate methods for the common functionality.
A new metric implemented this way will automatically support significance testing.references argument at
initialization time to process and cache the references. Further evaluations
of different systems against the same references becomes faster this way
for example when using significance testing.word_order argument. Added test cases against chrF++.py.
Exposed it through the CLI (--chrf-word-order) (#124)--input/-i can now ingest multiple systems. For this reason, the positional
references should always preceed the -i flag.--help is printed.--format/-f flag. The single-system output mode is now json by default.
If you want to keep the old text format persistently, you can export SACREBLEU_FORMAT=text into your
shell.json falls back to plain text. latex output can only
be generated for multi-system mode.tabulate package, the results are
nicely rendered into a plain text table, LaTeX, HTML or RST (cf. --format/-f argument).
The systems can be either given as a list of plain text files to -i/--input or
as a tab-separated single stream redirected into STDIN. In the former case,
the basenames of the files will be automatically used as system names.--confidence flag)
as well as paired bootstrap resampling (--paired-bs) and paired approximate
randomization tests (--paired-ar) when evaluating multiple systems (#40 and #78).Fix extraction error for WMT18 extra test sets (test-ts)
Minor bugfix release:
Add missing __repr__() methods for BLEU and TER
1.5.0 (2021-01-15)
__repr__() methods for BLEU and TER--short is used (#131)floor smoothing is now 0.1 instead of 0.sacrebleu.sentence_bleu() now uses the exp smoothing method,
exactly the same as the CLI's --sentence-level behavior. This was mainly done
to make two methods behave the same.Added character-based tokenization (-tok char). Thanks to Christian Federmann.
-tok char).
Thanks to Christian Federmann.
-m ter). Thanks to Ales Tamchyna! (fixes #90)Make mecab3-python an extra dependency, adapt code to new mecab3-python. This fixes the recent Windows installation issues as well (#104) Japanese sup
1.4.13 (2020-07-30)
mecab3-python. This fixes the recent Windows installation issues as well (#104) Japanese support should now be explicitly installed through sacrebleu[ja] package.- Fix a deployment bug
Added Multi30k multimodal MT test set metadata
utils.pyBLEUSignature and CHRFSignature classesCleaned up deprecation warnings (thanks to Karthikeyan Singaravelan @tirkarthi)
<= 3.4, as it was integrated in the standard
library in Python 3.5 (thanks to Erwan de Lépinau @ErwanDL).Changed get_available_testsets() to return a list
get_available_testsets() to return a list
Fixed descriptions of some WMT19/google test sets
Added Google's extra wmt19/en-de refs (-t wmt19/google/{ar,arp,hqall,hqp,hqr,wmtp}) (Freitag, Grangier, & Caswell BLEU might be Guilty but References
Large internal reorganization as a module (thanks to Thamme Gowda @thammegowda)
Added Japanese MeCab tokenizer (-tok ja-mecab) (thanks to Makoto Morishita @MorinoseiMorizo)
-tok ja-mecab) (thanks to Makoto Morishita @MorinoseiMorizo)
Smoothing changes (Sebastian Nickels @sn1c)
--list now returns a list of all language pairs for a task when combined with -t
(e.g., sacrebleu -t wmt19 --list)Bugfix: handling of result object for CHRF
Tokenization variant omitted from the chrF signature; it is relevant only for BLEU (thanks to Martin Popel)
Added sentence-level scoring via -sl (--sentence-level)
Many thanks to Martin Popel for all the changes below!
-t wmt17,wmt18).
Works as long as they all have the same language pair.sacrebleu --origlang (both for evaluation on a subset and for --echo).
Note that while echoing prints just the subset, evaluation expects the complete
test set (and just skips the irrelevant parts).sacrebleu --detail for breakdown by domain-specific subsets of the test sets.
(Available for WMT19).sacrebleu -hsacrebleu --listos.makedirs(outdir, exist_ok=True) instead of if os.path.exists)Lazy loading of regexes cuts import time from ~1s to nearly nothing (thanks, @louismartin!)
--num-refs N to tell it to run the split.
Only works with a single reference file passed from the command line.Removed another f-string for Python 3.5 compatibility
Restored Python 3.5 compatibility
- Added MTNT 2019 test sets - Added a BLEU object
- Added WMT'19 test sets
Bugfix in test case (thanks to Adam Roberts, @adarob)
sentence_bleuChanged interface to some functions (backwards incompatible)
--smooth exp|floor|add-n|none) and the associated value (--smooth-value), when relevant.
Ctrl-M characters are now treated as normal characters, previously treated as newline.
Tokenization now defaults to "zh" when language pair is known
Updated checksum for wmt19/dev (seems to have changed)
Fixed checksum for wmt17/dev (copy-paste error)
Added kk-en and en-kk to wmt19/dev
Added gu-en and en-gu to wmt19/dev
Added MD5 checksumming of downloaded files for all datasets.
Added mtnt1.1/train mtnt1.1/valid mtnt1.1/test data from MTNT
Added 'wmt19/dev' task for 'lt-en' and 'en-lt' (development data for new tasks).
Now outputs only only digit after the decimal
Added a function for sentence-level, smoothed BLEU
Added wmt18 test set (with references)
Added zh-en, en-zh, tr-en, and en-tr datasets for wmt18/test-ts
Added wmt18/test-ts, the test sources (only) for WMT18
sacrebleu.py and the CHANGELOG into a separate filefixed another locale issue (with --echo)
-tok none from the command lineadded wmt17/ms (Microsoft's additional ZH-EN references). Try sacrebleu -t wmt17/ms --cite.
sacrebleu -t wmt17/ms --cite.
--echo ref now pastes together all references, if there is more than oneadded wmt18/dev datasets (en-et and et-en)
Nothing published for this version
metrics (-m) are now printed in the order requested
-m) are now printed in the order requested
Nothing published for this version
Nothing published for this version
Your coding agent can read these notes before it upgrades. Set up the MCP server →