NewYour coding agent can read the release notes before it upgrades.Set up the MCP server →
PyPI · #2310 most downloaded on PyPI
Hassle-free computation of shareable, comparable, and reproducible BLEU, chrF, and TER scores
Last release 8 months ago
12 Jan 2026
Release timing varies
gaps range from 2 weeks to 13 months
Nearly every release is documented
notes for 57 of the last 60 stable releases
Nothing withdrawn
no release was ever pulled
9 years old
73 releases · first in 2017
added the chrF metric (-m chrf or -m bleu chrf for both) See 'CHRF: character n-gram F-score for automatic MT evaluation' by Maja Popovic (WMT 2015) [
-m chrf or -m bleu chrf for both)
See 'CHRF: character n-gram F-score for automatic MT evaluation' by Maja Popovic (WMT 2015)
[http://www.statmt.org/wmt15/pdf/WMT49.pdf]
--cite to produce the citation for easy inclusion in papers--input (-i) to set input to a file instead of STDINcorpus_bleu() now raises an exception if input streams are different lengths
--tok intl (international tokenization)One column per quarter.
bugfix for tokenization warning
added -b option (only output the BLEU score)
added effective order for sentence-level BLEU computation
Factored code a bit to facilitate API:
Nothing published for this version
Nothing published for this version
Nothing published for this version
Contributions from Christian Federmann:
Nothing published for this version
Small bugfix affecting some versions of Python.
Nothing published for this version
Your coding agent can read these notes before it upgrades. Set up the MCP server →