NewYour coding agent can read the release notes before it upgrades.Set up the MCP server →
PyPI · #2668 most downloaded on PyPI
Look up the frequencies of words in many languages, based on many sources of data.
Last release 3 years ago
no release in 18 months
Release timing varies
gaps range from 3 weeks to 13 months
Most releases are documented
notes for 20 of 30 stable releases
Nothing withdrawn
no release was ever pulled
13 years old
31 releases · first in 2014
Stopped using the deprecated pkg_resources and implicitly depending on setuptools. We now use the locate package instead.
pkg_resources and implicitly depending on
setuptools. We now use the locate package instead.Nothing published for this version
wordfreq has changed from the MIT license to the Apache License 2.0.
wordfreq has changed from the MIT license to the Apache License 2.0.
The Apache license is equally permissive, but clarifies requirements about attributing the author (Robyn Speer), making this attribution readable, and maintaining credit for the included data resources.
Minor packaging changes to make sure that tox can test the library on Python 3.7 through 3.11.
One column per quarter.
Updated the range of allowable versions of regex. Versions before 2021.7.6 don't have the regex.Match class.
Updated the range of allowable versions of regex. Versions before 2021.7.6 don't have the regex.Match class.
Added the extras dependencies as optional dependencies in pyproject.toml.
move mypy to dev dependencies
move mypy to dev dependencies
mypy was accidentally listed as a full dependency; moved it to dev dependencies.
Nothing published for this version
Import ftfy and use its uncurl_quotes method to turn curly quotes into straight ones, providing consistency with multiple forms of apostrophes.
Import ftfy and use its uncurl_quotes method to turn curly quotes into
straight ones, providing consistency with multiple forms of apostrophes.
Set minimum version requierements on regex, jieba, and langcodes
so that tokenization will give consistent results.
Work around an inconsistency in the msgpack API around
strict_map_key=False.
Import ftfy and use its uncurl_quotes method to turn curly quotes into
straight ones, providing consistency with multiple forms of apostrophes.
Set minimum version requirements on regex, jieba, and langcodes
so that tokenization will give consistent results.
Workaround an inconsistency in the msgpack API around
strict_map_key=False.
Nothing published for this version
When tokenizing Japanese or Korean, MeCab's dictionaries no longer have to be installed separately as system packages. They can now be found via the P
When tokenizing Japanese or Korean, MeCab's dictionaries no longer have to
be installed separately as system packages. They can now be found via the
Python packages ipadic and mecab-ko-dic.
When the tokenizer had to infer word boundaries in languages without spaces, inputs that were too long (such as the letter 'l' repeated 800 times) were causing overflow errors. We changed the sequence of operations so that it no longer overflows, and such inputs simply get a frequency of 0.
Changed a log message to not try to call a language by name, to remove the dependency on a database of language names.
Relaxing the dependency on regex had an unintended consequence in 2.3.1: it could no longer get the frequency of French phrases such as "l'écran" beca
Relaxing the dependency on regex had an unintended consequence in 2.3.1: it could no longer get the frequency of French phrases such as "l'écran" because their tokenization behavior changed.
2.3.2 fixes this with a more complex tokenization rule that should handle apostrophes the same across these various versions of regex.
Fixed an incompatibility with newly-released msgpack 1.0.
Library change:
msgpack 1.0.Fixed calling msgpack.load with a deprecated parameter.
Library changes:
Relaxed the version requirement on the 'regex' dependency, allowing compatibility with spaCy.
The range of regex versions that wordfreq now allows is from 2017.07.11 to 2018.02.21. No changes to word boundary matching were made between these versions.
Fixed calling msgpack.load with a deprecated parameter.
Nothing published for this version
Nothing published for this version
Fixed edge cases that inserted spurious token boundaries when Japanese text is run through simple_tokenize, because of a few characters that don't mat
Fixed edge cases that inserted spurious token boundaries when Japanese text is
run through simple_tokenize, because of a few characters that don't match any
of our "spaceless scripts".
It is not a typical situation for Japanese text to be passed through
simple_tokenize, because Japanese text should instead use the
Japanese-specific tokenization in wordfreq.mecab.
However, some downstream uses of wordfreq have justifiable reasons to pass all
terms through simple_tokenize, even terms that may be in Japanese, and in
those cases we want to detect only the most obvious token boundaries.
In this situation, we no longer try to detect script changes, such as between kanji and katakana, as token boundaries. This particularly allows us to keep together Japanese words where ヶ appears between kanji, as well as words that use the iteration mark 々.
This change does not affect any word frequencies. (The Japanese word list uses
wordfreq.mecab for tokenization, not simple_tokenize.)
As a breaking change, this means that the tokenize function no longer has the combine_numbers option, because that's a postprocessing step. For the sa…
The big change in this version is that text preprocessing, tokenization, and postprocessing to look up words in a list are separate steps.
If all you need is preprocessing to make text more consistent, use
wordfreq.preprocess.preprocess_text(text, lang). If you need preprocessing
and tokenization, use wordfreq.tokenize(text, lang) as before. If you need
all three steps, use the new function wordfreq.lossy_tokenize(text, lang).
As a breaking change, this means that the tokenize function no longer has
the combine_numbers option, because that's a postprocessing step. For
the same behavior, use lossy_tokenize, which always combines numbers.
Similarly, tokenize will no longer replace Chinese characters with their
Simplified Chinese version, while lossy_tokenize will.
Other changes:
There's a new default wordlist for each language, called "best". This chooses the "large" wordlist for that language, or if that list doesn't exist, it falls back on "small".
The wordlist formerly named "combined" (this name made sense long ago) is now named "small". "combined" remains as a deprecated alias.
The "twitter" wordlist has been removed. If you need to compare word frequencies from individual sources, you can work with the separate files in exquisite-corpus.
Tokenizing Chinese will preserve the original characters, no matter whether they are Simplified or Traditional, instead of replacing them all with Simplified characters.
Different languages require different processing steps, and the decisions
about what these steps are now appear in the wordfreq.language_info module,
replacing a bunch of scattered and inconsistent if statements.
Tokenizing CJK languages while preserving punctuation now has a less confusing implementation.
The preprocessing step can transliterate Azerbaijani, although we don't yet have wordlists in this language. This is similar to how the tokenizer supports many more languages than the ones with wordlists, making future wordlists possible.
Speaking of that, the tokenizer will log a warning (once) if you ask to tokenize text written in a script we can't tokenize (such as Thai).
New source data from exquisite-corpus includes OPUS OpenSubtitles 2018.
Nitty gritty dependency changes:
Updated the regex dependency to 2018.02.21. (We would love suggestions on
how to coexist with other libraries that use other versions of regex,
without a >= requirement that could introduce unexpected data-altering
changes.)
We now depend on msgpack, the new name for msgpack-python.
Tokenization will always keep Unicode graphemes together, including complex emoji introduced in Unicode 10
Depend on langcodes 1.4, with a new language-matching system that does not depend on SQLite.
Depend on langcodes 1.4, with a new language-matching system that does not depend on SQLite.
This prevents silly conflicts where langcodes' SQLite connection was preventing langcodes from being used in threads.
This release of wordfreq gives word frequencies in 27 languages from a variety of data sources, which it checks against each other to mitigate outlier
This release of wordfreq gives word frequencies in 27 languages from a variety of data sources, which it checks against each other to mitigate outliers.
See CHANGELOG.md for more details on the version history.
Nothing published for this version
Nothing published for this version
Add large lists in English, German, Spanish, French, and Portuguese
zipf_frequency functionAdd Reddit comments as an English source
Better support for Chinese, using Jieba for tokenization, and mapping Traditional Chinese characters to Simplified
Use the 'regex' package to implement Unicode tokenization that's mostly consistent across languages
Create compact word frequency lists in English, Arabic, German, Spanish, French, Indonesian, Japanese, Malay, Dutch, Portuguese, and Russian
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Your coding agent can read these notes before it upgrades. Set up the MCP server →