NewYour coding agent can read the release notes before it upgrades.Set up the MCP server →
PyPI · #967 most downloaded on PyPI
Fixes mojibake and other problems with Unicode, after the fact
Last release 2 years ago
no release in 18 months
Release timing varies
gaps range from 2 weeks to 1.8 years
Most releases are documented
notes for 43 of 55 stable releases
Nothing withdrawn
no release was ever pulled
14 years old
55 releases · first in 2012
Fixed license metadata field in pyproject.toml.
license metadata field in pyproject.toml.hatchling sdist output.Switched packaging from poetry to uv.
chardata.py to be more human-readable and debuggable, instead of being full of keysmash-like character sets.See CHANGELOG.md for the full changelog.
Trusted Publishing is now supposed to create these releases on GitHub at the same time that it publishes to PyPI, following the user guide. It didn't, but it's supposed to.
I think I've fixed the problem (upgrading to sigstore/gh-action-sigstore-python@v3.0.0 from the broken v2.1.1), and maybe future releases really will be as simple as pushing a tag.
One column per quarter.
See CHANGELOG.md for version changes.
See CHANGELOG.md for version changes.
Can you tell that I'm creating these releases manually? I've set up a GitHub action that publishes to PyPI, which is reasonably well documented, but I can't find one that creates a release here on GitHub with the Python package included. Please let me know (or make a PR) if you know how.
Updated Read the Docs config so that docs might build again.
Updated setup.py and tox.ini to indicate support for Python 3.8 through 3.13.
Fixed a case where an en-dash and a space near other mojibake would be interpreted (probably incorrectly) as MacRoman mojibake.
Switched to the Apache 2.0 license.
update setup.py and README text about it
update setup.py and README text about it
Updated the heuristic to fix the letter ß in UTF-8/MacRoman mojibake, which had regressed since version 5.6.
Packaging fixes to pyproject.toml.
Nothing published for this version
Nothing published for this version
New function: ftfy.fix_and_explain() can describe all the transformations that happen when fixing a string. This is similar to what ftfy.fixes.fix_enc
Updates in 6.0.x:
Allow the keyword argument fix_entities as a deprecated alias for
unescape_html, raising a warning.
ftfy.formatting functions now disregard ANSI terminal escapes when
calculating text width.
The remove_terminal_escapes step was accidentally not being used. This version restores it.
The remove_terminal_escapes step was accidentally not being used. This
version restores it.
Specified in setup.py that ftfy 6 requires Python 3.6 or later.
Use a lighter link color when the docs are viewed in dark mode.
New function: ftfy.fix_and_explain() can describe all the transformations that happen when fixing a string. This is similar to what ftfy.fixes.fix_enc
New function: ftfy.fix_and_explain() can describe all the transformations
that happen when fixing a string. This is similar to what
ftfy.fixes.fix_encoding_and_explain() did in previous versions, but it
can fix more than the encoding.
fix_and_explain() and fix_encoding_and_explain() are now in the top-level
ftfy module.
Changed the heuristic entirely. ftfy no longer needs to categorize every Unicode character, but only characters that are expected to appear in mojibake.
Because of the new heuristic, ftfy will no longer have to release a new version for every new version of Unicode. It should also run faster and use less RAM when imported.
The heuristic ftfy.badness.is_bad(text) can be used to determine whether
there appears to be mojibake in a string. Some users were already using
the old function sequence_weirdness() for that, but this one is actually
designed for that purpose.
Instead of a pile of named keyword arguments, ftfy functions now take in a TextFixerConfig object. The keyword arguments still work, and become settings that override the defaults in TextFixerConfig.
Added support for UTF-8 mixups with Windows-1253 and Windows-1254.
Overhauled the documentation: https://ftfy.readthedocs.org
This version is brought to you by the letter à and the number 0xC3.
This version is brought to you by the letter à and the number 0xC3.
Tweaked the heuristic to decode, for example, "Ã " as the letter "à" more often.
This combines with the non-breaking-space fixer to decode "Ã " as "à" as well. However, in many cases, the text " Ã " was intended to be " à ", preserving the space -- the underlying mojibake had two spaces after it, but the Web coalesced them into one. We detect this case based on common French and Portuguese words, and preserve the space when it appears intended.
Thanks to @zehavoc for bringing to my attention how common this case is.
Improved detection of UTF-8 mojibake of Greek, Cyrillic, Hebrew, and Arabic scripts.
Improved detection of UTF-8 mojibake of Greek, Cyrillic, Hebrew, and Arabic scripts.
Fixed the undeclared dependency on setuptools by removing the use of
pkg_resources.
Updated the data file of Unicode character categories to Unicode 12.1, as used in Python 3.8. (No matter what version of Python you're on, ftfy uses t
Updated the data file of Unicode character categories to Unicode 12.1, as used in Python 3.8. (No matter what version of Python you're on, ftfy uses the same data.)
Corrected an omission where short sequences involving the ACUTE ACCENT character were not being fixed.
Nothing published for this version
The unescape_html function now supports all the HTML5 entities that appear in html.entities.html5, including those with long names such as ˝.
The unescape_html function now supports all the HTML5 entities that appear
in html.entities.html5, including those with long names such as
˝.
Unescaping of numeric HTML entities now uses the standard library's
html.unescape, making edge cases consistent.
(The reason we don't run html.unescape on all text is that it's not always
appropriate to apply, and can lead to false positive fixes. The text
"This&NotThat" should not have "&Not" replaced by a symbol, as
html.unescape would do.)
On top of Python's support for HTML5 entities, ftfy will also convert HTML
escapes of common Latin capital letters that are (nonstandardly) written
in all caps, such as Ñ for Ñ.
See CHANGELOG.md for release notes.
See CHANGELOG.md for release notes.
Added Python 3.7 support.
Updated the data file of Unicode character categories to Unicode 11, as used in Python 3.7.0. (No matter what version of Python you're on, ftfy uses the same data.)
Nothing published for this version
Fixed a bug in the setup.py metadata.
Fixed a bug in the setup.py metadata.
This bug was causing ftfy, a package that fixes encoding mismatches, to not install in some environments due to an encoding mismatch. (We were really putting the "meta" in "metadata" here.)
Nothing published for this version
Nothing published for this version
Nothing published for this version
These releases fix two unrelated problems with the tests, one in each version.
These releases fix two unrelated problems with the tests, one in each version.
v5.1.1: fixed the CLI tests (which are new in v5) so that they pass on Windows, as long as the Python output encoding is UTF-8.
v4.4.3: added the # coding: utf-8 declaration to two files that were
missing it, so that tests can run on Python 2.
Removed the dependency on html5lib by dropping support for Python 3.2.
Removed the dependency on html5lib by dropping support for Python 3.2.
We previously used the dictionary html5lib.constants.entities to decode
HTML entities. In Python 3.3 and later, that exact dictionary is now in the
standard library as html.entities.html5.
Moved many test cases about how particular text should be fixed into
test_cases.json, which may ease porting to other languages.
The functionality of this version remains the same as 5.0.2 and 4.4.2.
Added a MANIFEST.in that puts files such as the license file and this changelog inside the source distribution.
Added a MANIFEST.in that puts files such as the license file and this
changelog inside the source distribution.
The unescape_html fixer will decode entities between € and Ÿ as what they would be in Windows-1252, even without the help of fix_encoding.
Bug fix:
The unescape_html fixer will decode entities between € and Ÿ
as what they would be in Windows-1252, even without the help of
fix_encoding.
This better matches what Web browsers do, and fixes a regression that version
4.4 introduced in an example that uses … as an ellipsis.
Dropped support for Python 2. If you need Python 2 support, you should get version 4.4, which has the same features as this version.
Breaking changes:
Dropped support for Python 2. If you need Python 2 support, you should get version 4.4, which has the same features as this version.
The top-level functions require their arguments to be given as keyword arguments.
Version 5.0 also now has tests for the command-line invocation of ftfy.
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
remove_control_chars was removing U+0D ('\r') prematurely. That's the job of fix_line_breaks.
Bug fix:
remove_control_chars was removing U+0D ('\r') prematurely. That's the
job of fix_line_breaks.Math symbols next to currency symbols are no longer considered 'weird' by the heuristic. This fixes a false positive where text that involved the mult
Heuristic changes:
Math symbols next to currency symbols are no longer considered 'weird' by the heuristic. This fixes a false positive where text that involved the multiplication sign and British pounds or euros (as in '5×£35') could turn into Hebrew letters.
A heuristic that used to be a bonus for certain punctuation now also gives a bonus to successfully decoding other common codepoints, such as the non-breaking space, the degree sign, and the byte order mark.
In version 4.0, we tried to "future-proof" the categorization of emoji (as a kind of symbol) to include codepoints that would likely be assigned to emoji later. The future happened, and there are even more emoji than we expected. We have expanded the range to include those emoji, too.
ftfy is still mostly based on information from Unicode 8 (as Python 3.5 is), but this expanded range should include the emoji from Unicode 9 and 10.
Emoji are increasingly being modified by variation selectors and skin-tone modifiers. Those codepoints are now grouped with 'symbols' in ftfy, so they fit right in with emoji, instead of being considered 'marks' as their Unicode category would suggest.
This enables fixing mojibake that involves iOS's new diverse emoji.
An old heuristic that wasn't necessary anymore considered Latin text with high-numbered codepoints to be 'weird', but this is normal in languages such as Vietnamese and Azerbaijani. This does not seem to have caused any false positives, but it caused ftfy to be too reluctant to fix some cases of broken text in those languages.
The heuristic has been changed, and all languages that use Latin letters should be on even footing now.
Bug fix: in the command-line interface, the -e option had no effect on Python 3 when using standard input. Now, it correctly lets you specify a differ
-e option had no effect on
Python 3 when using standard input. Now, it correctly lets you specify
a different encoding for standard input.ftfy can now deal with "lossy" mojibake. If your text has been run through a strict Windows-1252 decoder, such as the one in Python, it may contain th
Heuristic changes:
ftfy can now deal with "lossy" mojibake. If your text has been run through a strict Windows-1252 decoder, such as the one in Python, it may contain the replacement character � (U+FFFD) where there were bytes that are unassigned in Windows-1252.
Although ftfy won't recover the lost information, it can now detect this situation, replace the entire lossy character with �, and decode the rest of the characters. Previous versions would be unable to fix any string that contained U+FFFD.
As an example, text in curly quotes that gets corrupted “ like this �
now gets fixed to be “ like this �.
Updated the data file of Unicode character categories to Unicode 8.0, as used in Python 3.5.0. (No matter what version of Python you're on, ftfy uses the same data.)
Heuristics now count characters such as ~ and ^ as punctuation instead
of wacky math symbols, improving the detection of mojibake in some edge cases.
New features:
A new module, ftfy.formatting, can be used to justify Unicode text in a
monospaced terminal. It takes into account that each character can take up
anywhere from 0 to 2 character cells.
Internally, the utf-8-variants codec was simplified and optimized.
The remove_unsafe_private_use parameter has been removed entirely, after two versions of deprecation. The function name fix_bad_encoding is also gone.
Breaking changes:
The default normalization form is now NFC, not NFKC. NFKC replaces a large number of characters with 'equivalent' characters, and some of these replacements are useful, but some are not desirable to do by default.
The fix_text function has some new options that perform more targeted
operations that are part of NFKC normalization, such as
fix_character_width, without requiring hitting all your text with the huge
mallet that is NFKC.
fix_character_width=False.The remove_unsafe_private_use parameter has been removed entirely, after
two versions of deprecation. The function name fix_bad_encoding is also
gone.
New features:
Fixers for strange new forms of mojibake, including particularly clear cases of mixed UTF-8 and Windows-1252.
New heuristics, so that ftfy can fix more stuff, while maintaining approximately zero false positives.
The command-line tool trusts you to know what encoding your input is in,
and assumes UTF-8 by default. You can still tell it to guess with the -g
option.
The command-line tool can be configured with options, and can be used as a pipe.
Recognizes characters that are new in Unicode 7.0, as well as emoji from Unicode 8.0+ that may already be in use on iOS.
Deprecations:
fix_text_encoding is being renamed again, for conciseness and consistency.
It's now simply called fix_encoding. The name fix_text_encoding is
available but emits a warning.Pending deprecations:
Python 2.6 support is largely coincidental.
Python 2.7 support is on notice. If you use Python 2, be sure to pin a version of ftfy less than 5.0 in your requirements.
ftfy.fixes.fix_surrogates will fix all 16-bit surrogate codepoints, which would otherwise break various encoding and output functions.
New features:
ftfy.fixes.fix_surrogates will fix all 16-bit surrogate codepoints,
which would otherwise break various encoding and output functions.Deprecations:
remove_unsafe_private_use emits a warning, and will disappear in the
next minor or major version.Certain symbols are marked as "ending punctuation" that may naturally occur after letters. When they follow an accented capital letter and look like m
Heuristic changes:
New features:
ftfy.explain_unicode is a diagnostic function that shows you what's going
on in a Unicode string. It shows you a table with each code point in
hexadecimal, its glyph, its name, and its Unicode category.
ftfy.fixes.decode_escapes adds a feature missing from the standard library:
it lets you decode a Unicode string with backslashed escape sequences in it
(such as "\u2014") the same way that Python itself would.
ftfy.streamtester is a release of the code that I use to test ftfy on
an endless stream of real-world data from Twitter. With the new heuristics,
the false positive rate of ftfy is about 1 per 6 million tweets. (See
the "Accuracy" section of the documentation.)
Deprecations:
Python 2.6 is no longer supported.
remove_unsafe_private_use is no longer needed in any current version of
Python. This fixer will disappear in a later version of ftfy.
fix_line_breaks fixes three additional characters that are considered line breaks in some environments, such as Javascript, and Python's "codecs" libr
fix_line_breaks fixes three additional characters that are considered line
breaks in some environments, such as Javascript, and Python's "codecs"
library. These are all now replaced with \n:
U+0085 <control>, with alias "NEXT LINE"
U+2028 LINE SEPARATOR
U+2029 PARAGRAPH SEPARATOR
Fix utf-8-variants so it never outputs surrogate codepoints, even on Python 2 where that would otherwise be possible.
utf-8-variants so it never outputs surrogate codepoints, even on
Python 2 where that would otherwise be possible.Fix bug in 3.1.1 where strings with backslashes in them could never be fixed
Add the ftfy.bad_codecs package, which registers new codecs that can decoding things that Python may otherwise refuse to decode:
Add the ftfy.bad_codecs package, which registers new codecs that can
decoding things that Python may otherwise refuse to decode:
utf-8-variants, which decodes CESU-8 and its Java lookalike
sloppy-windows-*, which decodes character-map encodings while treating
unmapped characters as Latin-1
Simplify the code using ftfy.bad_codecs.
Nothing published for this version
Fix the arguments to fix_file, because they were totally wrong.
fix_file, because they were totally wrong.Restore compatibility with Python 2.6.
Fixed an ugly regular expression bug that prevented ftfy from importing on a narrow build of Python.
Basically, 3.0.1 was too eager to treat text as MacRoman or cp437 when three consecutive characters coincidentally decoded as UTF-8. Increased the cos
Fixed some false positives.
Basically, 3.0.1 was too eager to treat text as MacRoman or cp437 when three consecutive characters coincidentally decoded as UTF-8. Increased the cost of those encodings so that they have to successfully decode multiple UTF-8 characters.
See tests/test_real_tweets.py for the new test cases that were added as a
result.
Fix bug in fix_java_encoding that led to only the first instance of CESU-8 badness per line being fixed
fix_java_encoding that led to only the first instance of
CESU-8 badness per line being fixedUnderstands more encodings and more kinds of mistakes
fix_text_encoding) will be
consistently skipped on long lines, but all other fixes will applyFix breaking up of long lines, so it can't go into an infinite loop
- Restored Python 2.6 support
Use fast Python built-ins to speed up fixes
Made into its own package with no dependencies, instead of a part of metanl
metanlYour coding agent can read these notes before it upgrades. Set up the MCP server →