NewYour coding agent can read the release notes before it upgrades.Set up the MCP server →
PyPI · #8 most downloaded on PyPI
The Real First Universal Charset Detector. Open, modern and actively maintained alternative to Chardet.
Last release 4 days ago
30 Sep 2026
Release timing varies
gaps range from 9 days to 11 months
Nearly every release is documented
notes for 53 of 57 stable releases
1 version withdrawn
withdrawn after publishing
7 years old
66 releases · first in 2019
Raised the Cython upper bound to <3.4 for native builds. The bound remains <3.3 for abi3 builds to preserve compatibility with the Python 3.7 Limited
<3.4 for native builds. The bound remains <3.3 forabi3 builds to preserve compatibility with the Python 3.7 Limited API.Raised upper bound of setuptools to v84
One column per quarter.
Explicit support for Python 3.15
Regression in our fallback path leading to a decode error. ( #771 ) We've yanked 3.4.8 as a result of that bug.
Wall import time due to cascade codec imports for our multibyte first sort of iana supported codecs
Pre-built optimized version using mypy[c] v1.20.
setuptools constraint to setuptools>=68,<82.1.Flattened the logic in charset_normalizer.md for higher performance. Removed eligible(..) and feed(...) in favor of feed_info(...) .
charset_normalizer.md for higher performance. Removed eligible(..) and feed(...)feed_info(...).UNICODE_RANGES_COMBINED using Unicode blocks v17.--normalize writing to wrong path when passing multiple files in. (#702)Update setuptools constraint to setuptools>=68,<=82 .
setuptools constraint to setuptools>=68,<=82.query_yes_no function (inside CLI) to avoid using ambiguous licensed code.cd.py submodule into mypyc optional compilation to reduce further the performance impact.Warning
mypyc changed the usual binary output for the optimized wheel. Beware, especially if using PyInstaller or alike. See #714
Bound setuptools to a specific constraint setuptools>=68,<=81 .
setuptools to a specific constraint setuptools>=68,<=81.setuptools-scm as a build dependency.dev-requirements.txt and created ci-requirements.txt for security purposes.multiple.intoto.jsonl in GitHub releases in addition to individual attestation file per wheel.mypy(c) is no longer a required dependency at build time if CHARSET_NORMALIZER_USE_MYPYC isn't set to 1 .
CHARSET_NORMALIZER_USE_MYPYC isn't set to 1. (#595) (#583)detect output legacy function. (#391)Addressed the DeprecationWarning in our CLI regarding argparse.FileType by backporting the target class into the package.
argparse.FileType by backporting the target class into the package. (#591)Deprecation warning "'count' is passed as positional argument" when converting to Unicode bytes on Python 3.13+
Did you know that Internet Explorer 11 shipped with an optional HTTP/2 support back in 2013? also libcurl did ship it in 2014[...] Using Requests today is the rough equivalent of using EOL Windows 8! We promptly invite Python developers to look at the first drop-in replacement for Requests, namely Niquests. Ship with native WebSocket, SSE, Happy Eyeballs, DNS over HTTPS, and so on[...] All of this while remaining compatible with all Requests prior plug-ins / add-ons.
It leverages charset-normalizer in a better way! Check it out, you will gain up to being 3X faster and get a real/respectable support with it.
pyproject.toml instead of setup.cfg using setuptools as the build backend.build-requirements.txt as per using pyproject.toml native build configuration.bin/integration.py and bin/serve.py in favor of downstream integration test (see noxfile).setup.cfg in favor of pyproject.toml metadata configuration.utils.range_scan function.utf_8 instead of preferred utf-8. (#572)Did you know that Internet Explorer 11 shipped with an optional HTTP/2 support back in 2013? also libcurl did ship it in 2014[...] All of this while o
Did you know that Internet Explorer 11 shipped with an optional HTTP/2 support back in 2013? also libcurl did ship it in 2014[...] All of this while our community is still struggling to make a firm advancement in HTTP clients. Now, many of you use Requests as the defacto http client, now, and for many years now, Requests has been frozen. Being left in a vegetative state and not evolving, this blocked millions of developers from using more advanced features.
We promptly invite Python developers to look at the drop-in replacement for Requests, namely Niquests. It leverage charset-normalizer in a better way! Check it out, you will be positively surprised! Don't wait another decade.
We are thankful to @microsoft and involved parties for funding our work through the Microsoft FOSS Fund program.
--no-preemptive in the CLI to prevent the detector to search for hints.Unintentional memory usage regression when using large payloads that match several encodings
Optional mypyc compilation upgraded to version 1.6.1 for Python >= 3.8
Allow to execute the CLI (e.g. normalizer) through python -m charset_normalizer.cli or python -m charset_normalizer
python -m charset_normalizer.cli or python -m charset_normalizerencoding.aliases as they have no alias (#323)Typehint for function from_path no longer enforce PathLike as its first argument
from_path no longer enforce PathLike as its first argumentis_binary that relies on main capabilities, and is optimized to detect binariesenable_fallback argument throughout from_bytes, from_path, and from_fp that allow a deeper control over the detection (default True)Argument should_rename_legacy for legacy function detect and disregard any new arguments without errors (PR #262)
should_rename_legacy for legacy function detect and disregard any new arguments without errors (PR #262)Multi-bytes cutter/chunk generator did not always cut correctly (PR #233)
Extend the capability of explain=True when cp_isolation contains at most two entries (min one), will log in details of the Mess-detector results
language_threshold in from_bytes, from_path and from_fp to adjust the minimum expected coherence rationormalizer --version now specify if the current version provides extra speedup (meaning mypyc compilation whl)md.py can be compiled using Mypyc to provide an extra speedup up to 4x faster than v2.1first() and best() from CharsetMatchnormalizechaos_secondary_pass, coherence_non_latin and w_counter from CharsetMatchunicodedata2This is the last version (3.0.x) to support Python 3.6 We plan to drop it for 3.1.x
language_threshold in from_bytes, from_path and from_fp to adjust the minimum expected coherence rationormalizer --version now specify if current version provide extra speedup (meaning mypyc compilation whl)md.py can be compiled using Mypyc to provide an extra speedup up to 4x faster than v2.1first() and best() from CharsetMatchnormalizechaos_secondary_pass, coherence_non_latin and w_counter from CharsetMatchunicodedata2_This is the last pre-release. If everything goes well, I will publish the stable tag._
This is the last pre-release. If everything goes well, I will publish the stable tag.
language_threshold in from_bytes, from_path and from_fp to adjust the minimum expected coherence ratiolanguage_threshold in from_bytes, from_path and from_fp to adjust the minimum expected coherence rationormalizer --version now specify if current version provide extra speedup (meaning mypyc compilation whl)
normalizer --version now specify if current version provide extra speedup (meaning mypyc compilation whl)first() and best() from CharsetMatchOptional: Module md.py can be compiled using Mypyc to provide an extra speedup up to 4x faster than v2.1
md.py can be compiled using Mypyc to provide an extra speedup up to 4x faster than v2.1normalizechaos_secondary_pass, coherence_non_latin and w_counter from CharsetMatchunicodedata2Function normalize scheduled for removal in 3.0
normalize scheduled for removal in 3.0Output the Unicode table version when running the CLI with --version (PR #194)
--version (PR #194)unicodedata2 as Python is quickly catching up, scheduled for removal in 3.0 (PR #194)ASCII miss-detection on rare cases (PR #170)
Explicit support for Python 3.11 (PR #164)
Fallback match entries might lead to UnicodeDecodeError for large bytes sequence (PR #154)
Moderating the logging impact (since 2.0.8) for specific environments (PR #147)
explain to True (PR #146)Improvement over Vietnamese detection (PR #126)
NullHandler by default from @nmaynes (PR #135)explain to True will add provisionally (bounded to function lifespan) a specific stream handler (PR #135)set_logging_handler to configure a specific StreamHandler from @nmaynes (PR #135)CHANGELOG.md entries, format is based on Keep a Changelog (PR #141)We arrived in a pretty stable state.
We arrived in a pretty stable state.
Changes:
SyntaxError (Not about ASCII decoding error) for those trying to install this package using a non-supported Python versionThis version pushes forward the detection-coverage to 98%! https://github.com/Ousret/charset_normalizer/runs/3863881150 The great filter (cannot be better than) shall be 99% in conjunction with the current dataset. In future releases.
Bugfix: :bug: Unforeseen regression with the loss of the backward-compatibility with some older minor of Python 3.5.x #100
Changes:
Internal: :art: The project now comply with: flake8, mypy, isort and black to ensure a better overall quality #81 Internal: :art: The MANIFEST.in was
Changes:
Internal: :art: The project now comply with: flake8, mypy, isort and black to ensure a better overall quality #81
Internal: :art: The MANIFEST.in was not exhaustive #78
Improvement: :sparkles: The BC-support with v1.x was improved, the old staticmethods are restored #82
Remove: :fire: The project no longer raise warning on tiny content given for detection, will be simply logged as warning instead #92
Improvement: :sparkles: The Unicode detection is slightly improved, see #93
Bugfix: :bug: In some rare case, the chunks extractor could cut in the middle of a multi-byte character and could mislead the mess detection #95
Bugfix: :bug: Some rare 'space' characters could trip up the UnprintablePlugin/Mess detection #96
Improvement: :art: Add syntax sugar __bool__ for results CharsetMatches list-container see #91
This release push further the detection coverage to 97 % !
Improvement: ❇️ Adjust the MD to lower the sensitivity, thus improving the global detection reliability (#69 #76)
Changes:
Improvement: ✨ Part of the detection mechanism has been improved to be less sensitive, resulting in more accurate detection results. Especially ASCII.
Changes:
Be assured that this project is disposed to listen to any of your concerns you may have. I know the vast majority did not expect requests to switch Chardet to Charset-Normalizer. I am inclined to make this change worth it, only together that we can achieve great things. Do not hesitate to leave feedback or bug report, will answer them all!
Bugfix: 🐛 Empty/Too small JSON payload miss-detection fixed (#59) Thanks @tseaver for the report
Changes:
Bugfix: :bug: Make it work where there isn't a filesystem available, dropping assets frequencies.json #54 #55 original report by @sethmlarson
Minor bug fixes release.
Changes:
cp_isolation and cp_exclusion arguments #47explain=False permanently disable the verbose output in the current runtime #47explain=True #47normalize default args values were not aligned with from_bytes #53Deprecation: 🔴 Methods coherence_non_latin, w_counter, chaos_secondary_pass of the class CharsetMatch are now deprecated and scheduled for removal in…
This package is reaching its two years of existence, now is a good time for a nice refresh.
Changes: See PR #45
detect function returns an identical charset name whenever possible.Performance, Backward-Compatibility with Chardet, and Detection Coverage in addition to currents tests. (+CodeQL)cached_property)--version argument to CLIcoherence_non_latin, w_counter, chaos_secondary_pass of the class CharsetMatch are now deprecated and scheduled for removal in v3.0utf_7 detection has been reinstated.After much consideration, this release won't drop Python 3.5 in v2.
Improvement: :art: Logger configuration/usage no longer conflict with others #44
Changes :
Thanks to @potiuk for his tests/ideas that permitted us to improve the quality of this project.
Changes :
Thanks to @potiuk for his tests/ideas that permitted us to improve the quality of this project.
logging instead of using the package loguru.nose test framework in favor of the maintained pytest.dragonmapper package to help with gibberish Chinese/CJK text.cached_property only for Python 3.5 due to constraint. Dropping for every other interpreter version.CharsetNormalizerMatch instance could be False in rare cases even if obviously present. Due to the sub-match factoring process.frequencies.json.This project no longer requires anything except for python 3.5. It is still supported even if passed EOL. Version 2.x will require Python 3.6+
Bugfix: :bug: In some very rare cases, you may end up getting encode/decode errors due to a bad bytes payload #40
Changes :
Bugfix: :bug: Empty given payload for detection may cause an exception if trying to access the alphabets property. #39
Changes :
alphabets property. #39Bugfix: :bug: The legacy detect function should return UTF-8-SIG if sig is present in the payload. #38
Changes :
detect function should return UTF-8-SIG if sig is present in the payload. #38Amend the previous release to allow prettytable 2.0 Thanks to @jayvdb #35
Amend the previous release to allow prettytable 2.0
Thanks to @jayvdb #35
Miscellaneous: 🔧 Dependencies refactor, add python 3.9 and 3.10 to the supported interpreters
Changes :
Small refresh to keep the project up and running until further dev. (Upcoming version 1.4.0) Thanks to the many adopters.
Improvement/Bugfix : False positive when searching for successive upper, lower char. (ProbeChaos)
Changes :
Improvement : Noticeable better detection for jp #30
Changes :
Bugfix : Passing zero-length bytes to from_bytes
Changes :
Improvement : Expose __version__ in package
Changes :
Feature : Now support unicodedata2 backport. To benefit from it install using pip install charset-normalizer[UnicodeDataBackport]. Python 3.7 have Uni
Changes :
<img src="https://media.tenor.com/images/181fd3b2bcb8c61f2eeda4d5f3d3fd3b/tenor.gif" width="90"/>
unicodedata2 backport. To benefit from it install using pip install charset-normalizer[UnicodeDataBackport]. Python 3.7 have UnicodeData v11. You could upgrade it to v12.preemptive_behaviour. Default to True. Does not take declared encoding for it, testing it first.CharsetDetector; EncodingDetector and CharsetDoctor.Feature : Added has_submatch, percent_chaos and percent_coherence properties on single match object.
Changes :
<img src="https://media.tenor.com/images/181fd3b2bcb8c61f2eeda4d5f3d3fd3b/tenor.gif" width="90"/>
has_submatch, percent_chaos and percent_coherence properties on single match object.best() method of CharsetNormalizerMatches has been rewritten for better readability.explain boolean positional parameter to print out what actually happen when searching for a match.cp_exclusion. List of str. for from_bytes from_path and from_fp.cp_isolation. List of str. for from_bytes from_path and from_fp.import charset_normalizer is enough to provide additional help when you encounter UnicodeDecodeError exception.Bugfix : from_bytes parameters *steps* and *chunk_size* were not adapted to sequence len if provided values were not fitted to content. Therefore coul
Changes :
from_bytes parameters steps and chunk_size were not adapted to sequence len if provided values were not fitted to content. Therefore could lead to misdetection on small content.Bugfix : Sequence having lenght bellow 10 chars was not checked by ProbeChaos at all.
Changes :
detect method inspired by chardet was not returning intended result when having no result. (#14)Nothing published for this version
Nothing published for this version
- Bugfix on UTF 7 support - Legacy detect(byte_str) method
Chaos prober improved on small text
RC 5
Probe Chaos: Code cleanup, performance review and accuracy improved
RC 4
Nothing published for this version
Close file after reading them in CLI mode
RC 2
Your coding agent can read these notes before it upgrades. Set up the MCP server →