NewYour coding agent can read the release notes before it upgrades.Set up the MCP server →
PyPI · #166 most downloaded on PyPI
Universal character encoding detector
Last release 1 months ago
14 Aug 2026
Ships unpredictably
gaps range from 8 days to 3.5 years
Nearly every release is documented
notes for 30 of 30 stable releases
Nothing withdrawn
no release was ever pulled
20 years old
31 releases · first in 2006
Big release: a Cython scoring kernel joins mypyc in compiled wheels, every model retrained on a deduplicated corpus, UTF-7 fixed in both directions, a
Big release: a Cython scoring kernel joins mypyc in compiled wheels, every model retrained on a deduplicated corpus, UTF-7 fixed in both directions, and a guarantee that detect() never returns an encoding that can't decode your complete input.
_kernel.py stays plain Python (PyPy and pure wheels run it interpreted, unchanged), and detection output is bit-identical. The kernel declares itself safe without the GIL, so free-threaded CPython scales instead of silently re-enabling the GIL on import: 3.14t runs the whole suite in ~340ms across 8 threads, the fastest configuration measured. Compiled builds now need both hooks: HATCH_BUILD_HOOK_ENABLE_MYPYC=true HATCH_BUILD_HOOK_ENABLE_CUSTOM=true.detect() no longer returns an encoding that cannot decode the input it was given (#380, thanks @yarikoptic). When the whole input has been examined and the winner's only multi-byte evidence is an incomplete trailing sequence, the best candidate that decodes the input completely wins instead. Genuinely truncated data keeps its answer.|NAME,+LAY| misdetecting as UTF-7 (#371 follow-up, thanks @agreenburg). The whole buffer must now actually decode as UTF-7, and a lone shifted character must land in a plausible script range.Full Changelog: 7.5.1...7.6.0
One column per quarter.
Performance:
Compiled wheels now score bigram profiles through a small Cython kernel alongside mypyc, and the pair is 4.7x faster than the pure wheel on CPython 3.14. _kernel.py stays plain Python (PyPy and pure wheels run it interpreted, unchanged), _kernel.pxd adds C types at build time and ships nothing, and detection output is bit-identical. The kernel declares itself safe without the GIL, so free-threaded CPython scales instead of silently re-enabling the GIL on import: 3.14t runs the whole suite in ~340ms across 8 threads, the fastest configuration measured. Compiled builds now need both hooks:
HATCH_BUILD_HOOK_ENABLE_MYPYC = true HATCH_BUILD_HOOK_ENABLE_CUSTOM = true uv build
( Dan Blanchard via Claude)
Added support for CPython 3.15, including the free-threaded build. No code changes were needed. ( Dan Blanchard via Claude)
Bug Fixes:
Fixed delimited ASCII data like |NAME,+LAY| misdetecting as UTF-7, a follow-up to #371 . Two new checks: the whole buffer must actually decode as UTF-7 ( +| is an illegal shift, so tabular data fails immediately), and a block encoding a single code unit must land in a script range where a lone shifted character plausibly occurs. +LAY decodes to U+2C06, Glagolitic; no genuine lone block in the corpus lands anywhere like it, while em dashes, ellipses, kanji, and accented letters all pass. ( Dan Blanchard via Claude)
Signed UTF-7 no longer reads as ASCII. The BOM stage recognizes the four UTF-7 signature prefixes ( +/v8- and friends) when the rest of the buffer decodes as UTF-7 — the prefix alone is ordinary ASCII (a diff of V8 source paths starts with +/v8 ). This is a deliberate divergence from WHATWG’s browser-security exclusion of UTF-7: chardet already detects the unsigned form, so refusing only the signed one made no sense. ( Dan Blanchard via Claude)
detect() no longer returns an encoding that cannot decode the input it was given ( #380 ). When the whole input has been examined and the winner’s only multi-byte evidence is an incomplete trailing sequence, the best candidate that decodes the input completely wins instead. Genuinely truncated data keeps its answer: CJK cut mid-character, or input sliced at max_bytes . ( Dan Blanchard via Claude)
Short apostrophe-heavy English is no longer labeled Scottish Gaelic or Breton. A rare-language label on an input under 128 bytes now needs a 0.03 lead over the best mainstream language, mirroring the encoding-side arbitration in ADR-0005. ( Dan Blanchard via Claude)
Hungarian text no longer loses to a Czech reading: confusion rescoring compares tied pairs only under language models both encodings have. ( Dan Blanchard via Claude)
Space-padded text no longer matches a degenerate Serbian model at high confidence. Statistical scoring now skips repeated-whitespace bigrams, matching the whitespace collapse training already applies. Fixed 21 files plus a long-standing GB2312 known failure. ( Dan Blanchard via Claude)
EBCDIC text is no longer invisible to the early pipeline. The binary stage treats EBCDIC’s 0x05/0x15 tab and newline as whitespace when the data is high-byte-dominated, and the markup stage reads charset declarations through a cp037 decode of the head. ( Dan Blanchard via Claude)
Fixed the last two EBCDIC sibling misdetections: a letter reading beating punctuation is no longer evidence by itself, and low-confidence near-ties scan deeper but need the category vote and the bigram rescore to agree. ( Dan Blanchard via Claude)
Fixed training normalization gaps that starved ISO-8859-16 (legacy cedilla forms) and the 26 pre-euro encodings (the euro sign) at exactly the bytes that discriminate them from their siblings. Cedilla folds to comma-below and the euro to the currency sign wherever the target encoding cannot represent them. ( Dan Blanchard via Claude)
Improvements:
Retrained every bigram model on a refreshed, deduplicated corpus with the ADR-0004 hardening: whitespace collapses after encoding, retention guards hard-fail mostly unencodable corpora, Serbian gets real Latin-script text, CP1006 gains its sixteen missing Urdu letters, and wiki markup is stripped before bigram counting. Training provenance is now recorded per model, so test data added after a retrain is detectable as such. ( Dan Blanchard via Claude)
New ANSI-art model. cp437 detection now includes a profile trained on 16,621 text-mode art files from 16colo.rs , keyed under the zxx pseudo-language and reported with language=None . ( Dan Blanchard via Claude)
Rare-language arbitration: a low-confidence statistical winner from a language with no documented legacy-encoding population (Scottish Gaelic, Welsh, Irish, Breton) yields to a near-tied mainstream candidate. Genuine Celtic text wins by landslides and is unaffected; eight boundary sentinels in the test suite guard the gate. Design and evidence in ADR-0005. ( Dan Blanchard via Claude)
Confusion-group resolution is context-aware: votes are counted per occurrence, letter readings with no word shape are demoted, and art-model wins are exempt. Fixed twelve EBCDIC and Latin files that were riding single-byte coin flips. ( Dan Blanchard via Claude)
Statistical dead heats no longer resolve by candidate enumeration order. Three tiebreaks: prefer the Windows superset, prefer the more prevalent era when there is no high-byte evidence, and prefer a classic-Mac candidate when line endings are bare \r . ( Dan Blanchard via Claude)
Training pipeline hardening after a cache-loss post-mortem: retrains that would silently drop a model now abort loudly, caches filter in place instead of being deleted wholesale, and the artpack fetcher builds into a temporary directory. ( Dan Blanchard via Claude)
Patch release: three detection fixes found while benchmarking against charset-normalizer's char-dataset.
Patch release: three detection fixes found while benchmarking against charset-normalizer's char-dataset.
Shift_JIS but using CP932 extension characters (like ①) came back as SHIFT_JIS, which fails .decode() on those same bytes. Superset promotion (CP932, CP949) now always fires when the reported name can't decode the data but the superset can.<meta charset="iso-8859-1"> came back as ISO-8859-1, which decodes to mojibake. Valid multi-byte UTF-8 now wins over a conflicting declaration.Full Changelog: 7.5.0...7.5.1
Bug Fixes:
Fixed markup-declared encodings being reported under a name that can’t decode the input. A page declaring Shift_JIS but using CP932 extension characters (like ①) came back as SHIFT_JIS , which fails .decode() on those same bytes. Superset promotion ( CP932 , CP949 ) now always fires when the reported name can’t decode the data but the superset can. ( Dan Blanchard via Claude)
Fixed a lying charset declaration beating genuine UTF-8 content. A UTF-8 page declaring <meta charset="iso-8859-1"> came back as ISO-8859-1, which decodes to mojibake. Valid multi-byte UTF-8 now wins over a conflicting declaration; pure ASCII and real single-byte content still honor it. ( Dan Blanchard via Claude)
Fixed BOM-less UTF-16 byte-order detection for pure-CJK text. With no ASCII in the sample, the only null bytes come from the low byte of characters like U+4E00 (一), which sit in the wrong parity position, so short Chinese UTF-16 samples came back with reversed endianness at full confidence. Byte order is now chosen by decoding both ways and comparing text quality, with the null signal breaking near-ties. Found by scoring chardet against charset-normalizer’s char-dataset. ( Dan Blanchard via Claude)
chardet.equivalences is now a deprecation shim: evaluation predicates moved to chardet.evaluation , output-name remapping to chardet.output_names . Re…
Accuracy and speed release: truncation-proof byte validity, statistical pruning worth ~2.9x, and half the peak memory.
max_bytes and _SCAN_LIMIT slicing. Validity checks now decode incrementally with final=False. (#376, thanks @aadsm)compat_names (the default) leaking internal Python codec names for seven encodings (ISO-8859-2, ISO-8859-6, ISO-8859-13, Windows-1250, Windows-1256, Windows-1257, CP874). (#374, thanks @aadsm)compat_names leaking the internal cp932 codec name; detect() now returns CP932. (#375, thanks @uttam12331)bytes (native array indexing under mypyc) and the model blob is decompressed in chunks: peak process memory dropped from 53.9 to 27.4 MiB.bytes.translate prefilters instead of per-byte Python scans.prefer_superset=True is now documented as the recommended mode and will become the default in chardet 8.0. Callers that depend on subset names should start passing prefer_superset=False explicitly.chardet.equivalences is now a deprecation shim: evaluation predicates moved to chardet.evaluation, output-name remapping to chardet.output_names. Removal in 8.0.Full Changelog: 7.4.3...7.5.0
Bug Fixes:
Fixed multi-byte encodings being eliminated when the input ends in an incomplete character. Byte-validity filtering decoded with a one-shot strict decode, which cannot tell a truncated tail from corrupt data, so a single dangling lead byte dropped every CJK candidate and the result came down to input-length parity — a 184-byte GBK sample detected as GB18030 , the same sample minus one byte as Windows-1256 . This was also reachable on complete, well-formed files, because chardet slices its own input at max_bytes and at _SCAN_LIMIT in _validate_bytes() : a valid 14 kB GBK page with an honest <meta charset="gbk"> lost its declaration, and with it text/html and 0.95 confidence, whenever byte 4096 happened to split a character. Validity checks now decode incrementally with final=False , deferring a partial trailing character while still rejecting corruption anywhere before it. ( António Afonso via Claude, #376 )
Fixed compat_names (the default) leaking internal Python codec names for seven encodings. detect() now returns ISO-8859-2 , ISO-8859-6 , ISO-8859-13 , Windows-1250 , Windows-1256 , Windows-1257 , and CP874 instead of their lowercase codec spellings. These were absent from _COMPAT_NAMES after the 7.1.0 switch to codec-name canonicals, which made default output inconsistent with their siblings (e.g. cp1250 vs Windows-1251 ) and with the encoding-name table in Usage . ( António Afonso via Claude, #374 )
Fixed compat_names (the default) leaking the internal cp932 codec name. detect() now returns CP932 instead of cp932 , matching its Japanese siblings ( shift_jis_2004 → SHIFT_JIS ) and the value chardet 5.x/6.x returned. ( uttam12331 , #375 )
Performance:
Statistical scoring now skips single-byte models that provably can’t beat the current runner-up, using per-model row-maximum tables ( rowmax.bin ). Results are bit-identical: multi-byte models are always scored in full and detect_all() bypasses pruning. Mean detection time dropped ~2.9x with mypyc, with the largest gains on legacy CJK (p99 from 9.5ms to 3.3ms). ( Dan Blanchard via Claude)
Model tables are now bytes (native array indexing under mypyc, instead of boxed memoryview calls), and the model blob is decompressed in chunks rather than one shot: peak process memory dropped from 53.9 to 27.4 MiB. ( Dan Blanchard via Claude)
Confusion-group resolution and post-processing now use bytes.translate prefilters instead of per-byte Python scans, making near-tie resolution cheaper on large inputs. ( Dan Blanchard via Claude)
Improvements:
prefer_superset=True is now documented as the recommended mode and will become the default in chardet 8.0 . Detection examines at most max_bytes of input, so only the superset encoding is guaranteed to decode bytes beyond that window — the same reasoning behind the WHATWG/W3C Encoding Standard’s rule that browsers decode ascii and iso-8859-1 content as windows-1252 . Callers that depend on subset names should start passing prefer_superset=False explicitly. ( Dan Blanchard via Claude)
chardet.equivalences is now a deprecation shim. Accuracy-evaluation predicates ( is_correct , is_equivalent_detection , etc.) moved to chardet.evaluation ; public-API encoding-name remapping ( apply_compat_names , apply_preferred_superset ) moved to chardet.output_names . Existing imports keep working with a DeprecationWarning . chardet.equivalences will be removed in 8.0. ( Dan Blanchard via Claude)
Internal pipeline reorganization: language detection, markup-superset promotion, and post-processing rank corrections moved out of the orchestrator into pipeline/language.py , pipeline/markup.py , and pipeline/postprocess.py respectively. No behavior change. The two new modules are also added to the mypyc compilation list. ( Dan Blanchard via Claude)
Patch release: fixes a crash when input contains null bytes inside a <meta charset> declaration.
Patch release: fixes a crash when input contains null bytes inside a <meta charset> declaration.
ValueError: embedded null character crash when input contained a <meta charset> declaration with a null byte in the encoding name (e.g. b'<meta charset="\x00utf-8">'). codecs.lookup() raises ValueError on embedded nulls, and lookup_encoding() was only catching LookupError. Also added defensive ValueError catches in _validate_bytes() and _to_utf8() for completeness. (#369, thanks @DRMacIver for the report)Full Changelog: 7.4.2...7.4.3
Bug Fixes:
Fixed ValueError: embedded null character crash when input contained a <meta charset> declaration with a null byte in the encoding name (e.g. b'<meta charset="\x00utf-8">' ). codecs.lookup() raises ValueError on embedded nulls, and lookup_encoding() was only catching LookupError . Also added defensive ValueError catches in _validate_bytes() and _to_utf8() for completeness. ( Dan Blanchard via Claude, #369 )
Patch release: fixes a crash on short inputs and closes a bunch of WHATWG/IANA alias gaps.
Patch release: fixes a crash on short inputs and closes a bunch of WHATWG/IANA alias gaps.
RuntimeError: pipeline must always return at least one result on ~2% of all possible two-byte inputs (e.g. b"\xf9\x92"). Multi-byte encodings like CP932 and Johab could score above the structural confidence threshold on very short inputs, but then statistical scoring would return nothing, leaving an empty result list instead of falling through to the fallback. (#367, #368, thanks @jasonwbarnett)<meta charset> labels like x-cp1252, x-sjis, dos-874, csUTF8, and the cswindows* family all resolve correctly through the markup detection stage. Every alias was driven by a failing spec-compliance test, not speculative. (#366)Full Changelog: 7.4.1...7.4.2
Bug Fixes:
Fixed RuntimeError: pipeline must always return at least one result on ~2% of all possible two-byte inputs (e.g. b"\xf9\x92" ). Multi-byte encodings like CP932 and Johab could score above the structural confidence threshold on very short inputs, but then statistical scoring would return nothing, leaving the pipeline with an empty result list instead of falling through to the no_match_encoding fallback. ( Jason Barnett via Claude, #367 , #368 )
Improvements:
Added ~90 encoding aliases from the WHATWG Encoding Standard and IANA Character Sets registry so that <meta charset> labels like x-cp1252 , x-sjis , dos-874 , csUTF8 , and the cswindows* family all resolve correctly through the markup detection stage. Every alias was driven by a failing spec-compliance test. ( Dan Blanchard via Claude, #366 )
Added a spec-compliance test suite covering Python decode round-trips for all 86 registry encodings, WHATWG web-platform label resolution, IANA preferred MIME names, and Unicode/RFC conformance (BOM sniffing, UTF-8 boundary cases, UTF-16 surrogate pairs). This is the test suite that would have caught the 7.4.1 BOM bug before release. ( Dan Blanchard via Claude, #366 )
BOM-prefixed UTF-16/32 input now returns utf-16 / utf-32 instead of utf-16-le / utf-16-be / utf-32-le / utf-32-be . The endian-specific codecs don't s
utf-16/utf-32 instead of utf-16-le/utf-16-be/utf-32-le/utf-32-be. The endian-specific codecs don't strip the BOM on decode, so callers were getting a stray U+FEFF at the start of their text. BOM-less detection is unchanged. (#364, #365)Full Changelog: 7.4.0...7.4.1
Bug Fixes:
BOM-prefixed UTF-16 and UTF-32 input now reports utf-16 and utf-32 instead of the endian-specific variants. Python’s utf-16-le / utf-16-be / utf-32-le / utf-32-be codecs keep the BOM as a U+FEFF in the decoded string, while utf-16 / utf-32 strip it, so callers passing the detection result directly to .decode() were getting a stray BOM at the start of their text. BOM-less UTF-16/32 detection (via null-byte patterns) is unchanged and still returns the endian-specific name. ( Dan Blanchard via Claude, #364 , #365 )
Release 7.4.0.post2: fix Windows wheel versioning in release workflow
Release 7.4.0.post2: fix Windows wheel versioning in release workflow
Fix Windows mypyc wheel builds (MSVC C2026 string literal limit). No code changes from 7.4.0, only build configuration.
Fix Windows mypyc wheel builds (MSVC C2026 string literal limit).
No code changes from 7.4.0, only build configuration.
0BSD license — the project license has been changed from MIT to 0BSD, a maximally permissive license with no attribution requirement. All prior 7.x re
mime_type field to detection results — identifies file types for both binary (via magic number matching) and text content. Returned in all detect(), detect_all(), and UniversalDetector results. (#350)pipeline/magic.py module detects 40+ binary file formats including images, audio/video, archives, documents, executables, and fonts. ZIP-based formats (XLSX, DOCX, JAR, APK, EPUB, wheel, OpenDocument) are distinguished by entry filenames. (#350)dataclasses.replace() with direct DetectionResult construction on hot paths, eliminating ~354k function calls per full test suite runLicense:
0BSD license — the project license has been changed from MIT to 0BSD , a maximally permissive license with no attribution requirement. All prior 7.x releases should also be considered 0BSD licensed as of this release. ( Dan Blanchard via Claude)
Features:
Added mime_type field to detection results — identifies file types for both binary (via magic number matching) and text content. Returned in all detect() , detect_all() , and UniversalDetector results. ( Dan Blanchard via Claude, #350 )
New pipeline/magic.py module detects 40+ binary file formats including images, audio/video, archives, documents, executables, and fonts. ZIP-based formats (XLSX, DOCX, JAR, APK, EPUB, wheel, OpenDocument) are distinguished by entry filenames. ( Dan Blanchard via Claude, #350 )
Bug Fixes:
Fixed incorrect equivalence between UTF-16-LE and UTF-16-BE in accuracy testing — these are distinct encodings with different byte order, not interchangeable ( Dan Blanchard via Claude)
Performance:
Added 4 new modules to mypyc compilation (orchestrator, confusion, magic, ascii), bringing the total to 11 compiled modules ( Dan Blanchard via Claude)
Capped statistical scoring at 16 KB — bigram models converge quickly, so large files no longer score the full 200 KB. Worst-case detection time dropped from 62ms to 26ms with no accuracy loss. ( Dan Blanchard via Claude)
Replaced dataclasses.replace() with direct DetectionResult construction on hot paths, eliminating ~354k function calls per full test suite run ( Dan Blanchard via Claude)
Build:
Added riscv64 to the mypyc wheel build matrix — prebuilt wheels are now published for RISC-V Linux alongside existing architectures ( Bruno Verachten , #348 )
Added include_encodings and exclude_encodings parameters to detect(), detect_all(), and UniversalDetector — restrict or exclude specific encodings fro
include_encodings and exclude_encodings parameters to detect(), detect_all(), and UniversalDetector — restrict or exclude specific encodings from the candidate set, with corresponding -i/--include-encodings and -x/--exclude-encodings CLI flags (#343)no_match_encoding (default "cp1252") and empty_input_encoding (default "utf-8") parameters — control which encoding is returned when no candidate survives the pipeline or the input is empty, with corresponding CLI flags (#343)-l/--language flag to chardetect CLI — shows the detected language (ISO 639-1 code and English name) alongside the encoding (#342)Full changelog: https://chardet.readthedocs.io/en/latest/changelog.html
Features:
Added include_encodings and exclude_encodings parameters to ~chardet.detect, ~chardet.detect_all, and ~chardet.UniversalDetector — restrict or exclude specific encodings from the candidate set, with corresponding -i/--include-encodings and -x/--exclude-encodings CLI flags (Dan Blanchard via Claude, #343)
Added no_match_encoding (default "cp1252") and empty_input_encoding (default "utf-8") parameters — control which encoding is returned when no candidate survives the pipeline or the input is empty, with corresponding CLI flags (Dan Blanchard via Claude, #343)
Added -l/--language flag to chardetect CLI — shows the detected language (ISO 639-1 code and English name) alongside the encoding (Dan Blanchard via Claude, #342)
…import UniversalDetector works with a deprecation warning
# -*- coding: ... -*- and # coding=... declarations on lines 1–2 of Python source files are now recognized with confidence 0.95 (#249)chardet.universaldetector backward-compatibility stub so that from chardet.universaldetector import UniversalDetector works with a deprecation warning (#341)++ or +word patterns (#332)detect() call — model norms are now computed during loading instead of lazily iterating 21M entries (#333)detect() now returns chardet 5.x-compatible names by default (#338)iter_unpack yielded fewer tuples instead of raising)load_models()struct.iter_unpack for bulk entry extraction (eliminates ~305K individual unpack calls)compat_names parameter (default True) to detect(), detect_all(), and UniversalDetector — set to False to get raw Python codec names instead of chardet 5.x/6.x compatible display namesprefer_superset parameter (default False) — remaps legacy ISO/subset encodings to their modern Windows/CP superset equivalents (e.g., ASCII → Windows-1252, ISO-8859-1 → Windows-1252). This will default to True in the next major version (8.0).should_rename_legacy in favor of prefer_superset — a deprecation warning is emitted when used"utf-8" instead of "UTF-8"), with compat_names controlling the public output formatlookup_encoding() to registry for case-insensitive resolution of arbitrary encoding name input to canonical namesFull changelog: https://chardet.readthedocs.io/en/latest/changelog.html
Features:
Added PEP 263 encoding declaration detection — # -- coding: ... -- and # coding=... declarations on lines 1–2 of Python source files are now recognized with confidence 0.95 ( Dan Blanchard via Claude, #249 )
Added chardet.universaldetector backward-compatibility stub so that from chardet.universaldetector import UniversalDetector works with a deprecation warning ( Dan Blanchard via Claude, #341 )
Fixes:
Fixed false UTF-7 detection of ASCII text containing ++ or +word patterns ( Dan Blanchard , #332 , #335 )
Fixed 0.5s startup cost on first detect() call — model norms are now computed during loading instead of lazily iterating 21M entries ( Dan Blanchard via Claude, #333 , #336 )
Fixed undocumented encoding name changes between chardet 5.x and 7.0 — detect() now returns chardet 5.x-compatible names by default ( Dan Blanchard via Claude, #338 )
Improved ISO-2022-JP family detection — recognizes ESC sequences for ISO-2022-JP-2004 (JIS X 0213) and ISO-2022-JP-EXT (JIS X 0201 Kana) ( Dan Blanchard via Claude)
Fixed silent truncation of corrupt model data ( iter_unpack yielded fewer tuples instead of raising) ( Dan Blanchard via Claude)
Fixed incorrect date in LICENSE ( Dan Blanchard )
Performance:
5.5x faster first-detect time (~0.42s → ~0.075s) by computing model norms as a side-product of load_models() ( Dan Blanchard via Claude)
~40% faster model parsing via struct.iter_unpack for bulk entry extraction (eliminates ~305K individual unpack calls) ( Dan Blanchard via Claude)
New API parameters:
Added compat_names parameter (default True ) to detect() , detect_all() , and UniversalDetector — set to False to get raw Python codec names instead of chardet 5.x/6.x compatible display names ( Dan Blanchard via Claude)
Added prefer_superset parameter (default False ) — remaps legacy
ISO/subset encodings to their modern Windows/CP superset equivalents
(e.g., ASCII → Windows-1252, ISO-8859-1 → Windows-1252). This will default to True in the next major version (8.0). ( Dan Blanchard via Claude)
Deprecated should_rename_legacy in favor of prefer_superset — a deprecation warning is emitted when used ( Dan Blanchard via Claude)
Improvements:
Switched internal canonical encoding names to Python codec names (e.g., "utf-8" instead of "UTF-8" ), with compat_names controlling the public output format. See Usage for the full mapping table. ( Dan Blanchard via Claude)
Added lookup_encoding() to registry for case-insensitive resolution of arbitrary encoding name input to canonical names ( Dan Blanchard via Claude)
Achieved 100% line coverage across all source modules (+31 tests) ( Dan Blanchard via Claude)
Updated benchmark numbers: 98.2% encoding accuracy, 95.2% language accuracy on 2,510 test files ( Dan Blanchard via Claude)
Pinned test-data cloning to chardet release version tags for reproducible builds ( Dan Blanchard via Claude)
Fixed false UTF-7 detection of SHA-1 git hashes (#324, fixing #323) — requirements files with VCS pins (e.g., +4bafdea3...) were misdetected as UTF-7,
+4bafdea3...) were misdetected as UTF-7, breaking tools like tox_SINGLE_LANG_MAP missing aliases for single-language encoding lookup (e.g., big5 → big5hkscs)TypeError in UTF-7 codec handlingFixes:
Fixed false UTF-7 detection of SHA-1 git hashes ( Alex Rembish , #324 )
Fixed _SINGLE_LANG_MAP missing aliases for single-language encoding lookup (e.g., big5 → big5hkscs ) ( Dan Blanchard )
Fixed PyPy TypeError in UTF-7 codec handling ( Dan Blanchard )
Improvements:
Retrained bigram models — 24 previously failing test cases now pass ( Dan Blanchard via Claude)
Updated language equivalences for mutual intelligibility (Slovak/Czech, East Slavic + Bulgarian, Malay/Indonesian, Scandinavian languages) ( Dan Blanchard via Claude)
LanguageFilter is accepted but ignored (deprecation warning emitted)
Ground-up, MIT-licensed rewrite of chardet. Same package name, same public API — drop-in replacement for chardet 5.x/6.x. Just way faster and more accurate!
Highlights:
detect() and detect_all() with no measurable overhead; scales on free-threaded Python 3.13t+Breaking changes vs 6.0.0:
detect() and detect_all() now default to encoding_era=EncodingEra.ALL (6.0.0 defaulted to MODERN_WEB)LanguageFilter is accepted but ignored (deprecation warning emitted)chunk_size is accepted but ignored (deprecation warning emitted)Ground-up, 0BSD-licensed rewrite of chardet ( Dan Blanchard via Claude, #322 ). Same package name, same public API — drop-in replacement for chardet 5.x/6.x.
Highlights:
0BSD license (previous versions were LGPL)
96.8% accuracy on 2,179 test files (+2.3pp vs chardet 6.0.0, +7.7pp vs charset-normalizer)
41x faster than chardet 6.0.0 with mypyc ( 28x pure Python), 7.5x faster than charset-normalizer
Language detection for every result (90.5% accuracy across 49 languages)
99 encodings across six eras (MODERN_WEB, LEGACY_ISO, LEGACY_MAC, LEGACY_REGIONAL, DOS, MAINFRAME)
12-stage detection pipeline — BOM, UTF-16/32 patterns, escape sequences, binary detection, markup charset, ASCII, UTF-8 validation, byte validity, CJK gating, structural probing, statistical scoring, post-processing; the markup stage’s PEP 263 declaration sniffing was requested by patrikha in #249
Bigram frequency models trained on CulturaX multilingual corpus data for all supported language/encoding pairs
Optional mypyc compilation — 1.49x additional speedup on CPython
Thread-safe detect() and detect_all() with no measurable overhead; scales on free-threaded Python 3.13t+
Negligible import memory (96 B)
Zero runtime dependencies
Breaking changes vs 6.0.0:
detect() and detect_all() now default to encoding_era=EncodingEra.ALL (6.0.0 defaulted to MODERN_WEB )
Internal architecture is completely different (probers replaced by pipeline stages). Only the public API is preserved.
LanguageFilter is accepted but ignored (deprecation warning emitted)
chunk_size is accepted but ignored (deprecation warning emitted)
Nothing published for this version
Fixed version number in chardet/version.py still being set to 6.0.0dev0. Otherwise identical to 6.0.0.
6.0.0dev0. Otherwise identical to 6.0.0.Fixed version not being set correctly in the package ( Dan Blanchard )
Unified single-byte charset detection: Instead of only having trained language models for a handful of languages (Bulgarian, Greek, Hebrew, Hungarian,
Latin1Prober and MacRomanProber heuristics for Western encodings, chardet now treats all single-byte charsets the same way: every encoding gets proper language-specific bigram models trained on CulturaX corpus data. This means chardet can now accurately detect both the encoding and the language for all supported single-byte encodings.EncodingEra filtering: New encoding_era parameter to detect allows filtering by an EncodingEra flag enum (MODERN_WEB, LEGACY_ISO, LEGACY_MAC, LEGACY_REGIONAL, DOS, MAINFRAME, ALL) allows callers to restrict detection to encodings from a specific era. detect() and detect_all() default to MODERN_WEB. The new MODERN_WEB default should drastically improve accuracy for users who are not working with legacy data. The tiers are:
MODERN_WEB: UTF-8/16/32, Windows-125x, CP874, CJK multi-byte (widely used on the web)LEGACY_ISO: ISO-8859-x, KOI8-R/U (legacy but well-known standards)LEGACY_MAC: Mac-specific encodings (MacRoman, MacCyrillic, etc.)LEGACY_REGIONAL: Uncommon regional/national encodings (KOI8-T, KZ1048, CP1006, etc.)DOS: DOS/OEM code pages (CP437, CP850, CP866, etc.)MAINFRAME: EBCDIC variants (CP037, CP500, etc.)--encoding-era CLI flag: The chardetect CLI now accepts -e/--encoding-era to control which encoding eras are considered during detection.max_bytes and chunk_size parameters: detect(), detect_all(), and UniversalDetector now accept max_bytes (default 200KB) and chunk_size (default 64KB) parameters for controlling how much data is examined. (#314, @bysiber)chardet.metadata.charsets module provides structured metadata about all supported encodings, including their era classification and language filter.should_rename_legacy now defaults intelligently: When set to None (the new default), legacy renaming is automatically enabled when encoding_era is MODERN_WEB.SJISDistributionAnalysis discarding valid second-byte range >= 0x80. (#315, @bysiber)MIN_RATIO threshold alongside the existing EXPECTED_RATIO.get_charset crash: Resolved a crash when looking up unknown charset names.char_len_table: Corrected the character length table for GB18030 multi-byte sequences.detect_all() returning inactive probers: Results from probers that determined "definitely not this encoding" are now excluded.Latin1Prober and MacRomanProber: These special-case probers have been replaced by the unified model-based approach described above. Latin-1, MacRoman, and all other single-byte encodings are now detected by SingleByteCharSetProber with trained language models, giving better accuracy and language identification.LanguageFilter.NONE removed: Use specific language filters or LanguageFilter.ALL instead.InputState, ProbingState, MachineState, SequenceLikelihood, and CharacterCategory are now IntEnum (previously plain classes or Enum). LanguageFilter values changed from hardcoded hex to auto().detect() default behavior change: detect() now defaults to encoding_era=EncodingEra.MODERN_WEB and should_rename_legacy=None (auto-enabled for MODERN_WEB), whereas previously it defaulted to considering all encodings with no legacy renaming.hatch-vcs for version management.create_language_model.py training script was rewritten to use the CulturaX multilingual corpus instead of Wikipedia, producing higher quality bigram frequency models.Language class converted to frozen dataclass: The language metadata class now uses @dataclass(frozen=True) with num_training_docs and num_training_chars fields replacing wiki_start_pages.pytest-timeout and pytest-xdist for faster parallel test execution. Reorganized test data directories.Thank you to everyone who contributed to this release!
And a special thanks to @helour, whose earlier Latin-1 prober work from an abandoned PR helped inform the approach taken in this release.
Features:
Unified single-byte charset detection with proper language-specific bigram models for all single-byte encodings (replaces Latin1Prober and MacRomanProber heuristics) ( Dan Blanchard )
38 new languages: Arabic, Belarusian, Breton, Croatian, Czech, Danish, Dutch, English, Esperanto, Estonian, Farsi, Finnish, French, German, Icelandic, Indonesian, Irish, Italian, Kazakh, Latvian, Lithuanian, Macedonian, Malay, Maltese, Norwegian, Polish, Portuguese, Romanian, Scottish Gaelic, Serbian, Slovak, Slovene, Spanish, Swedish, Tajik, Ukrainian, Vietnamese, Welsh ( Dan Blanchard )
EncodingEra filtering via new encoding_era parameter ( Dan Blanchard )
max_bytes and chunk_size parameters for detect() , detect_all() , and UniversalDetector ; chunked processing was proposed by deedy5 in #284 ( Dan Blanchard )
-e / --encoding-era CLI flag ( Dan Blanchard via Claude)
EBCDIC detection (CP037, CP500) ( Dan Blanchard )
Direct GB18030 support (replaces redundant GB2312 prober) ( Dan Blanchard )
Binary file detection ( Dan Blanchard )
Python 3.12, 3.13, and 3.14 support ( Hugo van Kemenade , #283 )
GitHub Codespaces support ( oxygen dioxide , #312 )
Breaking changes:
Dropped Python 3.7, 3.8, and 3.9 (requires Python 3.10+)
Removed Latin1Prober and MacRomanProber
Removed EUC-TW support
Removed LanguageFilter.NONE
detect() default changed to encoding_era=EncodingEra.MODERN_WEB
Fixes:
Fixed CP949 state machine ( nenw* , #268 )
Fixed SJIS distribution analysis (second-byte range >= 0x80) ( Kadir Can Ozden , #315 )
Fixed max_bytes not being passed to UniversalDetector ( Kadir Can Ozden , #314 )
Fixed UTF-16/32 detection for non-ASCII-heavy text ( Dan Blanchard )
Fixed GB18030 char_len_table ( Dan Blanchard )
Fixed UTF-8 state machine ( Dan Blanchard )
Fixed detect_all() returning inactive probers ( Dan Blanchard )
Fixed early cutoff bug ( Dan Blanchard )
Updated LGPLv2.1 license text for remote-only FSF address ( Ben Beasley , #307 )
Adds support for running chardet CLI via python -m chardet (0e9b7bc20366163efcc221281201baff4100fe19, @dan-blanchard)
Adds support for running chardet CLI via python -m chardet (0e9b7bc20366163efcc221281201baff4100fe19, @dan-blanchard)
Added support for running the CLI via python -m chardet ( Dan Blanchard )
Add should_rename_legacy argument to most functions, which will rename older encodings to their more modern equivalents (e.g., GB2312 becomes GB18030)
should_rename_legacy argument to most functions, which will rename older encodings to their more modern equivalents (e.g., GB2312 becomes GB18030) (#264, @dan-blanchard)--minimal flag to chardetect command (#214, @dan-blanchard)Added should_rename_legacy argument to remap legacy encoding names to modern equivalents ( Dan Blanchard , #264 )
Added MacRoman encoding prober ( Elia Robyn Lake )
Added --minimal flag to chardetect CLI ( Dan Blanchard , #214 )
Added type annotations and mypy CI ( Jon Dufresne , #261 )
Added support for Python 3.11 ( Hugo van Kemenade , #274 )
Added ISO-8859-15 capital letter sharp S handling ( Simon Waldherr , #222 )
Clarified LGPL version in license trove classifier ( Ben Beasley , #255 )
Removed support for Python 3.6 ( Jon Dufresne , #260 )
⚠️ This release is the first release of chardet that no longer supports Python < 3.6 ⚠️
⚠️ This release is the first release of chardet that no longer supports Python < 3.6 ⚠️
In addition to that change, it features the following user-facing changes:
SingleByteCharSetProber confidence to match latest uchardet (#209)detect_all return child prober confidences (#210)Added Johab Korean prober ( grizlupo , #172 , #207 )
Added UTF-16/32 BE/LE probers ( Jason Zavaglia , #109 , #206 )
Added test data for Croatian, Czech, Hungarian, Polish, Slovak, Slovene, Greek, Turkish ( Dan Blanchard )
Improved XML tag filtering ( Dan Blanchard , #208 )
Made detect_all return child prober confidences ( Dan Blanchard , #210 )
Added support for Python 3.10 ( Hugo van Kemenade , #232 )
Slight performance increase ( deedy5 , #252 )
Dropped Python 2.7, 3.4, 3.5 (requires Python 3.6+)
Remove deprecated license_file from setup.cfg (#182) @jdufresne
⚠️ This will be the last release of chardet to support Python 2.7. chardet 5.0 will only support 3.6+ ⚠️
This release is multiple years in the making, and provides some quality of life improvements to chardet. The primary user-facing changes are:
CharsetGroupProber class now properly short-circuits when one of the probers in the group is considered a definite match. This lead to a substantial speedup.chardet.detect_all function that returns a list of possible encodings for the input with associated confidences.The changes in this release have also laid the groundwork for retraining the models to make them more accurate, and to support some more encodings/languages (see #99 for progress). This is our main focus for chardet 5.0 (beyond dropping Python 2 support).
Running on a MacBook Pro (15-inch, 2018) with 2.2GHz 6-core i7 processor and 32GB RAM
Benchmarking chardet 3.0.4 on CPython 3.7.5 (default, Sep 8 2020, 12:19:42)
[Clang 11.0.3 (clang-1103.0.32.62)]
--------------------------------------------------------------------------------
Calls per second for each encoding:
ascii: 25559.439366240098
big5: 7.187002209518091
cp932: 4.71090956645177
cp949: 2.937256786994428
euc-jp: 4.870580412090848
euc-kr: 6.6910755971933416
euc-tw: 87.71098043480079
gb2312: 6.614302607154443
ibm855: 27.595893549680685
ibm866: 29.93483661732791
iso-2022-jp: 3379.5052775763434
iso-2022-kr: 26181.67290886392
iso-8859-1: 120.63424740403983
iso-8859-5: 32.65106262196898
iso-8859-7: 62.480089080556084
koi8-r: 13.72481001727257
maccyrillic: 33.018537255804496
shift_jis: 4.996013583677438
tis-620: 14.323112928341818
utf-16: 166771.53081510935
utf-32: 198782.18009478672
utf-8: 13.966236809766901
utf-8-sig: 193732.28637413395
windows-1251: 23.038910006925768
windows-1252: 99.48409117053738
windows-1255: 6.336261495718825
Total time: 357.05358052253723s (10.054513372323958 calls per second)
Benchmarking chardet 4.0.0 on CPython 3.7.5 (default, Sep 8 2020, 12:19:42)
[Clang 11.0.3 (clang-1103.0.32.62)]
--------------------------------------------------------------------------------
.......................................................................................................................................................................................................................................................................................................................................................................
Calls per second for each encoding:
ascii: 38176.31067961165
big5: 12.86915132656389
cp932: 4.656400877065864
cp949: 7.282976434315926
euc-jp: 4.329381447610525
euc-kr: 8.16386823884839
euc-tw: 90.230745070368
gb2312: 14.248865889128146
ibm855: 33.30225548069821
ibm866: 44.181691968506
iso-2022-jp: 3024.2295767539117
iso-2022-kr: 25055.57945041816
iso-8859-1: 59.25262902122995
iso-8859-5: 39.7069713674529
iso-8859-7: 61.008422013862194
koi8-r: 41.21560517643845
maccyrillic: 31.402474369805002
shift_jis: 4.9091652743515155
tis-620: 14.408875278821073
utf-16: 177349.00634249471
utf-32: 186413.51111111112
utf-8: 108.62174360115105
utf-8-sig: 181965.46637744035
windows-1251: 43.16933400329809
windows-1252: 211.27653358317968
windows-1255: 16.15113643694104
Total time: 268.0230791568756s (13.394368915143872 calls per second)
Thank you to @aaaxx, @edumco, @hrnciar, @hroncok, @jdufresne, @mdamien, @saintamh , @xeor for submitting pull requests, to all of our users for being patient with how long this release has taken.
Added detect_all() function returning all candidate encodings ( Damien , #111 )
Converted single-byte charset probers to nested dicts (performance) ( Dan Blanchard , #121 )
CharsetGroupProber now short-circuits on definite matches (performance) ( Dan Blanchard , #203 )
Added language field to detect_all output ( Dan Blanchard )
Switched from Travis to GitHub Actions ( Dan Blanchard , #204 )
Dropped Python 2.6, 3.4, 3.5
This minor bugfix release just fixes some packaging and documentation issues:
This minor bugfix release just fixes some packaging and documentation issues:
setup.py where pytest_runner was always being installed. (PR #119, thanks @zmedico)test.py is included in the manifest (PR #118, thanks @zmedico)Fixed packaging issue with pytest_runner ( Zac Medico , #119 )
Included test.py in source distribution ( Zac Medico , #118 )
Updated old URLs in README and docs ( Qi Fan , #123 ; Jon Dufresne , #129 )
This release fixes a crash when debugging logging was enabled. (Issue #115, PRs #117 and #125)
This release fixes a crash when debugging logging was enabled. (Issue #115, PRs #117 and #125)
Fixed crash when debug logging was enabled ( Dan Blanchard , #117 )
Fixes an issue where detect would sometimes return None instead of a dict with the keys encoding, language, and confidence (Issue #113, PR #114).
Fixes an issue where detect would sometimes return None instead of a dict with the keys encoding, language, and confidence (Issue #113, PR #114).
Fixed detect sometimes returning None instead of a result dict ( Dan Blanchard , #114 )
This bugfix release fixes a crash in the EUC-TW prober when it encountered certain strings (Issue #67).
This bugfix release fixes a crash in the EUC-TW prober when it encountered certain strings (Issue #67).
Fixed crash in EUC-TW prober with certain strings ( Dan Blanchard )
This release is long overdue, but still mostly serves as a placeholder for the impending 4.0.0 release, which will have retrained models for better ac
This release is long overdue, but still mostly serves as a placeholder for the impending 4.0.0 release, which will have retrained models for better accuracy. For now, this release will get the following improvements up on PyPI:
'rb' for chardetect CLI. (PR #38, thanks @lpsinger)chardetect crash with non-ascii file names (PR #39, thanks @nkanaev)mTypicalPositiveRatio, and instead typical_positive_ratio)filter_without_english_words to filter_international_words and make it match current Mozilla implementation (PR #44, thanks @rsnair2)filter_english_letters to match C implementation (c6654595)hypotheis-based test (PR #66, thanks @DRMacIver)bytes.decode() (PR #73, thanks @snoack)chardetect when encoding is detected instead of looping through entire file (PR #103, thanks @jpz)bytearray objects internally instead of wrap_ord calls, which provides a nice performance boost across the board (PR #106)language property to probers and UniversalDetector results (PR #180)Added Turkish ISO-8859-9 detection ( queeup )
Modernized naming conventions ( typical_positive_ratio instead of mTypicalPositiveRatio ) ( Dan Blanchard , #107 )
Added language property to probers and results ( Dan Blanchard , #108 )
Switched from Travis to GitHub Actions ( Dan Blanchard )
Fixed CharsetGroupProber.state not being set to FOUND_IT ( Dan Blanchard )
Added Hypothesis-based fuzz testing ( David R. MacIver , #66 )
Don’t indicate byte order for UTF-16/32 with given BOM, for compatibility with decode() ( Sebastian Noack , #73 )
Stop reading file immediately when file type is known ( Jason Zavaglia , #103 )
Added support for CP932 detection (thanks to @hashy).
In this release, we:
chardetect to use argparse for argument parsing.gh-pages branch. You can now access them at http://chardet.github.io.Added CP932 detection ( hashy )
Fixed UTF-8 BOM not detected as UTF-8-SIG ( atbest , #32 )
Switched chardetect to use argparse ( Dan Blanchard )
Fix missing paren in chardetect.py
Fix missing paren in chardetect.py
Fixed missing parenthesis in chardetect.py ( Owen , #12 )
chardet 2.1.1 (2012-10-01)
chardet 2.1.1 (2012-10-01)
Bumped version past Mark Pilgrim’s last release
chardetect can now read from stdin ( Erik Rose )
Fixed BOM byte strings for UCS-4-2143 and UCS-4-3412 ( Toshio Kuratomi )
Restored Mark Pilgrim’s original docs and COPYING file ( Toshio Kuratomi )
- Added chardetect CLI tool ( Erik Rose )
Added chardetect CLI tool ( Erik Rose )
Fixed utf8prober crash when character is out of range ( David Cramer )
Cleaned up detection logic to fail gracefully ( David Cramer )
Fixed feed encoding errors ( David Cramer )
- Version fix ( Ian Cordasco )
Version fix ( Ian Cordasco )
- Initial release: Python 2 port of Mozilla’s universal charset detector ( Mark Pilgrim )
Initial release: Python 2 port of Mozilla’s universal charset detector ( Mark Pilgrim )
On this page
Changelog
Unreleased
7.6.0 (2026-08-14)
7.5.1 (2026-08-06)
7.5.0 (2026-08-05)
7.4.3 (2026-04-13)
7.4.2 (2026-04-12)
7.4.1 (2026-04-07)
7.4.0 (2026-03-26)
7.3.0 (2026-03-24)
7.2.0 (2026-03-17)
7.1.0 (2026-03-11)
7.0.1 (2026-03-04)
7.0.0 (2026-03-02)
6.0.0.post1 (2026-02-22)
6.0.0 (2026-02-22)
5.2.0 (2023-08-01)
5.1.0 (2022-12-01)
5.0.0 (2022-06-25)
4.0.0 (2020-12-10)
3.0.4 (2017-06-08)
3.0.3 (2017-05-16)
3.0.2 (2017-04-12)
3.0.1 (2017-04-11)
3.0.0 (2017-04-11)
chardet 2.3.0 (2014-10-07)
chardet 2.2.1 (2013-12-18)
chardet 2.2.0 (2013-12-16)
charade 1.0.3 (2013-01-18)
charade 1.0.2 (2013-01-18)
charade 1.0.1 (2012-12-03)
charade 1.0.0 (2012-12-02)
chardet 2.1.1 (2012-10-01)
chardet 1.1 (2012-07-27)
chardet 1.0.1 (2008-04-19)
chardet 1.0 (2006-12-23)
Your coding agent can read these notes before it upgrades. Set up the MCP server →