NewYour coding agent can read the release notes before it upgrades.Set up the MCP server →
PyPI · #3383 most downloaded on PyPI
Python version of Sudachi, the Japanese Morphological Analyzer
Last release 10 days ago
24 Sep 2026
Release timing varies
gaps range from 2 weeks to 1.3 years
Some releases are documented
notes for 13 of 42 stable releases
Nothing withdrawn
no release was ever pulled
7 years old
43 releases · first in 2019
This release introduced many API update including some breaking changes. See migration guide and CHANGELOG .
The v0.7.* is intended as an intermediate release series before the v1.
Please pin the exact version when using this series, as breaking behavioral changes may be introduced even in patch releases.
This release introduced new dictionary format V1.
Because of this change, you need to update both of the system dictionary and your user dictionaries.
See migration guide for details.
This release introduced many API update including some breaking changes.
See migration guide and CHANGELOG.
Full Changelog: v0.6.11...v0.7.0
See migration guide for details.
JapaneseDictionary::entries,
entries_subset, lookup_all_entries, and lookup_all_entries_subset. (#336)Dictionary.oov_morpheme to create OOV morpheme.(#344)MorphemeList::lookup to normalize queries with dictionary input-text
plugins before indexed lookup, matching Java Sudachi behavior. (#336)synonym_group_id to synonym_group_ids. (#357)fetch_dictionary.sh targets V1 dictionary by default. (#362)One column per quarter.
Support CPython 3.14 and 3.14t (free-threaded)
Support py3.13t (wheels are provided only for linux-amd64 and macos-univarsal2)
Remove Python 3.7 and 3.8 support as it reaches its end of life ( https://devguide.python.org/versions/ ) ( #249 , #281 )
fetch_dictionary.sh targets latest dictionary by default (#240)SplitMode (#245)sudachipy.Config and sudachipy.errors.SudachiError to default import (#260)-s (system dictionary path) of sudachi ubuild command is now required (#239)-d option of sudachi cli (which is no-op) now warns (#278)sudachi dump subcommand (#277)see rust changelog and python changelog for more.
fetch_dictionary.sh targets latest dictionary by default (#240)For chiTra compatibility SudachiPy can now directly produce different tokens in the surface field.
Morheme.raw_surface() methodProvide binary wheels for Python 3.11
Dictionary.lookup() method which allows you to enumerate morphemes from the dictionary without performing analysis.Add boundary matching mode to regex oov handler
Fixed invalid POS tags which appeared when using user-defined POS tags both in user dictionaries and OOV handlers. You are not affected by this bug if
Remove Python 3.6 support which reached end-of-life status on 2021-12-23
maxLength setting defines maximum length in unicode codepoints, not in utf-8 bytes as in Java (will be changed to codepoints later)Fixed path resolution algorithm for resources. They are now resolved in the following order (first existing file wins):
Dictionary constructorDictionary now has __repr__() function which displays absolute paths to dictionaries in use.Dictionary now has pos_of() function which returns a POS tuple for a given POS id.PosMatcher supports set operations
m1 | m2)m1 & m2)m1 - m2)~m1)Dictionary constructorDictionary now has __repr__() function which displays absolute paths to dictionaries in use.Dictionary now has pos_of() function which returns a POS tuple for a given POS id.PosMatcher supports set operations
m1 | m2)m1 & m2)m1 - m2)~m1)Fixed analysis differences from 0.5.4
Added Fuzzing (see sudachi-fuzz subdirectory), Sudachi.rs seems to be pretty robust towards arbitrary inputs (no crashes and panics)
sudachi-fuzz subdirectory), Sudachi.rs seems to be pretty robust towards arbitrary inputs (no crashes and panics)
Morpheme.part_of_speech method now returns Tuple of POS components instead of a list.Dictionary.create(), Dictionary.pre_tokenizer()Dictionary.pre_tokenizer()out parameters which accept MorphemeListsTokenizer.tokenize(), Morpheme.split()len(Morpheme) now returns the length of the morpheme in Unicode codepoints. Use it instead of len(m.surface())Morpheme.split() has new add_single parameter, which can be used to check whether the split has produced anything
if m.split(SplitMode.A, out=res, add_single=False): handle_splits(res)add_single=True, returning the list with the current morpheme is the current behaviorMorpheme/MorphemeList now have readable __repr__ and __str__
Full feature parity with Java version
--split-sentences=nosudachipy build and sudachipy ubuild should work once more
List of deprecated SudachiPy API:
MorphemeList.empty(dict: Dictionary)
Morpheme.split(mode: SplitMode)Morpheme.get_word_info()Dictionary.grammar, Dictionary.lexicon.
sudachipy build and sudachipy ubuild will not work, please use 0.5.3 in another virtual environment for the time being until the feature is implemented: #13Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Your coding agent can read these notes before it upgrades. Set up the MCP server →