NewYour coding agent can read the release notes before it upgrades.Set up the MCP server →
PyPI · #1780 most downloaded on PyPI
Chinese Words Segmentation Utilities
Last release 7 years ago
no release in 18 months
Ships unpredictably
gaps range from 2 weeks to 2.3 years
Most releases are documented
notes for 20 of 32 stable releases
Nothing withdrawn
no release was ever pulled
14 years old
32 releases · first in 2012
修复setup.py在python2.7版本无法工作的问题
修复setup.py在python2.7版本无法工作的问题
修复paddle模式空字符串coredump问题 @JesseyXujin
One column per quarter.
1. 开启paddle模式更友好 2. 修复cut_all模式不支持中英混合词的bug
支持基于paddle的深度学习分词模式(use_paddle=True); by @JesseyXujin , @xyzhou-puck
del_word支持强行拆开词语; by @gumblex , @fxsjy
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
新增ChineseAnalyzer,用于支持whoosh搜索引擎 3)添加了更多的中英混合词汇 4)修改了一些py文件的加载方法,从而支持py2exe,cxfree打包为exe
1) 优化了viterbi算法的代码,分词速度提升15% 2) 去除了词典中的一些低质词
1) 提升了finalseg子模块命名体识别的准确度 2) 修正了一些badcase
add wraps decorator, by @cloudaice
1) 修正了临时cache文件生成在pypy解析器下出错的问题
1) 修正了initialize函数默认参数绑定的bug.
Nothing published for this version
新增词典lazy load功能,用户可以在'import jieba'后再改变词典的路径. 感谢hermanschaaf
修正了“的”字频过高引起的bug;修正了对小数点和下划线的处理
Nothing published for this version
改进了关键词提取的功能jieba.analyse.extract_tags;
1)支持繁体中文的分词 2)修正了多python进程时生成cache文件失败的bug
1)支持繁体中文的分词 2)修正了多python进程时生成cache文件失败的bug
解决了没有标点的长句子分词效果差的问题,问题在于连续的小概率乘法可能会导致浮点下溢或为0.
1) 修复了之前版本不能识别中英混合词语的问题
新增jieba.cut_for_search方法, 该方法在精确分词的基础上对“长词”进行再次切分,适用于搜索引擎领域的分词,比精确分词模式有更高的召回率。
用户自定义词典函数load_userdict支持file-like object作为输入
1) 新增词性标注功能
Your coding agent can read these notes before it upgrades. Set up the MCP server →