NewYour coding agent can read the release notes before it upgrades.Set up the MCP server →
A Japanese tokenizer based on recurrent neural networks
Last release 3 months ago
06 Jul 2026
Release timing varies
gaps range from 2 weeks to 2.2 years
Some releases are documented
notes for 14 of 26 stable releases
Nothing withdrawn
no release was ever pulled
9 years old
27 releases · first in 2018
nagisa 0.3.0 incorporates the following changes:
nagisa 0.3.0 incorporates the following changes:
DyNet is now only required for training your own models with nagisa.fit. The wakati/tagging/postagging inference path runs entirely on NumPy and a small Cython extension (nagisa_utils), loading the trained weights directly from DyNet's text-format model dump. Tokenization output is unchanged; only the execution engine underneath was replaced.
This makes inference roughly 3x faster on a single thread and uses about 25% less memory than v0.2.12:
<p align="center"> <img width="480" alt="Single-thread inference is about 3x faster in v0.3.0" src="https://github.com/user-attachments/assets/9894c60a-9012-4298-b3ab-8fdfc5464e61" /> </p>
The new kernels release the GIL and hold no shared mutable state, so nagisa now scales correctly with threads, including under Python 3.14's free-threaded (no-GIL) build:
<p align="center"> <img width="480" alt="v0.3.0 scales with threads, while v0.2.12 got slower" src="https://github.com/user-attachments/assets/3e62fece-9085-48ed-89e7-cc3e2186902b" /> </p>
setup.py to pyproject.toml https://github.com/taishi-i/nagisa/commit/b4ff68a0301d310610e1a6edeb11fa7dae1ef6a1tagger.py by using raw strings https://github.com/taishi-i/nagisa/commit/e33d62ef8f5b1a6073eef31aca57a9bc3969ed89Known issues
ValueError: Not a DyNet text-format model file while loading nagisa_v001.model. The bundled model file has always had stray \r\n bytes, which the old DyNet loader tolerated but the new byte-exact dynet_loader.py parser doesn't. nagisa/data/* isn't marked binary in .gitattributes, so Windows' core.autocrlf can mangle it further on checkout. Doesn't affect pip install nagisa from PyPI. Fix: mark data files as binary in .gitattributes.One column per quarter.
nagisa 0.2.12 incorporates the following changes:
nagisa 0.2.12 incorporates the following changes:
This happened when the input text ended with a partially formed word (i.e., a sequence starting with a BEGIN tag but missing a corresponding END tag). In such cases, the segmenter_for_bmes function discarded the incomplete word instead of including it in the output.
Before (v0.2.11):
import nagisa
print(nagisa.tagging("いい写真やちゃ").words)
# ['いい', '写真', 'や']
After (v0.2.12):
import nagisa
print(nagisa.tagging("いい写真やちゃ").words)
# ['いい', '写真', 'や', 'ちゃ']
Nagisa (v0.2.12+) includes a built-in list of Japanese stopwords, accessible via nagisa.stopwords. This makes it easy to filter out common function words from your tokenized results without needing an external stopword list.
import nagisa
print(nagisa.stopwords)
# ['~', '\u3000', 'あ', 'あっ', 'あり', 'ある', 'あれ', 'い', 'いう', 'いずれ', 'いつ', 'いる', 'う', 'うち', 'え', 'お', 'および', 'おり', 'か', 'かつて', 'から', 'が', 'き', 'くらい', 'けど', 'こ', 'ここ', 'こちら', 'こと', 'この', 'これ', 'ご', 'ごと', 'さ', 'さらに', 'し', 'しか', 'しかし', 'じゃ', 'す', 'する', 'ず', 'せ', 'せる', 'そう', 'そこ', 'そして', 'その', 'それ', 'それぞれ', 'た', 'たい', 'たく', 'ただ', 'ただし', 'たち', 'ため', 'たら', 'たり', 'だ', 'だけ', 'だっ', 'だろう', 'ち', 'って', 'つ', 'て', 'てる', 'で', 'でき', 'できる', 'でし', 'でしょう', 'です', 'と', 'とき', 'ところ', 'とも', 'どこ', 'どちら', 'どれ', 'な', 'ない', 'なお', 'なかっ', 'ながら', 'なく', 'なけれ', 'なっ', 'など', 'なに', 'なら', 'なり', 'なる', 'なん', 'なんて', 'に', 'にて', 'ぬ', 'ね', 'の', 'のみ', 'は', 'ば', 'ぶり', 'へ', 'べき', 'ほか', 'ほとんど', 'ほど', 'ほぼ', 'ま', 'まし', 'ましょう', 'ます', 'ませ', 'また', 'まで', 'まま', 'み', 'も', 'もの', 'や', 'よ', 'よう', 'より', 'ら', 'られ', 'られる', 'る', 'れ', 'れる', 'を', 'ん', '及び']
import nagisa
text = "日本語のストップワードを簡単に利用できます。"
tokens = nagisa.tagging(text)
print(tokens.words)
# ['日本', '語', 'の', 'ストップ', 'ワード', 'を', '簡単', 'に', '利用', 'でき', 'ます', '。']
# Filter out stopwords
words = [word for word in tokens.words if word not in nagisa.stopwords]
print(words)
# ['日本', '語', 'ストップ', 'ワード', '簡単', '利用', '。']
nagisa 0.2.11 incorporates the following changes:
nagisa 0.2.11 incorporates the following changes:
This issue was caused by the lack of correct library dependencies in the tar.gz files registered on PyPI. Specifically, in versions prior to 0.2.10, the following information was not included in the PKG-INFO file of the tar.gz.
Requires-Dist: six
Requires-Dist: numpy
Requires-Dist: DyNet38
nagisa-0.2.10.tar.gz does not include the
Requires-Distin PKG-INFO.
The issue was caused by an outdated build environment when creating tar.gz. Therefore, by updating pip, wheel, and build to their latest versions and creating tar.gz, we were able to include the correct dependency information in the tar.gz, resolving this problem.
nagisa-0.2.11.tar.gz includes the
Requires-Distin PKG-INFO.
In Poetry, dependency library information is fetched from https://pypi.org/pypi/nagisa/json, which refers to tar.gz. Therefore, errors occurred in versions prior to 0.2.10.
import requests
version = "0.2.10"
url = f"https://pypi.org/pypi/nagisa/{version}/json"
response = requests.get(url)
data = response.json()
dependencies = data.get("info", {}).get("requires_dist", [])
print(f"Version {version}: {dependencies}")
# Version 0.2.10: ['six', 'numpy', 'DyNet']
With this update, it is now possible to obtain correct information about dependency libraries.
import requests
version = "0.2.11"
url = f"https://pypi.org/pypi/nagisa/{version}/json"
response = requests.get(url)
data = response.json()
dependencies = data.get("info", {}).get("requires_dist", [])
print(f"Version {version}: {dependencies}")
# Version 0.2.11: ['six', 'numpy', 'DyNet38']
Add Python wheels (3.6, 3.7, 3.8, 3.9, 3.10, 3.11, 3.12) to PyPI for Linux
Add Python wheels (3.6, 3.7, 3.8, 3.9, 3.10, 3.11) to PyPI for macOS Intel
Add Python wheels (3.6, 3.7, 3.8) to PyPI for Windows
Fix the macOS M1/2 installation error https://github.com/taishi-i/nagisa/issues/35 https://github.com/taishi-i/nagisa/issues/30 (Updated on Jun 15, 2024)
Add Python wheels (3.9, 3.10, 3.11, 3.12) to PyPI for macOS M1/M2 (Updated on Jun 15, 2024)
Fix the aarch64 installation error https://github.com/taishi-i/nagisa/issues/33 (Updated on Jun 16, 2024)
Add Python wheels (3.9, 3.10, 3.11, 3.12) to PyPI for aarch64 (Updated on Jun 16, 2024)
Add Python wheels (3.13) to PyPI for Linux (Updated on May 4, 2025)
Add Python wheels (3.13) to PyPI for macOS M1/M2 (Updated on May 4, 2025)
Add Python wheels (3.13) to PyPI for aarch64 (Updated on May 4, 2025)
Fix the Windows installation error that @ROBERT-MCDOWELL reported #37 (Updated on Dec 30, 2025)
Add Python wheels for Windows (3.9, 3.10, 3.11, 3.12, 3.13, 3.14) to PyPI for win_arm64 (Updated on Dec 30, 2025)
Add Python wheels (3.14) to PyPI for Linux, macOS, aarch64 (Updated on Dec 30, 2025)
Nothing published for this version
nagisa 0.2.10 incorporates the following changes:
nagisa 0.2.10 incorporates the following changes:
Fix hard-coding process for noun id conversion in tagger.py https://github.com/taishi-i/nagisa/commit/e638343e21559a072471c03ff672a67cf8196b77
Hidden the dynet log that appears when import nagisa is used in Python 3.8, 3.9, 3.10, 3.11, and 3,12 on Linux
Provide the nagisa-demo page in Hugging Face Spaces
Provide the stopwords for nagisa in Hugging Face Datasets
Update read the docs documents
Compatible with Python 3.12 on Linux
Add Python wheels (3.6, 3.7, 3.8, 3.9, 3.10, 3.11, 3,12) to PyPI for Linux
Add Python wheels (3.6, 3.7, 3.8, 3.9, 3.10, 3.11) to PyPI for macOS Intel
Add Python wheels (3.6, 3.7, 3.8) to PyPI for Windows
nagisa 0.2.9 incorporates the following changes:
nagisa 0.2.9 incorporates the following changes:
Until now, there was an issue where the processing time would slow down as the results analyzed by the following code increased in tagger.py.
tids = []
for w in words:
if w in self._word2postags:
w2p = self._word2postags[w]
else:
w2p = [0]
if self.use_noun_heuristic is True:
if w.isalnum() is True:
if w2p == [0]:
w2p = [self._pos2id[u'名詞']]
else:
# bottleneck is here!
w2p.append(self._pos2id[u'名詞'])
w2p = list(set(w2p))
tids.append(w2p)
By changing to the following code, we have resolved the issue of the processing slowing down.
tids = []
for w in words:
w2p = set(self._word2postags.get(w, [0]))
if self.use_noun_heuristic and w.isalnum():
if 0 in w2p:
w2p.remove(0)
w2p.add(2) # nagisa.tagger._pos2id["名詞"] = 2
tids.append(list(w2p))
[metadata]
description_file = README.md
nagisa 0.2.8 incorporates the following changes:
nagisa 0.2.8 incorporates the following changes:
AttributeError in nagisa_utils.pyx when tokenizing a text containing Latin capital letter I with dot above 'İ'When tokenizing a text containing 'İ', an AttributeError has occurred. This is because, as the following example shows, lowering 'İ' would have changed to the length of 2, and would not have been extracting features correctly.
>>> text = "İ" # [U+0130]
>>> print(len(text))
1
>>> text = text.lower() # [U+0069] [U+0307]
>>> print(text)
'i̇'
>>> print(len(text))
2
To avoid this error, the following preprocess was added to the source code modification 1, modification 2.
text = text.replace('İ', 'I')
nagisa 0.2.7 incorporates the following changes:
nagisa 0.2.7 incorporates the following changes:
AttributeError: module 'utils' to rename utils.pyx into nagisa_utils.pyx #14nagisa 0.2.6 incorporates the following changes:
nagisa 0.2.6 incorporates the following changes:
readFile(filename) in mecab_system_eval.py for windows usersnagisa-0.2.6-cp36-cp36m-win_amd64.whl and nagisa-0.2.6-cp37-cp37m-win_amd64.whl to PyPI to install nagisa without Build Tools for Windows users #23nagisa-0.2.6-*-manylinux1_i686.whl and nagisa-0.2.6-*-manylinux1_x86_64.whl to PyPI to install nagisa for Linux usersnagisa 0.2.5 incorporates the following changes:
nagisa 0.2.5 incorporates the following changes:
__version__ to __init__.pynagisa 0.2.4 incorporates the following changes:
nagisa 0.2.4 incorporates the following changes:
nagisa 0.2.3 incorporates the following changes:
nagisa 0.2.3 incorporates the following changes:
nagisa.tagging reduces wasteful memory and improves the speed in word segmentation.nagisa 0.2.2 incorporates the following changes:
nagisa 0.2.2 incorporates the following changes:
Nothing published for this version
nagisa 0.2.0 incorporates the following changes:
nagisa 0.2.0 incorporates the following changes:
nagisa 0.1.2 incorporates the following changes:
nagisa 0.1.2 incorporates the following changes:
nagisa.Tagger(single_word_list)Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Your coding agent can read these notes before it upgrades. Set up the MCP server →