NewYour coding agent can read the release notes before it upgrades.Set up the MCP server →
Lexical text utilities (tokenizer, stemmer, stopwords) for Dart and Flutter.
Last release 1 months ago
24 Aug 2026
Ships fairly regularly
a new release about every 5 weeks
Nearly every release is documented
notes for 1 of 1 stable releases
Nothing withdrawn
no release was ever pulled
3 months old
3 releases · first in 2026
One column per month.
Record what the first stable release actually contains: the API is unchanged from 0.1.0-dev.2, so point readers at the feature set and note the two co
Record what the first stable release actually contains: the API is
unchanged from 0.1.0-dev.2, so point readers at the feature set and note
the two constraint bumps (Dart 3.13, betto_icu ^0.1.0).
Co-Authored-By: Claude Opus 5 noreply@anthropic.com
First stable release. The public API is unchanged from 0.1.0-dev.2 — see the
entries below for the full feature set: tokenization (createDefaultTokenizer,
IcuTokenizer, BrowserTokenizer, RegExpTokenizer, OffsetTokenizer),
stemming in 28 languages, and stop-word sets for 58 languages.
sdk: ^3.13.0, up from ^3.12.0).betto_icu constraint moved to the stable ^0.1.0 (from ^0.1.0-dev.2).Stemmer previously wired up only English, throwing ArgumentError for every other locale, even though the underlying snowball_stemmer dependency alread
Stemmer previously wired up only English, throwing ArgumentError for
every other locale, even though the underlying snowball_stemmer
dependency already implements 28 languages. Extend the factory to map
all of them (excluding the generic porter variant, not a distinct
language). Each new mapping was verified against the actual stemmer
output for a representative inflected word per language, not just
type-checked.
Also bump the betto_icu dependency to 0.1.0-dev.2 and re-export its
new OffsetTokenizer/TokenSpan types, needed by kmdb's vault chunker
for position-aware tokenisation.
Co-Authored-By: Claude Sonnet 5 noreply@anthropic.com
Stemmer now supports 28 languages, up from just English: Arabic, Armenian,
Basque, Catalan, Danish, Dutch, English, Finnish, French, German, Greek,
Hindi, Hungarian, Indonesian, Irish, Italian, Lithuanian, Nepali, Norwegian,
Portuguese, Romanian, Russian, Serbian, Spanish, Swedish, Tamil, Turkish, and
Yiddish — every language package:snowball_stemmer implements (excluding its
generic porter variant, an alternate English algorithm rather than a
distinct language). Unsupported language codes still throw ArgumentError, as
before.OffsetTokenizer and TokenSpan from betto_icu (now pinned to
^0.1.0-dev.2) — a Tokenizer that also reports each token's character
offsets in the source text. Implemented by IcuTokenizer and
RegExpTokenizer.Co-Authored-By: Claude Sonnet 4.6 noreply@anthropic.com
Co-Authored-By: Claude Sonnet 4.6 noreply@anthropic.com
Initial development release.
createDefaultTokenizer() returns the best tokenizer for the current platform
at compile time — no runtime Platform checks.IcuTokenizer — UAX #29 word segmentation via the system ICU library (native:
macOS, Linux, Windows, Android, iOS).BrowserTokenizer — word segmentation via Intl.Segmenter (web).RegExpTokenizer — lightweight Latin-script tokenizer in pure Dart, available
on all platforms.Stemmer — Snowball-based stemmer. Construct with a Locale and call
stem(word) to reduce tokens to their base form. English (en) supported.getStopWords(Locale) — returns a Stopwords enum value containing the
stop-word set for the given locale. Throws ArgumentError for unsupported
language codes.Your coding agent can read these notes before it upgrades. Set up the MCP server →