NewYour coding agent can read the release notes before it upgrades.Set up the MCP server →
pub.dev
Pure Dart WordPiece tokenizer for BERT NLP models. Completely independent of Flutter.
Last release 1 months ago
12 Aug 2026
Ships on a steady schedule
a new release about every 2 weeks
Nearly every release is documented
notes for 7 of 7 stable releases
Nothing withdrawn
no release was ever pulled
4 months old
7 releases · first in 2026
One column per month.
Added batch processing (prepareBatch) and support for pre‑tokenized input (prepareFromTokens).
prepareBatch) and support for pre‑tokenized input (prepareFromTokens).TruncationStrategy) and custom segment ID override (per‑token or global).padTokenId, unkTokenId, etc.) for easy access.\r for cross‑platform compatibility).Upgraded the Example Code: Replaced the basic example script with a comprehensive, multi-step version that demonstrates raw string tokenization, ML in
Resolved regex syntax errors by utilizing raw triple quotes for the punctuation character class.
Resolved regex syntax errors by utilizing raw triple quotes for the punctuation character class.
Enforced strict WordPiece standard (entire word becomes [UNK] if any subword fails).
Added toLowerCase parameter for cased/uncased model flexibility.
Added pre-tokenization stripping of invisible control characters.
Robust line ending handling (\r\n and \n) for cross-platform vocab loading.
Fixed Truncation Crash: Rewrote the sequence length handling in prepareNerInput to prevent unhandled StateError exceptions when maxLength is set to sm
Fixed Truncation Crash: Rewrote the sequence length handling in prepareNerInput to prevent unhandled StateError exceptions when maxLength is set to small values (like 0 or 1).
Preserved Special Tokens: Adjusted the truncation logic to trim the raw word pieces before injecting the [CLS] and [SEP] tokens, ensuring the structural markers required by BERT models are never overwritten.
Enforced Minimum Bounds: Added an ArgumentError validation check to prepareNerInput ensuring maxLength is at least 2, which is the mathematical minimum required to contain the mandatory start and end tokens.
Added Initialization Validation: Updated fromStringContent to verify that both [UNK] and [PAD] are explicitly defined within the parsed vocabulary map, replacing late-stage runtime crashes (!) with a descriptive FormatException during setup.
* Update pub.dev visibility.
Fix documentation links and update repository visibility.
Initial release of pure Dart BERT tokenizer.
Your coding agent can read these notes before it upgrades. Set up the MCP server →