NewYour coding agent can read the release notes before it upgrades.Set up the MCP server →
pub.dev · #1857 most downloaded on pub.dev
A lightweight pure Dart SentencePiece tokenizer supporting BPE (Gemma), Unigram (Llama), and Hugging Face tokenizer.json pipelines.
Last release 1 months ago
06 Sep 2026
Release timing varies
gaps range from 1 weeks to 5 months
Nearly every release is documented
notes for 11 of 11 stable releases
Nothing withdrawn
no release was ever pulled
9 months old
11 releases · first in 2026
One column per month.
release: prepare v1.4.1 by @DenisovAV , @brody-0125 in #40
Full Changelog: v1.4.0...v1.4.1
Split pre-tokenizers preserve the 1.3.3 token IDs,
tokens, offsets, masks, sequence IDs, and decoded text, while now assigning
Hugging Face-compatible word ID 0 to their single pre-tokenized segment.tokenizer.json files that declare a no-op Split pre-tokenizer (#35).
U+0020) to SentencePiece whitespace (U+2581) before pre-tokenization.Split patterns, behaviors, and inverted matches remain rejected rather than being silently mis-tokenized.Split form and unsupported Split variants.chore: add Pub Monthly Downloads badge by @brody-0125 in #30
Full Changelog: v1.3.3...v1.4.0
tokenizer.json compatibility for SentencePiece-based BPE and Unigram tokenizers (#31).
Precompiled normalizers with ordered Replace operations.WhitespaceSplit and Metaspace pre-tokenizer pipelines, including current and legacy configuration fields.TemplateProcessing special tokens with declared BOS/EOS IDs and type IDs.fuse_unk) and declared added-token IDs.SpAddedToken for tokenizer JSON metadata interoperability.Encoding.int32 decoding without ZigZag conversion.feat: add Unicode normalization via precompiled charsmap (DARTS trie) by @brody-0125 in #23
Full Changelog: v1.3.2...v1.3.3
merges entries represented as [left, right] pairs (#26).This is a maintenance release with no API changes or breaking changes. Focus: CI infrastructure, code hygiene, and documentation clarity.
GitHub Actions CI Pipeline (#19)
Automated continuous integration is now configured for the project:
dart format consistency and dart analyze --fatal-infos with zero tolerance for warnings.contents: read) and concurrency groups to cancel stale runs.dart format to 23 source files for consistent code style across the codebase.dart analyze --fatal-infos issues:
avoid_returning_null_for_future lint rule.if statements, const constructors, and final local variables where required.google/sentencepiece proto spec compliance for default token IDs (unkId=0, bosId=1, eosId=2, padId=-1).This is a maintenance release with no API changes or breaking changes. Focus: CI infrastructure, code hygiene, and documentation clarity.
Full Changelog: v1.3.1...v1.3.2
analyze job - dart format --set-exit-if-changed and dart analyze --fatal-infostest job - Matrix testing across Dart stable and 3.10.7 (minimum supported version)contents: read) and concurrency group to cancel stale runsdart format to 23 files for consistent code styledart analyze --fatal-infos issues
avoid_returning_null_for_future lint ruleconst constructors, final local variablesgoogle/sentencepiece proto spec compliance for default token IDs (unkId=0, bosId=1, eosId=2, padId=-1) (#20)Load HuggingFace tokenizers directly from tokenizer.json — no conversion step required.
Load HuggingFace tokenizers directly from tokenizer.json — no conversion step required.
tokenizer.json Format SupportYou can now load any HuggingFace tokenizer.json file without converting it to SentencePiece .model format first. This makes it straightforward to use tokenizers published on the HuggingFace Hub.
// Load from file
final tokenizer = await HuggingFaceTokenizerLoader.fromJsonFile('tokenizer.json');
// Load from a pre-parsed map
final tokenizer = HuggingFaceTokenizerLoader.fromMap(jsonMap);
// Auto-detection — works transparently with TokenizerJsonLoader
final tokenizer = await TokenizerJsonLoader.fromJsonFile('tokenizer.json');Supported model types:
Automatic configuration inference:
unk, bos, eos, pad) are detected from the added_tokens sectionaddDummyPrefix, escapeWhitespaces) are inferred from the HuggingFace normalizer configaddBosToken, addEosToken) are parsed from TemplateProcessingFormat detection:
TokenizerJsonLoader.isHuggingFaceFormat() lets you check whether a JSON map uses the HuggingFace format. When you call TokenizerJsonLoader.fromJsonFile(), HuggingFace format is detected and delegated automatically — no code changes needed if you already use TokenizerJsonLoader.
dependencies:
dart_sentencepiece_tokenizer: ^1.3.1Full Changelog: https://github.com/brody-0125/dart_sentencepiece_tokenizer/blob/develop/CHANGELOG.md
tokenizer.json Format Support
HuggingFaceTokenizerLoader class for loading HuggingFace tokenizer.json files directly
fromJsonString() / fromMap() - Parse from JSON string or pre-parsed mapfromJsonFile() / fromJsonFileSync() - Load from file (async/sync)added_tokens sectionTokenizerJsonLoader.isHuggingFaceFormat() - Helper to detect HuggingFace formatTokenizerJsonLoader - Automatically delegates to HuggingFaceTokenizerLoader when HuggingFace format is detectedfeat: add HuggingFace TextStreamer compatible streaming API by @brody-0125 in #8
Full Changelog: 1.2.2...1.3.0
BaseStreamer - Abstract interface for streaming token decoders with put() and end() methodsTextStreamer - HuggingFace TextStreamer-compatible class for real-time LLM token decoding
put(int tokenId) - Add tokens as they are generatedend() - Signal end of generation and flush remaining contentonFinalizedText callback for custom text handlingskipSpecialTokens option to filter BOS/EOS/PAD tokensskipPrompt option to skip initial prompt tokenspromptLength option to skip multiple prompt tokensSentencePieceTokenizer.createTextStreamer() - Factory for TextStreamerSentencePieceTokenizer.decodeStream() - Stream-based token decodingSentencePieceTokenizer.decodeWithCallback() - Callback-based token decodingTextStreamer (HuggingFace-compatible):
final streamer = tokenizer.createTextStreamer();
for (final id in llmOutput) {
streamer.put(id);
}
streamer.end();
// With custom callback
final streamer = tokenizer.createTextStreamer(
onFinalizedText: (text, {required streamEnd}) {
myTextController.append(text);
if (streamEnd) myTextController.complete();
},
);
Stream-based decoding:
final textStream = tokenizer.decodeStream(llmTokenStream);
await for (final chunk in textStream) {
stdout.write(chunk);
}
Callback-based decoding:
tokenizer.decodeWithCallback(
tokenIds,
(chunk) => stdout.write(chunk),
);
feat: optimize memory usage and refactor tests for v1.2.2 by @brody-0125 in #7
Full Changelog: 1.2.1...1.2.2
Trie into shared _decodeCodePoint helpersequenceIds in Encoding to avoid O(n) recomputation on repeated accessBpeAlgorithm and BpeAlgorithmOptimized to prevent unbounded memory growthfillRange for padding initialization in Encoding.withPadding()feat: add JSON Serialization API, Dynamic Token Addition API and Optimized BPE Algorithm by @brody-0125 in #5
Full Changelog: 1.2.0...1.2.1
addTokens() to use single typed array allocation instead of per-token expansion (O(N) instead of O(N²))TokenizerJsonLoader)_kMaxInputLength constant declarationsJSON Serialization - HuggingFace-compatible tokenizer.json format
JSON Serialization - HuggingFace-compatible tokenizer.json format
toJson() - Serialize tokenizer to JSON stringsaveToJson() / saveToJsonSync() - Save to fileTokenizerJsonLoader.fromJsonString() - Load from JSON stringTokenizerJsonLoader.fromJsonFile() / fromJsonFileSync() - Load from fileDynamic Token Addition API
addTokens(List<String>) - Add new tokens to vocabularyaddSpecialTokens(Map<String, String>) - Add special tokens (pad, mask, etc.)getAddedVocab() - Get map of dynamically added tokensisAddedToken(String) - Check if token was added dynamicallygetVocab({withAddedTokens}) - Get full vocabulary as Map<String, int>HuggingFace-compatible Methods
tokenize(String) - Returns List<String> of tokenstokenizeBatch(List<String>) - Batch tokenizationOptimized BPE Algorithm (BpeAlgorithmOptimized)
SpVocabulary now uses growable list for dynamic token addition supportfeat: improve BPE Algorithm by @brody-0125 in #1
Full Changelog: 1.0.0...1.2.0
example/example.dart)Full Changelog : https://github.com/brody-0125/dart_sentencepiece_tokenizer/commits/1.0.0
SentencePieceTokenizer - Main tokenizer class
fromBytes() - Load from protobuf bytesfromModelFile() / fromModelFileSync() - Load from .model fileencode() - Encode single textencodeBatch() - Encode multiple textsencodeBatchParallel() - Parallel batch encoding using IsolatesencodePair() - Encode text pairs for sequence classificationencodePairBatch() - Batch encode text pairsdecode() / decodeBatch() - Decode token IDs back to textEncoding class with:
ids - Token IDs (Int32List)tokens - Token stringstypeIds - Segment type IDs (Uint8List)attentionMask - Attention mask (Uint8List)specialTokensMask - Special token indicators (Uint8List)offsets - Character offsets for each tokenwithPadding() / withTruncation() - Post-processing methodstruncatePair() - Static method for pair truncationPredefined configurations:
SentencePieceConfig.llama - Llama-style (BOS only)SentencePieceConfig.gemma - Gemma-style (BOS + EOS)Truncation strategies:
longestFirst - Truncate longer sequence firstonlyFirst - Only truncate first sequenceonlySecond - Only truncate second sequencedoNotTruncate - No truncationPadding options:
Your coding agent can read these notes before it upgrades. Set up the MCP server →