NewYour coding agent can read the release notes before it upgrades.Set up the MCP server →
PyPI · #164 most downloaded on PyPI
Last release 13 days ago
21 Sep 2026
Ships fairly regularly
a new release about every 5 weeks
Some releases are documented
notes for 35 of the last 60 stable releases
2 versions withdrawn
withdrawn after publishing
7 years old
103 releases · first in 2019
Making tokenizers the sota tokenization library, same API, same token IDs, same standards
tokenizers the sota tokenization library, same API, same token IDs, same standardsWe are happy to finally bring tokenizers up to speed ⚡ !
tokenizers was released 7 years ago, in a time where running the model was always the bottleneck. The ecosystem has evolved so much that tokenizers became a the bottleneck.
As we saw many different libraries and project pop out, improving the tokenizer, we realized that our library's entire stack was too outdated for external contributors to contribute. This pushed us to refactor the codebase, and make sure the next developers come here to push the performances of tokenization.
The TLDR; is that we're a lot faster than before, crate size 6x smaller, peak memory is reduced.
We worked a lot to reduce the UTF-8 tax, and make sure the improvements span across all languages.
For more details on the release, read this.
Big Kudos to @SBrandeis and @McPatate for this sprint. Special thanks to @sebpop for his involvement as well.
One column per quarter.
Nothing published for this version
This is the last v0 release, we are moving to v1!!
This is the last v0 release, we are moving to v1!!
More details coming soon 👀
decode API by @SBrandeis in #2099Full Changelog: v0.23.1...v0.23.2
⚠️ As called out in breaking changes: stricter type info means previously-hidden type errors in user code may now surface under mypy --strict .
tokenizers 0.23.1 is the first proper stable release in the 0.23 line — 0.23.0 only ever shipped as rc0 because the release pipeline itself was broken (Node side hadn't shipped multi-platform binaries since 2023, Python side was on pyo3 0.27 without free-threaded support). 0.23.1 is the version where everything actually goes out the door together: full Node multi-platform wheels for the first time in years, Python 3.14 (regular and free-threaded 3.14t), full type hints for every Python class, and a stack of measurable perf wins on the BPE / added-vocab hot paths.
There is no functional 0.23.0 published — we tag 0.23.1 directly so users don't accidentally pull a never-shipped version.
requires-python = ">=3.10"; 3.9 users stay on 0.22.x.add_tokens normalizes content at insertion (#1995) — re-saved tokenizer.json may differ in the added_tokens block. Existing files load unchanged.Any now return real types; mypy --strict may surface previously-hidden errors. Stub layout also moved from tokenizers/<sub>/__init__.pyi to tokenizers/<sub>.pyi. This breaks the surface of some of the processors like RobertaProcessign's __init__ .PyResult<T> because of Arc<RwLock<Tokenizer>>; a poisoned lock surfaces as PyException instead of a panic.Run with cargo bench --bench <name> -- --save-baseline v0_22_2 on v0.22.2, then --baseline v0_22_2 on v0.23.1. Numbers are point-in-time wall clock on a single laptop; relative deltas are what matters, absolute numbers will differ on CI hardware.
bench: improve added_vocab_deserialize to reflect real-world workloads (#2000) is now representative of how transformers actually loads tokenizer.json files. The combined effect of daachorse for the matching automaton plus the normalize-on-insert refactor is enormous on this workload:
| benchmark | v0.22.2 | v0.23.1 | change |
|---|---|---|---|
| 100k tokens, special, no norm | ~410 ms | 248 ms | −40% |
| 100k tokens, non-special, no norm | ~7.1 s | 273 ms | −96% |
| 100k tokens, special, NFKC | ~395 ms | 235 ms | −40% |
| 100k tokens, non-special, NFKC | ~7.4 s | 290 ms | −96% |
| 400k tokens, special, no norm | ~15 s | 980 ms | −94% |
Real-world impact: loading a Llama-3-style tokenizer with a large set of added tokens dropped from "noticeable pause" to "instant".
| benchmark | v0.22.2 | v0.23.1 | change |
|---|---|---|---|
BPE GPT2 encode batch, no cache |
530 ms | 446 ms | −16% |
BPE GPT2 encode batch (cached) |
690 ms | 685 ms | noise |
BPE GPT2 encode (single) |
1.95 s | 1.94 s | noise |
BPE Train (small) |
32.6 ms | 31.5 ms | −3% |
BPE Train (big) |
1.01 s | 988 ms | −2% |
The BPE per-thread cache PR (#2028) shows much larger wins on highly-parallel workloads (+47–62% at 88+ threads on a server box, per the PR's own measurements on Vera). Single-thread batch numbers above are flat or slightly improved because cache-hit overhead was already low without contention.
| benchmark | v0.22.2 | v0.23.1 | change |
|---|---|---|---|
llama3-encode (single) |
2.10 s | 2.02 s | −4% |
llama3-batch |
438 ms | 408 ms | −7% |
llama3-offsets |
410 ms | 395 ms | −4% |
Right-direction truncation no longer pre-tokenizes past max_length. The new truncation_benchmark doesn't exist on v0.22.2 so there's no apples-to-apples here, but the PR's own measurements on the same machine showed −20–28% across a range of max_length values for right-truncation; left-truncation unchanged.
BPE::Builder::build no longer formats strings in a hot loop (#2010) — ~45% faster Tokenizer::from_file on Llama-3 in the PR's profile.The tokenizer.json format is forward-compatible: existing files load on 0.23 unchanged. Two things to know if you re-save:
added_tokens entries created via add_tokens(..., normalized=True) will have their content normalized at save time — see breaking-change note above.tokenizer.train(...) no longer keeps a redundant added_tokens/special_tokens Vec separate from the added_tokens_map_r. Public API surface unchanged; only the internal struct shape moved.bench: improve added_vocab_deserialize to reflect real-world workloads (#2000) lands a more realistic micro-benchmark for this surface; if you're tracking deserialize perf in your own CI, the new bench is the one to compare against.
Dedicated wheels for python3.14t (the free-threaded build introduced in PEP 703). The wheel:
Py_MOD_GIL_NOT_USED, so importing tokenizers does not force the GIL back on.abi3 cargo feature (free-threaded Python doesn't expose the limited API).Arc<RwLock<Tokenizer>> for the inner state so concurrent setters and encoders don't race PyO3's per-pyclass borrow check.A new stress-test module tests/test_freethreaded.py exercises N-encoder × M-setter races on a single Tokenizer and asserts no RuntimeError: Already borrowed, no RwLock poisoning, and that sys._is_gil_enabled() is False post-import.
For the regular CPython wheel everything is unchanged.
The npm package now ships 13 platforms (macOS x64/arm64/universal, Windows x64/i686/arm64, Linux x64/arm64/armv7 in both glibc and musl, Android arm64/armv7) — previous workflows only built 3 of those, leaving Apple Silicon / Linux ARM / Alpine users with package-not-found errors since 2023 (#1365, #1703, #1922). Fixed via #1970 + #2034, which also bumps @napi-rs/cli to v3 and switches cross-builds to cargo-zigbuild.
Every class in the python bindings now ships proper .pyi stubs — Tokenizer, AddedToken, Encoding, every decoder / model / normalizer / pre-tokenizer / processor / trainer. Editors and type checkers (mypy, pyright, ty) see real signatures with types and docstrings instead of falling back to Any.
The stubs are generated automatically from the compiled extension via tools/stub-gen (Rust binary using pyo3-introspection). Re-running make style regenerates them; CI guards against regenerated-vs-checked-in drift. If the generator ever returns 0 docstrings (e.g. because the [patch.crates-io] pin in .cargo/config.toml falls out of sync with the pyo3 dep version), it now hard-aborts with a precise diagnostic instead of silently emitting bare-bones stubs.
>>> from tokenizers import Tokenizer
>>> # IDEs now resolve every method, every kwarg, every return type
>>> Tokenizer.from_pretrained("bert-base-cased")⚠️ As called out in breaking changes: stricter type info means previously-hidden type errors in user code may now surface under mypy --strict.
models.Unigram now exposes alpha and nbest_size for subword regularization (parity with Google's implementation, #1994). Closes long-standing requests #730 and #849.Tokenizer (#1958) — useful for long-lived caches that don't want to keep tokenizers alive.ci_benchmark against the stored baseline and posts a comparison chart to the PR.EncodingVisualizer: unclosed annotation span fixed (#1911), HTML escape applied to output (#1937).__copy__ / __deepcopy__ (#1930).to_vec() from slice (#1964).wget / norvig URL with HF Hub downloads in test data fetch (#2018).uv support in the Python Makefile (#1977).Thanks to everyone who shipped commits between v0.22.2 and v0.23.1:
@ArthurZucker, @finnagin, @gordonmessmer, @jberg5, @kennethsible, @llukito, @MayCXC, @McPatate, @michaelfeil, @mrkm4ntr, @musicinmybrain, @ngoldbaum, @OhashiReon, @paulinebm, @podarok, @rtrompier, @sebpop, @Shivam-Bhardwaj, @threexc, @wheynelau, @xanderlent — plus @dependabot and @hf-security-analysis for keeping pins fresh.
Full Changelog: v0.22.2...v0.23.1
Nothing published for this version
@napi-rs/cli removed the --name option; the name is now read from the napi.name field in package.json. Replace with --dir ./artifacts so the command p
@napi-rs/cli removed the --name option; the name is now read from the
napi.name field in package.json. Replace with --dir ./artifacts so the
command picks up the downloaded per-target .node files.
Also pins the second PyO3/maturin-action reference (inside the build
job) that was missed by the earlier action-modernization commit.
Co-Authored-By: Claude Opus 4.7 (1M context) noreply@anthropic.com
Okay mostly doing the release for these PR:
Okay mostly doing the release for these PR:
Basically good typing with at least ty, and a lot fast (from 4 to 8x faster) loading vocab with a lot of added tokens and GIL free !?
ci: add support for building Win-ARM64 wheels by @MugundanMCW in #1869
Add cargo-semver-checks to Rust CI workflow by @haixuanTao in #1875
Update indicatif dependency by @gordonmessmer in #1867
Bump node-forge from 1.3.1 to 1.3.2 in /tokenizers/examples/unstable_wasm/www by @dependabot[bot] in #1889
Bump js-yaml from 3.14.1 to 3.14.2 in /bindings/node by @dependabot[bot] in #1892
fix: used normalize_str in BaseTokenizer.normalize by @ishitab02 in #1884
Remove runtime stderr warning from Python bindings by @Copilot in #1898
Mark immutable pyclasses as frozen by @ngoldbaum in #1861
DOCS: add add_prefix_space to processors.ByteLevel by @CloseChoice in #1878
Bump express from 4.21.2 to 4.22.1 in /tokenizers/examples/unstable_wasm/www by @dependabot[bot] in #1903
Full Changelog: v0.22.1...v0.22.2
Okay mostly doing the release for these PR:
<img width="2400" height="1200" alt="image" src="https://github.com/user-attachments/assets/0b974453-1fc6-4393-84ea-da99269e2b34" />
Basically good typing with at least ty, and a lot fast (from 4 to 8x faster) loading vocab with a lot of added tokens and GIL free !?
ci: add support for building Win-ARM64 wheels by @MugundanMCW in https://github.com/huggingface/tokenizers/pull/1869
Add cargo-semver-checks to Rust CI workflow by @haixuanTao in https://github.com/huggingface/tokenizers/pull/1875
Update indicatif dependency by @gordonmessmer in https://github.com/huggingface/tokenizers/pull/1867
Bump node-forge from 1.3.1 to 1.3.2 in /tokenizers/examples/unstable_wasm/www by @dependabot[bot] in https://github.com/huggingface/tokenizers/pull/1889
Bump js-yaml from 3.14.1 to 3.14.2 in /bindings/node by @dependabot[bot] in https://github.com/huggingface/tokenizers/pull/1892
fix: used normalize_str in BaseTokenizer.normalize by @ishitab02 in https://github.com/huggingface/tokenizers/pull/1884
[MINOR:TYPO] Update mod.rs by @cakiki in https://github.com/huggingface/tokenizers/pull/1883
Remove runtime stderr warning from Python bindings by @Copilot in https://github.com/huggingface/tokenizers/pull/1898
Mark immutable pyclasses as frozen by @ngoldbaum in https://github.com/huggingface/tokenizers/pull/1861
DOCS: add add_prefix_space to processors.ByteLevel by @CloseChoice in https://github.com/huggingface/tokenizers/pull/1878
Bump express from 4.21.2 to 4.22.1 in /tokenizers/examples/unstable_wasm/www by @dependabot[bot] in https://github.com/huggingface/tokenizers/pull/1903
Full Changelog: https://github.com/huggingface/tokenizers/compare/v0.22.1...v0.22.2
Nothing published for this version
Bump huggingface_hub upper version ( #1866 ) from @Wauplin
update version to 0.22.1.rc0
update version to 0.22.1.rc0
Bump on-headers and compression in /tokenizers/examples/unstable_wasm/www by @dependabot [bot] in #1827
from_bytes and read_bytes Methods in WordPiece Tokenizer for WebAssembly Compatibility by @sondalex in #1758EncodingVisualizer.calculate_label_colors by @Liam-DeVoe in #1853Full Changelog: v0.21.3...v0.22.0rc0
Nothing published for this version
No change, the 0.21.3 release failed, this is just a re-release.
Full Changelog: https://github.com/huggingface/tokenizers/compare/v0.21.3...v0.21.4
No change, the 0.21.3 release failed, this is just a re-release.
https://github.com/huggingface/tokenizers/releases/tag/v0.21.3
This release if focused around some performance optimization, enabling broader python no gil support, and fixing some onig issues!
This release if focused around some performance optimization, enabling broader python no gil support, and fixing some onig issues!
Full Changelog: https://github.com/huggingface/tokenizers/compare/v0.21.1...v0.21.2rc0
Nothing published for this version
Update dev version and pyproject.toml by @ArthurZucker in https://github.com/huggingface/tokenizers/pull/1693
Full Changelog: https://github.com/huggingface/tokenizers/compare/v0.21.0...v0.21.1
Update dev version and pyproject.toml by @ArthurZucker in https://github.com/huggingface/tokenizers/pull/1693
Full Changelog: https://github.com/huggingface/tokenizers/compare/v0.21.0...v0.21.1rc0
We also no longer support python 3.7 or 3.8 (similar to transformers) as they are deprecated.
We also no longer support python 3.7 or 3.8 (similar to transformers) as they are deprecated.
Full Changelog: https://github.com/huggingface/tokenizers/compare/v0.20.3...v0.21.0
Nothing published for this version
Nothing published for this version
Nothing published for this version
There was a breaking change in 0.20.3 for tuple inputs of encode_batch!
There was a breaking change in 0.20.3 for tuple inputs of encode_batch!
Full Changelog: https://github.com/huggingface/tokenizers/compare/v0.20.2...v0.20.3
Thanks a MILE to @diliop we now have support for python 3.13! 🥳
Thanks a MILE to @diliop we now have support for python 3.13! 🥳
set_var by @sftse in https://github.com/huggingface/tokenizers/pull/1664Full Changelog: https://github.com/huggingface/tokenizers/compare/v0.20.1...v0.20.2
The most awaited offset issue with Llama is fixed 🥳
The most awaited offset issue with Llama is fixed 🥳
ignore_merges] Fix offsets by @ArthurZucker in https://github.com/huggingface/tokenizers/pull/1640Full Changelog: https://github.com/huggingface/tokenizers/compare/v0.20.0...v0.20.1
[BREAKING CHANGE] Ignore added_tokens (both special and not) in the decoder by @Narsil in https://github.com/huggingface/tokenizers/pull/1513
This release is focused on performances and user experience.
First off, we did a bit of benchmarking, and found some place for improvement for us!
With a few minor changes (mostly #1587) here is what we get on Llama3 running on a g6 instances on AWS https://github.com/huggingface/tokenizers/blob/main/bindings/python/benches/test_tiktoken.py :
We shipped better deserialization errors in general, and support for __str__ and __repr__ for all the object. This allows for a lot easier debugging see this:
>>> from tokenizers import Tokenizer;
>>> tokenizer = Tokenizer.from_pretrained("bert-base-uncased");
>>> print(tokenizer)
Tokenizer(version="1.0", truncation=None, padding=None, added_tokens=[{"id":0, "content":"[PAD]", "single_word":False, "lstrip":False, "rstrip":False, ...}, {"id":100, "content":"[UNK]", "single_word":False, "lstrip":False, "rstrip":False, ...}, {"id":101, "content":"[CLS]", "single_word":False, "lstrip":False, "rstrip":False, ...}, {"id":102, "content":"[SEP]", "single_word":False, "lstrip":False, "rstrip":False, ...}, {"id":103, "content":"[MASK]", "single_word":False, "lstrip":False, "rstrip":False, ...}], normalizer=BertNormalizer(clean_text=True, handle_chinese_chars=True, strip_accents=None, lowercase=True), pre_tokenizer=BertPreTokenizer(), post_processor=TemplateProcessing(single=[SpecialToken(id="[CLS]", type_id=0), Sequence(id=A, type_id=0), SpecialToken(id="[SEP]", type_id=0)], pair=[SpecialToken(id="[CLS]", type_id=0), Sequence(id=A, type_id=0), SpecialToken(id="[SEP]", type_id=0), Sequence(id=B, type_id=1), SpecialToken(id="[SEP]", type_id=1)], special_tokens={"[CLS]":SpecialToken(id="[CLS]", ids=[101], tokens=["[CLS]"]), "[SEP]":SpecialToken(id="[SEP]", ids=[102], tokens=["[SEP]"])}), decoder=WordPiece(prefix="##", cleanup=True), model=WordPiece(unk_token="[UNK]", continuing_subword_prefix="##", max_input_chars_per_word=100, vocab={"[PAD]":0, "[unused0]":1, "[unused1]":2, "[unused2]":3, "[unused3]":4, ...}))
>>> tokenizer
Tokenizer(version="1.0", truncation=None, padding=None, added_tokens=[{"id":0, "content":"[PAD]", "single_word":False, "lstrip":False, "rstrip":False, "normalized":False, "special":True}, {"id":100, "content":"[UNK]", "single_word":False, "lstrip":False, "rstrip":False, "normalized":False, "special":True}, {"id":101, "content":"[CLS]", "single_word":False, "lstrip":False, "rstrip":False, "normalized":False, "special":True}, {"id":102, "content":"[SEP]", "single_word":False, "lstrip":False, "rstrip":False, "normalized":False, "special":True}, {"id":103, "content":"[MASK]", "single_word":False, "lstrip":False, "rstrip":False, "normalized":False, "special":True}], normalizer=BertNormalizer(clean_text=True, handle_chinese_chars=True, strip_accents=None, lowercase=True), pre_tokenizer=BertPreTokenizer(), post_processor=TemplateProcessing(single=[SpecialToken(id="[CLS]", type_id=0), Sequence(id=A, type_id=0), SpecialToken(id="[SEP]", type_id=0)], pair=[SpecialToken(id="[CLS]", type_id=0), Sequence(id=A, type_id=0), SpecialToken(id="[SEP]", type_id=0), Sequence(id=B, type_id=1), SpecialToken(id="[SEP]", type_id=1)], special_tokens={"[CLS]":SpecialToken(id="[CLS]", ids=[101], tokens=["[CLS]"]), "[SEP]":SpecialToken(id="[SEP]", ids=[102], tokens=["[SEP]"])}), decoder=WordPiece(prefix="##", cleanup=True), model=WordPiece(unk_token="[UNK]", continuing_subword_prefix="##", max_input_chars_per_word=100, vocab={"[PAD]":0, "[unused0]":1, "[unused1]":2, ...}))
The pre_tokenizer.Sequence and normalizer.Sequence are also more accessible now:
from tokenizers import normalizers
norm = normalizers.Sequence([normalizers.Strip(), normalizers.BertNormalizer()])
norm[0]
norm[1].lowercase=False
USED_PARALLELISM atomic by @nathaniel-daniel in https://github.com/huggingface/tokenizers/pull/1532cached_download to hf_hub_download in tests by @Wauplin in https://github.com/huggingface/tokenizers/pull/1547dropout = 0.0 as an equivalent to none in BPE by @mcognetta in https://github.com/huggingface/tokenizers/pull/1550None to reset pre_tokenizers and normalizers, and index sequences by @ArthurZucker in https://github.com/huggingface/tokenizers/pull/1590Full Changelog: https://github.com/huggingface/tokenizers/compare/v0.19.1...v0.20.0rc1
add serialization for ignore_merges by @ArthurZucker in https://github.com/huggingface/tokenizers/pull/1504
ignore_merges by @ArthurZucker in https://github.com/huggingface/tokenizers/pull/1504Full Changelog: https://github.com/huggingface/tokenizers/compare/v0.19.0...v0.19.1
🚨🚨 BREAKING CHANGE 🚨🚨: (add_prefix_space dropped everything is using prepend_scheme enum instead) Refactor metaspace by @ArthurZucker in https://githu…
remove black] And use ruff by @ArthurZucker in https://github.com/huggingface/tokenizers/pull/1436AddedVocabulary. by @eaplatanios in https://github.com/huggingface/tokenizers/pull/1443Full Changelog: https://github.com/huggingface/tokenizers/compare/v0.15.2...v0.19.0
Big shoutout to @rlrs for the fast replace normalizers PR. This boosts the performances of the tokenizers: !image
Big shoutout to @rlrs for the fast replace normalizers PR. This boosts the performances of the tokenizers:
Full Changelog: https://github.com/huggingface/tokenizers/compare/v0.15.1...v0.15.2rc1
udpate to version = "0.15.1-dev0" by @ArthurZucker in https://github.com/huggingface/tokenizers/pull/1390
Clone on Tokenizer, add Encoding.into_tokens() method by @epwalsh in https://github.com/huggingface/tokenizers/pull/1381Full Changelog: https://github.com/huggingface/tokenizers/compare/v0.15.0...v0.15.1
fix a clerical error in the comment by @tiandiweizun in https://github.com/huggingface/tokenizers/pull/1356
huggingface_hub<1.0 by @Wauplin in https://github.com/huggingface/tokenizers/pull/1385pre_tokenizers] Fix sentencepiece based Metaspace by @ArthurZucker in https://github.com/huggingface/tokenizers/pull/1357Full Changelog: https://github.com/huggingface/tokenizers/compare/v0.14.1...v0.15.0
Fix conda release by @ArthurZucker in https://github.com/huggingface/tokenizers/pull/1211
decode and decode_batch work on borrowed content. by @mfuntowicz in https://github.com/huggingface/tokenizers/pull/1251expect() for disabling truncation by @boyleconnor in https://github.com/huggingface/tokenizers/pull/1316safetensors. + Rewritten node bindings. by @Narsil in https://github.com/huggingface/tokenizers/pull/1331Full Changelog: https://github.com/huggingface/tokenizers/compare/v0.13.3...v0.14.1
⚠️ Reworks the release pipeline. Other breaking changes ⚠️ :
⚠️ Reworks the release pipeline. Other breaking changes ⚠️ :
is_special_token rename to special for consistencyOFF by default, and depends on hf-hub instead of cached_path (updated cache directory, better sync implementation)decode and decode_batch work on borrowed content. by @mfuntowicz in https://github.com/huggingface/tokenizers/pull/1251expect() for disabling truncation by @boyleconnor in https://github.com/huggingface/tokenizers/pull/1316safetensors. + Rewritten node bindings. by @Narsil in https://github.com/huggingface/tokenizers/pull/1331Full Changelog: https://github.com/huggingface/tokenizers/compare/v0.13.3...v0.14.0
Update pr docs actions by @mishig25 in https://github.com/huggingface/tokenizers/pull/1101
Tokenizer clone. by @Narsil in https://github.com/huggingface/tokenizers/pull/1152from_pretrained on invalid ids (better error message). by @Narsil in https://github.com/huggingface/tokenizers/pull/1153tokenizers. by @Narsil in https://github.com/huggingface/tokenizers/pull/1183datasets train example by @lhoestq in https://github.com/huggingface/tokenizers/pull/1192Replace to decoder (to undo the Replace Normalizer for Metaspace split). by @Narsil in https://github.com/huggingface/tokenizers/pull/1195normalizers.Prepend (To be used instead of Metaspace). by @Narsil in https://github.com/huggingface/tokenizers/pull/1194content to Strip decoder to allow decoding mid tokens. by @Narsil in https://github.com/huggingface/tokenizers/pull/1199Full Changelog: https://github.com/huggingface/tokenizers/compare/v0.13.2...v0.13.3
Python 3.11 support (Python only modification)
Python 3.11 support (Python only modification)
[#1072] Fixing Roberta type ids.
[#1008] Decoder is now a composable trait, but without being backward incompatible
unstable_wasm feature to support building on Wasm (it's unstable !)Decoder is now a composable trait, but without being backward incompatibleProcessor is now a composable trait, but without being backward incompatibleBoth trait changes warrant a "major" number since, despite best efforts to not break backward compatibility, the code is different enough that we cannot be exactly sure.
[#938] Reverted breaking change. https://github.com/huggingface/transformers/issues/16520
Bump minor version because of a breaking change.
Bump minor version because of a breaking change.
The breaking change was causing more issues upstream in transformers than anticipated:
https://github.com/huggingface/transformers/pull/16537#issuecomment-1085682657
The decision was to rollback on that breaking change, and figure out a different way later to do this modification
[#938] Breaking change. Decoder trait is modified to be composable. This is only breaking if you are using decoders on their own. tokenizers should be error free.
[#939] Making the regex in ByteLevel pre_tokenizer optional (necessary for BigScience)
[#952] Fixed the vocabulary size of UnigramTrainer output (to respect added tokens)
[#954] Fixed not being able to save vocabularies with holes in vocab (ConvBert). Yell warnings instead, but stop panicking.
[#961] Added link for Ruby port of tokenizers
[#960] Feature gate for cli and its clap dependency
Nothing published for this version
Nothing published for this version
Nothing published for this version
[#919] Fixing single_word AddedToken. (regression from 0.11.2)
added_tokens by loading them in batch.[#919] Fixing single_word AddedToken. (regression from 0.11.2)
added_tokens by loading them in batch.[#882] Fixing Punctuation deserialize without argument.
[#236]: Fix a bug with offsets being shifted when there are sub-sequences (Usually with special tokens and/or added tokens in the sequence).
File::open in count_wordsEncoding. Previous mappings
were misleading and only providing offsets. New ones provide methods to easily convert between
char or word (input space) and token (output space)AddedToken with special options like rstrip will keep the matched whitespaces
in the textual representation of the token, exposed in tokens on the Encoding. The ID stays
the same as usual. This fixes the offsets for said tokens.add_prefix_space attribute to determine how to
trim offsets.TruncationError to handle cases where provided max length is too low.encode and encode_batch input has been greatly improved, and it now also accept
pre-tokenized inputs.TruncationError to handle cases where provided max length is too low.onig for byte-level pre-tokenization to remove all the differences with the original
implementation from GPT-2normalized, controlling whether a token should be extracted from the normalized version of the
input text.strip_accents is not specified.Tokenizer and all the parts (PreTokenizer, Normalizer, ...)
using serde. It is now easy to save/load an entire tokenizer.TOKENIZERS_PARALLELISM environment
variable.TemplateProcessing PostProcessor.XXX_to_YYY_offsets() method call by any of the new ones.add_prefix_space and trim_offsets options on RobertaProcessing if you don't
want the offsets trimmed out.PostProcessor now handles offsets relative to the original string (as opposed to the
normalized one).Nothing published for this version
Nothing published for this version
[#226]: Fix the word indexes when there are special tokens
[#222]: All Tokenizer's subparts must now be Send + Sync
Send + SyncTokenizer & ModelBPEDecoderNothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Only one progress bar while reading files during training. This is better for use-cases with a high number of files as it avoids having too many progr
encode and encode_batch now take a new argument, specifying whether we should add the
special tokensNormalizedString has been removed from the Encoding. It is now possible to
retrieve it by calling normalize on the Tokenizer. This brings a reduction of 70% of the memory
footprintNormalizedString API has been improved. It is now possible to retrieve parts of both
strings using both "normalized" or "original" offsetsEncoding are now relative to the original string, and not the
normalized one anymoreAddedToken are now used for both add_special_tokens and add_tokens. Also, these AddedToken
have more options to allow various behaviors.impl PostProcessor for ByteLevel: Handles trimming the offsets if activated. This avoids
the unintuitive inclusion of the whitespaces in the produced offsets, even if these whitespaces are
part of the actual tokenEncoding.post_process can be called on the TokenizerByteLevel BPE:
add_prefix_space is activatedByteLevel PostProcessor to your byte-level BPE tokenizers if relevant.Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Your coding agent can read these notes before it upgrades. Set up the MCP server →