NewYour coding agent can read the release notes before it upgrades.Set up the MCP server →
PyPI · #1879 most downloaded on PyPI
TF.Text is a TensorFlow library of text related ops, modules, and subgraphs.
Last release 26 days ago
08 Sep 2026
Release timing varies
gaps range from 2 weeks to 10 months
Nearly every release is documented
notes for 35 of 35 stable releases
Nothing withdrawn
no release was ever pulled
7 years old
66 releases · first in 2019
Exclude stubdata and i18n BUILD.bazel from ICU 77 repository archive
This release contains contributions from many people at Google, as well as:
C. Antonio Sánchez
Update TensorFlow Text WORKSPACE for latest Xcode.
use_simple_allocator option.This release contains contributions from many people at Google, as well as:
C. Antonio Sánchez
One column per quarter.
Fixed a memory safety bug in FastWordpieceTokenizer concerning StringVocab lifetime. This prevents temporary copies that were previously invalidating
This release contains contributions from many people at Google, as well as:
C. Antonio Sánchez
iterate over Unicode White_Space directly, rather than testing each of 1.1M code points
cleanup of deprecated test methods
srcs_version and python_version attributes, as they already default to "PY3"python_wheel_version_suffix_repository for TF dependency in tf-text project after tensorflow/tensorflow@805775f.python_wheel_version_suffix_repository for TF dependency in tf-text project after tensorflow/tensorflow@805775f."This release contains contributions from many people at Google, as well as:
C. Antonio Sánchez
cleanup of deprecated test methods
srcs_version and python_version attributes, as they already default to "PY3"python_wheel_version_suffix_repository for TF dependency in tf-text project after tensorflow/tensorflow@805775f.python_wheel_version_suffix_repository for TF dependency in tf-text project after tensorflow/tensorflow@805775f."This release contains contributions from many people at Google, as well as:
C. Antonio Sánchez
Major Features and Improvements
Fix out-of-bounds read in whitespace tokenizer
Fix out-of-bounds read in whitespace tokenizer
negative sampling excludes positive class
tensorflow-macos from setup.pyThis release contains contributions from many people at Google, as well as:
Alex Shroyer, C. Antonio Sánchez, Maggie Zhang
negative sampling excludes positive class
tensorflow-macos from setup.pyThis release contains contributions from many people at Google, as well as:
Alex Shroyer, C. Antonio Sánchez, Maggie Zhang
Support resource manager scoped Sentencepiece resources.
use_unique_shared_resource_name.This release contains contributions from many people at Google, as well as:
Raviteja Gorijala
Support resource manager scoped Sentencepiece resources.
use_unique_shared_resource_name.This release contains contributions from many people at Google, as well as:
Raviteja Gorijala
Update TF versions and scripts to allow consistently building against tf-nightly.
Update TF versions and scripts to allow consistently building against tf-nightly.
This release contains contributions from many people at Google, as well as:
nallave
Fix nullptr dereference issue in UnicodeScriptTokenizeWithOffsetOp.
coerce_to_valid_utf8_op_test test on macThis release contains contributions from many people at Google, as well as:
Add @tensorflow/docs-team to CODEOWNERS
coerce_to_valid_utf8_op_test test on macUpdate word_embeddings.ipynb with Time Series as default dashboard.
Update word_embeddings.ipynb with Time Series as default dashboard.
Replace usage of the tsl::Status constructor with a tsl::{error, errors}::Code.
This release contains contributions from many people at Google, as well as:
Raviteja Gorijala
Deprecate PY37 support for TF-Text
concatenate_segments op.Thanks to our Contributors This release contains contributions from many people at Google, as well as: synandi, tilakrayal
Deprecate PY37 support for TF-Text
BoiseTagsTpOffsets opThis release contains contributions from many people at Google, as well as:
synandi, tilakrayal
Added op for converting to/from BOISE labels to offsets
OffsetsToBoiseTags opPositionalEmbedding layer.tensorflow::Status::OK() with tensorflow::OkStatus().Model.fit.This release contains contributions from many people at Google, as well as:
satojkovic
Moving logging.h and bitmap from tf/core to tf/tsl.
OffsetsToBoiseTags opPositionalEmbedding layer.tensorflow::Status::OK() with tensorflow::OkStatus().Model.fit.This release contains contributions from many people at Google, as well as:
satojkovic
New ByteSplitter which tokenizes strings into bytes.
Trimmer ABC to be usable as tf_text.Trimmerfine_tune_bert.This release contains contributions from many people at Google, as well as:
gadagashwini, mnahinkhan, Steve R. Sun, synandi
New ByteSplitter which tokenizes strings into bytes.
Trimmer ABC to be usable as tf_text.Trimmerfine_tune_bert.This release contains contributions from many people at Google, as well as:
gadagashwini, mnahinkhan, Steve R. Sun, synandi
Added FastSentencepieceTokenizer which is convertible to TF Lite. Please note the op name in the graph will change, so any models trained with this ve
New FastBertNormalizer that improves speed for BERT normalization and is convertible to TF Lite.
:all.This release contains contributions from many people at Google, as well as:
Aflah, Connor Brinton, devnev39, Janak Ramakrishnan, Martin, Nathan Luehr, Pierre Dulac, Rabin Adhikari, gadagashwini, mohantym, rtg0795
New FastBertNormalizer that improves speed for BERT normalization and is convertible to TF Lite.
This release contains contributions from many people at Google, as well as:
Aflah, Connor Brinton, devnev39, Janak Ramakrishnan, Martin, Nathan Luehr, Pierre Dulac, Rabin Adhikari
New FastBertNormalizer that improves speed for BERT normalization and is convertible to TF Lite.
This release contains contributions from many people at Google, as well as:
Aflah, Connor Brinton, devnev39, Janak Ramakrishnan, Martin, Nathan Luehr, Pierre Dulac, Rabin Adhikari
📦️ Fix macOS packaging so it works with package managers like Poetry
This release contains contributions from many people at Google, as well as:
Connor Brinton
Upgrade Sentencepiece to v0.1.96
This release contains contributions from many people at Google, as well as:
Abhijeet Manhas, chunduriv, Dean Wyatte, Feiteng, jaymessina3, Mao, Olivier Bacs, RenuPatelGoogle, Steve R. Sun, Stonepia, sun1638650145, Tharaka De Silva, thuang513, Xiaoquan Kong, devnev39, Janak Ramakrishnan, Pierre Dulac
Upgrade Sentencepiece to v0.1.96
ShrinkLongestTrimmerThis release contains contributions from many people at Google, as well as:
Abhijeet Manhas, chunduriv, Dean Wyatte, Feiteng, jaymessina3, Mao, Olivier Bacs, RenuPatelGoogle, Steve R. Sun, Stonepia, sun1638650145, Tharaka De Silva, thuang513, Xiaoquan Kong
Fixed broken packages for MacOS & Windows
Added new tokenizer: FastWordpieceTokenizer that is considerably faster than the original WordpieceTokenizer
This release contains contributions from many people at Google, as well as:
Aaron Siddhartha Mondal, Abhijeet Manhas, Dominik Schlösser, jaymessina3, Mao, Xiaoquan Kong, Yasir Modak, Olivier Bacs, Tharaka De Silva
Added new tokenizer: FastWordpieceTokenizer that is considerably faster than the original WordpieceTokenizer
This release contains contributions from many people at Google, as well as:
Aaron Siddhartha Mondal, Abhijeet Manhas, Dominik Schlösser, jaymessina3, Mao, Xiaoquan Kong, Yasir Modak
WhitespaceTokenizer was rewritten to increase speed and smaller kernel size
This release contains contributions from many people at Google, as well as:
Aaron Siddhartha Mondal, Dominik Schlösser, Xiaoquan Kong, Yasir Modak
Update __init__.py: Added a __version__ variable
__init__.py: Added a __version__ variableThis release contains contributions from many people at Google, as well as:
8bitmp3, akiprasad, bongbonglemon, Jules Gagnon-Marchand, Stonepia
Update __init__.py: Added a __version__ variable
__init__.py: Added a __version__ variableThis release contains contributions from many people at Google, as well as:
8bitmp3, akiprasad, bongbonglemon, Jules Gagnon-Marchand, Stonepia
Update the sentence_breaking_ops docstring to indicate that it's deprecated.
We want to particularly point out that guides, tutorials, and API docs are currently being published to http://tensorflow.org/text ! This should make it easier for users to find our documentation. We worked hard on improving docs across the board, so feel free to let us know if further clarification is needed.
Replacing use of TFT's deprecated dataset_schema.from_feature_spec with its replacement schema_utils.schema_from_feature_spec.
BertTokenizer and WordpieceTokenizer.shape attribute to the ToDense Keras layer.unselectable_ids shape check in ItemSelector.text.BertTokenizertensorflow_text pip package.tools pip package inclusion.This release contains contributions from many people at Google, as well as:
Rens, Samuel Marks, thuang513
Fix export as saved model of hub_module_splitter
This release contains contributions from many people at Google, as well as:
fsx950223
We are now building a nightly package - tensorflow-text-nightly. This is available for Linux immediately, with other platforms to be added soon.
tensorflow-text-nightly. This is available for Linux immediately, with other platforms to be added soon.New APIs proposed in RFC: End-to-end text preprocessing with TF.Text #283 have been added, including:
Splitter
RegexSplitterStateBasedSentenceBreakerTrimmer
WaterfallTrimmerRoundRobinTrimmerItemSelector
RandomItemSelectorFirstNItemSelectorMaskValuesChoosermask_language_model()combine_segments()pad_model_inputs()Spliter / SplitterWithOffsets abstract base classes. These are meant to replace the current Tokenizer / TokenizerWithOffsets base classes. The Tokenizer base classes will continue to work and will implement these new Splitter base classes. The reasoning behind the change is to prevent confusion when future splitting operations that also use this interface do not tokenize into words (sentences, subwords, etc).offset_end is a positional value rather than a length.HubModuleSplitter that helps handle ragged tensor input and outputs for hub modules which implement the Splitter class.SplitMergeFromLogitsTokenizer which is a narrowly focused tokenizer that splits text based on logits from a model. This is used with the newly released Chinese segmentation model.normalize_utf8_with_offsets and find_source_offsets ops.normalization_form that will be ignored.This release contains contributions from many people at Google, as well as:
Pranay Joshi, Siddharths8212376, Vincent Bodin
Released our first TF Hub module for Chinese segmentation! Please visit the hub module page here for more info including instructions on how to use th
Spliter / SplitterWithOffsets abstract base classes. These are meant to replace the current Tokenizer / TokenizerWithOffsets base classes. The Tokenizer base classes will continue to work and will implement these new Splitter base classes. The reasoning behind the change is to prevent confusion when future splitting operations that also use this interface do not tokenize into words (sentences, subwords, etc).offset_end is a positional value rather than a length.HubModuleSplitter that helps handle ragged tensor input and outputs for hub modules which implement the Splitter class.SplitMergeFromLogitsTokenizer which is a narrowly focused tokenizer that splits text based on logits from a model. This is used with the newly released Chinese segmentation model.normalize_utf8_with_offsets and find_source_offsets ops.normalization_form that will be ignored.This release contains contributions from many people at Google, as well as:
Pranay Joshi, Siddharths8212376, Vincent Bodin
Released our first TF Hub module for Chinese segmentation! Please visit the hub module page here for more info including instructions on how to use th
Spliter / SplitterWithOffsets abstract base classes. These are meant to replace the current Tokenizer / TokenizerWithOffsets base classes. The Tokenizer base classes will continue to work and will implement these new Splitter base classes. The reasoning behind the change is to prevent confusion when future splitting operations that also use this interface do not tokenize into words (sentences, subwords, etc).offset_end is a positional value rather than a length.HubModuleSplitter that helps handle ragged tensor input and outputs for hub modules which implement the Splitter class.SplitMergeFromLogitsTokenizer which is a narrowly focused tokenizer that splits text based on logits from a model. This is used with the newly released Chinese segmentation model.normalize_utf8_with_offsets and find_source_offsets ops.normalization_form that will be ignored.This release contains contributions from many people at Google, as well as:
Pranay Joshi, Siddharths8212376, Vincent Bodin
Please note that this is a pre-release and meant to run with TF v2.3.x. We wanted to give access to some of the features we were adding to 2.4.x, but
Please note that this is a pre-release and meant to run with TF v2.3.x. We wanted to give access to some of the features we were adding to 2.4.x, but did not want to wait for the TF release.
Spliter / SplitterWithOffsets abstract base classes. These are meant to replace the current Tokenizer / TokenizerWithOffsets base classes. The Tokenizer base classes will continue to work and will implement these new Splitter base classes. The reasoning behind the change is to prevent confusion when future splitting operations that also use this interface do not tokenize into words (sentences, subwords, etc).offset_end is a positional value rather than a length.HubModuleSplitter that helps handle ragged tensor input and outputs for hub modules which implement the Splitter class.SplitMergeFromLogitsTokenizer which is a narrowly focused tokenizer that splits text based on logits from a model. This is used with the newly released Chinese segmentation model.Added UnicodeCharacterTokenizer
Added UnicodeCharacterTokenizer
Python 3.8 release builds added
# Release 2.2 ## Major Features and Improvements ## Breaking Changes ## Bug Fixes and Other Changes * Update version ## Thanks to our Contributors
Force MacOS builds to build for OSX 10.9 so they can be installed to a wider range of MacOS versions.
Add op for solving max-spanning-tree (MST) problems. The code here is intended for NLP applications, but attempts to remain agnostic to particular NLP
This release contains contributions from many people at Google, as well as:
Hyunwoo Cho
BertTokenizer to accept a string tensor for the vocab_lookup_table.
Add support for token offsets to BertTokenizer.
Missing regex character class used in BertTokenizer
Fixes a bug in case_fold_utf8 and normalize_utf8 ops where they were unable to locate the ICU data file.
Please note that moving forward our releases and branches will match the major & minor versions of core TensorFlow. This should prevent future confusi
Please note that moving forward our releases and branches will match the major & minor versions of core TensorFlow. This should prevent future confusion. As such, this (previously 1.0) release is 2.0, and we will be skiping straight to 1.15 for the next 1.x release to support TF 1.15.
Major Updates:
Minor Updates:
Missing regex character class used in BertTokenizer
Fixes a bug in case_fold_utf8 and normalize_utf8 ops where they were unable to locate the ICU data file.
Your coding agent can read these notes before it upgrades. Set up the MCP server →