NewYour coding agent can read the release notes before it upgrades.Set up the MCP server →
PyPI · #4757 most downloaded on PyPI
tensorflow/datasets is a library of datasets ready to use with TensorFlow.
Last release 4 months ago
08 May 2026
Release timing varies
gaps range from 8 days to 11 months
Nearly every release is documented
notes for 37 of 39 stable releases
Nothing withdrawn
no release was ever pulled
8 years old
39 releases · first in 2019
Increment TFDS version to 4.9.10.
Increment TFDS version to 4.9.10.
PiperOrigin-RevId: 912492881
apache-beam version is pinned at <2.65.0 until related tests are fixed, see issue 11055 .
apache-beam version is pinned at <2.65.0 until related tests are fixed,One column per quarter.
New Beam writer NoShuffleBeamWriter that doesn't shuffle, which speeds up dataset generation significantly, but does not have deterministic order guar
NoShuffleBeamWriter that doesn't shuffle, which speeds up--nondeterministic_order.CroissantBuilder 's API to generate TFDS datasets from Croissant files.
CroissantBuilder's API to generate TFDS datasets from Croissant files.Added Full support for Python 3.12.
Support to download and prepare datasets using the Parquet data format.
Support to download and prepare datasets using the
Parquet data format.
builder = tfds.builder('fashion_mnist', file_format='parquet')
builder.download_and_prepare()
ds = builder.as_dataset(split='train')
print(next(iter(ds)))tfds.data_source
is pickable, thus working smoothly with
PyGrain. Learn more by following the
tutorial.
TFDS plays nicely with
Croissant. Learn more by
following the
recipe.
A new CroissantBuilder which initializes a DatasetBuilder based on a Croissant metadata file.
HuggingfaceDatasetBuilder.Segment Anything (SA-1B) dataset.
None values for any features. TFDS has notfds.features.Optional, so None values are converted to default values.0 and 0.0 for int and float. Now, it's-inf as defined by NumPy (e.g., np.iinfo(np.int32).min ornp.finfo(np.float32).min). This avoids ambiguous values when 0 and 0.0tfds.features.Optional.[Experimental] A list of freeform text tags can now be attached to a BuilderConfig . For example:
BuilderConfig. For example:
BUILDER_CONFIGS = [
tfds.core.BuilderConfig(name="foo", tags=["foo", "live"]),
tfds.core.BuilderConfig(name="bar", tags=["bar", "old"]),
]builder.info.config_tags # ["foo", "live"]The installation on macOS now works (see issues 4805 and 4852 ). The ArrayRecord dependency is lazily loaded, so the TensorFlow-less path is not possi
Native support for JAX and PyTorch. TensorFlow is no longer a dependency for reading datasets. See the documentation.
tensorflow=2.12.Flag ignore_verifications from Hugging Face's datasets.load_dataset is deprecated, and used to cause errors in tfds.load(huggingface:foo).
ignore_verifications from Hugging Face's datasets.load_dataset is
deprecated, and used to cause errors in tfds.load(huggingface:foo).Python 3.7 support: this is the last version of TFDS supporting Python 3.7. Future versions will use Python 3.8.
tfds new and tfds build better support the new recommended datasets
organization, where individual datasets have their own package under
datasets/, builder class is called Builder and is defined within module
${dsname}_dataset_builder.py.Added file valid_tags.txt to not break builds.
valid_tags.txt to not break builds.tf.bool: np.bool_tf.string: np.str_tf.int64, tf.int32, etc: np.int64, np.int32, etctf.float64, tf.float32, etc: np.float64, np.float32, etc[API] DatasetBuilder's description and citations can be specified in dedicated README.md and CITATIONS.bib files, within the dataset package (see http
DatasetBuilder's description and citations can be specified in
dedicated README.md and CITATIONS.bib files, within the dataset package
(see https://www.tensorflow.org/datasets/add_dataset).TAGS.txt file. For
now, they are only used in the generated documentation.ViewBuilder to define datasets as transformations
of existing datasets. Also adds tfds.transform with functionality to apply
transformations.tfds.as_numpy(...), base Logger class has a
new corresponding method.tfds.core.DatasetBuilder can have a default limit for the number of
simultaneous downloads. tfds.download.DownloadConfig can override it.tfds.features.Audio supports storing raw audio data for lazy decoding.builder.download_and_prepare(download_config=tfds.download.DownloadConfig(num_shards=42)).
Alternatively, you can configure the min and max shard size if you want TFDS
to compute the number of shards for you, but want to have control over the
shard sizes.[API] Added TfDataBuilder that is handy for storing experimental ad hoc TFDS datasets in notebook-like environments such that they can be versioned, d
tfds.beam.inc_counter to reduce beam.metrics.Metrics.counter boilerplatetfds build my_dataset --config='{"name": "my_custom_config", "description": "Abc"}')20220620 snapshot.Logger class expects more information to be passed to the as_dataset method. This should only be relevant to people who have implemented and registered custom Logger class(es).DEFAULT_BUILDER_CONFIG_NAME in a DatasetBuilder to change the default config if it shouldn't be the first builder config defined in BUILDER_CONFIGS.~) directory, a new ~ directory is not created in the current directory (fixes #4117).Support for community datasets on GCS.
tfds.builder_from_directory and tfds.builder_from_directories, see
https://www.tensorflow.org/datasets/external_tfrecord#directly_from_folder.file_format argument to download_and_prepare method, allowing user
to specify an alternative file format to store prepared data (e.g. "riegeli").file_format to DatasetInfo string representation.tf_example_spec to public.doc kwarg on Features, to describe a feature.num_shards is now optional in the shard name.etils.epath (see
https://github.com/google/etils).DatasetNotFoundError messages.deterministic on a global level but locally in interleave, so it
only apply to interleave and not all transformations.As always, thank you to all contributors!
Release notes: * Fix import bug on Windows (#3709) * Updated documentation
Release notes:
split=tfds.split_for_jax_process('train') (alias of
tfds.even_splits('train', n=jax.process_count())[jax.process_index()]).Add split=tfds.split_for_jax_process('train') (alias of tfds.even_splits('train', n=jax.process_count())[jax.process_index()])
Release notes:
split=tfds.split_for_jax_process('train') (alias of tfds.even_splits('train', n=jax.process_count())[jax.process_index()])This is the last version of TFDS supporting 3.6. Future version will use 3.7
This is the last version of TFDS supporting 3.6. Future version will use 3.7
Better split API:
split='train[3shard]'split='train[:500_000]'split='all'tfds.even_splits is more precise and flexible:
tfds.even_splits('train', n=3, drop_remainder=True)tfds.even_splits('train[:75%]', n=3) or even nestedtfds.even_splits('train', n=3)[0] + 'test'FeatureConnectors:
serialize_example / deserialize_example methods to encode/decode example to proto: example_bytes = features.serialize_example(example_data)Audio now supports encoding='zlib' for better compressionBetter testing:
Documentation update:
RLDS:
Misc:
tfds build --file_format=tfrecordtfds.typingtfds.ReadConfig has a new assert_cardinality=False to disable cardinality.release_notesAnd of course, new datasets, bug fixes,...
Thank you to all our contributors for improving TFDS!
split='train[3shard]'.split='train[:500_000]'.split='all'.tfds.even_splits
is more precise and flexible:tfds.even_splits('train', n=3, drop_remainder=True).tfds.even_splits('train[:75%]', n=3) or even
nested.tfds.even_splits('train', n=3)[0] + 'test'.serialize_example / deserialize_example methods on features to
encode/decode example to proto: example_bytes = features.serialize_example(example_data).Audio feature now supports encoding='zlib' for better compression.tfds build --file_format=tfrecord.tfds.typing.tfds.ReadConfig has a new assert_cardinality=False argument to
disable cardinality.tfds.display_progress_bar(True) for functional control..release_notes.Add `PartialDecoding` support, to decode only a subset of the features (for performances)
API:
PartialDecoding support, to decode only a subset of the features (for performances)tfds.as_numpy supports datasets with Nonedisable_shuffling=True are now read in generation order.tfds.features.FeatureConnectortfds.testing.mock_data now supports
tf.stringbuilder_from_files and path-based community datasetstfds.builder(..., file_format=)).Dataset creation:
tfds.features.LabeledImage for semantic segmentation (like image but with additional info.features['image_label'].name label metadata)tfds.features.Image (e.g. for depth map)None dimension anywhere (previously restricted to the first position).tfds.features.Tensor() can have arbitrary number of dynamic dimension (Tensor(..., shape=(None, None, 3, None)))tfds.features.Tensor can now be serialised as bytes, instead of float/int values (to allow better compression): Tensor(..., encoding='zlib')Thank you all for your support and contribution!
PartialDecoding support,
to decode only a subset of the features (for performances).tfds.features.LabeledImage for semantic segmentation (like image but
with additional info.features['image_label'].name label metadata).tfds.features.Image (e.g. for depth map).tfds.features.FeatureConnector.None dimension anywhere
(previously restricted to the first position).tfds.features.Tensor() can have arbitrary number of dynamic
dimension (Tensor(..., shape=(None, None, 3, None))).tfds.features.Tensor can now be serialised as bytes, instead of
float/int values (to allow better compression): Tensor(..., encoding='zlib').None in tfds.as_numpy.tfds.testing.mock_data now supports:
tf.string;builder_from_files and path-based community datasets.disable_shuffling=True are now read in
generation order.tfds.builder(..., file_format=)).Add dataset.info.splits['train'].num_shards to expose the number of shards to the user
API:
• Add dataset.info.splits['train'].num_shards to expose the number of shards to the user
• Add tfds.features.Dataset to have a field containing sub-datasets (e.g. used in RL datasets)
• Add dtype and tf.uint16 supports for tfds.features.Video
• Add DatasetInfo.license field to add redistributing information
• Better tfds.benchmark(ds) (compatible with any iterator, not just tf.data, better colab representation)
Other
• Faster tfds.as_numpy() (avoid extra tf.Tensor <> np.array copy)
• Better tfds.as_dataframe visualisation (Sequence, ragged tensor, semantic masks with use_colormap)
• (experimental) community datasets support. To allow dynamically import datasets defined outside the TFDS repository.
• (experimental) Add a hugging-face compatibility wrapper to use Hugging-face datasets directly in TFDS.
• (experimental) Riegelli format support
• (experimental) Add DatasetInfo.disable_shuffling to force examples to be read in generation order.
• Add .copy, .format methods to GPath objects
• Many bug fixes
Testing:
• Supports custom BuilderConfig in DatasetBuilderTest
• DatasetBuilderTest now has a dummy_data class property which can be used in setUpClass
• Add add_tfds_id and cardinality support to tfds.testing.mock_data
And of course, many new datasets and datasets updates.
We would like to thank all the TFDS contributors!
dataset.info.splits['train'].num_shards to expose the number of
shards to the user.tfds.features.Dataset to have a field containing sub-datasets (e.g.
used in RL datasets).tf.uint16 support in tfds.features.Video.DatasetInfo.license field to add redistributing information..copy, .format methods to GPath objects.tfds.benchmark(ds) (compatible with any iterator, not just
tf.data, better colab representation).tfds.as_numpy() (avoid extra tf.Tensor <>
np.array copy).BuilderConfig in DatasetBuilderTest.DatasetBuilderTest now has a dummy_data class property which
can be used in setUpClass.add_tfds_id and cardinality support to tfds.testing.mock_data.tfds.as_dataframe visualisation (Sequence, ragged
tensor, semantic masks with use_colormap).DatasetInfo.disable_shuffling to force examples to be read
in generation order.Add tfds build to the CLI. See documentation.
API:
tfds build to the CLI. See documentation.tfds.as_numpy are compatible with len(ds)tfds.features.Dataset to represent nested datasetstfds.ReadConfig(add_tfds_id=True) to add a unique id to the example ex['tfds_id'] (e.g. b'train.tfrecord-00012-of-01024__123')num_parallel_calls option to tfds.ReadConfig to overwrite to default AUTOTUNE optiontfds.ImageFolder now support tfds.decode.SkipDecodertfds.features.Audiotfds.as_dataframe visualization (ffmpeg video if installed, bounding boxes,...)try_gcs to tfds.builder(..., try_gcs=True)BuilderConfig definition: class VERSION and RELEASE_NOTES are applied to all BuilderConfig. Config description is now optional.Breaking compatibility changes:
multi_nli/plain_text -> multi_nli.str, bytes and int). New errors likely indicates an issue in the dataset implementation.tfds.core.benchmark now returns a pd.DataFrame (instead of a dict)tfds.units is not visible anymore from the public APIBug fixes:
max_examples_per_splits=0 in tfds build --max_examples_per_splits=0 to test _split_generators only (without _generate_examples).And of course, many new datasets and datasets updates.
Thank you the community for their many valuable contributions and to supporting us in this project!!!
tfds build to the CLI. See
documentation.tfds.features.Dataset to represent nested datasets.tfds.ReadConfig(add_tfds_id=True) to add a unique id to the example
ex['tfds_id'] (e.g. b'train.tfrecord-00012-of-01024__123').num_parallel_calls option to tfds.ReadConfig to overwrite to
default AUTOTUNE option.tfds.ImageFolder support for tfds.decode.SkipDecoder.tfds.features.Audio.try_gcs to tfds.builder(..., try_gcs=True)tfds.as_dataframe visualization (ffmpeg video if installed,
bounding boxes,...).max_examples_per_splits=0 in tfds build --max_examples_per_splits=0 to test _split_generators only (without
_generate_examples).BuilderConfig definition: class VERSION and
RELEASE_NOTES are applied to all BuilderConfig. Config description is
now optional.str, bytes and int). New
errors likely indicates an issue in the dataset implementation.tfds.core.benchmark now returns a pd.DataFrame (instead of a
dict).tfds.units is not visible anymore from the public API.multi_nli/plain_text -> multi_nli.tfds.as_numpy are compatible with len(ds).tfds.core.SplitGenerator, tfds.core.BeamBasedBuilder are deprecated and will be removed in future version.
When generating a dataset, if download fails for any reason, it is now possible to manually download the data. See doc.
Simplification of the dataset creation API.
_split_generators should now returns {'split_name': self._generate_examples(), ...} (but current datasets are backward compatible).tfds.core.GeneratorBasedBuilder. Converting a dataset to beam now only require changing _generate_examples (see example and doc).tfds.core.SplitGenerator, tfds.core.BeamBasedBuilder are deprecated and will be removed in future version.Better pathlib.Path, os.PathLike compatibility:
dl_manager.manual_dir now returns a pathlib-Like object. Example:text = (dl_manager.manual_dir / 'downloaded-text.txt').read_text()
dl_manager.download, .extract,... will return pathlib-like objects in future versionsFeatureConnector,... and most functions should accept PathLike objects. Let us know if some functions you need are missing.tfds.core.as_path to create pathlib.Path-like objects compatible with GCS (e.g. tfds.core.as_path('gs://my-bucket/labels.csv').read_text()).Other bug fixes and improvement. E.g.
verify_ssl= option to tfds.download.DownloadConfig to disable SSH certificate during download.BuilderConfig are now compatible with Beam datasets #2348--record_checksums now assume the new dataset-as-folder modeltfds.features.Images can accept encoded bytes images directly (useful when used with img_name, img_bytes = dl_manager.iter_archive('images.zip')).imagenet2012 with only a single split (e.g. only the validation data). Other split will be skipped if not present.And of course new datasets
Thank you to all our contributors for improving TFDS!
tfds.core.as_path to create pathlib.Path-like objects compatible with GCS
(e.g. tfds.core.as_path('gs://my-bucket/labels.csv').read_text()).verify_ssl= option to tfds.download.DownloadConfig to disable SSH
certificate during download.tfds.core.GeneratorBasedBuilder. Converting a
dataset to beam now only require changing _generate_examples (see
example and doc)._split_generators should now returns {'split_name': self._generate_examples(), ...} (but current datasets are backward
compatible).pathlib.Path, os.PathLike compatibility: dl_manager.manual_dir
now returns a pathlib-Like object. Example: python text = (dl_manager.manual_dir / 'downloaded-text.txt').read_text() Note: Other
dl_manager.download, .extract,... will return pathlib-like objects in
future versions. FeatureConnector,... and most functions should accept
PathLike objects. Let us know if some functions you need are missing.--record_checksums now assume the new dataset-as-folder model.tfds.core.SplitGenerator, tfds.core.BeamBasedBuilder are deprecated and
will be removed in a future version.BuilderConfig are now compatible with Beam datasets #2348tfds.features.Images can accept encoded bytes images directly (useful
when used with img_name, img_bytes = dl_manager.iter_archive('images.zip')).imagenet2012 with only a single split (e.g. only the
validation data). Other split will be skipped if not present.Fix tfds.load when generation code isn't present
tfds.load when generation code isn't presentThanks @carlthome for reporting and fixing the issue.
tfds.load when generation code isn't present.Rename tfds.features.text.Xyz -> tfds.deprecated.text.Xyz
API changes, new features:
tfds.load can now load dataset without using the generation class. So tfds.load('my_dataset:1.0.0') can work even if MyDataset.VERSION == '2.0.0' (See #2493).tfds.testing.mock_data does not require metadata files anymore!tfds.as_dataframe(ds, ds_info) with custom visualisation (example)tfds.even_splits to generate subsplits (e.g. tfds.even_splits('train', n=3) == ['train[0%:33%]', 'train[33%:67%]', ...]DatasetBuilder.RELEASE_NOTES propertytfds.ImageFolder now supports custom shape, dtypeMyDataset.url_infosskip_prefetch option to tfds.ReadConfigas_supervised=True support for tfds.show_examples, tfds.as_dataframeBreaking compatible changes:
tfds.as_numpy() now returns an iterable which can be iterated multiple times. To migrate next(ds) -> next(iter(ds))tfds.features.text.Xyz -> tfds.deprecated.text.XyzDatasetBuilder.IN_DEVELOPMENT propertytfds.core.disallow_positional_args (should use Py3 *, instead)FeatureConnector.to_json_content to support this feature.Other bug fixes:
dl_manager.download_and_extracttfds.__version__ in TFDS nightly to be PEP440 compliantAnd of course, new datasets, datasets updates.
A gigantic thanks to our community which has helped us debugging issues and with the implementation of many features, especially vijayphoenix@ for being a major contributor.
tfds.load can now load dataset without using the generation class. So
tfds.load('my_dataset:1.0.0') can work even if MyDataset.VERSION == '2.0.0' (See #2493).tfds.testing.mock_data does not require metadata files anymore!tfds.as_dataframe(ds, ds_info) with custom visualisation
(example).tfds.even_splits to generate subsplits (e.g. tfds.even_splits('train', n=3) == ['train[0%:33%]', 'train[33%:67%]', ...].DatasetBuilder.RELEASE_NOTES property.tfds.features.Image now supports PNG with 4-channels.tfds.ImageFolder now supports custom shape, dtype.MyDataset.url_infos.skip_prefetch option to tfds.ReadConfig.as_supervised=True support for tfds.show_examples, tfds.as_dataframe.FeatureConnector.to_json_content to support this feature.tfds.as_numpy() now returns an iterable which can be iterated multiple
times. To migrate: next(ds) -> next(iter(ds)).tfds.features.text.Xyz -> tfds.deprecated.text.Xyz.DatasetBuilder.IN_DEVELOPMENT property.tfds.core.disallow_positional_args (should use Py3 *, instead).dl_manager.download_and_extract.tfds.__version__ in TFDS nightly to be PEP440 compliantFix an issue with GCS on Windows.
The tfds.features.text encoding API is deprecated. Please use tensorflow_text instead.
Future breaking change:
tfds.features.text encoding API is deprecated. Please use tensorflow_text instead.New features
API:
tfds.ImageFolder and tfds.TranslateFolder to easily create custom datasets with your custom data.tfds.ReadConfig(input_context=) to shard dataset, for better multi-worker compatibility (#1426).data_dir can be controlled by the TFDS_DATA_DIR environment variable.tfds.show_statistics(ds_info) to display FACETS OVERVIEW. Note: This require the dataset to have been generated with the statistics.Documentation:
tfds-nightly <span class="material-icons">nights_stay</span>Breaking compatibility change:
tfds.load('image_label_folder') in favor of the more user-friendly tfds.ImageFolderOther:
__slot__, fix parallelisation bug in tf.data.TFRecordReader,...)Thanks to all our contributors who help improving the state of dataset for the entire research community!
tfds.ImageFolder and tfds.TranslateFolder to easily create custom
datasets with your custom data.tfds.ReadConfig(input_context=) to shard dataset, for better
multi-worker compatibility (#1426).data_dir can be controlled by the TFDS_DATA_DIR
environment variable.tfds-nightly
<span class="material-icons">nights_stay</span>.tfds.show_statistics(ds_info) to display
FACETS OVERVIEW. Note: This require
the dataset to have been generated with the statistics.tfds.features.text encoding API. Please use
tensorflow_text
instead.tfds.load('image_label_folder') in favor of the more user-friendly
tfds.ImageFolder.__slot__, fix parallelisation bug in tf.data.TFRecordReader, ...).Invert ds, ds_info argument orders for tfds.show_examplesFuture breaking change:
Beaking compatibility change:
tfds.core.NamedSplit, tfds.core.SplitBase -> tfds.Split. Now tfds.Split.TRAIN,... are instance of tfds.Splitnum_shards argument from tfds.core.SplitGenerator. This argument was ignored as shards are automatically computed.Future breaking compatibility changes:
interleave_parallel_reads -> interleave_cycle_length for tfds.ReadConfig.tfds.show_examplesFuture breaking change:tfds.features.text encoding API is deprecated. Please use tensorflow_text instead.Other changes:
tfds.testing.mock_datatfds-nightlytfds.builder_cls(name) to access a DatasetBuilder class by nameinfo.split['train'].filenames for access to the tf-record files.tfds.core.add_data_dir to register an additional data dirds.with_options which where applied by TFDS. Now use tf.data default.Thank you all for your contributions, and helping us make TFDS better for everyone!
tfds.builder_cls(name) to access a DatasetBuilder class by nameinfo.split['train'].filenames for access to the tf-record files.tfds.core.add_data_dir to register an additional data dir.tfds.testing.mock_data.tfds-nightly.tfds.core.NamedSplit, tfds.core.SplitBase -> tfds.Split. Now
tfds.Split.TRAIN,... are instance of tfds.Split.interleave_parallel_reads -> interleave_cycle_length for
tfds.ReadConfig.tfds.show_examples.tfds.features.text encoding API. Please use tensorflow_text instead.num_shards argument from tfds.core.SplitGenerator. This argument was
ignored as shards are automatically computed.ds.with_options which where applied by TFDS. Now use tf.data
default.The tfds.features.text encoding API is deprecated. Please use tensorflow_text instead.
Breaking changes:
tfds.experiment.S3 has been removedimage_classification section. Some datasets have been move there from images.in_memory argument has been removed from as_dataset/tfds.load (small datasets are now auto-cached).DownloadConfig do not append the dataset name anymore (manual data should be in <manual_dir>/ instead of <manual_dir>/<dataset_name>/)dl_manager.download urls has registered checksums. To opt-out, add SKIP_CHECKSUMS = True to your DatasetBuilderTestCase.tfds.load now always returns tf.compat.v2.Dataset. If you're using still using tf.compat.v1:
tf.compat.v1.data.make_one_shot_iterator(ds) rather than ds.make_one_shot_iterator()isinstance(ds, tf.compat.v2.Dataset) instead of isinstance(ds, tf.data.Dataset)tfds.Split.ALL has been removed from the API.Future breaking change:
tfds.features.text encoding API is deprecated. Please use tensorflow_text instead.num_shards argument of tfds.core.SplitGenerator is currently ignored and will be removed in the next version.Features:
DownloadManager is now pickable (can be used inside Beam pipelines)tfds.features.Audio:
info.features['audio'].sample_rateThank you to all our contributors for helping us make TFDS better for everyone!
DownloadManager is now pickable (can be used inside Beam pipelines).tfds.features.Audio:
info.features['audio'].sample_rate.image_classification section. Some datasets have been move there from
images.DownloadConfig does not append the dataset name anymore (manual data
should be in <manual_dir>/ instead of <manual_dir>/<dataset_name>/).dl_manager.download urls has registered
checksums. To opt-out, add SKIP_CHECKSUMS = True to your
DatasetBuilderTestCase.tfds.load now always returns tf.compat.v2.Dataset. If you're using still
using tf.compat.v1:
tf.compat.v1.data.make_one_shot_iterator(ds) rather than
ds.make_one_shot_iterator().isinstance(ds, tf.compat.v2.Dataset) instead of isinstance(ds, tf.data.Dataset).tfds.features.text encoding API is deprecated. Please use
tensorflow_text
instead.num_shards argument of tfds.core.SplitGenerator is currently ignored and
will be removed in the next version.tfds.experiment.S3 has been removedin_memory argument has been removed from as_dataset/tfds.load (small
datasets are now auto-cached).tfds.Split.ALL.Auto-caching small datasets. in_memory argument is deprecated and will be removed in a future version.
New features:
info.dataset_size and info.download_size. All datasets generated with 2.1.0 cannot be loaded with previous version (previous datasets can be read with 2.1.0 however).in_memory argument is deprecated and will be removed in a future version.num_examples = tf.data.experimental.cardinality(ds) (Requires tf-nightly or TF >= 2.2.0)info.splits['train[70%:]'].num_examplesinfo.dataset_size and info.download_size.num_examples = tf.data.experimental.cardinality(ds) (Requires tf-nightly or TF >= 2.2.0)info.splits['train[70%:]'].num_examples2.1.0 however).in_memory argument is deprecated and will be removed in a future version.The previous split API is still available, but is deprecated. If you wrote DatasetBuilders outside the TFDS repository, please make sure they do not u…
DatasetBuilders outside the TFDS repository, please make sure they do not use experiments={tfds.core.Experiment.S3: False}. This will be removed in the next version, as well as the num_shards kwargs from SplitGenerator.shuffle_files defaults to False so that dataset iteration is deterministic by default. You can customize the reading pipeline, including shuffling and interleaving, through the new read_config parameter in tfds.load.urls kwargs renamed homepage in DatasetInfotfds.features.Sequence and tf.RaggedTensorFeatureConnectors can override the decode_batch_example method for efficient decoding when wrapped inside a tfds.features.Sequence(my_connector)tfds.core.BeamMetadataDict to store additional metadata computed as part of the Beam pipeline._split_generators accepts an additional pipeline kwargs to define a pipeline shared between all splits.Nothing published for this version
Nothing published for this version
Bug fixes and performance improvements.
Bug fixes and performance improvements.
Add shuffle_files argument to tfds.load function. The semantic is the same as in builder.as_dataset function, which for now means that by default, fil
shuffle_files argument to tfds.load function. The semantic is the same as in builder.as_dataset function, which for now means that by default, files will be shuffled for TRAIN split, and not for other splits. Default behaviour will change to always be False at next release.Add in_memory option to cache small dataset in RAM.
in_memory option to cache small dataset in RAM.tfds.core.DatasetInfo
which will be stored/restored with the dataset. See tfds.core.Metadata.decoders kwargs to override the default feature decoding
(guide).More datasets added:
Add direct GCS access for MNIST (with tfds.load('mnist', try_gcs=True))
tfds.load('mnist', try_gcs=True))tfds.disable_progress_bar())Thanks to all external contributors for raising issues, their feedback and their pull request.
tfds.load('mnist', try_gcs=True)).tfds.disable_progress_bar()).Fixes bug #52 that was putting the process in Eager mode by default
celeb_a_hqYour coding agent can read these notes before it upgrades. Set up the MCP server →