NewYour coding agent can read the release notes before it upgrades.Set up the MCP server →
PyPI · #4428 most downloaded on PyPI
Grain: A library for loading and transforming data for ML training.
Last release 3 months ago
17 Jun 2026
Release timing varies
gaps range from 2 weeks to 3 months
Some releases are documented
notes for 10 of 18 stable releases
1 version withdrawn
withdrawn after publishing
3 years old
20 releases · first in 2024
Added PyGrain telemetry metrics (bytes read, read latency, and prefetch depth) for OSS via Prometheus.
InterleaveDatasetIterator now overrides start_prefetch() such that it indeed starts prefetching (it did not do anything before).New features:
Bug fixes:
InterleaveDatasetIterator now overrides start_prefetch() such that it
indeed starts prefetching (it did not do anything before).Adds automated batching into shared memory when multiprocess prefetch is used to save one data copy.
One column per month.
New features:
grain_enable_multiprocess_worker_profiling=true and add "profile_subprocesses" = True in advanced profiler options.Breaking changes:
Bug fixes:
New features:
ShapeDtypeStructProtocol and ShapeDtypeStruct to
represent dataset element specs.grain_enable_multiprocess_worker_profiling=true and add
"profile_subprocesses" = True in advanced profiler options.Breaking changes:
Deprecations:
Bug fixes:
Deprecates grain.python.experimental.MultiprocessPrefetchIterDataset , use the graduated version instead: grain.IterDataset.mp_prefetch .
New features:
get_next_index and set_next_index to fetchgrain.DatasetIterator to the given produced element index.IterDataset.mp_prefetch when free-threaded Python is detected.grain.DataLoaderIterator can now asynchronously start processing elementsstart_prefetch call.SharedMemoryArrayMetadata in a public API as a metadata descriptor for SharedMemoryArray.Breaking changes:
RandomAccessDataSource should accept int__getitem__. Legacy paths that handle SupportsIndex will stillgrain.RandomAccessDataSource and callsuper().__getitem__ with Supportsindex you may see a type checkingint to fix it.Deprecations:
grain.python.experimental.MultiprocessPrefetchIterDataset,grain.IterDataset.mp_prefetch.grain.python.experimental.ConcatenateMapDataset, use thegrain.MapDataset.concatenate.Bug fixes:
WindowShuffleIterDataset checkpointing.New features:
get_next_index and set_next_index to fetch
and advance a grain.DatasetIterator to the given produced element index.IterDataset.mp_prefetch when free-threaded Python is detected.grain.DataLoaderIterator can now asynchronously start processing elements
in background with start_prefetch call.SharedMemoryArrayMetadata in a public API as a metadata descriptor
for SharedMemoryArray.ParquetIterDataset can read from multiple string paths interleaving reads.ElasticIterDatasetIterator for scaling up and down the number of shards between checkpoints.Breaking changes:
RandomAccessDataSource should accept int
index in __getitem__. Legacy paths that handle SupportsIndex will still
work at runtime, but depending on the type checker in use, if you're
directly inheriting from grain.RandomAccessDataSource and call
super().__getitem__ with Supportsindex you may see a type checking
error. Switch to int to fix it.Deprecations:
grain.python.experimental.MultiprocessPrefetchIterDataset,
use the graduated version instead: grain.IterDataset.mp_prefetch.grain.python.experimental.ConcatenateMapDataset, use the
graduated version instead: grain.MapDataset.concatenate.Bug fixes:
WindowShuffleIterDataset checkpointing.Adds public DatasetIterator.close API as an explicit blocking alternative to a cleanup during GC.
DatasetIterator.close API as an explicit blocking alternativeDatasetIterator.start_prefetch now propagates to the first asynchronousNotImplemented. This API is useful forgrain.experimental.multithread_prefetch as an{Map|Iter}Dataset elementIterDataset.mix components and weights after a checkpoint.Deprecates Python 3.10 support.
New features
dm-tree dependency with pure Python implementation for Pytreejax is not installed. If jaxjax.tree_util instead.Breaking changes:
grain[testing] PyPi build. It is an implementation detail andmanylinux_2_28.Deprecations:
grain.python.experimental.visualize_dataset. Use visualizationGraduates grain.experimental.apply_transformations to grain.{MapDataset|IterDataset}.apply . The experimental API will soon be deprecated.
New features
reseed_each_epoch option to MapDataset.repeat that allows to replaygrain.experimental.RebatchIterDataset for efficient rebatch.max_sequences_per_bin to packing transformations to limit the numbergrain.experimental.RepeatIterDataset.grain.DataLoader.grain.experimental.FlatMapTransform support to grain.DataLoader.Breaking changes:
Deprecations:
grain.experimental.apply_transformations tograin.{MapDataset|IterDataset}.apply. The experimental API will soon beBug fixes
ThreadPrefetchDatasetIterator deletion.Allow passing read_kwargs to ParquetIterDataset for configuring parquet file reading.
New features:
read_kwargs to ParquetIterDataset for configuring parquetThreadPrefetchDatasetIterator now supports non-Grain iterators thatgrain.experimental.device_put() forIterDataset, finds number of processes for mp_prefetchPrefetchDatasetIterator.reader_options to ArrayRecordDataSource for configuringgrain.experimental.batch_and_pad for padding a partial batch toBreaking changes:
array_record and protobuf.New features:
read_kwargs to ParquetIterDataset for configuring parquet
file reading.ThreadPrefetchDatasetIterator now supports non-Grain iterators that
support checkpointing.grain.experimental.device_put() for
easy CPU and device prefetching.IterDataset, finds number of processes for mp_prefetch
and buffer size for PrefetchDatasetIterator.reader_options to ArrayRecordDataSource for configuring
array record file reading.grain.experimental.batch_and_pad for padding a partial batch to
avoid dropping batch remainder data.MultiprocessPrefetchIterDataset. New slicing allows each worker process to
read unique file shards and thus improving performance.Breaking changes:
array_record and protobuf.Deprecations:
Bug fixes
Automatic publishing releases to PyPI via GitHub actions.
Disable PyPi attestations to make the upload work with nested workflow.
Disable PyPi attestations to make the upload work with nested workflow.
PiperOrigin-RevId: 772107796
Nothing published for this version
Grain 0.2.9 release. PiperOrigin-RevId: 753168774
Grain 0.2.9 release.
PiperOrigin-RevId: 753168774
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Your coding agent can read these notes before it upgrades. Set up the MCP server →