NewYour coding agent can read the release notes before it upgrades.Set up the MCP server →
PyPI · #4169 most downloaded on PyPI
Streaming lets users create PyTorch compatible datasets that can be streamed from cloud-based object stores
Last release 1 years ago
15 Jul 2025
Ships fairly regularly
a new release about every 5 weeks
Nearly every release is documented
notes for 30 of 30 stable releases
2 versions withdrawn
withdrawn after publishing
4 years old
35 releases · first in 2022
One column per quarter.
Fix typing by @dakinggg in https://github.com/mosaicml/streaming/pull/907
Full Changelog: https://github.com/mosaicml/streaming/compare/v0.12.0...v0.13.0
We've added support for Python 3.12 and deprecated Python 3.9 support.
We've added support for Python 3.12 and deprecated Python 3.9 support.
Full Changelog: https://github.com/mosaicml/streaming/compare/v0.11.0...v0.12.0
Streaming v0.11.0 is released! Install via pip:
Streaming v0.11.0 is released! Install via pip:
pip install --upgrade mosaicml-streaming==0.11.0
StreamingDataset can now be used with custom Stream implementations via a registry. See the documentation page for example usage.
simulation module import paths (@srstevenson)S3Downloader serialization issues (@wouterzwerink)simulation module by @srstevenson in https://github.com/mosaicml/streaming/pull/838epoch_seed_change attribute on SimulationDataset by @srstevenson in https://github.com/mosaicml/streaming/pull/840Full Changelog: https://github.com/mosaicml/streaming/compare/v0.10.0...v0.11.0
The py1b shuffle algorithm has now been deprecated. Please use the improved py1e (default) or the py1br shuffle algorithms instead.
Streaming v0.10.0 is released! Install via pip:
pip install --upgrade mosaicml-streaming==0.10.0
py1b shuffle algorithm deprecation (https://github.com/mosaicml/streaming/pull/837)py1b shuffle algorithm has now been deprecated. Please use the improved py1e (default) or the py1br shuffle algorithms instead.Full Changelog: https://github.com/mosaicml/streaming/compare/v0.9.1...v0.10.0
Nothing published for this version
Streaming v0.9.1 is released! Install via pip:
Streaming v0.9.1 is released! Install via pip:
pip install --upgrade mosaicml-streaming==0.9.1
Full Changelog: https://github.com/mosaicml/streaming/compare/v0.9.0...v0.9.1
Streaming v0.9.0 is released! Install via pip:
Streaming v0.9.0 is released! Install via pip:
pip install --upgrade mosaicml-streaming==0.9.0
It is now possible to have columns including a map type successfully convert to JSON in an MDS file if the given type for the column is specified as 'json', and allows the JSON encoder to handle ndarray types.
Full Changelog: https://github.com/mosaicml/streaming/compare/v0.8.1...v0.9.0
Patching conf.py due to Sphinx deprecating config manipulation by @snarayan21 in https://github.com/mosaicml/streaming/pull/746
Streaming v0.8.1 is released! Install via pip:
pip install --upgrade mosaicml-streaming==0.8.1
Dataloader hanging between epochs has now been resolved! We've seen training time improvements of up to 40% for some many-epoch training jobs. If this was impacting your runs and has now been fixed, please let us know!
shuffle=False by @snarayan21 in https://github.com/mosaicml/streaming/pull/750Full Changelog: https://github.com/mosaicml/streaming/compare/v0.8.0...v0.8.1
Streaming now supports streaming data from HF file system! This adds another popular backend as an option to host your data.
Streaming now supports streaming data from HF file system! This adds another popular backend as an option to host your data.
py1e for improperly written datasets by @snarayan21 in https://github.com/mosaicml/streaming/pull/673batch_size typo for Stream object in docs by @snarayan21 in https://github.com/mosaicml/streaming/pull/682remote is specified by @snarayan21 in https://github.com/mosaicml/streaming/pull/683replication for World object by @snarayan21 in https://github.com/mosaicml/streaming/pull/685Spanner object instead of ValueError by @snarayan21 in https://github.com/mosaicml/streaming/pull/701drop_first checking in partitioning to account for world_size divisibility by @snarayan21 in https://github.com/mosaicml/streaming/pull/706dbfs: prefix from error message by @vanshcsingh in https://github.com/mosaicml/streaming/pull/712Full Changelog: https://github.com/mosaicml/streaming/compare/v0.7.6...v0.8.0
Streaming v0.7.6 is released! Install via pip:
Streaming v0.7.6 is released! Install via pip:
pip install --upgrade mosaicml-streaming==0.7.6
device_per_stream batching methodUsers can now construct batches such that each device sees only samples from a single stream. This is very useful in cases where different data sources have samples/tensors of different sizes, but the model should still see samples from these different data sources at each optimizer step.
device_per_stream batching by @snarayan21 in https://github.com/mosaicml/streaming/pull/661ndarray type for Spark dataframes.Enable parsing Spark's ArrayType (of ShortType, LongType, IntegerType, FloatType, DoubleType) when converting a Spark dataframe to MDS.
Adds support for Alipan, Alibaba's cloud storage service.
Full Changelog: https://github.com/mosaicml/streaming/compare/v0.7.5...v0.7.6
Add support for Python 3.11 and deprecate Python 3.8 by @karan6181 in https://github.com/mosaicml/streaming/pull/586
Streaming v0.7.5 is released! Install via pip:
pip install --upgrade mosaicml-streaming==0.7.5
Using the replication argument, easily share data samples across multiple ranks, enabling sequence or tensor parallelism.
New and improved streaming documentation can be found here -- please submit issues with any feedback.
batch_size is now required for StreamingDatasetAs we have seen multiple errors and performance degradations from users not setting the batch_size argument to StreamingDataset, we are making it a requirement to iterate over the dataset.
allow_unsafe_types=True by @snarayan21 in https://github.com/mosaicml/streaming/pull/647Full Changelog: https://github.com/mosaicml/streaming/compare/v0.7.4...v0.7.5
Streaming v0.7.4 is released! Install via pip:
Streaming v0.7.4 is released! Install via pip:
pip install --upgrade mosaicml-streaming==0.7.4
load_state_dict multiple times. by @snarayan21 in https://github.com/mosaicml/streaming/pull/593tempfile.gettempdir() instead of a hardcoded temp root. by @knighton in https://github.com/mosaicml/streaming/pull/570load_state_dict multiple times. by @snarayan21 in https://github.com/mosaicml/streaming/pull/593Full Changelog: https://github.com/mosaicml/streaming/compare/v0.7.3...v0.7.4
Streaming v0.7.3 is released! Install via pip:
Streaming v0.7.3 is released! Install via pip:
pip install --upgrade mosaicml-streaming==0.7.3
Full Changelog: https://github.com/mosaicml/streaming/compare/v0.7.2...v0.7.3
The pickle serialization format, one of the available MDS encodings, is a potential security vulnerability. We added a boolean flag allow_unsafe_types…
Streaming v0.7.2 is released! Install via pip:
pip install --upgrade mosaicml-streaming==0.7.2
Add support for the Canned ACL using the environment variable S3_CANNED_ACL for AWS S3. Checkout Canned ACL document on how to use it.
The pickle serialization format, one of the available MDS encodings, is a potential security vulnerability. We added a boolean flag allow_unsafe_types in the StreamingDataset class to allow or reject datasets containing Pickle.
Full Changelog: https://github.com/mosaicml/streaming/compare/v0.7.1...v0.7.2
Streaming v0.7.1 is released! Install via pip:
Streaming v0.7.1 is released! Install via pip:
pip install --upgrade mosaicml-streaming==0.7.1
simulator is fixed (#499)Full Changelog: https://github.com/mosaicml/streaming/compare/v0.7.0...v0.7.1
Streaming v0.7.0 is released! Install via pip:
Streaming v0.7.0 is released! Install via pip:
pip install --upgrade mosaicml-streaming==0.7.0
StreamingDataset (#479)StreamingDataset have been updated to be more performant and are applicable for most use cases, detailed below:| Parameter | Old Value | New Value | Benefit |
|---|---|---|---|
shuffle_algo |
py1s |
py1e |
Better shuffle and balanced downloading |
num_canonical_nodes |
64 * physical nodes |
if py1s or py2s, 64 * physical_nodes, otherwise physical_nodes |
Consistently good shuffle for all shuffle algos |
shuffle_block_size |
262,144 |
4,000,000 / num_canonical_nodes |
Consistently good shuffle for all num_canonical_nodes values |
predownload |
max(batch_size, 256 * batch_size // num_canonical_nodes) |
8 * batch_size |
Better balanced downloading |
partition_algo |
orig |
relaxed |
More flexible deterministic resumptions on nodes |
simulator in your terminal to open the simulation interface.num_canonical_nodes parameter had to divide or be a multiple of the number of physical nodes for determinism.Full Changelog: https://github.com/mosaicml/streaming/compare/v0.6.1...v0.7.0
Streaming v0.6.1 is released! Install via pip:
Streaming v0.6.1 is released! Install via pip:
pip install --upgrade mosaicml-streaming==0.6.1
merge_index() utility method to merge subdirectories index files from an MDS dataset. The subdirectories can be local or any supported cloud provider URL path.retry in Writer.from streaming import MDSWriter
with MDSWriter(
...,
retry=3) as out:
for sample in dataset:
out.write(sample)
py1e shuffling algorithm by varying shard sample ranges, helping to reduce throughput drops at scale. (#442)Full Changelog: https://github.com/mosaicml/streaming/compare/v0.6.0...v0.6.1
The py1br algorithm is a replacement for the py1b algorithm, which will be deprecated soon.
Streaming v0.6.0 is released! Install via pip:
pip install --upgrade mosaicml-streaming==0.6.0
Support for reading and writing data from and to the Databricks File System (DBFS) and Unity Catalog (UC) Volumes. This means that you can now use DBFS and UC Volumes as a source or sink for your streaming data pipelines or model training. Below is the path structure:
Databricks File System (DBFS)
DBFS path structure is a hierarchical namespace that is organized into directories and files. The DBFS prefix must starts with dbfs:/.
UC Volumes
The path structure for UC Volumes is similar to the path structure for DBFS, but with a few key differences.
The root of the UC Volumes namespace is dbfs:/Volumes/<catalog>/<schema>/<volume>, where:
<catalog> is the name of the catalog where the volume is created.<schema> is the name of the schema where the volume is created.<volume> is the name of the volume.Hence, use a dbfs://Volumes prefix to specify a UC Volumes path.
Introducing the new DataFrameToMDS API, empowering users to effortlessly leverage Spark's capabilities for handling diverse datasets in various formats. This API enables seamless conversion of Spark DataFrames into MDS datasets, with the flexibility to specify output locations to both local and cloud storage. Index files are optionally merged. Additionally, users can add data preprocessing steps by defining custom iterator functions and arguments. All these features are seamlessly bundled into a single Spark job, ensuring an efficient and streamlined workflow for data transformation. An example notebook is provided to help users get started.
The new py1br shuffle algorithm helps mitigate download spikes that occur when using the py1b algorithm. With py1b, shuffle blocks are all the same size, so when progressing through training, nodes will have to download many shards at the same time. In contrast, with py1br, shuffle blocks are offset from each other and are variably sized. This results in more balanced downloads over time. The py1br algorithm is a replacement for the py1b algorithm, which will be deprecated soon.
from streaming import StreamingDataset
dataset = StreamingDataset(
shuffle_algo='py1br',
...
)
The new py1e shuffle algorithm helps reduce the minimum cache limit needed for training, and results in much smoother downloads than both py1br and py1e. However, its shuffle quality is slightly lower. Rather than shuffling all samples in blocks of size shuffle_block_size, it instead spreads the samples of each shard over a range of maximum size shuffle_block_size, retaining most of the shuffle quality from py1b and py1br while reducing download spikes across the duration of training.
from streaming import StreamingDataset
dataset = StreamingDataset(
shuffle_algo='py1e',
...
)
Users are now able to ensure that each batch comes has samples from only a single stream. You can now set the new parameter batching_method to per_stream to access this functionality. Per-stream batching will still take into account upsampling and downsampling of streams, set by proportion, repeat, or choose. To make batches contain only samples from a group of streams, merge streams’ index.json files to create a single one for each group.
from streaming import StreamingDataset
dataset = StreamingDataset(
batching_method='per_stream',
...
)
Users are now able to ensure that each batch has a consistent number of samples from every stream. Previously, stream proportions were satisfied in the aggregate but not at the batch level. You can now set the new parameter batching_method to stratified to access this functionality. Stratified batching will still take into account upsampling and downsampling of streams, set by proportion, repeat, or choose.
from streaming import StreamingDataset
dataset = StreamingDataset(
batching_method='stratified',
...
)
Previous versions of StreamingDataset implement downsampling/upsampling by giving each sample equal probability of being selected (plus or minus one due when sampling is fractional), without regard to what shard a sample is on. This means that no matter how small your desired downsampling is, StreamingDataset will still use each shard at as equal a rate as possible. This is problematic for downloading performance.
In this version of Streaming, we have added a new optional StreamingDataset argument sampling_granularity which can be used to configure how sampling is done. It is an integer, defaulting to 1, that determines how many samples are to be drawn at a time from a single random shard until we have enough samples.
Note that the default setting of 1 is equivalent to the old non-shard-aware behavior. Setting it high, e.g. the number of samples in a full shard or more, means it will draw all the samples in a randomly chosen (without replacement) shard until it has enough samples, which is much more download-effiicient but results in the samples of each shard always being seen close together in training, which may have implications to convergence depending on your workload. Setting sampling granularity to half a shard means, roughly speaking, you'll see half the samples of a shard at a time during training.
from streaming import StreamingDataset
dataset = StreamingDataset(
sampling_granularity=1,
...
)
Users can now instantiate more than one StreamingDataset with same local directory and remote=None. This would be useful if there is a high-speed storage mounted on a node and multiple folks are trying to read the dataset directly from mount storage on the same node without having to copy the data on local disk.
from streaming import StreamingDataset
local = '<local disk directory or a mount point directory>'
dataset_0 = StreamingDataset(local=local, remote=None)
dataset_1 = StreamingDataset(local=local, remote=None)
cache_limit is lower than the size of a single shard file to avoid deadlock. (#420)predownload value to zero issue where users can now provide predownload=0 in StreamingDataset. (#383)index.json exists locally before downloading to avoid duplicate downloads (#372).Full Changelog: https://github.com/mosaicml/streaming/compare/v0.5.2...v0.6.0
Streaming v0.5.2 is released! Install via pip:
Streaming v0.5.2 is released! Install via pip:
pip install --upgrade mosaicml-streaming==0.5.2
Full Changelog: https://github.com/mosaicml/streaming/compare/v0.5.1...v0.5.2
Improved shard eviction test execution time by @karan6181 in https://github.com/mosaicml/streaming/pull/291
Full Changelog: https://github.com/mosaicml/streaming/compare/v0.5.0...v0.5.1
Streaming v0.5.0 is released! Install via pip:
Streaming v0.5.0 is released! Install via pip:
pip install --upgrade mosaicml-streaming==0.5.0
Dynamically delete least recently used shards in order to keep disk usage under a specified limit. This is enabled by setting the StreamingDataset argument cache_limit. See the shuffling guide for more details.
from streaming import StreamingDataset
dataset = StreamingDataset(
cache_limit='100gb',
...
)
Users can now randomly access samples using NumPy-style indexing with StreamingDataset. For example,
import numpy as np
from streaming import StreamingDataset
dataset = StreamingDataset(local=local, remote=remote)
dataset[0] # Fetch sample 0
dataset[-1] # Fetch last sample
dataset[[10, 20]] # Fetch sample 10 and 20
dataset[slice(1, 10, 2)] # Fetch sample 1, 3, 5, 7, and 9
dataset[5:0:-1] # Fetch sample 5, 4, 3, 2, 1
dataset[np.array([4, 7])] # Fetch sample 4 and 7
Support of any S3 compatible object stores, meaning, an object store which uses the S3 API to communicate with any connected device or system. Some of the S3 compatible object stores are Cloudflare R2, Coreweave, Backblaze b2, etc. User needs to provide an environment variable S3_ENDPOINT_URL based on the object store that you are using. Details on how to configure credentials can be found here.
Support of Azure cloud blob storage. Details on how to configure credentials can be found here.
samples_per_epoch has been renamed to epoch_size in StreamingDatasetto better distinguish the actual number of underlying samples as serialized and the number of observed samples when iterating (which may be different due to weighting sub-datasets).samples has been renamed to choose in Stream to better distinguish the underlying sample vs resampled data.keep_raw has been removed in StreamingDataset in the process of finalizing the design for shard eviction (see the newly-added cache_limit parameter).predownload in StreamingDataset was updated; it is now derived using batch size and number of canonical nodes instead of previous constant value of 100_000. This is to prevent predownloaded shards from getting evicted before ever being used.num_canonical_nodes in StreamingDataset was updated to 64 times the number of nodes of the initial run instead of number of nodes of the initial run to increase data source diversity and improve convergence.shuffle_algo in StreamingDataset was changed from py1b to py1s as it requires less shards to be downloaded during iteration. More details about different shuffling algorithms can be found here.pile.py link by @ouhenio in https://github.com/mosaicml/streaming/pull/259Stream usage example to README by @hanlint in https://github.com/mosaicml/streaming/pull/266Full Changelog: https://github.com/mosaicml/streaming/compare/v0.4.1...v0.5.0
Streaming v0.4.1 is released! Install via pip:
Streaming v0.4.1 is released! Install via pip:
pip install --upgrade mosaicml-streaming==0.4.1
Full Changelog: https://github.com/mosaicml/streaming/compare/v0.4.0...v0.4.1
Streaming v0.4.0 is released! Install via pip:
Streaming v0.4.0 is released! Install via pip:
pip install --upgrade mosaicml-streaming==0.4.0
streams parameter which takes one or more sub-datasets and it intelligently fetches samples across sub-datasets. You can mix (upsample or downsample) datasets by defining each either relatively (proportion) or absolutely (repeat or samples or none of them to sample 1:1).Full Changelog: https://github.com/mosaicml/streaming/compare/v0.3.0...v0.4.0
Streaming v0.3.0 is released! Install via pip:
Streaming v0.3.0 is released! Install via pip:
pip install --upgrade mosaicml-streaming==0.3.0
Now, you can automatically upload shards to cloud storage on the fly by providing a cloud path to MDSWriter. Track the progress of individual uploads with progress_bar=True, and tune background upload workers with max_workers=4.
User can choose to upload a output shard files automatically to a supported cloud (AWS S3, GCP, OCI) by providing a out parameter as a cloud provider bucket location as part of Writer class. Below is the example to upload output files to AWS S3 bucket
output_dir = 's3://bucket/dir/path'
with MDSWriter(out=output_dir, ...) as out:
for sample in samples:
pass
User can choose to keep a output shard files locally by providing a local directory path as part of Writer. For example,
output_dir = '/tmp/mds'
with MDSWriter(out=output_dir, ...) as out:
for sample in samples:
pass
User can see the progress of the cloud upload file by setting progress_bar=True as part of Writer. For example,
output_dir = 's3://bucket/dir/path'
with MDSWriter(out=output_dir, progress_bar=True, ...) as out:
for sample in samples:
pass
User can control the number of background upload threads via parameter max_workers as part of Writer who is responsible for uploading the shard files to a remote location if provided. One thread is responsible for one file upload. For example, if max_workers=4, maximum 4 threads would be active at a same time uploading one shard file each.
output_dir = 's3://bucket/dir/path'
with MDSWriter(out=output_dir, max_workers=4, ...) as out:
for sample in samples:
pass
We’ve added a new shuffling algorithm py1s which is twice as fast on typical workloads. You can toggle which shuffling algorithm is used by overriding shuffle_algo (old behavior: py2s). You will experience this as faster epoch starts and faster mid-epoch resumption for large datasets.
We’ve also reimplemented how shards/samples are assigned to nodes/devices/dataloader workers to run about twice as fast on typical workloads while giving identical results. This is exposed as the partition_algo argument to StreamingDataset. You will experience this as faster start and resumption for large datasets.
We provide examples of modifying StreamingDataset to stream from a dataset of links to external data sources. In our examples, using the WebVid dataset, each sample points to a video file which exists outside of the shards in its original format and is downloaded separately. Benchmarking is included.
Class Writer and its derived classes (MDSWriter, XSVWriter, TSVWriter, CSVWriter, and JSONWriter) parameter has been changed from dirname to out with the following advanced functionalities:
out is a local directory, shard files are saved locally. For example, out=/tmp/mds/.out is a remote directory, a local temporary directory is created to cache the shard files and then the shard files are uploaded to a remote location. At the end, the temp directory is deleted once shards are uploaded. For example, out=s3://bucket/dir/path.out is a tuple of (local_dir, remote_dir), shard files are saved in the
local_dir and also uploaded to a remote location. For example, out=('/tmp/mds/', 's3://bucket/dir/path').Given the complexity of their arguments, and the need to be able to safely upgrade them over time, we have updated the APIs of Writer and its subclasses (like MDSWriter) and StreamingDataset to require kwargs.
Full Changelog: https://github.com/mosaicml/streaming/compare/v0.2.5...v0.3.0
Streaming v0.2.5 is released! Install via pip:
Streaming v0.2.5 is released! Install via pip:
pip install --upgrade mosaicml-streaming==0.2.5
Full Changelog: https://github.com/mosaicml/streaming/compare/v0.2.4...v0.2.5
Streaming v0.2.4 is released! Install via pip:
🚀 Streaming v0.2.4
Streaming v0.2.4 is released! Install via pip:
pip install --upgrade mosaicml-streaming==0.2.4
Full Changelog: https://github.com/mosaicml/streaming/compare/v0.2.3...v0.2.4
Streaming v0.2.3 is released! Install via pip:
Streaming v0.2.3 is released! Install via pip:
pip install --upgrade mosaicml-streaming==0.2.3
StreamingDataset[sample_id] block to download the given sample's shard if it is not present, so that the dataset can be used lazily (https://github.com/mosaicml/streaming/pull/118)Full Changelog: https://github.com/mosaicml/streaming/compare/v0.2.2...v0.2.3
Streaming v0.2.2 is released! Install via pip:
Streaming v0.2.2 is released! Install via pip:
pip install --upgrade mosaicml-streaming==0.2.2
Full Changelog: https://github.com/mosaicml/streaming/compare/v0.2.1...v0.2.2
Streaming v0.2.1 is released! Install via pip:
Streaming v0.2.1 is released! Install via pip:
pip install --upgrade mosaicml-streaming==0.2.1
Full Changelog: https://github.com/mosaicml/streaming/compare/v0.2.0...v0.2.1
Streaming v0.2.0 is released! Install via pip:
Streaming v0.2.0 is released! Install via pip:
pip install --upgrade mosaicml-streaming==0.2.0
Elastic world size deterministic shuffle
Shuffled or not, StreamingDataset now collectively traverses the samples in identical order across all the devices, given a seed and a canonical number of nodes. This ordering holds true even if you checkpoint and resume training of the same epoch on a different number of nodes.
Instant Mid-Epoch Resumption
Waiting while your data loader spins to resume from where you left off can be costly! StreamingDataset now lets you resume immediately.
NEW StreamingDataLoader
A StreamingDataLoader is a drop-in replacement for your PyTorch DataLoader with a Mid-Epoch Resumption functionality where it resumes from where you left off without spinning the dataloader.
Support for Oracle Cloud Infrastructure (OCI) blob storage
Streaming now supports OCI blob storage as a storage backend for streaming. One can pass the OCI blob storage as either oci://<bucket_name>@<namespace>/<folder_name>/<filename> or oci://<bucket_name>/<folder_name>/<filename> to a StreamingDataset class. For example:
from streaming import StreamingDataset
remote = 'oci://<bucket>@<namespace>/<path>'
local = '/tmp/dataset/'
train_dataset = StreamingDataset(local=local, remote=remote, split='train')
Streaming expects the credentials to be present in ~/.oci/config path.
Support for public AWS S3 buckets
Streaming now supports AWS S3 buckets which are public resources that can be accessed without credentials, apart from the already supported private AWS S3 buckets. One can instantiate the StreamingDataset class with an AWS S3 bucket as follows
from streaming import StreamingDataset
remote = 's3://<bucket>/<path>'
local = '/tmp/dataset/'
train_dataset = StreamingDataset(local=local, remote=remote, split='train')
Dataset has been renamed as class StreamingDataset (https://github.com/mosaicml/streaming/pull/37).
C4 renamed as StreamingC4EnWiki renamed as StreamingEnWikiPile renamed as StreamingEnWikiADE20K renamed as StreamingADE20KCIFAR10 renamed as StreamingCIFAR10COCO renamed as StreamingCOCOImageNet renamed as StreamingImageNetprefetch in class Dataset has been renamed as predownload in class StreamingDataset (https://github.com/mosaicml/streaming/pull/37).retry in class Dataset has been renamed as download_retry in class StreamingDataset (https://github.com/mosaicml/streaming/pull/37).timeout in class Dataset has been renamed as download_timeout in class StreamingDataset (https://github.com/mosaicml/streaming/pull/37).hash in class Dataset has been renamed as validate_hash in class StreamingDataset (https://github.com/mosaicml/streaming/pull/37).Full Changelog: https://github.com/mosaicml/streaming/compare/v0.1.2...v0.2.0
Streaming v0.1.2 is released! Install via pip:
Streaming v0.1.2 is released! Install via pip:
pip install --upgrade mosaicml-streaming==0.1.2
Full Changelog: https://github.com/mosaicml/streaming/compare/v0.1.1...v0.1.2
Streaming v0.1.1 is released! Install via pip:
Streaming v0.1.1 is released! Install via pip:
pip install --upgrade mosaicml-streaming==0.1.1
Full Changelog: https://github.com/mosaicml/streaming/commits/v0.1.1
Nothing published for this version
Nothing published for this version
Nothing published for this version
Your coding agent can read these notes before it upgrades. Set up the MCP server →