NewYour coding agent can read the release notes before it upgrades.Set up the MCP server →
PyPI · #3605 most downloaded on PyPI
Composable data loading modules for PyTorch
Last release 2 years ago
no release in 18 months
Release timing varies
gaps range from 5 weeks to 9 months
Nearly every release is documented
notes for 14 of 14 stable releases
1 version withdrawn
withdrawn after publishing
5 years old
16 releases · first in 2021
One column per quarter.
max_concurrent instance check and error message fix
Unbatcher node (#1416 )max_concurrent instance check and error message fix (#1420 )CYCLE_FOREVER stop criterion for MultiNodeWeightedSampler node (#1424 )generator initial seed setting in RandomSampler (#1441 )StatefulDataloader restart behavior when reloaded from a state_dict saved at end of epoch with num_workers=0 (#1439 )Full Changelog: https://github.com/pytorch/data/compare/v0.10.1...v0.11.0 @andrewkho @divyanshk @scotts @atalman
Unbatcher node (#1416 )max_concurrent instance check and error message fix (#1420 )CYCLE_FOREVER stop criterion for MultiNodeWeightedSampler node (#1424 )generator initial seed setting in RandomSampler (#1441 )StatefulDataloader restart behavior when reloaded from a state_dict saved at end of epoch with num_workers=0 (#1439 )Full Changelog: v0.10.1...v0.11.0
@andrewkho @divyanshk @scotts @atalman
PyTorch's official conda channel is deprecated . TorchData has removed its conda builds as well. TorchData will be available for installation through…
This release introduces 3 major changes:
Introducing torchdata.nodes, a library of extensible and composable iterators that lets you chain together common dataloading and pre-proc operations! This initial release includes the following features, with more on the way:
torchdata.nodes docs for more details.This release drops support for DataPipes and DataLoader2. Release v0.9 was the last stable release which includes them. Please see this issue for more details.
PyTorch's official conda channel is deprecated. TorchData has removed its conda builds as well. TorchData will be available for installation through pip, on PyPI and download.pytorch.org.
Full Changelog: v0.9.0...v0.10.1
add new torchdata.nodes doc file
add new torchdata.nodes doc file
update main readme.md, move content of nodes readme.md to torchdata.nodes.rst
main readme
udpate readme
This was a relatively small release compared to previous. This will notably be the last stable release to feature DataPipes and DataLoader2!
This was a relatively small release compared to previous. This will notably be the last stable release to feature DataPipes and DataLoader2!
Full Changelog: https://github.com/pytorch/data/commits/v0.9.0
This was a relatively small release compared to previous. This will notably be the last stable release to feature DataPipes and DataLoader2!
Full Changelog: https://github.com/pytorch/data/commits/v0.9.0
In release torchdata==0.8.0 (July 2024) they will be marked as deprecated, and in 0.9.0 (Oct 2024) they will be deleted. Existing users are advised to…
We are excited to announce the release of TorchData 0.8.0. This first release of StatefulDataLoader, which is a drop-in replacement for torch.utils.data.DataLoader, offering state_dict/load_state_dict methods for handling mid-epoch checkpointing.
⚠️ June 2024 Status Update: Removing DataPipes and DataLoader V2
We are re-focusing the torchdata repo to be an iterative enhancement of torch.utils.data.DataLoader. We do not plan on continuing development or maintaining the [DataPipes] and [DataLoaderV2] solutions, and they will be removed from the torchdata repo. We'll also be revisiting the DataPipes references in pytorch/pytorch. In release torchdata==0.8.0 (July 2024) they will be marked as deprecated, and in 0.9.0 (Oct 2024) they will be deleted. Existing users are advised to pin to torchdata==0.8.0 or an older version until they are able to migrate away. Subsequent releases will not include DataPipes or DataLoaderV2. The old version of this README is available here. Please reach out if you suggestions or comments (please use https://github.com/pytorch/data/issues/1196 for feedback).
Full Changelog: https://github.com/pytorch/data/commits/v0.8.0
Current status ⚠️ As of July 2023, we have paused active development on TorchData and have paused new releases. We have learnt a lot from building it
Current status ⚠️ As of July 2023, we have paused active development on TorchData and have paused new releases. We have learnt a lot from building it and hearing from users, but also believe we need to re-evaluate the technical design and approach given how much the industry has changed since we began the project. During the rest of 2023 we will be re-evaluating our plans in this space. Please reach out if you suggestions or comments (please use https://github.com/pytorch/data/issues/1196 for feedback).
This is a patch release, which is compatible with PyTorch 2.1.1. There are no new features added.
:warning: As of July 2023, we have paused active development on TorchData and have paused new releases. We have learnt a lot from building it and hear
:warning: As of July 2023, we have paused active development on TorchData and have paused new releases. We have learnt a lot from building it and hearing from users, but also believe we need to re-evaluate the technical design and approach given how much the industry has changed since we began the project. During the rest of 2023 we will be re-evaluating our plans in this space. Please reach out if you suggestions or comments (please use #1196 for feedback).
This minor release is aligned with PyTorch 2.0.1 and primarily fixes bugs that are introduced in the 0.6.0 release. We sincerely thank our users and c
This minor release is aligned with PyTorch 2.0.1 and primarily fixes bugs that are introduced in the 0.6.0 release. We sincerely thank our users and contributors for spotting various bugs and helping us to fix them.
seed = 0 bug (#1098)
seed = 0 was passed into DataLoader2, the seed value in DataLoader2 would not be set and the seed would be unused. This change fixes that and allow seed = 0 to be used normally.worker_init_fn to update DataPipe graph and move worker prefetch to the end of Worker pipeline (#1100)pin_memory_fn to support namedtuple (#1086)portalocker at import time (#1099)FullSync operation when world_size == 1 (#1065)setup.py for display on PyPI (#1094)This library is currently in the Beta stage and currently does not have a fully stable release. The API may change based on user feedback or performance. We are committed to bring this library to stable release, but future changes may not be completely backward compatible. If you install from source or use the nightly version of this library, use it along with the PyTorch nightly binaries. If you have suggestions on the API or use cases you'd like to be covered, please open a GitHub issue. We'd love to hear thoughts and feedback. As always, we welcome new contributors to our repo.
Remove previously deprecated FileLoaderDataPipe
We are excited to announce the release of TorchData 0.6.0. This release is composed of about 130 commits since 0.5.0, made by 27 contributors. We want to sincerely thank our community for continuously improving TorchData.
TorchData 0.6.0 updates are primarily focused on DataLoader2. We graduate some of its APIs from the prototype stage and introduce additional features. Highlights include:
MultiProcessingReadingService from prototype to beta
ReadingService that we expect most users to use; it closely aligns with the functionalities of old DataLoader with improvementsReadingServices at the same timeDataLoader2 and its subcomponentsMultiProcessingReadingService as well as the internal implementation have changed. Overall, this should provide a better user experience.<p align="center"> <table align="center"> <tr><th>0.5.0</th><th>0.6.0</th></tr> <tr valign="top"> <td><sub> It previously took the following arguments: <pre lang="python"> MultiProcessingReadingService( num_workers: int = 0, pin_memory: bool = False, timeout: float = 0, worker_init_fn: Optional[Callable[[int], None]] = None, multiprocessing_context=None, prefetch_factor: Optional[int] = None, persistent_workers: bool = False, ) </pre></sub></td> <td><sub> The new version takes these arguments: <pre lang="python"> MultiProcessingReadingService( num_workers: int = 0, multiprocessing_context: Optional[str] = None, worker_prefetch_cnt: int = 10, main_prefetch_cnt: int = 10, worker_init_fn: Optional[Callable[[DataPipe, WorkerInfo], DataPipe]] = None, worker_reset_fn: Optional[Callable[[DataPipe, WorkerInfo, SeedGenerator], DataPipe]] = None, ) </pre></sub></td> </tr> </table> </p>
DataLoader2 initialization (#746)
DataLoader2, a deep copy of the passed-in ReadingService object is created during initialization and will be subsequently used.DataLoader2s from accidentally sharing states when the same ReadingService object is passed into them.<p align="center"> <table align="center"> <tr><th>0.5.0</th><th>0.6.0</th></tr> <tr valign="top"> <td><sub> Previously, a ReadingService object that is used in multiple DataLoader2 shared state among them. <pre lang="python">
dp = IterableWrapper([0, 1, 2, 3, 4]) rs = MultiProcessingReadingService(num_workers=2) dl1 = DataLoader2(dp, reading_service=rs) dl2 = DataLoader2(dp, reading_service=rs) next(iter(dl1)) print(f"Number of processes that exist in
dl1's RS after initializingdl1: {len(dl1.reading_service._worker_processes)}")
dl1's RS after initializing dl1: 2next(iter(dl2))
dl1.read_service belowprint(f"Number of processes that exist in
dl1's RS after initializingdl2: {len(dl1.reading_service._worker_processes)}")
dl1's RS after initializing dl1: 4 </pre></sub></td>
<td><sub> DataLoader2 now deep copies the ReadingService object during initialization and the ReadingService state is no longer shared.
<pre lang="python">
dp = IterableWrapper([0, 1, 2, 3, 4]) rs = MultiProcessingReadingService(num_workers=2) dl1 = DataLoader2(dp, reading_service=rs) dl2 = DataLoader2(dp, reading_service=rs) next(iter(dl1)) print(f"Number of processes that exist in
dl1's RS after initializingdl1: {len(dl1.reading_service._worker_processes)}")
dl1's RS after initializing dl1: 2next(iter(dl2))
dl1.read_service belowprint(f"Number of processes that exist in
dl1's RS after initializingdl2: {len(dl1.reading_service._worker_processes)}")
dl1's RS after initializing dl1: 2 </pre></sub></td>
</tr>
</table> </p>
In PyTorch Core
FileLoaderDataPipe (#89794)torch.utils.data.datapipes.iter.grouping as deprecated (#94527)TorchData
worker_init_fn and worker_reset_fn to MultiProcessingReadingService (#907)reset_iterator when all loops have received reset request in the dispatching process (#994)DataLoader2 (#801)limit, pause, resume operations to halt DataPipes in DataLoader2 (#879)ShardExpander IterDataPipe (#405)RoundRobinDemux IterDataPipe (#903)PinMemory IterDataPipe (#1014)In PyTorch Core
apply_sharding to accept one sharding_filter per branch (#90769)TorchData
load_state_dict() signature to align with TorchSnapshot (#887)DistInfo and ExtraInfo (used to store distributed metadata) (#916)DataLoader2Iterator's __getattr__ method (#1004)DataLoader2 exit with them (#1003)In PyTorch Core
sharding_filter (#88424)keep_key option to Grouper (#92532)TorchData
limit=None a no-op (#908)fsspec DataPipe to be compatible with the latest version of fsspec (#957)MapKeyZipper (#1042)Slicer (#1041)In PyTorch Core
In PyTorch Core
__len__ of datapipes dynamic (#88302)Zipper (#89974)TorchData
to_graph DataPipeGraph visualization function (#872)max_token_bucketize to accept incomparable data (#883)S3FileLoader local file clobbering (#895)fsspec DataPipe for paths starting with az:// (#849)DataLoader2 training loop example (#863)DataLoader2 content (#979)DataLoader2 Tutorial (#980)DataLoader2 (#1034)pin_memory to documentation (#1046)In PyTorch Core
TorchData
sphinx doctest (#850)portalocker optional dependency (#1007)For DataLoader2, we are actively developing new features such as the checkpointing and the ability to execute part of the DataPipe graph on a single process before dispatching the outputs to worker processes. You may begin to see some of these features in nightly builds. We expect them to be part of the next release.
We welcome feedback and feature requests (let us know your use cases!). We always welcome potential contributors.
This library is currently in the Beta stage and currently does not have a fully stable release. The API may change based on user feedback or performance. We are committed to bring this library to stable release, but future changes may not be completely backward compatible. If you install from source or use the nightly version of this library, use it along with the PyTorch nightly binaries. If you have suggestions on the API or use cases you'd like to be covered, please open a GitHub issue. We'd love to hear thoughts and feedback.
This is a minor release to update PyTorch dependency from 1.13.0 to 1.13.1. Please check the release note of TorchData 0.5.0 major release for more de
This is a minor release to update PyTorch dependency from 1.13.0 to 1.13.1. Please check the release note of TorchData 0.5.0 major release for more detail.
We are excited to announce the release of TorchData 0.5.0. This release is composed of about 236 commits since 0.4.1, including ones from PyTorch Core
We are excited to announce the release of TorchData 0.5.0. This release is composed of about 236 commits since 0.4.1, including ones from PyTorch Core since 1.12.1, made by more than 35 contributors. We want to sincerely thank our community for continuously improving TorchData.
TorchData 0.5.0 updates are focused on consolidating the DataLoader2 and ReadingService APIs and benchmarking. Highlights include:
DataLoader2 and provided a few ReadingServices, with detailed documentation now available hereDataPipe operations, e.g., random_split, repeat, set_length, and prefetch.MapDataPipe.shuffle to an IterDataPipe (https://github.com/pytorch/pytorch/pull/83202)IterDataPipe is used to to preserve data order
<p align="center"> <table align="center"> <thead> <tr> <th colspan="2">MapDataPipe.shuffle</th> </tr> </thead> <tr><th>0.4.1</th><th>0.5.0</th></tr> <tr valign="top"> <td><sub><pre lang="python">
from torch.utils.data import IterDataPipe, MapDataPipe from torch.utils.data.datapipes.map import SequenceWrapper dp = SequenceWrapper(list(range(10))).shuffle() isinstance(dp, MapDataPipe) True isinstance(dp, IterDataPipe) False </pre></sub></td> <td><sub><pre lang="python"> from torch.utils.data import IterDataPipe, MapDataPipe from torch.utils.data.datapipes.map import SequenceWrapper dp = SequenceWrapper(list(range(10))).shuffle() isinstance(dp, MapDataPipe) False isinstance(dp, IterDataPipe) True </pre></sub></td> </tr> </table> </p>
on_disk_cache now doesn’t accept generator functions for the argument of filename_fn (https://github.com/pytorch/data/pull/810)<p align="center"> <table align="center"> <thead> <tr> <th colspan="2">on_disk_cache</th> </tr> </thead> <tr><th>0.4.1</th><th>0.5.0</th></tr> <tr valign="top"> <td><sub><pre lang="python">
url_dp = IterableWrapper(["https://path/to/filename", ]) def filepath_gen_fn(url): … yield from [url + f”/{i}” for i in range(3)] cache_dp = url_dp.on_disk_cache(filepath_fn=filepath_gen_fn) </pre></sub></td> <td><sub><pre lang="python"> url_dp = IterableWrapper(["https://path/to/filename", ]) def filepath_gen_fn(url): … yield from [url + f”/{i}” for i in range(3)] cache_dp = url_dp.on_disk_cache(filepath_fn=filepath_gen_fn)
</pre></sub></td>
</tr>
</table>
DataLoader2 (https://github.com/pytorch/data/pull/700)<p align="center"> <table align="center"> <thead> <tr> <th colspan="2">DataLoader2 with a single iterator</th> </tr> </thead> <tr><th>0.4.1</th><th>0.5.0</th></tr> <tr valign="top"> <td><sub><pre lang="python">
dl = DataLoader2(IterableWrapper(range(10))) it1 = iter(dl) print(next(it1)) 0 it2 = iter(dl) # No reset here print(next(it2)) 1 print(next(it1)) 2 </pre></sub></td> <td><sub><pre lang="python"> dl = DataLoader2(IterableWrapper(range(10))) it1 = iter(dl) print(next(it1)) 0 it2 = iter(dl) # DataLoader2 resets with the creation of a new iterator print(next(it2)) 0 print(next(it1))
</pre></sub></td>
</tr>
</table> </p>
DataPipe during DataLoader2 initialization or restoration (https://github.com/pytorch/data/pull/786, https://github.com/pytorch/data/pull/833)Previously, if a DataPipe is being passed to multiple DataLoaders, the DataPipe's state can be altered by any of those DataLoaders. In some cases, that may raise an exception due to the single iterator constraint; in other cases, some behaviors can be changed due to the adapters (e.g. shuffling) of another DataLoader.
<p align="center"> <table align="center"> <thead> <tr> <th colspan="2">Deep copy DataPipe during DataLoader2 constructor</th> </tr> </thead> <tr><th>0.4.1</th><th>0.5.0</th></tr> <tr valign="top"> <td><sub><pre lang="python">
dp = IterableWrapper([0, 1, 2, 3, 4]) dl1 = DataLoader2(dp) dl2 = DataLoader2(dp) for x, y in zip(dl1, dl2): … print(x, y)
</pre></sub></td>
<td><sub><pre lang="python">
dp = IterableWrapper([0, 1, 2, 3, 4]) dl1 = DataLoader2(dp) dl2 = DataLoader2(dp) for x, y in zip(dl1, dl2): … print(x, y) 0 0 1 1 2 2 3 3 4 4 </pre></sub></td> </tr> </table> </p>
traverse function and only_datapipe argument (https://github.com/pytorch/pytorch/pull/85667)Please use traverse_dps with the behavior the same as only_datapipe=True. (https://github.com/pytorch/data/pull/793)
<p align="center"> <table align="center"> <thead> <tr> <th colspan="2">DataPipe traverse function</th> </tr> </thead> <tr><th>0.4.1</th><th>0.5.0</th></tr> <tr valign="top"> <td><sub><pre lang="python">
dp_graph = torch.utils.data.graph.traverse(datapipe, only_datapipe=False) </pre></sub></td> <td><sub><pre lang="python"> dp_graph = torch.utils.data.graph.traverse(datapipe, only_datapipe=False) FutureWarning:
traversefunction and only_datapipe argument will be removed after 1.13. </pre></sub></td> </tr> </table> </p>
IterDataPipe to trace DataFrames operations (https://github.com/pytorch/pytorch/pull/71931,DataFrameMakerIterDataPipe to accept dtype_generator to solve unserializable dtype (https://github.com/pytorch/data/pull/537)IterDataPipe (https://github.com/pytorch/pytorch/pull/79479, https://github.com/pytorch/pytorch/pull/79657)drop operation for IterDataPipe to drop column(s) (https://github.com/pytorch/data/pull/725)FullSyncIterDataPipe to synchronize distributed shards (https://github.com/pytorch/data/pull/713)slice and flatten operations for IterDataPipe (https://github.com/pytorch/data/pull/730)repeat operation for IterDataPipe (https://github.com/pytorch/data/pull/748)LengthSetterIterDataPipe (https://github.com/pytorch/data/pull/747)RandomSplitter (without buffer) (https://github.com/pytorch/data/pull/724)padden_tokens to max_token_bucketize to bucketize samples based on total padded token length (https://github.com/pytorch/data/pull/789)PrefetcherIterDataPipe (https://github.com/pytorch/data/pull/770, https://github.com/pytorch/data/pull/818, https://github.com/pytorch/data/pull/826, https://github.com/pytorch/data/pull/842)CacheTimeout Adapter to redefine cache timeout of the DataPipe graph (https://github.com/pytorch/data/pull/571)DistribtuedReadingService to support uneven data sharding (https://github.com/pytorch/data/pull/727)PrototypeMultiProcessingReadingService
__next__ for IterDataPipe (https://github.com/pytorch/pytorch/pull/79757)TarArchiveLoader (https://github.com/pytorch/data/pull/653)
Improved GDrive 'content-disposition' error message (https://github.com/pytorch/data/pull/654)as_tuple argument for CSVParserIterDataPipe` to convert output from list to tuple (https://github.com/pytorch/data/pull/646)HTTPReader get 404 Response (#160) (https://github.com/pytorch/data/pull/569)flatmap (https://github.com/pytorch/data/pull/749)input_col with the provided map function for DataPipe (https://github.com/pytorch/pytorch/pull/80267, https://github.com/pytorch/data/pull/755, https://github.com/pytorch/pytorch/pull/84279)ShufflerIterDataPipe support snapshotting (#83535)in_batch_shuffle with shuffle for IterDataPipe (https://github.com/pytorch/data/pull/745)IterDataPipe.to_map_datapipe loading data lazily (https://github.com/pytorch/data/pull/765)kwargs to open files for FSSpecFileLister and FSSpecSaver (https://github.com/pytorch/data/pull/804)FileLister (#86497)DataPipes with set_shuffle API https://github.com/pytorch/pytorch/pull/83741)DataPipe (https://github.com/pytorch/pytorch/pull/80509, https://github.com/pytorch/data/pull/559)finalize and finalize_iteration are called during shutdown or exception (https://github.com/pytorch/data/pull/846)fork and unzip operations for the case of a single child (https://github.com/pytorch/pytorch/pull/81502)ShufflerMapDataPipe (https://github.com/pytorch/pytorch/pull/82666)unzip when columns_to_skip is specified (https://github.com/pytorch/data/pull/658)TarArchiveLoader to skip open for opened TarFile stream (https://github.com/pytorch/data/pull/679)IterDataPipe (https://github.com/pytorch/pytorch/pull/84676)DataLoader2
DataPipe
AIStoreDataPipe (https://github.com/pytorch/data/pull/582)DataPipe
DataPipe converters (https://github.com/pytorch/data/pull/710)S3 DataPipe (https://github.com/pytorch/data/pull/784)FileOpenerIterDataPipe (https://github.com/pytorch/pytorch/pull/81407)buffer_size for MaxTokenBucketizer (https://github.com/pytorch/data/pull/834)Prefetcher (https://github.com/pytorch/data/pull/835)generate_csv (https://github.com/pytorch/data/pull/675)random_split example (https://github.com/pytorch/data/pull/843)DataLoader2 (https://github.com/pytorch/data/pull/581, https://github.com/pytorch/data/pull/817)DataLoader2, ReadingService and DataPipe (https://github.com/pytorch/data/pull/563, https://github.com/pytorch/data/pull/664, https://github.com/pytorch/data/pull/670, https://github.com/pytorch/data/pull/787)We will continue benchmarking over datasets on local disk and cloud storage using TorchData. And, we will continue making DataLoader2 and related ReadingService more stable and provide more features like snapshotting the data pipeline and restoring it from the serialized state. Stay tuned and welcome any feedback.
This library is currently in the Beta stage and currently does not have a stable release. The API may change based on user feedback or performance. We are committed to bring this library to stable release, but future changes may not be completely backward compatible. If you install from source or use the nightly version of this library, use it along with the PyTorch nightly binaries. If you have suggestions on the API or use cases you'd like to be covered, please open a GitHub issue. We'd love to hear thoughts and feedback.
Fixed DataPipe working with DataLoader in the distributed environment
DataPipe working with DataLoader in the distributed environment (https://github.com/pytorch/pytorch/pull/80348, https://github.com/pytorch/pytorch/pull/81071, https://github.com/pytorch/pytorch/pull/81071)torchdata binaries for arm64 Apple Silicon (#692)
>>> dp = dp.open_file_by_fsspec() FutureWarning: FSSpecFileOpener()'s functional API .open_file_by_fsspec() is deprecated since 0.4.0 and will be remo…
We are excited to announce the release of TorchData 0.4.0. This release is composed of about 120 commits since 0.3.0, made by 23 contributors. We want to sincerely thank our community for continuously improving TorchData.
TorchData 0.4.0 updates are focused on consolidating the DataPipe APIs and supporting more remote file systems. Highlights include:
DataLoader regarding dynamic sharding and shuffle determinism in single-process, multiprocessing, and distributed environments. Please check the tutorial here.AWSSDK is integrated to support listing/loading files from AWS S3.TFRecord and Hugging Face Hub.DataLoader2 became available in prototype mode. For more details, please check our future plans.Multiplexer (functional API mux) to stop merging multiple DataPipes whenever the shortest one is exhausted (https://github.com/pytorch/pytorch/pull/77145)Please use MultiplexerLongest (functional API mux_longgest) to achieve the previous functionality.
<p align="center"> <table align="center"> <tr><th>0.3.0</th><th>0.4.0</th></tr> <tr valign="top"> <td><sub><pre lang="python">
dp1 = IterableWrapper(range(3)) dp2 = IterableWrapper(range(10, 15)) dp3 = IterableWrapper(range(20, 25)) output_dp = dp1.mux(dp2, dp3) list(output_dp) [0, 10, 20, 1, 11, 21, 2, 12, 22, 3, 13, 23, 4, 14, 24] len(output_dp) 13 </pre></sub></td> <td><sub><pre lang="python"> dp1 = IterableWrapper(range(3)) dp2 = IterableWrapper(range(10, 15)) dp3 = IterableWrapper(range(20, 25)) output_dp = dp1.mux(dp2, dp3) list(output_dp) [0, 10, 20, 1, 11, 21, 2, 12, 22] len(output_dp) 9 </pre></sub></td> </tr> </table> </p>
IterDataPipes w/wo multiple outputs https://github.com/pytorch/pytorch/pull/70479, (https://github.com/pytorch/pytorch/pull/75995)If you need to reference the same IterDataPipe multiple times, please apply .fork() on the IterDataPipe instance.
<p align="center"> <table align="center"> <thead> <tr> <th colspan="2">IterDataPipe with a single output</th> </tr> </thead> <tr><th>0.3.0</th><th>0.4.0</th></tr> <tr valign="top"> <td><sub><pre lang="python">
source_dp = IterableWrapper(range(10)) it1 = iter(source_dp) list(it1) [0, 1, ..., 9] it1 = iter(source_dp) next(it1) 0 it2 = iter(source_dp) next(it2) 0 next(it1) 1
source_dp = IterableWrapper(range(10)) zip_dp = source_dp.zip(source_dp) list(zip_dp) [(0, 0), ..., (9, 9)] </pre></sub></td> <td><sub><pre lang="python"> source_dp = IterableWrapper(range(10)) it1 = iter(source_dp) list(it1) [0, 1, ..., 9] it1 = iter(source_dp) # This doesn't raise any warning or error next(it1) 0 it2 = iter(source_dp) next(it2) # Invalidates
it10 next(it1) RuntimeError: This iterator has been invalidated because another iterator has been created from the same IterDataPipe: IterableWrapperIterDataPipe(deepcopy=True, iterable=range(0, 10)) This may be caused multiple references to the same IterDataPipe. We recommend using.fork()if that is necessary. For feedback regarding this single iterator per IterDataPipe constraint, feel free to comment on this issue: https://github.com/pytorch/data/issues/45.
source_dp = IterableWrapper(range(10)) zip_dp = source_dp.zip(source_dp) list(zip_dp) RuntimeError: This iterator has been invalidated because another iterator has been createdfrom the same IterDataPipe: IterableWrapperIterDataPipe(deepcopy=True, iterable=range(0, 10)) This may be caused multiple references to the same IterDataPipe. We recommend using
.fork()if that is necessary. For feedback regarding this single iterator per IterDataPipe constraint, feel free to comment on this issue: https://github.com/pytorch/data/issues/45. </pre></sub></td> </tr> </table> </p>
<p align="center"> <table align="center"> <thead> <tr> <th colspan="2">IterDataPipe with multiple outputs</th> </tr> </thead> <tr><th>0.3.0</th><th>0.4.0</th></tr> <tr valign="top"> <td><sub><pre lang="python">
source_dp = IterableWrapper(range(10)) cdp1, cdp2 = source_dp.fork(num_instances=2) it1, it2 = iter(cdp1), iter(cdp2) list(it1) [0, 1, 2, 3, 4, 5, 6, 7, 8, 9] list(it2) [0, 1, 2, 3, 4, 5, 6, 7, 8, 9] it1, it2 = iter(cdp1), iter(cdp2) it3 = iter(cdp1)
it1cdp1 hasn't been read since resetnext(it1) 0 next(it2) 0 next(it3) 1
cdp2 has started readingit4 = iter(cdp2) next(it3) 0 list(it4) [0, 1, 2, 3, 4, 5, 6, 7, 8, 9] </pre></sub></td> <td><sub><pre lang="python"> source_dp = IterableWrapper(range(10)) cdp1, cdp2 = source_dp.fork(num_instances=2) it1, it2 = iter(cdp1), iter(cdp2) list(it1) [0, 1, 2, 3, 4, 5, 6, 7, 8, 9] list(it2) [0, 1, 2, 3, 4, 5, 6, 7, 8, 9] it1, it2 = iter(cdp1), iter(cdp2) it3 = iter(cdp1) # This invalidates
it1andit2next(it1) RuntimeError: This iterator has been invalidated, because a new iterator has been created from one of the ChildDataPipes of _ForkerIterDataPipe(buffer_size=1000, num_instances=2). For feedback regarding this single iterator per IterDataPipe constraint, feel free to comment on this issue: https://github.com/pytorch/data/issues/45. next(it2) RuntimeError: This iterator has been invalidated, because a new iterator has been created from one of the ChildDataPipes of _ForkerIterDataPipe(buffer_size=1000, num_instances=2). For feedback regarding this single iterator per IterDataPipe constraint, feel free to comment on this issue: https://github.com/pytorch/data/issues/45. next(it3) 0
cdp2 after it2 was invalidatedit4 = iter(cdp2) next(it3) 1 list(it4) [0, 1, 2, 3, 4, 5, 6, 7, 8, 9] </pre></sub></td> </tr> </table> </p>
open_file_by_fsspec and open_file_by_iopath for IterDataPipe (https://github.com/pytorch/pytorch/pull/78970, https://github.com/pytorch/pytorch/pull/79302)Please use open_files_by_fsspec and open_files_by_iopath
<p align="center"> <table align="center"> <tr><th>0.3.0</th><th>0.4.0</th></tr> <tr valign="top"> <td><sub><pre lang="python">
dp = IterableWrapper([file_path, ]) dp = dp.open_file_by_fsspec() # No Warning dp = IterableWrapper([file_path, ]) dp = dp.open_file_by_iopath() # No Warning </pre></sub></td> <td><sub><pre lang="python"> dp = IterableWrapper([file_path, ]) dp = dp.open_file_by_fsspec() FutureWarning:
FSSpecFileOpener()'s functional API.open_file_by_fsspec()is deprecated since 0.4.0 and will be removed in 0.6.0. See https://github.com/pytorch/data/issues/163 for details. Please use.open_files_by_fsspec()instead. dp = IterableWrapper([file_path, ]) dp = dp.open_file_by_iopath() FutureWarning:IoPathFileOpener()'s functional API.open_file_by_iopath()is deprecated since 0.4.0 and will be removed in 0.6.0. See https://github.com/pytorch/data/issues/163 for details. Please use.open_files_by_iopath()instead. </pre></sub></td> </tr> </table> </p>
drop_empty_batches of Filter (functional API filter) is deprecated and going to be removed in the future release (https://github.com/pytorch/pytorch/pull/76060)<p align="center"> <table align="center"> <tr><th>0.3.0</th><th>0.4.0</th></tr> <tr valign="top"> <td><sub><pre lang="python">
dp = IterableWrapper([(1, 1), (2, 2), (3, 3)]) dp = dp.filter(lambda x: x[0] > 1, drop_empty_batches=True) </pre></sub></td> <td><sub><pre lang="python"> dp = IterableWrapper([(1, 1), (2, 2), (3, 3)]) dp = dp.filter(lambda x: x[0] > 1, drop_empty_batches=True) FutureWarning: The argument
drop_empty_batchesofFilterIterDataPipe()is deprecated since 1.12 and will be removed in 1.14. See https://github.com/pytorch/data/issues/163 for details. </pre></sub></td> </tr> </table> </p>
DataPipe graphs (https://github.com/pytorch/data/pull/330)Bz2FileLoader with functional API of load_from_bz2 (https://github.com/pytorch/data/pull/312)BatchMapper (functional API: map_batches) and FlatMapper (functional API: flat_map) (https://github.com/pytorch/data/pull/359)MultiplexerLongest with functional API of mux_longest (https://github.com/pytorch/data/pull/372)ZipperLongest with functional API of zip_longest (https://github.com/pytorch/data/pull/373)MaxTokenBucketizer with functional API of max_token_bucketize (https://github.com/pytorch/data/pull/283)S3FileLister (functional API: list_files_by_s3) and S3FileLoader (functional API: load_files_by_s3) integrated with the native AWSSDK (https://github.com/pytorch/data/pull/165)HuggingFaceHubReader (https://github.com/pytorch/data/pull/490)TFRecordLoader with functional API of load_from_tfrecord (https://github.com/pytorch/data/pull/308)UnZipper with functional API of unzip (https://github.com/pytorch/data/pull/325)MapToIterConverter with functional API of to_iter_datapipe (https://github.com/pytorch/data/pull/327)InMemoryCacheHolder with functional API of in_memory_cache (https://github.com/pytorch/data/pull/328)pip install –pre torchdata -f https://download.pytorch.org/whl/nightly/cpuconda install -c pytorch-nightly torchdataDataPipes. See: README
encoding argument to FileOpener (https://github.com/pytorch/pytorch/pull/72715)BucketBatcher argument to avoid name collision (https://github.com/pytorch/data/pull/304)ShufflerIterDataPipe (https://github.com/pytorch/pytorch/pull/74370)DataPipe iterator (https://github.com/pytorch/pytorch/pull/75275)input_col argument to flatmap for applying fn to the specific column(s) (https://github.com/pytorch/data/pull/363)IterDataPipe (https://github.com/pytorch/pytorch/pull/75618)DataPipes (https://github.com/pytorch/pytorch/pull/76134)StreamReader (functional API: open_files) and FileOpener (functional API: read_from_stream) (https://github.com/pytorch/pytorch/pull/76233)MapDataPipe (https://github.com/pytorch/pytorch/pull/74851)input_col argument to filter for applying filter_fn to the specific column(s) (https://github.com/pytorch/pytorch/pull/76060)OnlineReaders (https://github.com/pytorch/data/pull/369)
HTTPReaderIterDataPipe: read_from_httpGDriveReaderDataPipe: read_from_gdriveOnlineReaderIterDataPipe: read_from_remoteDataPipe during __del__ (https://github.com/pytorch/pytorch/pull/76345)IterDataPipe's pyi interface (https://github.com/pytorch/data/pull/326)IterDataPipe from iterator to instance (self) (https://github.com/pytorch/data/pull/388)DataPipe serialization:
ForkerIterDataPipe (https://github.com/pytorch/pytorch/pull/73118)DataPipe serialization with dill (https://github.com/pytorch/pytorch/pull/72896)demux and added cache to graph traverse (https://github.com/pytorch/pytorch/pull/75034)DataPipes (https://github.com/pytorch/pytorch/pull/74984)IterDataPipe buffers from iter to instance (self) (#76999)Multiplexer from __iter__ to instance (self) (https://github.com/pytorch/pytorch/pull/77775)GDriveReader handling Virus Scan Warning (https://github.com/pytorch/data/pull/442)**kwargs arguments to HttpReader to specify extra parameters for HTTP requests (https://github.com/pytorch/data/pull/392)FSSpecFileLister and IoPathFileLister to support multiple root paths and updated FSSpecFileLister to support S3 urls (https://github.com/pytorch/data/pull/383)filelock to IoPathSaver to prevent racing condition (https://github.com/pytorch/data/pull/413)on_disk_cache downloading twice https://github.com/pytorch/data/pull/409)DataPipes (https://github.com/pytorch/data/pull/479)list_file functional API to FSSpecFileLister and IoPathFileLister (https://github.com/pytorch/data/pull/463)list_files functional API to FileLister (https://github.com/pytorch/pytorch/pull/78419)DataPipes to accept extra keyword arguments (https://github.com/pytorch/data/pull/495)kwargs to json.loads call in JsonParse (https://github.com/pytorch/data/pull/518)dill to pass DataPipes in multiprocessing (https://github.com/pytorch/pytorch/pull/77288))DataLoader automatically apply sharding to DataPipe graph in single-process, multi-process and distributed environments (https://github.com/pytorch/pytorch/pull/78762, https://github.com/pytorch/pytorch/pull/78950, https://github.com/pytorch/pytorch/pull/79041, https://github.com/pytorch/pytorch/pull/79124, https://github.com/pytorch/pytorch/pull/79524)ShufflerDataPipe deterministic with DataLoader in single-process, multi-process and distributed environments (https://github.com/pytorch/pytorch/pull/77741, https://github.com/pytorch/pytorch/pull/77855, https://github.com/pytorch/pytorch/pull/78765, https://github.com/pytorch/pytorch/pull/79829)DataLoader for DataPipe (https://github.com/pytorch/pytorch/pull/75505)requirements.txt as the single source of truth for TorchData version (https://github.com/pytorch/data/pull/414)IterDataPipe by default (https://github.com/pytorch/pytorch/pull/78674)
DataPipes when their internal operations are very simple (e.g. IterableWrapper)DataPipe naming guidelines (https://github.com/pytorch/data/pull/428)DataSet to PyTorch Dataset (https://github.com/pytorch/data/pull/292)DataPipe (https://github.com/pytorch/data/pull/337)DataPipe (https://github.com/pytorch/data/pull/340)DataPipe (https://github.com/pytorch/data/pull/354)sharding_filter (https://github.com/pytorch/data/pull/487)IterToMapConverter, S3FileLister and S3FileLoader (https://github.com/pytorch/data/pull/381)MapDataPipe (https://github.com/pytorch/data/pull/379)
IterDataPipe and MapDataPipe, we encourage users to use the built-in functionalities of IterDataPipe and use the converter to MapDataPipe as needed.DataPipe working with DataLoader (https://github.com/pytorch/data/pull/458)For DataLoader2, we are introducing new ways to interact between DataPipes, DataLoading API, and backends (aka ReadingServices). Feature is stable in terms of API, but functionally not complete yet. We welcome early adopters and feedback, as well as potential contributors.
This library is currently in the Beta stage and currently does not have a stable release. The API may change based on user feedback or performance. We are committed to bring this library to stable release, but future changes may not be completely backward compatible. If you install from source or use the nightly version of this library, use it along with the PyTorch nightly binaries. If you have suggestions on the API or use cases you'd like to be covered, please open a GitHub issue. We'd love to hear thoughts and feedback.
We are delighted to present the Beta release of TorchData. This is a library of common modular data loading primitives for easily constructing flexibl
We are delighted to present the Beta release of TorchData. This is a library of common modular data loading primitives for easily constructing flexible and performant data pipelines. Based on community feedback, we have found that the existing DataLoader bundled too many features together and can be difficult to extend. Moreover, different use cases often have to rewrite the same data loading utilities over and over again. The goal here is to enable composable data loading through Iterable-style and Map-style building blocks called “DataPipes” that work well out of the box with the PyTorch’s DataLoader.
We are releasing DataPipes - there are Iterable-style DataPipe (IterDataPipe) and Map-style DataPipe (MapDataPipe).
Early on, we observed widespread confusion between the PyTorch DataSets which represented reusable loading tooling (e.g. TorchVision's ImageFolder), and those that represented pre-built iterators/accessors over actual data corpora (e.g. TorchVision's ImageNet). This led to an unfortunate pattern of siloed inheritance of data tooling rather than composition.
DataPipe is simply a renaming and repurposing of the PyTorch DataSet for composed usage. A DataPipe takes in some access function over Python data structures, __iter__ for IterDataPipes and __getitem__ for MapDataPipes , and returns a new access function with a slight transformation applied. For example, take a look at this JsonParser, which accepts an IterDataPipe over file names and raw streams, and produces a new iterator over the filenames and deserialized data:
import json
class JsonParserIterDataPipe(IterDataPipe):
def __init__(self, source_datapipe, **kwargs) -> None:
self.source_datapipe = source_datapipe
self.kwargs = kwargs
def __iter__(self):
for file_name, stream in self.source_datapipe:
data = stream.read()
yield file_name, json.loads(data)
def __len__(self):
return len(self.source_datapipe)
You can see in this example how DataPipes can be easily chained together to compose graphs of transformations that reproduce sophisticated data pipelines, with streamed operation as a first-class citizen.
Under this naming convention, DataSet simply refers to a graph of DataPipes, and a dataset module like ImageNet can be rebuilt as a factory function returning the requisite composed DataPipes.
In this example, we have a compressed TAR archive file stored in Google Drive and accessible via an URL. We demonstrate how you can use DataPipes to download the archive, cache the result, decompress the archive, filter for specific files, parse and return the CSV content. The full example with detailed explanation is included in the example folder.
url_dp = IterableWrapper([URL])
cache_compressed_dp = GDriveReader(cache_compressed_dp)
# cache_decompressed_dp = ... # See source file for full code example
# Opens and loads the content of the TAR archive file.
cache_decompressed_dp = FileOpener(cache_decompressed_dp, mode="b").load_from_tar()
# Filters for specific files based on the file name.
cache_decompressed_dp = cache_decompressed_dp.filter(
lambda fname_and_stream: _EXTRACTED_FILES[split] in fname_and_stream[0]
)
# Saves the decompressed file onto disk.
cache_decompressed_dp = cache_decompressed_dp.end_caching(mode="wb", same_filepath_fn=True)
data_dp = FileOpener(cache_decompressed_dp, mode="b")
# Parses content of the decompressed CSV file and returns the result line by line. return
return data_dp.parse_csv().map(fn=lambda t: (int(t[0]), " ".join(t[1:])))
[Beta] IterDataPipe
We have implemented over 50 Iterable-style DataPipes across 10 different categories. They cover different functionalities, such as opening files, parsing texts, transforming samples, caching, shuffling, and batching. For users who are interested in connecting to cloud providers (such as Google Drive or AWS S3), the fsspec and iopath DataPipes will allow you to do so. The documentation provides detailed explanations and usage examples of each IterDataPipe.
[Beta] MapDataPipe
Similar to IterDataPipe, we have various, but a more limited number of MapDataPipe available for different functionalities. More MapDataPipes support will come later. If the existing ones do not meet your needs, you can write a custom DataPipe.
The documentation for TorchData is now live. It contains a tutorial that covers how to use DataPipes, use them with DataLoader, and implement custom ones.
In this release, some of the PyTorch domain libraries have migrated their datasets to use DataPipes. In TorchText, the popular datasets provided by the library are implemented using DataPipes and a section of its SST-2 binary text classification tutorial demonstrates how you can use DataPipes to preprocess data for your model. There also are other prototype implementations of datasets with DataPipes in TorchVision (available in nightly releases) and in TorchRec. You can find more specific examples here.
There will be a new version of DataLoader in the next release. At the high level, the plan is that DataLoader V2 will only be responsible for multiprocessing, distributed, and similar functionalities, not data processing logic. All data processing features, such as the shuffling and batching, will be moved out of DataLoader to DataPipe. At the same time, the current/old version of DataLoader should still be available and you can use DataPipes with that as well.
This library is currently in the Beta stage and currently does not have a stable release. The API may change based on user feedback or performance. We are committed to bring this library to stable release, but future changes may not be completely backward compatible. If you install from source or use the nightly version of this library, use it along with the PyTorch nightly binaries. If you have suggestions on the API or use cases you'd like to be covered, please open a GitHub issue. We'd love to hear thoughts and feedback.
Nothing published for this version
Nothing published for this version
Your coding agent can read these notes before it upgrades. Set up the MCP server →