NewYour coding agent can read the release notes before it upgrades.Set up the MCP server →
PyPI · #331 most downloaded on PyPI
HuggingFace community-driven open-source library of datasets
Last release 2 months ago
28 Jul 2026
Ships fairly regularly
a new release about every 4 weeks
Nearly every release is documented
notes for 60 of the last 60 stable releases
Nothing withdrawn
no release was ever pulled
11 years old
119 releases · first in 2015
Fix version string in init .py by @qgallouedec in #8244
Full Changelog: 5.0.0...5.0.1
One column per quarter.
Note 3: This is a BREAKING CHANGE: the default shuffling mechanism now uses multiple input shards. You can get the old mechanism by passing max_buffer…
Parse Agent traces messages for SFT using teich by @lhoestq in #8232
teich library (new optional dependency), traces are parsed to messages to enable training on traces using e.g. trl>>> from datasets import load_dataset
>>> ds = load_dataset("lhoestq/agent-traces-example", split="train")
>>> ds[0]["messages"]
[{'role': 'user', 'content': 'Download a random dataset from Hugging Face, use DuckDB to inspect it, and come back with a short report about it. Be concise and include: dataset name, what files/format you found, row count or rough size if you can determine it,...'
...]trl sft --dataset-name lhoestq/agent-traces-example ...Use multiple input shards for shuffle buffer by @lhoestq in #8194
ds = load_dataset(..., streaming=True)
ds = ds.shuffle(seed=42)
# or configure local buffer shuffling manually, default is:
ds = ds.shuffle(seed=42, buffer_size=1000, max_buffer_input_shards=10)toy example comparison
from datasets import IterableDataset
ds = IterableDataset.from_dict({"i": range(123_456_789)}, num_shards=1024)
ds = ds.shuffle(seed=42)
print("Cold start ids:")
print(list(ds.take(10)["i"]))
print("Nominal regime ids:")
print(list(ds.skip(10_000).take(10)["i"]))before👎:
Cold start ids:
[6148853, 6149537, 6149418, 6149202, 6149197, 6149622, 6148849, 6149461, 6148965, 6148858]
Nominal regime ids:
[6149537, 6149418, 6149202, 6149197, 6149622, 6148849, 6149461, 6148965, 6148858, 6149290]
after✨:
Cold start ids:
[7836668, 9283505, 95847927, 482299, 9283471, 482341, 112003312, 59920157, 43764666, 95847871]
Nominal regime ids:
[9283505, 95847927, 482299, 9283471, 482341, 112003312, 59920157, 43764666, 95847871, 16758448]
Note: ds.state_dict() and ds.load_state_dict() are still supported for this improved shuffling :) enabling dataset checkpointing
Note 2: it uses threads to fetch the first examples in parallel from the input shards
Note 3: This is a BREAKING CHANGE: the default shuffling mechanism now uses multiple input shards. You can get the old mechanism by passing max_buffer_input_shards=1 to IterableDataset.shuffle()
Add batch(by_column=...) by @lhoestq in #8172
from datasets import Dataset
ds = Dataset.from_dict({"episode": [0] * 10 + [1] * 10, "frame": list(range(10)) * 2})
# ds = ds.to_iterable_dataset()
ds = ds.batch(by_column="episode")
for x in ds:
print(x)
# {'episode': [0, 0, 0, 0, 0, 0, 0, 0, 0, 0], 'frame': [0, 1, 2, 3, 4, 5, 6, 7, 8, 9]}
# {'episode': [1, 1, 1, 1, 1, 1, 1, 1, 1, 1], 'frame': [0, 1, 2, 3, 4, 5, 6, 7, 8, 9]}.conll / .conllu dataset format loader (CoNLL-2003 / 2000 / U) by @CrypticCortex in #8219num_proc argument to Dataset.to_sql by @EricSaikali in #7791ClassLabel docs: Correct value for unknown labels by @l-uuz in #7645Full Changelog: 4.8.5...5.0.0
fix: decode Json() values before calling DataFrame.to_json() ( #8116 ) by @Brianzhengca in #8122
Full Changelog: 4.8.4...4.8.5
Support latest torchvision by @lhoestq in #8087
Full Changelog: 4.8.3...4.8.4
Fix split_dataset_by_node step by @lhoestq in #8081
Full Changelog: 4.8.2...4.8.3
Json type for empty struct by @lhoestq in #8074
Fix formatted iter arrow double yield by @HaukurPall in #8063
Full Changelog: 4.8.0...4.8.1
Read (and write) from HF Storage Buckets : load raw data, process and save to Dataset Repos by @lhoestq in #8064
Read (and write) from HF Storage Buckets: load raw data, process and save to Dataset Repos by @lhoestq in #8064
from datasets import load_dataset
# load raw data from a Storage Bucket on HF
ds = load_dataset("buckets/username/data-bucket", data_files=["*.jsonl"])
# or manually, using hf:// paths
ds = load_dataset("json", data_files=["hf://buckets/username/data-bucket/*.jsonl"])
# process, filter
ds = ds.map(...).filter(...)
# publish the AI-ready dataset
ds.push_to_hub("username/my-dataset-ready-for-training")This also fixes multiprocessed push_to_hub on macos that was causing segfault (now it uses spawn instead of fork).
And it bumps dill and multiprocess versions to support python 3.14
Datasets streaming iterable packaged improvements and fixes by @Michael-RDev in #8068
max_shard_size to IterableDataset.push_to_hub (but requires iterating twice to know the full dataset twice - improvements are welcome)zip://*.jsonl::hf://datasets/username/dataset-name/data.zipFull Changelog: 4.7.0...4.8.0
Add Json() type by @lhoestq in #8027
Json() type by @lhoestq in #8027
Json()type is used to store such data that would normally not be supported in Arrow/ParquetJson() type in Features() for any dataset, it is supported in any functions that accepts features=like load_dataset(), .map(), .cast(), .from_dict(), .from_list()on_mixed_types="use_json" to automatically set the Json() type on mixed types in .from_dict(), .from_list() and .map()Examples:
You can use on_mixed_types="use_json" or specify features= with a [Json] type:
>>> ds = Dataset.from_dict({"a": [0, "foo", {"subfield": "bar"}]})
Traceback (most recent call last):
...
File "pyarrow/error.pxi", line 92, in pyarrow.lib.check_status
pyarrow.lib.ArrowInvalid: Could not convert 'foo' with type str: tried to convert to int64
>>> features = Features({"a": Json()})
>>> ds = Dataset.from_dict({"a": [0, "foo", {"subfield": "bar"}]}, features=features)
>>> ds.features
{'a': Json()}
>>> list(ds["a"])
[0, "foo", {"subfield": "bar"}]This is also useful for lists of dictionaries with arbitrary keys and values, to avoid filling missing fields with None:
>>> ds = Dataset.from_dict({"a": [[{"b": 0}, {"c": 0}]]})
>>> ds.features
{'a': List({'b': Value('int64'), 'c': Value('int64')})}
>>> list(ds["a"])
[[{'b': 0, 'c': None}, {'b': None, 'c': 0}]] # missing fields are filled with None
>>> features = Features({"a": List(Json())})
>>> ds = Dataset.from_dict({"a": [[{"b": 0}, {"c": 0}]]}, features=features)
>>> ds.features
{'a': List(Json())}
>>> list(ds["a"])
[[{'b': 0}, {'c': 0}]] # OKAnother example with tool calling data and the on_mixed_types="use_json" argument (useful to not have to specify features= manually):
>>> messages = [
... {"role": "user", "content": "Turn on the living room lights and play my electronic music playlist."},
... {"role": "assistant", "tool_calls": [
... {"type": "function", "function": {
... "name": "control_light",
... "arguments": {"room": "living room", "state": "on"}
... }},
... {"type": "function", "function": {
... "name": "play_music",
... "arguments": {"playlist": "electronic"} # mixed-type here since keys ["playlist"] and ["room", "state"] are different
... }}]
... },
... {"role": "tool", "name": "control_light", "content": "The lights in the living room are now on."},
... {"role": "tool", "name": "play_music", "content": "The music is now playing."},
... {"role": "assistant", "content": "Done!"}
... ]
>>> ds = Dataset.from_dict({"messages": [messages]}, on_mixed_types="use_json")
>>> ds.features
{'messages': List({'role': Value('string'), 'content': Value('string'), 'tool_calls': List(Json()), 'name': Value('string')})}
>>> ds[0][1]["tool_calls"][0]["function"]["arguments"]
{"room": "living room", "state": "on"}Full Changelog: 4.6.1...4.7.0
Remove tmp file in push to hub by @lhoestq in #8030
Support Image, Video and Audio types in Lance datasets
Support Image, Video and Audio types in Lance datasets
>>> from datasets import load_dataset
>>> ds = load_dataset("lance-format/Openvid-1M", streaming=True, split="train")
>>> ds.features
{'video_blob': Video(),
'video_path': Value('string'),
'caption': Value('string'),
'aesthetic_score': Value('float64'),
'motion_score': Value('float64'),
'temporal_consistency_score': Value('float64'),
'camera_motion': Value('string'),
'frame': Value('int64'),
'fps': Value('float64'),
'seconds': Value('float64'),
'embedding': List(Value('float32'), length=1024)}
Push to hub now supports Video types
>>> from datasets import Dataset, Video
>>> ds = Dataset.from_dict({"video": ["path/to/video.mp4"]})
>>> ds = ds.cast_column("video", Video())
>>> ds.push_to_hub("username/my-video-dataset")
Write image/audio/video blobs as is in parquet (PLAIN) in push_to_hub() by @lhoestq in https://github.com/huggingface/datasets/pull/7976
<p align="center"> <a href="https://huggingface.co/docs/hub/en/xet/deduplication"> <img height="200" alt="image" src="https://github.com/user-attachments/assets/dd0de6a2-24a1-4945-8d25-44b763c1151e" /> </a> </p>
Add IterableDataset.reshard() by @lhoestq in https://github.com/huggingface/datasets/pull/7992
Reshard the dataset if possible, i.e. split the current shards further into more shards. This increases the number of shards and the resulting dataset has num_shards >= previous_num_shards. Equality may happen if no shard can be split further.
The resharding mechanism depends on the dataset file format:
>>> from datasets import load_dataset
>>> ds = load_dataset("fancyzhx/amazon_polarity", split="train", streaming=True)
>>> ds
IterableDataset({
features: ['label', 'title', 'content'],
num_shards: 4
})
>>> ds.reshard()
IterableDataset({
features: ['label', 'title', 'content'],
num_shards: 3600
})
transformers v5 and huggingface_hub v1 by @hanouticelina in https://github.com/huggingface/datasets/pull/7989Full Changelog: https://github.com/huggingface/datasets/compare/4.5.0...4.6.0
Add lance format support by @eddyxu in https://github.com/huggingface/datasets/pull/7913
Add lance format support by @eddyxu in https://github.com/huggingface/datasets/pull/7913
from datasets import load_dataset
ds = load_dataset("lance-format/fineweb-edu", streaming=True)
for example in ds["train"]:
...
revision in load_dataset by @Scott-Simmons in https://github.com/huggingface/datasets/pull/7929Full Changelog: https://github.com/huggingface/datasets/compare/4.4.2...4.5.0
Fix embed storage nifti by @CloseChoice in https://github.com/huggingface/datasets/pull/7853
Full Changelog: https://github.com/huggingface/datasets/compare/4.4.1...4.4.2
Better streaming retries (504 and 429) by @lhoestq in https://github.com/huggingface/datasets/pull/7847
Full Changelog: https://github.com/huggingface/datasets/compare/4.4.0...4.4.1
Add nifti support by @CloseChoice in https://github.com/huggingface/datasets/pull/7815
Add nifti support by @CloseChoice in https://github.com/huggingface/datasets/pull/7815
ds = load_dataset("username/my_nifti_dataset")
ds["train"][0] # {"nifti": <nibabel.nifti1.Nifti1Image>}
files = ["/path/to/scan_001.nii.gz", "/path/to/scan_002.nii.gz"]
ds = Dataset.from_dict({"nifti": files}).cast_column("nifti", Nifti())
ds["train"][0] # {"nifti": <nibabel.nifti1.Nifti1Image>}
Add num channels to audio by @CloseChoice in https://github.com/huggingface/datasets/pull/7840
# samples have shape (num_channels, num_samples)
ds = ds.cast_column("audio", Audio()) # default, use all channels
ds = ds.cast_column("audio", Audio(num_channels=2)) # use stereo
ds = ds.cast_column("audio", Audio(num_channels=1)) # use mono
_batch_setitems() by @sghng in https://github.com/huggingface/datasets/pull/7817Full Changelog: https://github.com/huggingface/datasets/compare/4.3.0...4.4.0
Enable large scale distributed dataset streaming:
Enable large scale distributed dataset streaming:
These improvements require huggingface_hub>=1.1.0 to take full effect
from_generator by @simonreise in https://github.com/huggingface/datasets/pull/7533Full Changelog: https://github.com/huggingface/datasets/compare/4.2.0...4.3.0
Sample without replacement option when interleaving datasets by @radulescupetru in https://github.com/huggingface/datasets/pull/7786
Sample without replacement option when interleaving datasets by @radulescupetru in https://github.com/huggingface/datasets/pull/7786
ds = interleave_datasets(datasets, stopping_strategy="all_exhausted_without_replacement")
Parquet: add on_bad_files argument to error/warn/skip bad files by @lhoestq in https://github.com/huggingface/datasets/pull/7806
ds = load_dataset(parquet_dataset_id, on_bad_files="warn")
Add parquet scan options and docs by @lhoestq in https://github.com/huggingface/datasets/pull/7801
ds = load_dataset(parquet_dataset_id, columns=["col_0", "col_1"])
ds = load_dataset(parquet_dataset_id, filters=[("col_0", "==", 0)])
fragment_scan_options = pyarrow.dataset.ParquetFragmentScanOptions(cache_options=pyarrow.CacheOptions(prefetch_limit=1, range_size_limit=128 << 20))
ds = load_dataset(parquet_dataset_id, streaming=True, fragment_scan_options=fragment_scan_options)
Full Changelog: https://github.com/huggingface/datasets/compare/4.1.1...4.2.0
fix iterate nested field by @lhoestq in https://github.com/huggingface/datasets/pull/7775
Full Changelog: https://github.com/huggingface/datasets/compare/4.1.0...4.1.1
feat: use content defined chunking by @kszucs in https://github.com/huggingface/datasets/pull/7589
feat: use content defined chunking by @kszucs in https://github.com/huggingface/datasets/pull/7589
Parquet datasets are now Optimized Parquet ! <img width="462" height="103" alt="image" src="https://github.com/user-attachments/assets/43703a47-0964-421b-8f01-1a790305de79" />
internally uses use_content_defined_chunking=True when writing Parquet files
this enables fast deduped uploads to Hugging Face !
# Now faster thanks to content defined chunking
ds.push_to_hub("username/dataset_name")
write_page_index=True is also used to enable fast random access for the Dataset Viewer and tools that need itConcurrent push_to_hub by @lhoestq in https://github.com/huggingface/datasets/pull/7708
Concurrent IterableDataset push_to_hub by @lhoestq in https://github.com/huggingface/datasets/pull/7710
HDF5 support by @klamike in https://github.com/huggingface/datasets/pull/7690
ds = load_dataset("username/dataset-with-hdf5-files")
train_test_split by @qgallouedec in https://github.com/huggingface/datasets/pull/7736Full Changelog: https://github.com/huggingface/datasets/compare/4.0.0...4.1.0
Add IterableDataset.push_to_hub() by @lhoestq in https://github.com/huggingface/datasets/pull/7595
Add IterableDataset.push_to_hub() by @lhoestq in https://github.com/huggingface/datasets/pull/7595
# Build streaming data pipelines in a few lines of code !
from datasets import load_dataset
ds = load_dataset(..., streaming=True)
ds = ds.map(...).filter(...)
ds.push_to_hub(...)
Add num_proc= to .push_to_hub() (Dataset and IterableDataset) by @lhoestq in https://github.com/huggingface/datasets/pull/7606
# Faster push to Hub ! Available for both Dataset and IterableDataset
ds.push_to_hub(..., num_proc=8)
New Column object
# Syntax:
ds["column_name"] # datasets.Column([...]) or datasets.IterableColumn(...)
# Iterate on a column:
for text in ds["text"]:
...
# Load one cell without bringing the full column in memory
first_text = ds["text"][0] # equivalent to ds[0]["text"]
Torchcodec decoding by @TyTodd in https://github.com/huggingface/datasets/pull/7616
# Don't download full audios/videos when it's not necessary
# Now with torchcodec it only streams the required ranges/frames:
from datasets import load_dataset
ds = load_dataset(..., streaming=True)
for example in ds:
video = example["video"]
frames = video.get_frames_in_range(start=0, stop=6, step=1) # only stream certain frames
torch>=2.7.0 and FFmpeg >= 4datasets<4.0AudioDecoder:audio = dataset[0]["audio"] # <datasets.features._torchcodec.AudioDecoder object at 0x11642b6a0>
samples = audio.get_all_samples() # or use get_samples_played_in_range(...)
samples.data # tensor([[ 0.0000e+00, 0.0000e+00, 0.0000e+00, ..., 2.3447e-06, -1.9127e-04, -5.3330e-05]]
samples.sample_rate # 16000
# old syntax is still supported
array, sr = audio["array"], audio["sampling_rate"]
VideoDecoder:video = dataset[0]["video"] <torchcodec.decoders._video_decoder.VideoDecoder object at 0x14a61d5a0>
first_frame = video.get_frame_at(0)
first_frame.data.shape # (3, 240, 320)
first_frame.pts_seconds # 0.0
frames = video.get_frames_in_range(0, 6, 1)
frames.data.shape # torch.Size([5, 3, 240, 320])
Remove scripts altogether by @lhoestq in https://github.com/huggingface/datasets/pull/7592
trust_remote_code is no longer supportedTorchcodec decoding by @TyTodd in https://github.com/huggingface/datasets/pull/7616
Replace Sequence by List by @lhoestq in https://github.com/huggingface/datasets/pull/7634
List typefrom datasets import Features, List, Value
features = Features({
"texts": List(Value("string")),
"four_paragraphs": List(Value("string"), length=4)
})
Sequence was a legacy type from tensorflow datasets which converted list of dicts to dicts of lists. It is no longer a type but it becomes a utility that returns a List or a dict depending on the subfeaturefrom datasets import Sequence
Sequence(Value("string")) # List(Value("string"))
Sequence({"texts": Value("string")}) # {"texts": List(Value("string"))}
Dataset.map to reuse cache files mapped with different num_proc by @ringohoffman in https://github.com/huggingface/datasets/pull/7434RepeatExamplesIterable by @SilvanCodes in https://github.com/huggingface/datasets/pull/7581_dill.py to use co_linetable for Python 3.10+ in place of co_lnotab by @qgallouedec in https://github.com/huggingface/datasets/pull/7609Full Changelog: https://github.com/huggingface/datasets/compare/3.6.0...4.0.0
Enable xet in push to hub by @lhoestq in https://github.com/huggingface/datasets/pull/7552
aiohttp from direct dependencies by @akx in https://github.com/huggingface/datasets/pull/7294Full Changelog: https://github.com/huggingface/datasets/compare/3.5.1...3.6.0
support pyarrow 20 by @lhoestq in https://github.com/huggingface/datasets/pull/7540
TypeError: ArrayExtensionArray.to_pylist() got an unexpected keyword argument 'maps_as_pydicts'Full Changelog: https://github.com/huggingface/datasets/compare/3.5.0...3.5.1
Introduce PDF support (#7318) by @yabramuvdi in https://github.com/huggingface/datasets/pull/7325
>>> from datasets import load_dataset, Pdf
>>> repo = "path/to/pdf/folder" # or username/dataset_name on Hugging Face
>>> dataset = load_dataset(repo, split="train")
>>> dataset[0]["pdf"]
<pdfplumber.pdf.PDF at 0x1075bc320>
>>> dataset[0]["pdf"].pages[0].extract_text()
...
Full Changelog: https://github.com/huggingface/datasets/compare/3.4.1...3.5.0
Fix data_files filtering by @lhoestq in https://github.com/huggingface/datasets/pull/7459
Full Changelog: https://github.com/huggingface/datasets/compare/3.4.0...3.4.1
/!\ Breaking change: we replaced decord with torchvision to read videos, since decord is not maintained anymore and isn't available for recent python…
Faster folder based builder + parquet support + allow repeated media + use torchvideo by @lhoestq in https://github.com/huggingface/datasets/pull/7424
decord with torchvision to read videos, since decord is not maintained anymore and isn't available for recent python versions, see the video dataset loading documentation here for more details. The Video type is still marked as experimental is this versionfrom datasets import load_dataset, Video
dataset = load_dataset("path/to/video/folder", split="train")
dataset[0]["video"] # <torchvision.io.video_reader.VideoReader at 0x1652284c0>
metadata.parquet in addition to metadata.csv or metadata.jsonl for the metadata of the image/audio/video filesAdd IterableDataset.decode with multithreading by @lhoestq in https://github.com/huggingface/datasets/pull/7450
dataset = dataset.decode(num_threads=num_threads)
Add with_split to DatasetDict.map by @jp1924 in https://github.com/huggingface/datasets/pull/7368
string_to_dict to return None if there is no match instead of raising ValueError by @ringohoffman in https://github.com/huggingface/datasets/pull/7435ds.set_epoch(new_epoch) by @lhoestq in https://github.com/huggingface/datasets/pull/7451Full Changelog: https://github.com/huggingface/datasets/compare/3.3.2...3.4.0
Attempt to fix multiprocessing hang by closing and joining the pool before termination by @dakinggg in https://github.com/huggingface/datasets/pull/74
Full Changelog: https://github.com/huggingface/datasets/compare/3.3.1...3.3.2
Fix filter speed regression by @lhoestq in https://github.com/huggingface/datasets/pull/7408
Full Changelog: https://github.com/huggingface/datasets/compare/3.3.0...3.3.1
Support async functions in map() by @lhoestq in https://github.com/huggingface/datasets/pull/7384
Support async functions in map() by @lhoestq in https://github.com/huggingface/datasets/pull/7384
prompt = "Answer the following question: {question}. You should think step by step."
async def ask_llm(example):
return await query_model(prompt.format(question=example["question"]))
ds = ds.map(ask_llm)
Add repeat method to datasets by @alex-hh in https://github.com/huggingface/datasets/pull/7198
ds = ds.repeat(10)
Support faster processing using pandas or polars functions in IterableDataset.map() by @lhoestq in https://github.com/huggingface/datasets/pull/7370
ds = load_dataset("ServiceNow-AI/R1-Distill-SFT", "v0", split="train", streaming=True)
ds = ds.with_format("polars")
expr = pl.col("solution").str.extract("boxed\\{(.*)\\}").alias("value_solution")
ds = ds.map(lambda df: df.with_columns(expr), batched=True)
Apply formatting after iter_arrow to speed up format -> map, filter for iterable datasets by @alex-hh in https://github.com/huggingface/datasets/pull/7207
Full Changelog: https://github.com/huggingface/datasets/compare/3.2.0...3.3.0
Faster parquet streaming + filters with predicate pushdown by @lhoestq in https://github.com/huggingface/datasets/pull/7309
from datasets import load_dataset
filters = [('date', '>=', '2023')]
ds = load_dataset("HuggingFaceFW/fineweb-2", "fra_Latn", streaming=True, filters=filters)
ClassLabel by @sergiopaniego in https://github.com/huggingface/datasets/pull/7293Full Changelog: https://github.com/huggingface/datasets/compare/3.1.0...3.2.0
Video support by @lhoestq in https://github.com/huggingface/datasets/pull/7230 ```python
>>> from datasets import Dataset, Video, load_dataset
>>> ds = Dataset.from_dict({"video":["path/to/Screen Recording.mov"]}).cast_column("video", Video())
>>> # or from the hub
>>> ds = load_dataset("username/dataset_name", split="train")
>>> ds[0]["video"]
<decord.video_reader.VideoReader at 0x105525c70>
>>> from datasets import load_dataset
>>> full_ds = load_dataset("amphion/Emilia-Dataset", split="train", streaming=True)
>>> full_ds.num_shards
2360
>>> ds = full_ds.shard(num_shards=ds.num_shards, index=0)
>>> ds.num_shards
1
>>> ds = full_ds.shard(num_shards=8, index=0)
>>> ds.num_shards
295
Full Changelog: https://github.com/huggingface/datasets/compare/3.0.2...3.1.0
fix unbatched arrow map for iterable datasets by @alex-hh in https://github.com/huggingface/datasets/pull/7204
Full Changelog: https://github.com/huggingface/datasets/compare/3.0.1...3.0.2
Modify add_column() to optionally accept a FeatureType as param by @varadhbhatnagar in https://github.com/huggingface/datasets/pull/7143
Full Changelog: https://github.com/huggingface/datasets/compare/3.0.0...3.0.1
Remove deprecated code by @albertvillanova in https://github.com/huggingface/datasets/pull/6996
.map()
Allow Polars as valid output type by @psmyth94 in https://github.com/huggingface/datasets/pull/6762
Example:
>>> from datasets import load_dataset
>>> ds = load_dataset("lhoestq/CudyPokemonAdventures", split="train").with_format("polars")
>>> cols = [pl.col("content").str.len_bytes().alias("length")]
>>> ds_with_length = ds.map(lambda df: df.with_columns(cols), batched=True)
>>> ds_with_length[:5]
shape: (5, 5)
┌─────┬───────────────────────────────────┬───────────────────────────────────┬───────────────────────┬────────┐
│ idx ┆ title ┆ content ┆ labels ┆ length │
│ --- ┆ --- ┆ --- ┆ --- ┆ --- │
│ i64 ┆ str ┆ str ┆ str ┆ u32 │
╞═════╪═══════════════════════════════════╪═══════════════════════════════════╪═══════════════════════╪════════╡
│ 0 ┆ The Joyful Adventure of Bulbasau… ┆ Bulbasaur embarked on a sunny qu… ┆ joyful_adventure ┆ 180 │
│ 1 ┆ Pikachu's Quest for Peace ┆ Pikachu, with his cheeky persona… ┆ peaceful_narrative ┆ 138 │
│ 2 ┆ The Tender Tale of Squirtle ┆ Squirtle took everyone on a memo… ┆ gentle_adventure ┆ 135 │
│ 3 ┆ Charizard's Heartwarming Tale ┆ Charizard found joy in helping o… ┆ heartwarming_story ┆ 112 │
│ 4 ┆ Jolteon's Sparkling Journey ┆ Jolteon, with his zest for life,… ┆ celebratory_narrative ┆ 111 │
└─────┴───────────────────────────────────┴───────────────────────────────────┴───────────────────────┴────────┘
huggingface_hub cache by @lhoestq in https://github.com/huggingface/datasets/pull/7105
huggingface_hub cache for files downloaded from HF, by default at ~/.cache/huggingface/hubdatasets cache, by default at ~/.cache/huggingface/datasetsuse_auth_token, fs or ignore_verificationsload_metric, please use the evaluate library insteadtask argument in load_dataset() .prepare_for_task() method, datasets.tasks modulecache_dir from cache_file_name by @ringohoffman in https://github.com/huggingface/datasets/pull/7096Full Changelog: https://github.com/huggingface/datasets/compare/2.21.0...3.0.0
Support pyarrow large_list by @albertvillanova in https://github.com/huggingface/datasets/pull/7019
import polars as pl
from datasets import Dataset
df1 = pl.from_dict({"col_1": [[1, 2], [3, 4]]}
df2 = Dataset.from_polars(df).to_polars()
assert df1.equals(df2)
HF_HUB_OFFLINE instead of HF_DATASETS_OFFLINE by @Wauplin in https://github.com/huggingface/datasets/pull/6968Full Changelog: https://github.com/huggingface/datasets/compare/2.20.0...2.21.0
Update tqdm >= 4.66.3 to fix vulnerability by @albertvillanova in https://github.com/huggingface/datasets/pull/6870
trust_remote_code=True by @lhoestq in https://github.com/huggingface/datasets/pull/6954
trust_remote_code=True to be usedcheckpoint and resume an iterable dataset (e.g. when streaming):
>>> iterable_dataset = Dataset.from_dict({"a": range(6)}).to_iterable_dataset(num_shards=3)
>>> for idx, example in enumerate(iterable_dataset):
... print(example)
... if idx == 2:
... state_dict = iterable_dataset.state_dict()
... print("checkpoint")
... break
>>> iterable_dataset.load_state_dict(state_dict)
>>> print(f"restart from checkpoint")
>>> for example in iterable_dataset:
... print(example)
Returns:
{'a': 0}
{'a': 1}
{'a': 2}
checkpoint
restart from checkpoint
{'a': 3}
{'a': 4}
{'a': 5}
.pth support for torch tensors by @lhoestq in https://github.com/huggingface/datasets/pull/6920dataset_module_factory by @Wauplin in https://github.com/huggingface/datasets/pull/6959Full Changelog: https://github.com/huggingface/datasets/compare/2.19.0...2.20.0
Update requests >=2.32.1 to fix vulnerability by @albertvillanova in https://github.com/huggingface/datasets/pull/6909
Full Changelog: https://github.com/huggingface/datasets/compare/2.19.1...2.19.2
Fix download for dict of dicts of URLs by @albertvillanova in https://github.com/huggingface/datasets/pull/6871
Full Changelog: https://github.com/huggingface/datasets/compare/2.19.0...2.19.1
Deprecate Beam API and download from HF GCS bucket by @mariosasko in https://github.com/huggingface/datasets/pull/6474
.to_polars();import polars as pl
from datasets import load_dataset
ds = load_dataset("DIBT/10k_prompts_ranked", split="train")
ds.to_polars() \
.groupby("topic") \
.agg(pl.len(), pl.first()) \
.sort("len", descending=True)
ds = ds.with_format("polars")
ds[:10].group_by("kind").len()
fsspec support for to_json, to_csv, and to_parquet by @alvarobartt in https://github.com/huggingface/datasets/pull/6096
ds.to_json("hf://datasets/username/my_json_dataset/data.jsonl")
ds.to_csv("hf://datasets/username/my_csv_dataset/data.csv")
ds.to_parquet("hf://datasets/username/my_parquet_dataset/data.parquet")
mode parameter to Image feature by @mariosasko in https://github.com/huggingface/datasets/pull/6735
dataset = dataset.cast_column("image", Image(mode="RGB"))
datasets-cli convert_to_parquet <dataset_id>
ds = ds.take(10) # take only the first 10 examples
remove_columns/rename_columns doc fixes by @mariosasko in https://github.com/huggingface/datasets/pull/6772uv in CI by @mariosasko in https://github.com/huggingface/datasets/pull/6779_check_legacy_cache2 by @lhoestq in https://github.com/huggingface/datasets/pull/6792DatasetBuilder._split_generators incomplete type annotation by @JonasLoos in https://github.com/huggingface/datasets/pull/6799CachedDatasetModuleFactory and Cache by @izhx in https://github.com/huggingface/datasets/pull/6754os.path.relpath in resolve_patterns by @mariosasko in https://github.com/huggingface/datasets/pull/6815Dataset.__getitem__ by @mariosasko in https://github.com/huggingface/datasets/pull/6817Full Changelog: https://github.com/huggingface/datasets/compare/2.18.0...2.19.0
Silence ruff deprecation messages by @mariosasko in https://github.com/huggingface/datasets/pull/6707
num_workers could lead to incorrect shards assignments to workers and cause errorsxlistdir by @mariosasko in https://github.com/huggingface/datasets/pull/6698Full Changelog: https://github.com/huggingface/datasets/compare/2.17.1...2.18.0
Remove deprecated verbose parameter from CSV builder by @albertvillanova in https://github.com/huggingface/datasets/pull/6672
arrow_writer.py from #6636 by @bryant1410 in https://github.com/huggingface/datasets/pull/6664Full Changelog: https://github.com/huggingface/datasets/compare/2.17.0...2.17.1
[WebDataset] Audio support and bug fixes by @lhoestq in https://github.com/huggingface/datasets/pull/6573
drop_last_batchin map after shuffling or sharding by @lhoestq in https://github.com/huggingface/datasets/pull/6575setup.cfg to pyproject.toml by @mariosasko in https://github.com/huggingface/datasets/pull/6619tqdm bars in non-interactive environments by @mariosasko in https://github.com/huggingface/datasets/pull/6627with_rank param to Dataset.filter by @mariosasko in https://github.com/huggingface/datasets/pull/6608Full Changelog: https://github.com/huggingface/datasets/compare/2.16.1...2.17.0
Fix dl_manager.extract returning FileNotFoundError by @lhoestq in https://github.com/huggingface/datasets/pull/6543
cache_dir to load_datasetload_dataset("ted_talks_iwslt", language_pair=("ja", "en"), year="2015")Full Changelog: https://github.com/huggingface/datasets/compare/2.16.0...2.16.1
Fix deprecation warning when building conda package by @albertvillanova in https://github.com/huggingface/datasets/pull/6425
https://hf.co/datasets/<repo_id>. A warning is shown to let the user know about the custom code, and they can avoid this message in future by passing the argument trust_remote_code=True.trust_remote_code=True will be mandatory to load these datasets from the next major release of datasets.HF_DATASETS_TRUST_REMOTE_CODE=0 you can already disable custom code by default without waiting for the next release of datasetshttps://hf.co/datasets/<repo_id>/tree/refs%2Fconvert%2Fparquetload_dataset step that lists the data files of big repositories (up to x100) but requires huggingface_hub 0.20 or newerload_dataset that used to reload data from cache even if the dataset was updated on Hugging Face~/.cache/huggingface/datasets/username___dataset_name/config_name/version/commit_shadatasets 2.15 (using the old scheme) are still reloaded from cache_get_data_files_patterns by @lhoestq in https://github.com/huggingface/datasets/pull/6343usedforsecurity=False in hashlib methods (FIPS compliance) by @Wauplin in https://github.com/huggingface/datasets/pull/6414ruff for formatting by @mariosasko in https://github.com/huggingface/datasets/pull/6434tqdm wrapper by @mariosasko in https://github.com/huggingface/datasets/pull/6433Table.__getstate__ and Table.__setstate__ by @LZHgrla in https://github.com/huggingface/datasets/pull/6444filelock package for file locking by @mariosasko in https://github.com/huggingface/datasets/pull/6445** by @mariosasko in https://github.com/huggingface/datasets/pull/6449dill logic by @mariosasko in https://github.com/huggingface/datasets/pull/6454push_to_hub by @mariosasko in https://github.com/huggingface/datasets/pull/6461__repr__ by @lhoestq in https://github.com/huggingface/datasets/pull/6480torch.Generator objects by @mariosasko in https://github.com/huggingface/datasets/pull/6502list_files_info with list_repo_tree in push_to_hub by @mariosasko in https://github.com/huggingface/datasets/pull/6510Full Changelog: https://github.com/huggingface/datasets/compare/2.15.0...2.16.0
Replace deprecated license_file in setup.cfg by @albertvillanova in https://github.com/huggingface/datasets/pull/6332
dl_manager.iter_files when they are given as input by @mariosasko in https://github.com/huggingface/datasets/pull/6230audio.py by @mariosasko in https://github.com/huggingface/datasets/pull/6241apache_beam import in BeamBasedBuilder._save_info by @mariosasko in https://github.com/huggingface/datasets/pull/6265tensorflow maximum version by @mariosasko in https://github.com/huggingface/datasets/pull/6301jax maximum version by @mariosasko in https://github.com/huggingface/datasets/pull/6300push_to_hub by @mariosasko in https://github.com/huggingface/datasets/pull/6269fsspec version to the datasets-cli env command output by @mariosasko in https://github.com/huggingface/datasets/pull/6356Dataset.map docstring by @bryant1410 in https://github.com/huggingface/datasets/pull/6373Image by @mariosasko in https://github.com/huggingface/datasets/pull/6379Full Changelog: https://github.com/huggingface/datasets/compare/2.14.7...2.15.0
Support pyarrow 14.0.1 and fix vulnerability CVE-2023-47248 by @albertvillanova in https://github.com/huggingface/datasets/pull/6404
Full Changelog: https://github.com/huggingface/datasets/compare/2.14.6...2.14.7
Ignore dataset_info.json in data files resolution by @mariosasko in https://github.com/huggingface/datasets/pull/6224
Full Changelog: https://github.com/huggingface/datasets/compare/2.14.5...2.14.6
Deprecate Dataset.export by @mariosasko in https://github.com/huggingface/datasets/pull/6081
iter_files for hidden files by @mariosasko in https://github.com/huggingface/datasets/pull/6092columns by @mariosasko in https://github.com/huggingface/datasets/pull/6160datasets_info.json but no README by @clefourrier in https://github.com/huggingface/datasets/pull/6164revision argument by @qgallouedec in https://github.com/huggingface/datasets/pull/6191Dataset.export by @mariosasko in https://github.com/huggingface/datasets/pull/6081download_custom by @mariosasko in https://github.com/huggingface/datasets/pull/6093select_columns to guide by @unifyh in https://github.com/huggingface/datasets/pull/6119to_iterable_dataset by @stevhliu in https://github.com/huggingface/datasets/pull/6158image_load doc by @mariosasko in https://github.com/huggingface/datasets/pull/6181huggingface/documentation-images by @mariosasko in https://github.com/huggingface/datasets/pull/6177hf-internal-testing repos for hosting test dataset repos by @mariosasko in https://github.com/huggingface/datasets/pull/6180Full Changelog: https://github.com/huggingface/datasets/compare/2.14.4...2.14.5
Fix authentication issues by @albertvillanova in https://github.com/huggingface/datasets/pull/6127
Full Changelog: https://github.com/huggingface/datasets/compare/2.14.3...2.14.4
Fix deprecation of use_auth_token in file_utils by @albertvillanova in https://github.com/huggingface/datasets/pull/6107
Full Changelog: https://github.com/huggingface/datasets/compare/2.14.2...2.14.3
Fix deprecation of use_auth_token in DownloadConfig by @albertvillanova in https://github.com/huggingface/datasets/pull/6094
Full Changelog: https://github.com/huggingface/datasets/compare/2.14.1...2.14.2
Remove README link to deprecated Colab notebook by @mariosasko in https://github.com/huggingface/datasets/pull/6080
Overview.ipynb & detach Jupyter Notebooks from datasets repository by @alvarobartt in https://github.com/huggingface/datasets/pull/5902Full Changelog: https://github.com/huggingface/datasets/compare/2.14.0...2.14.1
Deprecate errors param in favor of encoding_errors in text builder by @mariosasko in https://github.com/huggingface/datasets/pull/5974
datasets>=2.14.0 may not be reloaded from cache using older version of datasets (and therefore re-downloaded).Support for multiple configs via metadata yaml info by @polinaeterna in https://github.com/huggingface/datasets/pull/5331
---
configs:
- config_name: default
data_files:
- split: train
path: data.csv
- split: test
path: holdout.csv
---
---
configs:
- config_name: main_data
data_files: main_data.csv
- config_name: additional_data
data_files: additional_data.csv
---
Support for multiple configs via metadata yaml info by @polinaeterna in https://github.com/huggingface/datasets/pull/5331
push_to_hub() additional dataset configurationsds.push_to_hub("username/dataset_name", config_name="additional_data")
# reload later
ds = load_dataset("username/dataset_name", "additional_data")
Support returning dataframe in map transform by @mariosasko in https://github.com/huggingface/datasets/pull/5995
errors param in favor of encoding_errors in text builder by @mariosasko in https://github.com/huggingface/datasets/pull/5974huggingface_hub's RepoCard API by @mariosasko in https://github.com/huggingface/datasets/pull/5949joblib to avoid joblibspark test failures by @mariosasko in https://github.com/huggingface/datasets/pull/6000column_names type check with type hint in sort by @mariosasko in https://github.com/huggingface/datasets/pull/6001use_auth_token in favor of token by @mariosasko in https://github.com/huggingface/datasets/pull/5996ClassLabel min max check for None values by @mariosasko in https://github.com/huggingface/datasets/pull/6023task_templates in IterableDataset when they are no longer valid by @mariosasko in https://github.com/huggingface/datasets/pull/6027HfFileSystem and deprecate S3FileSystem by @mariosasko in https://github.com/huggingface/datasets/pull/6052Dataset.from_list docstring by @mariosasko in https://github.com/huggingface/datasets/pull/6062features are specified by @mariosasko in https://github.com/huggingface/datasets/pull/6045Full Changelog: https://github.com/huggingface/datasets/compare/2.13.1...2.14.0
Do not filter out .zip extensions from no-script datasets by @albertvillanova in https://github.com/huggingface/datasets/pull/6208
Full Changelog: https://github.com/huggingface/datasets/compare/2.13.1...2.13.2
Fix JSON generation in benchmarks CI by @mariosasko in https://github.com/huggingface/datasets/pull/5966
list_datasets by @mariosasko in https://github.com/huggingface/datasets/pull/5964encoding and errors params to JSON loader by @mariosasko in https://github.com/huggingface/datasets/pull/5969Full Changelog: https://github.com/huggingface/datasets/compare/2.13.0...2.13.1
Add IterableDataset.from_spark by @maddiedawson in https://github.com/huggingface/datasets/pull/5770
Add IterableDataset.from_spark by @maddiedawson in https://github.com/huggingface/datasets/pull/5770
from datasets import IterableDataset
from torch.utils.data import DataLoader
ids = IterableDataset.from_spark(df)
ids = ids.map(...).filter(...).with_format("torch")
for batch in DataLoader(ids, batch_size=16, num_workers=4):
...
IterableDataset formatting for PyTorch, TensorFlow, Jax, NumPy and Arrow:
from datasets import load_dataset
ids = load_dataset("c4", "en", split="train", streaming=True)
ids = ids.map(...).with_format("torch") # to get PyTorch tensors - also works with tf, np, jax etc.
Add IterableDataset.from_file to load local dataset as iterable by @mariusz-jachimowicz-83 in https://github.com/huggingface/datasets/pull/5893
from datasets import IterableDataset
ids = IterableDataset.from_file("path/to/data.arrow")
Arrow dataset builder to be able to load and stream Arrow datasets by @mariusz-jachimowicz-83 in https://github.com/huggingface/datasets/pull/5944
from datasets import load_dataset
ds = load_dataset("arrow", data_files={"train": "train.arrow", "test": "test.arrow"})
stopping_strategy of shuffled interleaved dataset (random cycling case) by @mariosasko in https://github.com/huggingface/datasets/pull/5816BuilderConfig by @Laurent2916 in https://github.com/huggingface/datasets/pull/5824accelerate as metric's test dependency to fix CI error by @mariosasko in https://github.com/huggingface/datasets/pull/5848date_format param to the CSV reader by @mariosasko in https://github.com/huggingface/datasets/pull/5845fn_kwargs to map and filter of IterableDataset and IterableDatasetDict by @yuukicammy in https://github.com/huggingface/datasets/pull/5810FixedSizeListArray casting by @mariosasko in https://github.com/huggingface/datasets/pull/5897DatasetBuilder.as_dataset when file_format is not "arrow" by @mariosasko in https://github.com/huggingface/datasets/pull/5915flatten_indices to DatasetDict by @maximxlss in https://github.com/huggingface/datasets/pull/5907batch_size optional, and minor improvements in Dataset.to_tf_dataset by @alvarobartt in https://github.com/huggingface/datasets/pull/5883to_numpy when None values in the sequence by @qgallouedec in https://github.com/huggingface/datasets/pull/5933Full Changelog: https://github.com/huggingface/datasets/compare/2.12.0...zef
Add Dataset.from_spark by @maddiedawson in https://github.com/huggingface/datasets/pull/5701
Add Dataset.from_spark by @maddiedawson in https://github.com/huggingface/datasets/pull/5701
>>> from datasets import Dataset
>>> ds = Dataset.from_spark(df)
Support streaming Beam datasets from HF GCS preprocessed data by @albertvillanova in https://github.com/huggingface/datasets/pull/5689
>>> from datasets import load_dataset
>>> ds = load_dataset("wikipedia", "20220301.de", streaming=True)
>>> next(iter(ds["train"]))
{'id': '1', 'url': 'https://de.wikipedia.org/wiki/Alan%20Smithee', 'title': 'Alan Smithee', 'text': 'Alan Smithee steht als Pseudonym für einen fiktiven Regisseur...}
Implement sharding on merged iterable datasets by @Hubert-Bonisseur in https://github.com/huggingface/datasets/pull/5735
>>> from datasets import load_dataset, interleave_datasets
>>> from torch.utils.data import DataLoader
>>> wiki = load_dataset("wikipedia", "20220301.en", split="train", streaming=True)
>>> c4 = load_dataset("c4", "en", split="train", streaming=True)
>>> merged = interleave_datasets([wiki, c4], probabilities=[0.1, 0.9], seed=42, stopping_strategy="all_exhausted")
>>> dataloader = DataLoader(merged, num_workers=4)
Consistent ArrayND Python formatting + better NumPy/Pandas formatting by @mariosasko in https://github.com/huggingface/datasets/pull/5751
Full Changelog: https://github.com/huggingface/datasets/compare/2.11.0...2.12.0
Deprecated batch_size on Dataset.to_dict()
batch_size on Dataset.to_dict()download_and_prepare() a datasetload_dataset():
from_dict by @mariosasko in https://github.com/huggingface/datasets/pull/5643ffmpeg system package installation on Colab by @polinaeterna in https://github.com/huggingface/datasets/pull/5558datasets.load_from_disk, DatasetDict.load_from_disk and Dataset.load_from_disk by @alvarobartt in https://github.com/huggingface/datasets/pull/5529huggingface_hub version to env cli command by @mariosasko in https://github.com/huggingface/datasets/pull/5578save_to_disk by @mariosasko in https://github.com/huggingface/datasets/pull/5588sort with indices mapping by @mariosasko in https://github.com/huggingface/datasets/pull/5587datasets-cli test by @lhoestq in https://github.com/huggingface/datasets/pull/5603verification_mode values by @polinaeterna in https://github.com/huggingface/datasets/pull/5607ruff by @polinaeterna in https://github.com/huggingface/datasets/pull/5636Features by @mariosasko in https://github.com/huggingface/datasets/pull/5646fsspec.open when using an HTTP proxy by @bryant1410 in https://github.com/huggingface/datasets/pull/5656Full Changelog: https://github.com/huggingface/datasets/compare/2.10.0...2.11.0
Fix sort with indices mapping by @mariosasko https://github.com/huggingface/datasets/pull/5587
IndexError when doing ds.filter(...).sort(...) or ds.select(...).sort(...)Full Changelog: https://github.com/huggingface/datasets/compare/2.10.0...2.10.1
Avoid saving sparse ChunkedArrays in pyarrow tables by @marioga in https://github.com/huggingface/datasets/pull/5542
.flatten_indices() (x2) + save/load_from_disk (x100) on selected/shuffled datasetsverification_mode you can pass to `load_dataset()):.map() in multiprocessing.to_iterable_dataset() to get a IterableDataset from a DatasetIterableDataset in the documentation about the differences between Dataset and IterableDataset.select_column() to return a dataset only containing the requested columnsds = ds.sort(['col_1', 'col_2'], reverse=[True, False])ds = ds.with_format("jax", device=device)nyu_depth_v2 dataset by @awsaf49 in https://github.com/huggingface/datasets/pull/5484load_from_cache_file arg from Dataset.shard() docstring by @polinaeterna in https://github.com/huggingface/datasets/pull/5493NumpyFormatter by @alvarobartt in https://github.com/huggingface/datasets/pull/5530load_from_cache_file type and logic by @HallerPatrick in https://github.com/huggingface/datasets/pull/5515ruff by @mariosasko in https://github.com/huggingface/datasets/pull/5519Full Changelog: https://github.com/huggingface/datasets/compare/2.9.0...ef
Fix deprecation warning when use_auth_token passed to download_and_prepare by @albertvillanova in https://github.com/huggingface/datasets/pull/5409
Parallel implementation of to_tf_dataset() by @Rocketknight1 in https://github.com/huggingface/datasets/pull/5377
num_workers= to .to_tf_dataset() to make your dataset faster with multiprocessingDistributed support by @lhoestq in https://github.com/huggingface/datasets/pull/5369
Dataset and IterableDataset (e.g. in streaming mode)import os
from datasets.distributed import split_dataset_by_node
rank = int(os.environ["RANK"])
world_size = int(os.environ["WORLD_SIZE"])
ds = split_dataset_by_node(ds, rank=rank, world_size=world_size)
Support streaming datasets with os.path.exists and Path.exists by @albertvillanova in https://github.com/huggingface/datasets/pull/5400
Tqdm progress bar for to_parquet by @zanussbaum in https://github.com/huggingface/datasets/pull/5456
ZIP files support in iter_archive with better compression type check by @Mehdi2402 in https://github.com/huggingface/datasets/pull/3379
Support other formats than uint8 for image arrays by @vigsterkr in https://github.com/huggingface/datasets/pull/5365
fs.open resource leaks by @tkukurin in https://github.com/huggingface/datasets/pull/5358cast_to_python_objects by @mariosasko in https://github.com/huggingface/datasets/pull/5384load_dataset docstring by @mariosasko in https://github.com/huggingface/datasets/pull/5389shard_size arg from .push_to_hub() by @polinaeterna in https://github.com/huggingface/datasets/pull/5469Full Changelog: https://github.com/huggingface/datasets/compare/2.8.0...2.9.0
Your coding agent can read these notes before it upgrades. Set up the MCP server →