NewYour coding agent can read the release notes before it upgrades.Set up the MCP server →
PyPI · #257 most downloaded on PyPI
HuggingFace community-driven open-source library of datasets
Last release 1 months ago
28 Jul 2026
Ships fairly regularly
a new release about every 4 weeks
Nearly every release is documented
notes for 60 of the last 60 stable releases
Nothing withdrawn
no release was ever pulled
11 years old
119 releases · first in 2015
Improve use_auth_token docstring and deprecate use_auth_token in download_and_prepare by @mariosasko in https://github.com/huggingface/datasets/pull/5…
datasets are not able to reload datasets pushed with this new model, so we encourage everyone to update.IterableDataset.map that lead to features=None by @alvarobartt in https://github.com/huggingface/datasets/pull/5287
features after column renaming or removalfeatures param to IterableDataset.map by @alvarobartt in https://github.com/huggingface/datasets/pull/5311num_shards or max_shard_size to ds.save_to_disk() or ds.push_to_hub()num_proc to use multiprocessing.from datasets import load_dataset
ds = load_dataset("c4", "en", streaming=True, split="train")
dataloader = DataLoader(ds, batch_size=32, num_workers=4)
max_shard_size docs by @lhoestq in https://github.com/huggingface/datasets/pull/5267from_generator docs by @mariosasko in https://github.com/huggingface/datasets/pull/5307wikipedia or natural_questionsArrowWriter.finalize before inference error by @mariosasko in https://github.com/huggingface/datasets/pull/5309num_proc for dataset download and generation by @mariosasko in https://github.com/huggingface/datasets/pull/5300IterableDataset.map param batch_size typing as optional by @alvarobartt in https://github.com/huggingface/datasets/pull/5336topdown parameter in xwalk by @mariosasko in https://github.com/huggingface/datasets/pull/5308use_auth_token docstring and deprecate use_auth_token in download_and_prepare by @mariosasko in https://github.com/huggingface/datasets/pull/5302.tar archives in the same way as for .tar.gz and .tgz in _get_extraction_protocol by @polinaeterna in https://github.com/huggingface/datasets/pull/5322Full Changelog: https://github.com/huggingface/datasets/compare/2.7.0...2.8.0
One column per quarter.
Remove YAML integer keys from class_label metadata by @albertvillanova in https://github.com/huggingface/datasets/pull/5277
Full Changelog: https://github.com/huggingface/datasets/compare/2.7.0...2.7.1
Deprecate num_proc parameter in DownloadManager.extract by @ayushthe1 in https://github.com/huggingface/datasets/pull/5142
from datasets import load_dataset
ds = load_dataset("imagenet-1k", num_proc=4)
map or filter that uses tensors or pipelines can now be cachedpyproject.toml for black by @mariosasko in https://github.com/huggingface/datasets/pull/5125tqdm zip bug by @david1542 in https://github.com/huggingface/datasets/pull/5120class_encode_column by @mariosasko in https://github.com/huggingface/datasets/pull/5130writer_batch_size by @mariosasko in https://github.com/huggingface/datasets/pull/5163co_filenames to remove by @gpucce in https://github.com/huggingface/datasets/pull/5169DownloadConfig.use_auth_token value by @alvarobartt in https://github.com/huggingface/datasets/pull/5205typer version in tests to <0.5 to fix Windows CI by @polinaeterna in https://github.com/huggingface/datasets/pull/5235Version hashable by @mariosasko in https://github.com/huggingface/datasets/pull/5238Full Changelog: https://github.com/huggingface/datasets/compare/2.6.1...2.7.0
Remove YAML integer keys from class_label metadata by @albertvillanova in https://github.com/huggingface/datasets/pull/5277
Full Changelog: https://github.com/huggingface/datasets/compare/2.6.1...2.6.2
Fix filter indices when batched by @albertvillanova in https://github.com/huggingface/datasets/pull/5113
filter could return examples with the wrong indicesmap with batch=True could return a dataset with less examplesFull Changelog: https://github.com/huggingface/datasets/compare/2.6.0...2.6.1
Add deprecation warning to multilingual_librispeech dataset card by @albertvillanova in https://github.com/huggingface/datasets/pull/5010
from datasets import Dataset
dataset = Dataset.from_sql("data_table", "sqlite:///sqlite_file.db")
from_sql + small doc improvement by @mariosasko in https://github.com/huggingface/datasets/pull/5091from datasets import Dataset
from sqlite3 import connect
con = connect(...)
dataset = Dataset.from_sql("SELECT text FROM table WHERE length(text) > 100 LIMIT 10", con)
from datasets import load_dataset
ds = load_dataset("imagenet-1k").with_format("torch") # or numpy/tf/jax
ds[0]["image"]
IterableDataset.from_generator by @hamid-vakilzadeh in https://github.com/huggingface/datasets/pull/5052kwargs to Dataset.from_generator by @mariosasko in https://github.com/huggingface/datasets/pull/5049converters in CsvBuilder by @mariosasko in https://github.com/huggingface/datasets/pull/5057load_from_disk by @asofiaoliveira in https://github.com/huggingface/datasets/pull/5073ClassLabel docstring example by @alvarobartt in https://github.com/huggingface/datasets/pull/5029flatten_indices with empty indices mapping by @mariosasko in https://github.com/huggingface/datasets/pull/5043hffs by @lhoestq in https://github.com/huggingface/datasets/pull/5101Full Changelog: https://github.com/huggingface/datasets/compare/2.5.1...2.6.0
Revert task removal in folder-based builders
Full Changelog: https://github.com/huggingface/datasets/compare/2.5.1...2.5.2
Revert input_columns change by @lhoestq in https://github.com/huggingface/datasets/pull/5006
Full Changelog: https://github.com/huggingface/datasets/compare/2.5.0...2.5.1
Deprecate metrics by @albertvillanova in https://github.com/huggingface/datasets/pull/4739
!pip install evaluate
import evaluate
metric = evaluate.load("accuracy")
Dataset.from_list by @sanderland in https://github.com/huggingface/datasets/pull/4890Dataset.from_generator by @mariosasko in https://github.com/huggingface/datasets/pull/4957input_colums in Dataset.map if input_columns are specified by @mariosasko in https://github.com/huggingface/datasets/pull/4971fn_kwargs param to IterableDataset.map by @mariosasko in https://github.com/huggingface/datasets/pull/4975language_bcp47 tag by @lhoestq in https://github.com/huggingface/datasets/pull/4753Full Changelog: https://github.com/huggingface/datasets/compare/2.4.0...2.5.0
Replace deprecated logging.warn with logging.warning by @hugovk in https://github.com/huggingface/datasets/pull/4539
concatenate_datasets for iterable datasets by @lhoestq in https://github.com/huggingface/datasets/pull/4500metadata.jsonl from parent directories in imagefolder @mariosasko in https://github.com/huggingface/datasets/pull/4576ArrowWriter.write_batch when batch is empty by @alvarobartt in https://github.com/huggingface/datasets/pull/4510batch_size parameter when calling add_faiss_index and add_faiss_index_from_external_arrays by @alvarobartt in https://github.com/huggingface/datasets/pull/4535load_dataset by @mariosasko in https://github.com/huggingface/datasets/pull/4577_arrow_to_datasets_dtype conversion by @mariosasko in https://github.com/huggingface/datasets/pull/4628assertEqual with assertTupleEqual in unit tests for verbosity by @alvarobartt in https://github.com/huggingface/datasets/pull/4496embed_storage on features inside lists/sequences by @mariosasko in https://github.com/huggingface/datasets/pull/4615from_pandas more robust by @mariosasko in https://github.com/huggingface/datasets/pull/4703DatasetInfo/Features by @mariosasko in https://github.com/huggingface/datasets/pull/4741Full Changelog: https://github.com/huggingface/datasets/compare/2.3.2...2.4.0
Fix double dots in data files by @lhoestq in https://github.com/huggingface/datasets/pull/4505
/../ is passed to data_files causing FileNotFoundErrorFull Changelog: https://github.com/huggingface/datasets/compare/2.3.1...2.3.2
Fix patching module that doesn't exist by @lhoestq in https://github.com/huggingface/datasets/pull/4495
DownloadConfig, DownloadMode, DownloadManagerFull Changelog: https://github.com/huggingface/datasets/compare/2.3.0...2.3.1
Update CI deprecated legacy image by @albertvillanova in https://github.com/huggingface/datasets/pull/4393
load_dataset without requiring a manual download !load_dataset("imagenet-1k", streaming=True)train_test_split by @nandwalritik in https://github.com/huggingface/datasets/pull/4322push_to_hub: skip identical files in push_to_hub instead of overwriting by @mariosasko in https://github.com/huggingface/datasets/pull/4402features in packaged loaders by @mariosasko in https://github.com/huggingface/datasets/pull/4364scene_parse_150 card by @mariosasko in https://github.com/huggingface/datasets/pull/4447new_fingerprint by @fxmarty in https://github.com/huggingface/datasets/pull/4326iter_files by @mariosasko in https://github.com/huggingface/datasets/pull/4412dataset_infos.json with new split info in Dataset.push_to_hub to avoid verification error by @mariosasko in https://github.com/huggingface/datasets/pull/4415inspect_dataset and inspect_metric by @mariosasko in https://github.com/huggingface/datasets/pull/4433_format_columns in remove_columns by @alvarobartt in https://github.com/huggingface/datasets/pull/4411Full Changelog: https://github.com/huggingface/datasets/compare/2.2.2...lol
Fix: irc_disentangle - fix checksum and bug dataset by @albertvillanova in https://github.com/huggingface/datasets/pull/4377
DatasetDict.push_to_hub by @mariosasko in https://github.com/huggingface/datasets/pull/4372transformers - pinning the version to <0.3.5 for nowFull Changelog: https://github.com/huggingface/datasets/compare/2.2.1...2.2.2
Fix cnn_dailymail (dm stories were ignored) by @lhoestq in https://github.com/huggingface/datasets/pull/4317
datasets 2.2.0 introduced a bug in cnn_dailymail and some examples were missing in the datasetFull Changelog: https://github.com/huggingface/datasets/compare/2.2.0...2.2.1
Deprecate shard_size in push_to_hub in favor of max_shard_size by @mariosasko in https://github.com/huggingface/datasets/pull/4190
path exists by @patrickvonplaten in https://github.com/huggingface/datasets/pull/4212imagefolder by @mariosasko in https://github.com/huggingface/datasets/pull/4069
metadata.jsonl, more info in the documentation on how to load an image datasetdata_dir parameter when loading datasets without script by @polinaeterna in https://github.com/huggingface/datasets/pull/4144
drop_last_batch to IterableDataset.map by @mariosasko in https://github.com/huggingface/datasets/pull/4215train-deval-index metadata to automate evaluation on your datasets based on their taskshuggingface_hub by @julien-c in https://github.com/huggingface/datasets/pull/4154shard_size in push_to_hub in favor of max_shard_size by @mariosasko in https://github.com/huggingface/datasets/pull/4190convert_file_size_to_int for kilobits and megabits by @mariosasko in https://github.com/huggingface/datasets/pull/4205faiss import to fix https://github.com/huggingface/datasets/issues/4287 by @alvarobartt in https://github.com/huggingface/datasets/pull/4288Full Changelog: https://github.com/huggingface/datasets/compare/2.1.0...2.2.0
Deprecated: Multilingual Librispeech - deprecate dataset in favor of facebook/multilingual_librispeechby @polinaeterna in https://github.com/huggingfa…
facebook/multilingual_librispeechby @polinaeterna in https://github.com/huggingface/datasets/pull/4060PIL.Image file handler in Image.decode_example by @mariosasko in https://github.com/huggingface/datasets/pull/3995map remove_columns on empty dataset by @lhoestq in https://github.com/huggingface/datasets/pull/4021push_to_hub by @lhoestq in https://github.com/huggingface/datasets/pull/4081cast_to_python_objects in TypedSequence by @mariosasko in https://github.com/huggingface/datasets/pull/4128Full Changelog: https://github.com/huggingface/datasets/compare/2.0.0...2.1.0
Remove deprecated methods/params (preparation for v2.0) by @mariosasko in https://github.com/huggingface/datasets/pull/3803
We're happy to announce that our new documentation is available at hf.co/docs/datasets !
imagefolder dataset loader:
push_to_hub:
Audio and Image feature in push_to_hub by @mariosasko in https://github.com/huggingface/datasets/pull/3685IterableDataset.filter by @lhoestq in https://github.com/huggingface/datasets/pull/3826IterableDataset (rename columns, cast, etc.) by @lhoestq in https://github.com/huggingface/datasets/pull/3862to_json by @bhavitvyamalik in https://github.com/huggingface/datasets/pull/3551FaissIndex by @rentruewang in https://github.com/huggingface/datasets/pull/3721map and shuffle for datasets loaded in streaming mode:
map when streaming: update instead of overwrite + add missing parameters by @lhoestq in https://github.com/huggingface/datasets/pull/3801IterableDataset.shuffle with Dataset.shuffle by @lhoestq in https://github.com/huggingface/datasets/pull/3842remove_columns param in filter by @mariosasko in https://github.com/huggingface/datasets/pull/3827module.builder_kwargs over defaults in TestCommand by @lvwerra in https://github.com/huggingface/datasets/pull/3672Dataset.select are within bounds by @mariosasko in https://github.com/huggingface/datasets/pull/3719push_to_hub by @mariosasko in https://github.com/huggingface/datasets/pull/3732ignore_verifications is True by @mariosasko in https://github.com/huggingface/datasets/pull/3796data_dir to data_files resolution and misc improvements to HfFileSystem by @mariosasko in https://github.com/huggingface/datasets/pull/3791predictions/references in Metric.compute by @mariosasko in https://github.com/huggingface/datasets/pull/3824ignore_verifications=True by @mariosasko in https://github.com/huggingface/datasets/pull/3868Full Changelog: https://github.com/huggingface/datasets/compare/1.18.3...0.0.0
Prioritize module.builder_kwargs over defaults in TestCommand #3672 (@lvwerra)
module.builder_kwargs over defaults in TestCommand #3672 (@lvwerra)Dataset.select are within bounds #3719 (@mariosasko)Full Changelog: https://github.com/huggingface/datasets/compare/1.18.3...1.18.4
Fix MP3 resampling when a dataset's audio files have different sampling rates by @lhoestq in https://github.com/huggingface/datasets/pull/3665
get_dataset_split_names by @mariosasko in https://github.com/huggingface/datasets/pull/3657Full Changelog: https://github.com/huggingface/datasets/compare/1.18.2...1.18.3
Fix streaming datasets that are not reset correctly by @lhoestq in https://github.com/huggingface/datasets/pull/3646
None by @mariosasko in https://github.com/huggingface/datasets/pull/3642add_column on datasets with indices mapping by @mariosasko in https://github.com/huggingface/datasets/pull/3647Full Changelog: https://github.com/huggingface/datasets/compare/1.18.1...1.18.2
Make decoding of Audio and Image feature optional by @mariosasko in https://github.com/huggingface/datasets/pull/3430
prepare_for_task() by @mariosasko in https://github.com/huggingface/datasets/pull/3614Full Changelog: https://github.com/huggingface/datasets/compare/1.18.0...1.18.1
Add VCTK dataset by @jaketae in https://github.com/huggingface/datasets/pull/3351
iter_files instead of str(Path(...) in image dataset by @mariosasko in https://github.com/huggingface/datasets/pull/3477ImageClassifcation task template by @mariosasko in https://github.com/huggingface/datasets/pull/3557DuplicatedKeysError and improve card by @mariosasko in https://github.com/huggingface/datasets/pull/3559gzip for to_json by @bhavitvyamalik in https://github.com/huggingface/datasets/pull/3492preserve_index to from_pandas by @Sorrow321 in https://github.com/huggingface/datasets/pull/3565str(Path(...)) conversion in streaming on Linux by @mariosasko in https://github.com/huggingface/datasets/pull/3472pretty_name for first 200 datasets by @bhavitvyamalik in https://github.com/huggingface/datasets/pull/3498pretty_name for all the other datasets by @bhavitvyamalik in https://github.com/huggingface/datasets/pull/3536Iterable.map call by @mariosasko in https://github.com/huggingface/datasets/pull/3556Full Changelog: https://github.com/huggingface/datasets/compare/1.17.0...1.18.0
Add The Pile dataset and PubMed Central subset by @albertvillanova in https://github.com/huggingface/datasets/pull/3287
None handling by @mariosasko in https://github.com/huggingface/datasets/pull/3195cast_column to IterableDataset by @mariosasko in https://github.com/huggingface/datasets/pull/3439Full Changelog: https://github.com/huggingface/datasets/compare/1.16.1...1.17.0
Fix import datasets on python 3.10 by @lhoestq in https://github.com/huggingface/datasets/pull/3326
datasets on python 3.10 by @lhoestq in https://github.com/huggingface/datasets/pull/3326Deprecate prepare_module by @albertvillanova in https://github.com/huggingface/datasets/pull/3166
Dataset and DatasetDict by @LysandreJik in https://github.com/huggingface/datasets/pull/3098:
push_to_hub() method !with_rank arg to pass process rank to map by @TevenLeScao in https://github.com/huggingface/datasets/pull/3314to_tf_dataset by @stevhliu in https://github.com/huggingface/datasets/pull/3175len(predictions) doesn't match len(references) in metrics by @mariosasko in https://github.com/huggingface/datasets/pull/3160fingerprint.py, search.py, arrow_writer.py and metric.py by @Ishan-Kumar2 in https://github.com/huggingface/datasets/pull/3305Full Changelog: https://github.com/huggingface/datasets/compare/1.15.1...1.16.0
Bump huggingface_hub to 0.1.0 by @lhoestq in https://github.com/huggingface/datasets/pull/3199
Fix numpy deprecation warning for ragged tensors by @lhoestq in https://github.com/huggingface/datasets/pull/3137
to_csv by @bhavitvyamalik in https://github.com/huggingface/datasets/pull/2896to_tf_dataset by @Rocketknight1 in https://github.com/huggingface/datasets/pull/3085zip_dict by @mariosasko in https://github.com/huggingface/datasets/pull/3170Update: LexGLUE and MultiEURLEX README - update dataset cards #3075 (@iliaschalkidis)
Update: Adapt all audio datasets #3081 (@patrickvonplaten)
Fix error related to huggingface_hub timeout parameter #3082 (@albertvillanova)
Fix loading a metric with internal import #3077 (@albertvillanova)
The script_version parameter in load_dataset is now deprecated, in favor of revision
to_tf_dataset method #2731 #2931 #2951 #2974 (@Rocketknight1)remove_columns to IterableDataset #3030 (@cccntu)get_dataset_split_names() to get a dataset config's split names #2906 (@severo)script_version parameter in load_dataset is now deprecated, in favor of revisionfilter several times in a row was not returning the right results in 1.12.0 and 1.12.1filter #2947 (@lhoestq)read_csv parameters #2960 (@SBrandeis)prepare_module function doesn't support the return_resolved_file_path and return_associated_base_path parameters. As an alternative, you may use the dataset_module_factory instead.Fix fsspec AbstractFileSystem access #2915 (@pierre-godard)
ArrowInvalid: Can only convert 1-dimensional array values errorsNew documentation structure #2718 (@stevhliu):
## New documentation - New documentation structure #2718 (@stevhliu):
New: Tutorials
New: Hot-to guides
New: Conceptual guides
Update: Reference
See the new documentation [here](https://huggingface.co/docs/datasets/) !
## Datasets changes - New: VIVOS dataset for Vietnamese ASR #2780 (@binh234) - New: The Pile books3 #2801 (@richarddwang) - New: The Pile stack exchange #2803 (@richarddwang) - New: The Pile openwebtext2 #2802 (@richarddwang) - New: Food-101 #2804 (@nateraw) - New: Beans #2809 (@nateraw) - New: cedr #2796 (@naumov-al) - New: cats_vs_dogs #2807 (@nateraw) - New: MultiEURLEX #2865 (@iliaschalkidis) - New: BIOSSES #2881 (@bwang482) - Update: TTC4900 - add download URL #2732 (@yavuzKomecoglu) - Update: Wikihow - Generate metadata JSON for wikihow dataset #2748 (@albertvillanova) - Update: lm1b - Generate metadata JSON #2752 (@albertvillanova) - Update: reclor - Generate metadata JSON #2753 (@albertvillanova) - Update: telugu_books - Generate metadata JSON #2754 (@albertvillanova) - Update: SUPERB - Add SD task #2661 (@albertvillanova) - Update: SUPERB - Add KS task #2783 (@anton-l) - Update: GooAQ - add train/val/test splits #2792 (@bhavitvyamalik) - Update: Openwebtext - update size #2857 (@lhoestq) - Update: timit_asr - make the dataset streamable #2835 (@lhoestq) - Fix: journalists_questions -fix key by recreating metadata JSON #2744 (@albertvillanova) - Fix: turkish_movie_sentiment - fix metadata JSON #2755 (@albertvillanova) - Fix: ubuntu_dialogs_corpus - fix metadata JSON #2756 (@albertvillanova) - Fix: CNN/DailyMail - typo #2791 (@omaralsayed) - Fix: linnaeus - fix url #2852 (@lhoestq) - Fix ToTTo - fix data URL #2864 (@albertvillanova) - Fix: wikicorpus - fix keys #2844 (@lhoestq) - Fix: COUNTER - fix bad file name #2894 (@albertvillanova) - Fix: DocRED - fix data URLs and metadata #2883 (@albertvillanova)
## Datasets features - Load Dataset from the Hub (NO DATASET SCRIPT) #2662 (@lhoestq) - Preserve dtype for numpy/torch/tf/jax arrays #2361 (@bhavitvyamalik) - add multi-proc in to_json #2747 (@bhavitvyamalik) - Optimize Dataset.filter to only compute the indices to keep #2836 (@lhoestq)
## Dataset streaming - better support for compression: - Fix streaming zip files #2798 (@albertvillanova) - Support streaming tar files #2800 (@albertvillanova) - Support streaming compressed files (gzip, bz2, lz4, xz, zst) #2786 (@albertvillanova) - Fix streaming zip files from canonical datasets #2805 (@albertvillanova) - Add url prefix convention for many compression formats #2822 (@lhoestq) - Support streaming datasets that use pathlib #2874 (@albertvillanova) - Extend support for streaming datasets that use pathlib.Path stem/suffix #2880 (@albertvillanova) - Extend support for streaming datasets that use pathlib.Path.glob #2876 (@albertvillanova)
## Metrics changes - Update: BERTScore - Add support for fast tokenizer #2770 (@mariosasko) - Fix: Sacrebleu - Fix sacrebleu tokenizers #2739 #2778 #2779 (@albertvillanova)
## Dataset cards - Updated dataset description of DaNE #2789 (@KennethEnevoldsen) - Update ELI5 README.md #2848 (@odellus)
## General improvements and bug fixes - Update release instructions #2740 (@albertvillanova) - Raise ManualDownloadError when loading a dataset that requires previous manual download #2758 (@albertvillanova) - Allow PyArrow from source #2769 (@patrickvonplaten) - fix typo (ShuffingConfig -> ShufflingConfig) #2766 (@daleevans) - Fix typo in test_dataset_common #2790 (@nateraw) - Fix type hint for data_files #2793 (@albertvillanova) - Bump tqdm version #2814 (@mariosasko) - Use packaging to handle versions #2777 (@albertvillanova) - Tiny typo fixes of "fo" -> "of" #2815 (@aronszanto) - Rename The Pile subsets #2817 (@lhoestq) - Fix IndexError by ignoring empty RecordBatch #2834 (@lhoestq) - Fix defaults in cache_dir docstring in load.py #2824 (@mariosasko) - Fix extraction protocol inference from urls with params #2843 (@lhoestq) - Fix caching when moving script #2854 (@lhoestq) - Fix windows CI CondaError #2855 (@lhoestq) - fix: 🐛 remove URL's query string only if it's ?dl=1 #2856 (@severo) - Update column_names showed as :func: in exploring.st #2851 (@ClementRomac) - Fix s3fs version in CI #2858 (@lhoestq) - Fix three typos in two files for documentation #2870 (@leny-mi) - Move checks from _map_single to map #2660 (@mariosasko) - fix regex to accept negative timezone #2847 (@jadermcs) - Prevent .map from using multiprocessing when loading from cache #2774 (@thomasw21) - Fix null sequence encoding #2900 (@lhoestq)
New: Add Russian SuperGLUE #2668 (@slowwavesleep)
tokenize_exemple #2726 (@shabie)The error message to tell which dataset config name to load was not displayed:
The error message to tell which dataset config name to load was not displayed:
Docstrings:
Fix minimum tqdm version and import on Colab #2697 (@nateraw)
Support remote data files #2616 (@albertvillanova) This allows to pass URLs of remote data files to any dataset loader: `python load_dataset("csv", da
load_dataset("csv", data_files={"train": [url_to_one_csv_file, url_to_another_csv_file...]})
This works for all these dataset loaders:
streaming=True. Main contributions:
filter with multiprocessing in case all samples are discarded #2601 (@mxschmdt)New: MasakhaNER #2465 (@dadelani)
Dataset.map #2540 (@lewtun)n>1M size tag #2527 (@lhoestq)New: Microsoft CodeXGlue Datasets #2357 (@madlag @ncoop57)
desc parameter in map for DatasetDict object #2423 (@bhavitvyamalik)Dataset.cast can now change the feature types of Sequence fieldskeep_in_memory=True when loading a dataset to load it in memoryNew: NLU evaluation data #2238 (@dkajtoch)
desc to tqdm in Dataset.map() #2374 (@bhavitvyamalik)key type and duplicates verification with hashing #2245 (@NikhilBartwal)Fix memory issue: don't copy recordbatches in memory during a table deepcopy #2291 (@lhoestq) This affected methods like concatenate_datasets, multipr
Fix memory issue: don't copy recordbatches in memory during a table deepcopy #2291 (@lhoestq)
This affected methods like concatenate_datasets, multiprocessed map and load_from_disk.
Breaking change:
Dataset.map with the input_columns parameter, the resulting dataset will only have the columns from input_columns and the columns added by the map functions. The other columns are discarded.Fix memory issue in multiprocessing: Don't pickle table index #2264 (@lhoestq)
Fix memory issue in multiprocessing: Don't pickle table index #2264 (@lhoestq)
Revert breaking change in cache_files property #2217 (@lhoestq)
seqeval metric #2204 (@marrodion)New: Europarl Bilingual #1874 (@lucadiliello)
Fix an issue #1981 with WMT downloads #1982 (@albertvillanova)
Fix an issue #1981 with WMT downloads #1982 (@albertvillanova)
The old in-place methods rename_column_, remove_columns_, flatten_ and cast_ are now deprecated.
ADD S3 support for downloading and uploading processed datasets
Fast start up (#1690): Importing datasets is now significantly faster.
datasets is now significantly faster.Moreover many dataset cards of datasets added during the sprint were updated ! Thanks to all the contributors :)
Intermediate release before v2.0.0 Includes all the datasets added during the datasets sprint of December 2020 (currently over 610 datasets).
Intermediate release before v2.0.0
Includes all the datasets added during the datasets sprint of December 2020 (currently over 610 datasets).
New: ASNQ - answer sentence selection
read_options, parse_options and convert_options are replaced with plain parameters like pandas.read_csvFix: text - use python read instead of pandas reader (#715):
data_files per splits using NamedSplit (#706)Nothing published for this version
fix numerous windows specific issues
dillUpdate: GLUE - update GLUE urls (now hosted on FB)
add multiprocessing to dataset dict
Fix:
Update now with ` pip install datasets `
Update now with
pip install datasets
map and filter (#552)input_column parameter in map and filter(#475)script_version with the env variable HF_SCRIPTS_VERSION (#584)HF_MODULES_CACHE (#574)map results (#601)select method for pyarrow < 1.0.0 (#585)Nothing published for this version
Your coding agent can read these notes before it upgrades. Set up the MCP server →