NewYour coding agent can read the release notes before it upgrades.Set up the MCP server →
PyPI · #89 most downloaded on PyPI
Client library to download and publish models, datasets and other repos on the huggingface.co hub
Last release 6 days ago
24 Sep 2026
Ships on a steady schedule
a new release about every 2 weeks
Nearly every release is documented
notes for 56 of the last 60 stable releases
7 versions withdrawn
withdrawn after publishing
6 years old
328 releases · first in 2020
One column per quarter.
On Python3.8, it is fairly easy to get a corrupted install of pydantic (more specificially, pydantic 2.x cannot run if tensorflow is installed because
On Python3.8, it is fairly easy to get a corrupted install of pydantic (more specificially, pydantic 2.x cannot run if tensorflow is installed because of an incompatible requirement on typing_extensions). Since pydantic is an optional dependency of huggingface_hub, we do not want to crash at huggingface_hub import time if pydantic install is corrupted. However this was the case because of how imports are made in huggingface_hub. This hot-fix releases fixes this bug. If pydantic is not correctly installed, we only raise a warning and continue as if it was not installed at all.
Related PR: https://github.com/huggingface/huggingface_hub/pull/1829
Full Changelog: https://github.com/huggingface/huggingface_hub/compare/v0.19.3...v0.19.4
Hot-fix release after https://github.com/huggingface/huggingface_hub/pull/1828.
Hot-fix release after https://github.com/huggingface/huggingface_hub/pull/1828.
In 0.19.0 we've loosen pydantic requirements to accept both 1.x and 2.x since huggingface_hub is compatible with both. However, it started to cause issues when installing both huggingface_hub[inference] and tensorflow in a Python3.8 environment. The problem comes from the fact that on Python3.8, Pydantic>=2.x and tensorflow don't seem to be compatible. Tensorflow depends on
typing_extension<=4.5.0 while pydantic 2.x requires typing_extensions>=4.6. This causes a ImportError: cannot import name 'TypeAliasType' from 'typing_extensions'. when importing huggingface_hub.
As a side note, tensorflow support for Python3.8 has been dropped since 2.14.0. Therefore this issue should affect less and less users over time.
Full Changelog: https://github.com/huggingface/huggingface_hub/compare/v0.19.2...v0.19.3
In https://github.com/huggingface/huggingface_hub/pull/1786 (already release in 0.19.0), we harmonized the environment variables in the HF ecosystem w
Not a hot-fix.
In https://github.com/huggingface/huggingface_hub/pull/1786 (already release in 0.19.0), we harmonized the environment variables in the HF ecosystem with the goal to propagate this harmonization to other HF libraries. In this work, we forgot to expose HF_HOME as a constant value that can be reused, especially by transformers or datasets. This release fixes this (see https://github.com/huggingface/huggingface_hub/pull/1825).
Full Changelog: https://github.com/huggingface/huggingface_hub/compare/v0.19.1...v0.19.2
Fixes a regression bug (PR https://github.com/huggingface/huggingface_hub/pull/1821) introduced in 0.19.0 that made looping over models with list_mode
Full Changelog: https://github.com/huggingface/huggingface_hub/compare/v0.19.0...v0.19.1.
Fixes a regression bug (PR https://github.com/huggingface/huggingface_hub/pull/1821) introduced in 0.19.0 that made looping over models with list_models fail. The problem came from the fact that we are now parsing the data returned by the server into Python objects. However for some models the metadata in the model card is not valid. This is usually checked by the server but some models created before we started to enforce correct metadata are not valid. This hot-fix fixes the issue by ignoring the corrupted data, if any.
Every change is meant to be backward compatible, meaning no breaking changes is expected. However, if you detect any inconsistency, please let us know…
(Discuss about the release in our Community Tab. Feedback welcome!! 🤗)
Inference Endpoints provides a secure solution to easily deploy models hosted on the Hub in a production-ready infrastructure managed by Huggingface. With huggingface_hub>=0.19.0 integration, you can now manage your Inference Endpoints programmatically. Combined with the InferenceClient, this becomes the go-to solution to deploy models and run jobs in production, either sequentially or in batch!
Here is an example how to get an inference endpoint, wake it up, wait for initialization, run jobs in batch and pause back the endpoint. All of this in a few lines of code! For more details, please check out our dedicated guide.
>>> import asyncio
>>> from huggingface_hub import get_inference_endpoint
# Get endpoint + wait until initialized
>>> endpoint = get_inference_endpoint("batch-endpoint").resume().wait()
# Run inference
>>> async_client = endpoint.async_client
>>> results = asyncio.gather(*[async_client.text_generation(...) for job in jobs])
# Pause endpoint
>>> endpoint.pause()
huggingface_hub is a library primarily used to transfer (huge!) files with the Huggingface Hub. Our goal is to keep improving the experience for this core part of the library. In this release, we introduce a more robust download mechanism for slow/limited connection while improving the UX for users with a high bandwidth available!
Getting a connection error in the middle of a download is frustrating. That's why we've implemented a retry mechanism that automatically reconnects if a connection get closed or a ReadTimeout error is raised. The download restart exactly where it stopped without having to redownload any bytes.
In addition to this, it is possible to configure huggingface_hub with higher timeouts thanks to @Shahafgo. This should help getting around some issues on slower connections.
hf_transferhf_transfer is a Rust-based library focused on improving upload and download speed on machines with a high bandwidth available. Once installed (pip install -U hf_transfer), it can transparently be used with huggingface_hub simply by setting HF_HUB_ENABLE_HF_TRANSFER=1 as environment variable. The counterpart of higher performances is the lack of some user-friendly features such as better error handling or a retry mechanism -meaning it is recommended only to power-users-. In this release we still ship a new feature to improve UX: progress bars. No need to update any existing code, a simple library upgrade is enough.
hf-transfer progress bar by @cbensimon in #1792huggingface-cli guidehuggingface-cli is the CLI tool shipped with huggingface_hub. It recently got some nice improvement, especially with commands to download and upload files directly from the terminal. All of this needed a guide, so here it is!
Environment variables are useful to configure how huggingface_hub should work. Historically we had some inconsistencies on how those variables were named. This is now improved, with a backward compatible approach. Please check the package reference for more details. The goal is to propagate those changes to the whole HF-ecosystem, making configuration easier for everyone.
HF_ENDPOINT environment variable by @Wauplin in #1799Hindi documentation landed on the Hub thanks to @aneeshd27! Checkout the Hindi version of the quickstart guide here.
[[autodoc]] for ModelStatus by @jamesbraza in #1758post and ModelStatus by @jamesbraza in #1740Legacy ModelSearchArguments and DatasetSearchArguments have been completely removed from huggingface_hub. This shouldn't cause problem as they were already not in use (and unusable in practice).
Classes containing details about a repo (ModelInfo, DatasetInfo and SpaceInfo) have been refactored by @mariosasko to be more Pythonic and aligned with the other classes in huggingface_hub. In particular those objects are now based the dataclass module instead of a custom ReprMixin class. Every change is meant to be backward compatible, meaning no breaking changes is expected. However, if you detect any inconsistency, please let us know and we will fix it asap.
ReprMixin with dataclasses by @mariosasko in #1788The legacy Repository and InferenceAPI classes are now deprecated but will not be removed before the next major release (v1.0).
Instead of the git-based Repository, we advice to use the http-based HfApi. Check out this guide explaining the reasons behind it. For InferenceAPI, we recommend to switch to InferenceClient which is much more feature-complete and will keep getting improved.
Repository class by @Wauplin in #1724InferenceClientInferenceClient.get_recommended_model by @jamesbraza in #1770pydantic<3 by @jamesbraza in #1727HfFileSystemNotImplementedError on transaction commits by @Wauplin in #1736HfFileSystemFile when init fails + improve error message by @Wauplin in #1805WEBHOOK_PAYLOAD_EXAMPLE deserialization by @jamesbraza in #1732/locks folder to prevent rare concurrency issue by @beeender in #1659@retry_endpoint a default for all test by @Wauplin in #1725InferenceClient.post by @jamesbraza in #1742The following contributors have made significant changes to the library over the last release:
Nothing published for this version
A breaking change has been introduced in CommitOperationAdd in order to implement preupload_lfs_files in a way that is convenient for the users. The m…
(Discuss about the release and provide feedback in the Community Tab!)
Collection API is now fully supported in huggingface_hub!
A collection is a group of related items on the Hub (models, datasets, Spaces, papers) that are organized together on the same page. Collections are useful for creating your own portfolio, bookmarking content in categories, or presenting a curated list of items you want to share. Check out this guide to understand in more detail what collections are and this guide to learn how to build them programmatically.
get_collectioncreate_collection: title, description, namespace, privateupdate_collection_metadata: title, description, position, private, themedelete_collectionadd_collection_item: item id, item type, noteupdate_collection_item: note, positiondelete_collection_item>>> from huggingface_hub import get_collection
>>> collection = get_collection("TheBloke/recent-models-64f9a55bb3115b4f513ec026")
>>> collection.title
'Recent models'
>>> len(collection.items)
37
>>> collection.items[0]
CollectionItem: {
{'_id': '6507f6d5423b46492ee1413e',
'id': 'TheBloke/TigerBot-70B-Chat-GPTQ',
'author': 'TheBloke',
'item_type': 'model',
'lastModified': '2023-09-19T12:55:21.000Z',
(...)
}}
>>> from huggingface_hub import create_collection
# Create collection
>>> collection = create_collection(
... title="ICCV 2023",
... description="Portfolio of models, papers and demos I presented at ICCV 2023",
... )
# Add item with a note
>>> add_collection_item(
... collection_slug=collection.slug, # e.g. "davanstrien/climate-64f99dc2a5067f6b65531bab"
... item_id="datasets/climate_fever",
... item_type="dataset",
... note="This dataset adopts the FEVER methodology that consists of 1,535 real-world claims regarding climate-change collected on the internet."
... )
url attribute to Collection class by @Wauplin in #1695Documentation is now available in both German and Korean thanks to community contributions! This is an important milestone for Hugging Face in its mission to democratize good machine learning.
(Disclaimer: this is a power-user usage. It is not expected to be used directly by end users.)
When using create_commit (or upload_file/upload_folder), the internal workflow has 3 main steps:
In this release, we introduce preupload_lfs_files to perform step 2 independently of step 3. This is useful for libraries like datasets that generate huge files "on-the-fly" and want to preupload them one by one before making one commit with all the files. For more details, please read this guide.
CommitOperationAdd's internal attributes by @mariosasko in #1716Similarly to list_user_likes (listing all likes of a user), we now introduce list_repo_likers to list all likes on a repo - thanks to @issamarabi.
>>> from huggingface_hub import list_repo_likers
>>> likers = list_repo_likers("gpt2")
>>> len(likers)
204
>>> likers
[User(username=..., fullname=..., avatar_url=...), ...]
Template for the Dataset Card has been updated to be more aligned with the Model Card template.
This release also adds a few QOL improvement for the users:
TimeoutError => asyncio.TimeoutError by @matthewgrossman in #1666refs/convert/parquet and PR revision correctly in hffs by @Wauplin in #1712A breaking change has been introduced in CommitOperationAdd in order to implement preupload_lfs_files in a way that is convenient for the users. The main change is that CommitOperationAdd is no longer a static object but is modified internally by preupload_lfs_files and create_commit. This means that you cannot reuse a CommitOperationAdd object once it has been committed to the Hub. If you do so, an explicit exception will be raised. You can still reuse the operation objects if the commit call failed and you retry it. We hope that it will not affect any users but please open an issue if you're encountering any problem.
fsspec to use default expand_path by @mariosasko in #16810.18.0.dev0 by @Wauplin in #1658HTTPError spec by @Wauplin in #1693The following contributors have made significant changes to the library over the last release:
Nothing published for this version
Fixing a bug when downloading files to a non-existent directory. In https://github.com/huggingface/huggingface_hub/pull/1590 we introduced a helper th
Full Changelog: https://github.com/huggingface/huggingface_hub/compare/v0.17.2...v0.17.3
Fixing a bug when downloading files to a non-existent directory. In https://github.com/huggingface/huggingface_hub/pull/1590 we introduced a helper that raises a warning if there is not enough disk space to download a file. A bug made the helper raise an exception if the folder doesn't exist yet as reported in https://github.com/huggingface/huggingface_hub/issues/1690. This hot-fix fixes it thanks to https://github.com/huggingface/huggingface_hub/pull/1692 which recursively checks the parent directories if the full path doesn't exist. If it keeps failing (for any OSError) we silently ignore the error and keep going. Not having the warning is worse than breaking the download of legit users.
Checkout those release notes to learn more about the v0.17 release.
Fixing a bug when uploading files to a Space repo using the CLI. The command was trying to create a repo (even if it already exists) and was failing b
Full Changelog: https://github.com/huggingface/huggingface_hub/compare/v0.17.1...v0.17.2
Fixing a bug when uploading files to a Space repo using the CLI. The command was trying to create a repo (even if it already exists) and was failing because space_sdk was not found in that case. More details in https://github.com/huggingface/huggingface_hub/pull/1669.
Also updated the user-agent when using huggingface-cli upload. See https://github.com/huggingface/huggingface_hub/pull/1664.
Checkout those release notes to learn more about the v0.17 release.
Nothing published for this version
…(1, sequence_length, hidden_size) which is the breaking change.
Thanks to a massive community effort, all inference tasks are now supported in InferenceClient. Newly added tasks are:
Documentation, including examples, for each of these tasks can be found in this table.
All those methods also support async mode using AsyncInferenceClient.
Sometimes knowing which models are available or not on the Inference API service can be useful. This release introduces two new helpers:
list_deployed_models aims to help users discover which models are currently deployed, listed by task.get_model_status aims to get the status of a specific model. That's useful if you already know which model you want to use.Those two helpers are only available for the Inference API, not Inference Endpoints (or any other provider).
>>> from huggingface_hub import InferenceClient
>>> client = InferenceClient()
# Discover zero-shot-classification models currently deployed
>>> models = client.list_deployed_models()
>>> models["zero-shot-classification"]
['Narsil/deberta-large-mnli-zero-cls', 'facebook/bart-large-mnli', ...]
# Get status for a specific model
>>> client.get_model_status("bigcode/starcoder")
ModelStatus(loaded=True, state='Loaded', compute_type='gpu', framework='text-generation-inference')
text_to_image and image_to_image parameters by @Wauplin in #1582This is a long-awaited feature finally implemented! huggingface-cli now offers two new commands to easily transfer file from/to the Hub. The goal is to use them as a replacement for git clone, git pull and git push. Despite being less feature-complete than git (no .git/ folder, no notion of local commits), it offers the flexibility required when working with large repositories.
Download
# Download a single file
>>> huggingface-cli download gpt2 config.json
/home/wauplin/.cache/huggingface/hub/models--gpt2/snapshots/11c5a3d5811f50298f278a704980280950aedb10/config.json
# Download files to a local directory
>>> huggingface-cli download gpt2 config.json --local-dir=./models/gpt2
./models/gpt2/config.json
# Download a subset of a repo
>>> huggingface-cli download bigcode/the-stack --repo-type=dataset --revision=v1.2 --include="data/python/*" --exclu
de="*.json" --exclude="*.zip"
Fetching 206 files: 100%|████████████████████████████████████████████| 206/206 [02:31<2:31, ?it/s]
/home/wauplin/.cache/huggingface/hub/datasets--bigcode--the-stack/snapshots/9ca8fa6acdbc8ce920a0cb58adcdafc495818ae7
Upload
# Upload single file
huggingface-cli upload my-cool-model model.safetensors
# Upload entire directory
huggingface-cli upload my-cool-model ./models
# Sync local Space with Hub (upload new files except from logs/, delete removed files)
huggingface-cli upload Wauplin/space-example --repo-type=space --exclude="/logs/*" --delete="*" --commit-message="Sync local Space with Hub"
Docs
For more examples, check out the documentation:
Some new features have been added to the Space API to:
>>> from huggingface_hub import HfApi
>>> api = HfApi()
>>> api.create_repo(
... repo_id=repo_id,
... repo_type="space",
... space_sdk="gradio",
... space_hardware="t4-medium",
... space_sleep_time="3600",
... space_storage="large",
... space_secrets=[{"key"="HF_TOKEN", "value"="hf_api_***"}, ...],
... space_variables=[{"key"="MODEL_REPO_ID", "value"="user/repo"}, ...],
... )
A special thank to @martinbrose who largely contributed on those new features.
A new section has been added to the upload guide with some tips about how to upload large models and datasets to the Hub and what are the limits when doing so.
:world_map: The documentation organization has been updated to support multiple languages. The community effort has started to translate the docs to non-English speakers. More to come in the coming weeks!
The behavior of InferenceClient.feature_extraction has been updated to fix a bug happening with certain models. The shape of the returned array for transformers models has changed from (sequence_length, hidden_size) to (1, sequence_length, hidden_size) which is the breaking change.
HfApi helpers:
Two new helpers have been added to check if a file or a repo exists on the Hub:
>>> from huggingface_hub import file_exists
>>> file_exists("bigcode/starcoder", "config.json")
True
>>> file_exists("bigcode/starcoder", "not-a-file")
False
>>> from huggingface_hub import repo_exists
>>> repo_exists("bigcode/starcoder")
True
>>> repo_exists("bigcode/not-a-repo")
False
Also, hf_hub_download and snapshot_download are now part of HfApi (keeping the same syntax and behavior).
hf_hub_download to HfApi by @Wauplin in #1580Download improvements:
missing_ok option in delete_repo by @Wauplin in #1640super_squash_history in HfApi by @Wauplin in #1639The following contributors have made significant changes to the library over the last release:
Nothing published for this version
Hotfix to avoid sharing requests.Session between processes. More information in https://github.com/huggingface/huggingface_hub/pull/1545. Internally,
Full Changelog: https://github.com/huggingface/huggingface_hub/compare/v0.16.3...v0.16.4
Hotfix to avoid sharing requests.Session between processes. More information in https://github.com/huggingface/huggingface_hub/pull/1545. Internally, we create a Session object per thread to benefit from the HTTPSConnectionPool (i.e. do not reopen connection between calls). Due to an implementation bug, the Session object from the main thread was shared if a fork of the main process happened. The shared Session gets corrupted in the process, leading to some random ConnectionErrors in rare occasions.
Check out these release notes to learn more about the v0.16 release.
Hotfix to print the request ID if any RequestException happen. This is useful to help the team debug users' problems. Request ID is a generated UUID,
Full Changelog: https://github.com/huggingface/huggingface_hub/compare/v0.16.2...v0.16.3
Hotfix to print the request ID if any RequestException happen. This is useful to help the team debug users' problems. Request ID is a generated UUID, unique for each HTTP call made to the Hub.
Check out these release notes to learn more about the v0.16 release.
ModelHubMixin got updated (after a deprecation cycle):
Introduced in the v0.15 release, the InferenceClient got a big update in this one. The client is now reaching a stable point in terms of features. The next updates will be focused on continuing to add support for new tasks.
Asyncio calls are supported thanks to AsyncInferenceClient. Based on asyncio and aiohttp, it allows you to make efficient concurrent calls to the Inference endpoint of your choice. Every task supported by InferenceClient is supported in its async version. Method inputs and outputs and logic are strictly the same, except that you must await the coroutine.
>>> from huggingface_hub import AsyncInferenceClient
>>> client = AsyncInferenceClient()
>>> image = await client.text_to_image("An astronaut riding a horse on the moon.")
Support for text-generation task has been added. It is focused on fully supporting endpoints running on the text-generation-inference framework. In fact, the code is heavily inspired by TGI's Python client initially implemented by @OlivierDehaene.
Text generation has 4 modes depending on details (bool) and stream (bool) values. By default, a raw string is returned. If details=True, more information about the generated tokens is returned. If stream=True, generated tokens are returned one by one as soon as the server generated them. For more information, check out the documentation.
>>> from huggingface_hub import InferenceClient
>>> client = InferenceClient()
# stream=False, details=False
>>> client.text_generation("The huggingface_hub library is ", max_new_tokens=12)
'100% open source and built to be easy to use.'
# stream=True, details=True
>>> for details in client.text_generation("The huggingface_hub library is ", max_new_tokens=12, details=True, stream=True):
>>> print(details)
TextGenerationStreamResponse(token=Token(id=1425, text='100', logprob=-1.0175781, special=False), generated_text=None, details=None)
...
TextGenerationStreamResponse(token=Token(
id=25,
text='.',
logprob=-0.5703125,
special=False),
generated_text='100% open source and built to be easy to use.',
details=StreamDetails(finish_reason=<FinishReason.Length: 'length'>, generated_tokens=12, seed=None)
)
Of course, the async client also supports text-generation (see docs):
>>> from huggingface_hub import AsyncInferenceClient
>>> client = AsyncInferenceClient()
>>> await client.text_generation("The huggingface_hub library is ", max_new_tokens=12)
'100% open source and built to be easy to use.'
InferenceClient now supports zero-shot-image-classification (see docs). Both sync and async clients support it. It allows to classify an image based on a list of labels passed as input.
>>> from huggingface_hub import InferenceClient
>>> client = InferenceClient()
>>> client.zero_shot_image_classification(
... "https://upload.wikimedia.org/wikipedia/commons/thumb/4/43/Cute_dog.jpg/320px-Cute_dog.jpg",
... labels=["dog", "cat", "horse"],
... )
[{"label": "dog", "score": 0.956}, ...]
Thanks to @dulayjm for your contribution on this task!
When using InferenceClient's task methods (text_to_image, text_generation, image_classification,...) you don't have to pass a model id. By default, the client will select a model recommended for the selected task and run on the free public Inference API. This is useful to quickly prototype and test models. In a production-ready setup, we strongly recommend to set the model id/URL manually, as the recommended model is expected to change at any time without prior notice, potentially resulting in different and unexpected results in your workflow. Recommended models are the ones used by default on https://hf.co/tasks.
It is now possible to configure headers and cookies to be sent when initializing the client: InferenceClient(headers=..., cookies=...). All calls made with this client will then use these headers/cookies.
The CommitScheduler is a new class that can be used to regularly push commits to the Hub. It watches changes in a folder and creates a commit every 5 minutes if it detected a file change. One intended use case is to allow regular backups from a Space to a Dataset repository on the Hub. The scheduler is designed to remove the hassle of handling background commits while avoiding empty commits.
>>> from huggingface_hub import CommitScheduler
# Schedule regular uploads every 10 minutes. Remote repo and local folder are created if they don't already exist.
>>> scheduler = CommitScheduler(
... repo_id="report-translation-feedback",
... repo_type="dataset",
... folder_path=feedback_folder,
... path_in_repo="data",
... every=10,
... )
Check out this guide to understand how to use the CommitScheduler. It comes with a Space to showcase how to use it in 4 practical examples.
CommitScheduler: upload folder every 5 minutes by @Wauplin in #1494The Hugging Face Hub offers nice support for Tensorboard data. It automatically detects when TensorBoard traces (such as tfevents) are pushed to the Hub and starts an instance to visualize them. This feature enable a quick and transparent collaboration in your team when training models. In fact, more than 42k models are already using this feature!
With the HFSummaryWriter you can now take full advantage of the feature for your training, simply by updating a single line of code.
>>> from huggingface_hub import HFSummaryWriter
>>> logger = HFSummaryWriter(repo_id="test_hf_logger", commit_every=15)
HFSummaryWriter inherits from SummaryWriter and acts as a drop-in replacement in your training scripts. The only addition is that every X minutes (e.g. 15 minutes) it will push the logs directory to the Hub. Commit happens in the background to avoid blocking the main thread. If the upload crashes, the logs are kept locally and the training continues.
For more information on how to use it, check out this documentation page. Please note that this is still an experimental feature so feedback is very welcome.
It is now possible to copy a file in a repo on the Hub. The copy can only happen within a repo and with an LFS file. File can be copied between different revisions. More information here.
ModelHubMixin got updated (after a deprecation cycle):
model_id as username/repo_name@revision in ModelHubMixin. Revision must be passed as a separate revision argument if needed.A x-request-id header is sent by default for every request made to the Hub. This should help debugging user issues.
3 PRs, 3 commits but in the end default timeout did not change. Problem has been solved server-side instead.
The following contributors have made significant changes to the library over the last release:
Nothing published for this version
Nothing published for this version
Some (announced) breaking changes have been introduced:
We introduce InferenceClient, a new client to run inference on the Hub. The objective is to:
summary = client.summarization("this is a long text"))Check out the Inference guide to get a complete overview.
>>> from huggingface_hub import InferenceClient
>>> client = InferenceClient()
>>> image = client.text_to_image("An astronaut riding a horse on the moon.")
>>> image.save("astronaut.png")
>>> client.image_classification("https://upload.wikimedia.org/wikipedia/commons/thumb/4/43/Cute_dog.jpg/320px-Cute_dog.jpg")
[{'score': 0.9779096841812134, 'label': 'Blenheim spaniel'}, ...]
The short-term goal is to add support for more tasks (here is the current list), especially text-generation and handle asyncio calls. The mid-term goal is to deprecate and replace InferenceAPI.
InferenceClient by @Wauplin in #1474It is now possible to run HfApi calls in the background! The goal is to make it easier to upload files periodically without blocking the main thread during a training. The was previously possible when using Repository but is now available for HTTP-based methods like upload_file, upload_folder and create_commit. If run_as_future=True is passed:
Future object is returned to check the job statusIn addition to this parameter, a run_as_future(...) method is available to queue any other calls to the Hub. More details in this guide.
>>> from huggingface_hub import HfApi
>>> api = HfApi()
>>> api.upload_file(...) # takes Xs
# URL to upload file
>>> future = api.upload_file(..., run_as_future=True) # instant
>>> future.result() # wait until complete
# URL to upload file
HfApi methods in the background (run_as_future) by @Wauplin in #1458Some (announced) breaking changes have been introduced:
list_models, list_datasets and list_spaces return an iterable instead of a list (lazy-loading of paginated results)cardData in list_datasets has been removed in favor of the parameter full.Both changes had a deprecation cycle for a few releases now.
New parameters in login() :
new_session : skip login if new_session=False and user is already logged inwrite_permission : write permission is required (login fails otherwise)Also added a new HfApi().get_token_permission() method that returns "read" or "write" (or None if not logged in).
New parameter to get more details when listing files: list_repo_files(..., expand=True).
API call is slower but lastCommit and security fields are returned as well.
ImportError when importing WebhooksServer and Gradio is not installed by @mariosasko in #1482_deprecation.py warning message for _deprecate_list_output() by @x11kjm in #1485Nothing published for this version
Nothing published for this version
Fixed an issue reported in `diffusers` impacting users downloading files from outside of the Hub. Expected download size now takes into account potent
Fixed an issue reported in diffusers impacting users downloading files from outside of the Hub. Expected download size now takes into account potential compression in the HTTP requests.
Full Changelog: https://github.com/huggingface/huggingface_hub/compare/v0.14.0...v0.14.1
No other breaking change expected in this release.
We introduce HfFileSystem, a pythonic filesystem interface compatible with fsspec. Built on top of HfApi, it offers typical filesystem operations like cp, mv, ls, du, glob, get_file and put_file.
>>> from huggingface_hub import HfFileSystem
>>> fs = HfFileSystem()
# List all files in a directory
>>> fs.ls("datasets/myself/my-dataset/data", detail=False)
['datasets/myself/my-dataset/data/train.csv', 'datasets/myself/my-dataset/data/test.csv']
>>> train_data = fs.read_text("datasets/myself/my-dataset/data/train.csv")
Its biggest advantage is to provide ready-to-use integrations with popular libraries like Pandas, DuckDB and Zarr.
import pandas as pd
# Read a remote CSV file into a dataframe
df = pd.read_csv("hf://datasets/my-username/my-dataset-repo/train.csv")
# Write a dataframe to a remote CSV file
df.to_csv("hf://datasets/my-username/my-dataset-repo/test.csv")
For a more detailed overview, please have a look to this guide.
hffs code to hfh by @mariosasko in #1420WebhooksServer allows to implement, debug and deploy webhook endpoints on the Hub without any overhead. Creating a new endpoint is as easy as decorating a Python function.
# app.py
from huggingface_hub import webhook_endpoint, WebhookPayload
@webhook_endpoint
async def trigger_training(payload: WebhookPayload) -> None:
if payload.repo.type == "dataset" and payload.event.action == "update":
# Trigger a training job if a dataset is updated
...
For more details, check out this twitter thread or the documentation guide.
Note that this feature is experimental which means the API/behavior might change without prior notice. A warning is displayed to the user when using it. As it is experimental, we would love to get feedback!
hf_transferIntegration with a Rust-based library to upload large files in chunks and concurrently. Expect x3 speed-up if your bandwidth allows it!
hf_transfer upload by @McPatate in #1395Uploading large folders at once might be annoying if any error happens while committing (e.g. a connection error occurs). It is now possible to upload a folder in multiple (smaller) commits. If a commit fails, you can re-run the script and resume the upload. Commits are pushed to a dedicated PR. Once completed, the PR is merged to the main branch resulting in a single commit in your git history.
upload_folder(
folder_path="local/checkpoints",
repo_id="username/my-dataset",
repo_type="dataset",
multi_commits=True, # resumable multi-upload
multi_commits_verbose=True,
)
Note that this feature is also experimental, meaning its behavior might be updated in the future.
create_commits_on_pr by @Wauplin in #1375Some more pre-validation done before committing files to the Hub. The .git folder is ignored in upload_folder (if any) + fail early in case of invalid paths.
path_in_repo validation when committing files by @Wauplin in #1382.git/ folder + ignore .git/ folder in upload_folder by @Wauplin in #1408Internal update to reuse the same HTTP session across huggingface_hub. The goal is to keep the connection open when doing multiple calls to the Hub which ultimately saves a lot of time. For instance, updating metadata in a README became 40% faster while listing all models from the Hub is 60% faster. This has no impact for atomic calls (e.g. 1 standalone GET call).
It is now possible to programmatically set a custom sleep time on your upgraded Space. After X seconds of inactivity, your Space will go to sleep to save you some $$$.
from huggingface_hub import set_space_sleep_time
# Put your Space to sleep after 1h of inactivity
set_space_sleep_time(repo_id=repo_id, sleep_time=3600)
sleep_time for Spaces by @Wauplin in #1438fsspec has been added as a main dependency. It's a lightweight Python library required for HfFileSystem.No other breaking change expected in this release.
A lot of effort has been invested in making huggingface_hub's cache system more robust especially when working with symlinks on Windows. Hope everything's fixed by now.
After a server-side configuration issue, we made huggingface_hub more robust when getting Hub's Etags to be more future-proof.
HUGGINGFACE_HEADER_X_LINKED_ETAG const by @julien-c in #1405Nothing published for this version
Nothing published for this version
Security patch to fix a vulnerability in huggingface_hub. In some cases, downloading a file with hf_hub_download or snapshot_download could lead to ov…
Security patch to fix a vulnerability in huggingface_hub. In some cases, downloading a file with hf_hub_download or snapshot_download could lead to overwriting any file on a Windows machine. With this fix, only files in the cache directory (or a user-defined directory) can be updated/overwritten.
Full Changelog: https://github.com/huggingface/huggingface_hub/compare/v0.13.3...v0.13.4
Patch to fix symlinks in the cache directory. Relative paths are used by default whenever possible. Absolute paths are used only on Windows when creat
Patch to fix symlinks in the cache directory. Relative paths are used by default whenever possible. Absolute paths are used only on Windows when creating a symlink betweenh 2 paths that are not on the same volume. This hot-fix reverts the logic to what it was in huggingface_hub<=0.12 given the issues that have being reported after the 0.13.2 release (https://github.com/huggingface/huggingface_hub/issues/1398, https://github.com/huggingface/diffusers/issues/2729 and https://github.com/huggingface/transformers/pull/22228)
Hotfix - use relative symlinks whenever possible https://github.com/huggingface/huggingface_hub/pull/1399 @Wauplin
Full Changelog: https://github.com/huggingface/huggingface_hub/compare/v0.13.2...v0.13.3
Patch to fix symlinks in the cache directory. All symlinks are now absolute paths.
Patch to fix symlinks in the cache directory. All symlinks are now absolute paths.
Full Changelog: https://github.com/huggingface/huggingface_hub/compare/v0.13.1...v0.13.2
Patch to fix upload_folder when passing path_in_repo=".". That was a breaking change compared to 0.12.1. Also added more validation around the path_in…
Patch to fix upload_folder when passing path_in_repo=".". That was a breaking change compared to 0.12.1. Also added more validation around the path_in_repo attribute to improve UX.
path_in_repo validation when committing files by @Wauplin in https://github.com/huggingface/huggingface_hub/pull/1382Full Changelog: https://github.com/huggingface/huggingface_hub/compare/v0.13.0...v0.13.1
This was problematic as it could lead to a loss of information. Fixing this is a breaking change but impact should be limited as the server is already…
It is now possible to download files from the Hub and move them to a specific folder!
Two behaviors are possible: either create symlinks or move the files from the cache. This can be controlled with the local_dir_use_symlinks input parameter. The default -and recommended- value is "auto" which will duplicate small files to ease user experience (no symlinks when editing a file) and create symlinks for big files (save disk usage).
from huggingface_hub import snapshot_download
# or "from huggingface_hub import hf_hub_download"
# Download and cache files + duplicate small files (<5MB) to "my-folder" + add symlinks for big files
snapshot_download(repo_id, local_dir="my-folder")
# Download and cache files + add symlinks in "my-folder"
snapshot_download(repo_id, local_dir="my-folder", local_dir_use_symlinks=True)
# Duplicate files already existing in cache and/or download missing files directly to "my-folder"
snapshot_download(repo_id, local_dir="my-folder", local_dir_use_symlinks=False)
Efforts to improve documentation have continued. The guides overview has been refactored to display which topics are covered (repository, upload, download, search, inference, community tab, cache, model cards, space management and integration).
The repository, upload and download guides have been revisited to showcase the different possibilities to manage a repository and upload/download files to/from it. The focus has been explicitly put on the HTTP endpoints rather than the git cli.
A new guide has been added on how to integrate any ML framework with the Hub. It explains what is meant by that and how to do it. Here is the summary table to remember:
It's now possible to duplicate a Space programmatically!
>>> from huggingface_hub import duplicate_space
# Duplicate a Space to your account
>>> duplicate_space("multimodalart/dreambooth-training")
RepoUrl('https://huggingface.co/spaces/nateraw/dreambooth-training',...)
delete_patterns in upload_folderNew input parameter delete_patterns for the upload_folder method. It allows to delete some remote files before pushing a folder to the Hub, in a single commit. Useful when you don't exactly know which files have already been pushed. Here is an example to upload log files while deleting existing logs on the Hub:
api.upload_folder(
folder_path="/path/to/local/folder/logs",
repo_id="username/trained-model",
path_in_repo="experiment/logs/",
allow_patterns="*.txt", # Upload all local text files
delete_patterns="*.txt", # Delete all remote text files before
)
Get the repo history (i.e. all the commits) for a given revision.
# Get initial commit on a repo
>>> from huggingface_hub import list_repo_commits
>>> initial_commit = list_repo_commits("gpt2")[-1]
# Initial commit is always a system commit containing the `.gitattributes` file.
>>> initial_commit
GitCommitInfo(
commit_id='9b865efde13a30c13e0a33e536cf3e4a5a9d71d8',
authors=['system'],
created_at=datetime.datetime(2019, 2, 18, 10, 36, 15, tzinfo=datetime.timezone.utc),
title='initial commit',
message='',
formatted_title=None,
formatted_message=None
)
huggingface-cli login--token and --add-to-git-credential option have been added to login directly from the CLI using an environment variable. Useful to login in a Github CI script for example.
huggingface-cli login --token $HUGGINGFACE_TOKEN --add-to-git-credential
Helper for external libraries to track usage of specific features of their package. Telemetry can be globally disabled by the user using HF_HUB_DISABLE_TELEMETRY.
from huggingface_hub.utils import send_telemetry
send_telemetry("gradio/local_link", library_name="gradio", library_version="3.22.1")
When loading a model card with an invalid model_index in the metadata, an error is explicitly raised. Previous behavior was to trigger a warning and ignore the model_index. This was problematic as it could lead to a loss of information. Fixing this is a breaking change but impact should be limited as the server is already rejecting invalid model cards. An optional ignore_metadata_errors argument (default to False) can be used to load the card with only a warning.
A few improvements in repo cards: expose RepoCard as top-level, dict-like methods for RepoCardData object (#1354), updated template and improved type annotation for metadata.
RepoCard at top level + few qol improvements by @Wauplin in #1354Nothing published for this version
Nothing published for this version
Hot-fix to remove authorization header when following redirection (using cached_download). Fix was already implemented for hf_hub_download but we forg
Hot-fix to remove authorization header when following redirection (using cached_download). Fix was already implemented for hf_hub_download but we forgot about this one. Has only a consequence when downloading LFS files from Spaces. Problem arose since a server-side change on how files are served. See https://github.com/huggingface/huggingface_hub/pull/1345.
Full Changelog: https://github.com/huggingface/huggingface_hub/compare/v0.12.0...v0.12.1
…depending on huggingface_hub. To help detect breaking changes that would affect third-party libraries, we built a framework to run simple end-to-end t…
Spaces support has been substantially enhanced. You can now:
# Assign hardware when creating the Space
api.create_repo(repo_id=repo_id, repo_type="space", space_sdk="gradio", space_hardware="cpu-upgrade")
# Configure some secrets
api.add_space_secret(repo_id=repo_id, key="HF_TOKEN", value="hf_api_***")
# Request hardware on the fly
api.request_space_hardware(repo_id=repo_id, hardware="t4-medium")
# Get Space runtime (state, hardware, sdk,...)
api.get_space_runtime(repo_id=repo_id)
Visit the docs for more details.
get_space_secrets endpoint by @Wauplin in #1264And in bonus: Spaces now support Dockerfile natively !
api.create_repo(repo_id=repo_id, repo_type="space", space_sdk="docker")
Check out the docs for more details.
HfApiIt's now possible to list branches/tags from a repo, getting exact ref and target_commit.
More details in the docs.
>>> api.list_repo_refs("bigcode/the-stack", repo_type='dataset')
GitRefs(
branches=[
GitRefInfo(name='main', ref='refs/heads/main', target_commit='18edc1591d9ce72aa82f56c4431b3c969b210ae3'),
GitRefInfo(name='v1.1.a1', ref='refs/heads/v1.1.a1', target_commit='f9826b862d1567f3822d3d25649b0d6d22ace714')],
converts=[],
tags=[
GitRefInfo(name='v1.0', ref='refs/tags/v1.0', target_commit='c37a8cd1e382064d8aced5e05543c5f7753834da')
]
)
New endpoints to like a repo, unlike it and list_liked_repos.
repo_id when creating a repoWhen using create_repo, one can provide a simple repo_name without specifying a namespace (example: "my-cool-model" instead of "Wauplin/my-cool-model"). This was annoying for as one could easily know if a namespace has been added to the provided repo_id. To facilitate this, the return value of create_repo is now an instance of RepoUrl which contains information like the endpoint, namespace, repo_id and repo_type.
By default, new branches start from main HEAD. It's now possible to specify any branch, tag or commit id to start from.
Modelcards module is getting some adjustments to better integrate with the Hub. The scope of this work is larger than "just" huggingface_hub and resulted in the launch of the HF Model Card Guidebook to help operationalize model cards in the ML community.
datasetcard_template: I think linking to a GH user does not make sense anymore now that dataset repos are fully on the Hub by @julien-c in #1257REGEX_YAML_BLOCK by @julien-c in #1285Quite some effort has been put into the documentation in the past few weeks:
huggingface_hub is getting more and more mature but you might still have some friction if you are maintainer of a library depending on huggingface_hub. To help detect breaking changes that would affect third-party libraries, we built a framework to run simple end-to-end tests in our CI. This is still preliminary work but the hope is make hfh ecosystem more and more robust over time. Check out our README for more details.
Goal is to download files faster. First step has been to increase the chunk size by which the files are uploaded. Second step has been to add an optional Rust extension. This is not officially documented for now as we are internally assessing its benefits and limits. Installing and activating hf_transfer is purely optional.
Repository "clone_from" feature do not create the remote repository if it doesn't exist on the Hub. Please use create_repo first before cloning it locally. The private init argument has also been removed as it was not needed anymore.allow_regex and ignore_regex have been removed from snapshot_download in favor allow_patterns and ignore_patterns.push_to_hub_fastai has been removed in favor of the HTTP-based approach. Same for ModelHubMixin, PyTorchModelHubMixin, push_to_hub_keras and KerasModelHubMixin.create_repo is now forced to use keyword-arguments. Same for metadata_eval_result.Not really some features, not really some fixes. But definitely a quality of life improvement for users 🙂
hf:// urls + raise ValueError if repo type is unknown by @Wauplin in #1298Nothing published for this version
Hot-fix to fix permission issues when downloading with hf_hub_download or snapshot_download. For more details, see https://github.com/huggingface/hugg
Hot-fix to fix permission issues when downloading with hf_hub_download or snapshot_download. For more details, see https://github.com/huggingface/huggingface_hub/pull/1220, https://github.com/huggingface/huggingface_hub/issues/1141 and https://github.com/huggingface/huggingface_hub/issues/1215.
Full changelog: https://github.com/huggingface/huggingface_hub/compare/v0.11.0...v0.11.1
In the future, listing models, datasets and spaces will be paginated on the Hub by default. To avoid breaking changes, huggingface_hub follows already…
HfApiHfApi is the central point to interact with the Hub API (manage repos, create commits,...). The goal is to propose more and more git-related features using HTTP endpoints to allow users to interact with the Hub without cloning locally a repo.
from huggingface_hub import create_branch, create_tag, delete_branch, delete_tag
create_tag(repo_id, tag="v0.11", tag_message="Release v0.11")
delete_tag(repo_id, tag="something") # If you created a tag by mistake
create_branch(repo_id, branch="experiment-154")
delete_branch(repo_id, branch="experiment-1") # Clean some old branches
create_tag method to create tags from the HTTP endpoint by @Wauplin in #1089delete_tag method to HfApi by @Wauplin in #1128Making a very large commit was previously tedious. Files are now processed by chunks which makes it possible to upload 25k files in a single commit (and 1Gb payload limitation if uploading only non-LFS files). This should make it easier to upload large datasets.
from huggingface_hub import CommitOperationDelete, create_commit, delete_folder
# Delete a single folder
delete_folder(repo_id=repo_id, path_in_repo="logs/")
# Alternatively, use the low-level `create_commit`
create_commit(
repo_id,
operations=[
CommitOperationDelete(path_in_repo="old_config.json") # Delete a file
CommitOperationDelete(path_in_repo="logs/") # Delete a folder
],
commit_message=...,
)
In the future, listing models, datasets and spaces will be paginated on the Hub by default. To avoid breaking changes, huggingface_hub follows already pagination. Output type is currently a list (deprecated), will become a generator in v0.14.
Authentication has been revisited to make it as easy as possible for the users.
login and logout methodsfrom huggingface_hub import login, logout
# `login` detects automatically if you are running in a notebook or a script
# Launch widgets or TUI accordingly
login()
# Now possible to login with a hardcoded token (non-blocking)
login(token="hf_***")
# If you want to bypass the auto-detection of `login`
notebook_login() # still available
interpreter_login() # to login from a script
# Logout programmatically
logout()
# Still possible to login from CLI
huggingface-cli login
HfApi sessionfrom huggingface_hub import HfApi
# Token will be sent in every request but not stored on machine
api = HfApi(token="hf_***")
use_auth_token in favor of token, everywheretoken parameter can now be passed to every method in huggingface_hub. use_auth_token is still accepted where it previously existed but the mid-term goal (~6 months) is to deprecate and remove it.
use_auth_token arg by token everywhere by @Wauplin in #1122Previously, token was stored in the git credential store. Can now be in any helper configured by the user -keychain, cache,...-.
# Dump all relevant information. To be used when reporting an issue.
➜ huggingface-cli env
Copy-and-paste the text below in your GitHub issue.
- huggingface_hub version: 0.11.0.dev0
- Platform: Linux-5.15.0-52-generic-x86_64-with-glibc2.35
- Python version: 3.10.6
...
x-error-message header if exists by @Wauplin in #1121Few improvements/fixes in the modelcard module:
model_name in metadata_update by @lvwerra in #1157New feature to provide a path in the cache where any downstream library can store assets (processed data, files from the web, extracted data, rendered images,...)
create_repoidentical_ok removed in upload_filevalidate_preupload_info, prepare_commit_payload, _upload_lfs_object (internal helpers for the commit API)huggingface_hub.snapshot_download is not exposed as a public module anymoreHfApi.move_repo(...) and complete tests by @Wauplin in #1136Nothing published for this version
Nothing published for this version
Hot-fix to force utf-8 encoding in modelcards. See https://github.com/huggingface/huggingface_hub/pull/1102 and https://github.com/skops-dev/skops/pul
Hot-fix to force utf-8 encoding in modelcards. See https://github.com/huggingface/huggingface_hub/pull/1102 and https://github.com/skops-dev/skops/pull/162#issuecomment-1263516507 for context.
Full Changelog: https://github.com/huggingface/huggingface_hub/compare/v0.10.0...v0.10.1
For consistency, the return type of create_commit has been modified. This is a breaking change, but we hope the return type of this method was never u…
Contribution from @nateraw to integrate the work done on Modelcards and DatasetCards (from nateraw/modelcards) directly in huggingface_hub.
>>> from huggingface_hub import ModelCard
>>> card = ModelCard.load('nateraw/vit-base-beans')
>>> card.data.to_dict()
{'language': 'en', 'license': 'apache-2.0', 'tags': ['generated_from_trainer', 'image-classification'],...}
modelcards repo by @nateraw in #940update_metadata by @Wauplin in #1061huggingface-cli scan-cache and huggingface-cli delete-cache)New commands in huggingface-cli to scan and delete parts of the cache. Goal is to manage the cache-system the same way for any dependent library that uses huggingface_hub. Only the new cache-system format is supported.
➜ huggingface-cli scan-cache
REPO ID REPO TYPE SIZE ON DISK NB FILES LAST_ACCESSED LAST_MODIFIED REFS LOCAL PATH
--------------------------- --------- ------------ -------- ------------- ------------- ------------------- -------------------------------------------------------------------------
glue dataset 116.3K 15 4 days ago 4 days ago 2.4.0, main, 1.17.0 /home/wauplin/.cache/huggingface/hub/datasets--glue
google/fleurs dataset 64.9M 6 1 week ago 1 week ago refs/pr/1, main /home/wauplin/.cache/
(...)
Done in 0.0s. Scanned 6 repo(s) for a total of 3.4G.
Got 1 warning(s) while scanning. Use -vvv to print details.
huggingface-cli delete-cache command by @Wauplin in #1046delete-cache by @Wauplin in #1063HTTP calls to the Hub have been harmonized to behave the same across the library.
Major differences are:
hf_raise_for_status (more informative error message)hf_hub_download.huggingface_hub by @Wauplin in #1019create_commit has been modified. This is a breaking change, but we hope the return type of this method was never used (quite recent and niche output type).repo_id is now validated using @validate_hf_hub_args (see below), a breaking change can be caused if repo_id was previously miused. A HFValidationError is now raised if repo_id is not valid.push_to_hub_fastai@validate_hf_hub_argshuggingface_hub by @Wauplin in #1029:warning: This is a breaking change if repo_id was previously misused :warning:
token in read-only methods of HfApi in favor of use_auth_token by @SBrandeis in #928/models/ path for api call to update settings by @Wauplin in #1049store in google colab by @Wauplin in #1053Nothing published for this version
Nothing published for this version
Nothing published for this version
Hot-fix error message on gated repositories (https://github.com/huggingface/huggingface_hub/pull/1015).
Hot-fix error message on gated repositories (https://github.com/huggingface/huggingface_hub/pull/1015).
Context: https://huggingface.co/CompVis/stable-diffusion-v1-4 has been widely shared in the last days but since it's a gated-repo, lots of users are getting confused by the Authentification error received. Error message is now more detailed.
Full Changelog: https://github.com/huggingface/huggingface_hub/compare/v0.9.0...v0.9.1
Huge work to programmatically interact with the community tab, thanks to @SBrandeis ! It is now possible to:
Huge work to programmatically interact with the community tab, thanks to @SBrandeis ! It is now possible to:
create_discussion, create_pull_request, merge_pull_request, change_discussion_status, rename_discussion)comment_discussion, edit_discussion_comment)get_repo_discussions, get_discussion_details)See full documentation for more details.
push_to_hub mixinspush_to_hub mixin and push_to_hub_keras have been refactored to leverage the http-endpoint. This means pushing to the hub will no longer require to first download the repo locally. Previous git-based version is planned to be supported until v0.12.
git by @LysandreJik in #847parent_commit argument for create_commit and related functions by @SBrandeis in #916files_metadata option to repo_info by @SBrandeis in #951upload_folderHF_HUB_DISABLE_PROGRESS_BARS env variable or using disable_progress_bars/enable_progress_bars helpers.try_to_load_from_cache to check if a file is locally cachedupload_file with upload_folder in upload_folder docstring by @mariosasko in #927hf_hub_download for a renamed repo by @Wauplin in #983path_in_repo optional in upload folder by @Wauplin in #988repocard_types.py by @julien-c in #931flake8-bugbear + adapt existing codebase by @Wauplin in #967Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Remove deprecations by @LysandreJik in #910
v0.8.1 introduces a new way of caching files from the Hugging Face Hub, to two methods: snapshot_download and hf_hub_download.
The new approach is extensively documented in the Documenting files guide and we recommend checking it out to get a better understanding of how caching works.
create_commit APIA new create_commit API allows users to upload and delete several files at once using HTTP-based methods. You can read more about it in this guide. The following convenience methods were also introduced:
upload_folder: Allows uploading a local directory to a repo.delete_file allows deleting a single file from a repo.upload_file now uses create_commit under the hood.
create_commit also allows creating pull requests with a create_pr=True flag.
None of the methods rely on Git locally.
create_commit API by @SBrandeis in #888All modules will now be lazy-loaded. This should drastically reduce the time it takes to import huggingface_hub as it will no longer load all soft dependencies.
hub-ci for tests by @SBrandeis in #898metadata_update: work on a copy of the upstream file, to not mess up the cache by @julien-c in #891Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
This PR adds a metadata_update function that allows the user to update the metadata in a repository on the hub. The function accepts a dict with metad
This PR adds a metadata_update function that allows the user to update the metadata in a repository on the hub. The function accepts a dict with metadata (following the same pattern as the YAML in the README) and behaves as follows for all top level fields except model-index.
Examples:
Starting from
existing_results = [{
'dataset': {'name': 'IMDb', 'type': 'imdb'},
'metrics': [{'name': 'Accuracy', 'type': 'accuracy', 'value': 0.995}],
'task': {'name': 'Text Classification', 'type': 'text-classification'}
}]
new_results = deepcopy(existing_results)
new_results[0]["metrics"][0]["value"] = 0.999
_update_metadata_model_index(existing_results, new_results, overwrite=True)
[{'dataset': {'name': 'IMDb', 'type': 'imdb'},
'metrics': [{'name': 'Accuracy', 'type': 'accuracy', 'value': 0.999}],
'task': {'name': 'Text Classification', 'type': 'text-classification'}}]
new_results = deepcopy(existing_results)
new_results[0]["metrics"][0]["name"] = "Recall"
new_results[0]["metrics"][0]["type"] = "recall"
[{'dataset': {'name': 'IMDb', 'type': 'imdb'},
'metrics': [{'name': 'Accuracy', 'type': 'accuracy', 'value': 0.995},
{'name': 'Recall', 'type': 'recall', 'value': 0.995}],
'task': {'name': 'Text Classification', 'type': 'text-classification'}}]
new_results = deepcopy(existing_results)
new_results[0]["dataset"] = {'name': 'IMDb-2', 'type': 'imdb_2'}
[{'dataset': {'name': 'IMDb', 'type': 'imdb'},
'metrics': [{'name': 'Accuracy', 'type': 'accuracy', 'value': 0.995}],
'task': {'name': 'Text Classification', 'type': 'text-classification'}},
{'dataset': ({'name': 'IMDb-2', 'type': 'imdb_2'},),
'metrics': [{'name': 'Accuracy', 'type': 'accuracy', 'value': 0.995}],
'task': {'name': 'Text Classification', 'type': 'text-classification'}}]
Nothing published for this version
*Disclaimer*: This release was initially released with advertised support for #844. It was not released in this release and will be in v0.7.
Disclaimer: This release was initially released with advertised support for #844. It was not released in this release and will be in v0.7.
v0.6.0 introduces downstream (download) and upstream (upload) support for the fastai libraries. It supports fastai versions above 2.4. The integration is detailed in the following blog.
RepositoryBinary files are now rejected by default by the Hub. v0.6.0 introduces automatic binary file tracking through the auto_lfs_track argument of the Repository.git_add method. It also introduces the Repository.auto_track_binary_files method which can be used independently of other methods.
skip_lfs_file is now added to mixinsThe parameter skip_lfs_files is now added to the different mixins. This will enable pushing files to the hub without first downloading the files above 10MB. This should drammatically reduce the time needed when updating a modelcard, a configuration file, and others.
The support for Keras model is greatly improved through several additions:
save_pretrained_keras method now accepts a list of tags that will automatically be added to the repository.hf_api a bit and add support for Spaces by @julien-c in #792Your coding agent can read these notes before it upgrades. Set up the MCP server →