NewYour coding agent can read the release notes before it upgrades.Set up the MCP server →
PyPI · #3046 most downloaded on PyPI
Download compute kernels
Last release 4 days ago
02 Oct 2026
Ships fairly regularly
a new release about every 3 weeks
Rarely documented
notes for 10 of 56 stable releases
Nothing withdrawn
no release was ever pulled
2 years old
59 releases · first in 2025
One column per month.
Fix fetch build fallback ( #848 ) by @SunMarc in #868
Full Changelog: v0.17.1...v0.17.2
Fix egg_info.writers entry point module path ( #829 ) by @jiqing-feng in #829 and #830
egg_info.writers entry point module path (#829) by @jiqing-feng in #829 and #830Full Changelog: v0.17.0...v0.17.1
doc: add a migration section for deprecated kernel functions by @danieldk in #679
get_kernel now raises an exception if the kernel does not support the architecture of the GPU in the machine. For instance, if a kernel is only compiled for Hopper (CUDA capability 9.0) is loaded on a Blackwell GPU with CUDA capability 12.0 an exception is raised. Previously, an incompatible kernel would load fine and then (loudly or quietly) fail at kernel launch time.
In some cases, an AOT kernel may support more architectures than the metadata declares, for instance when the kernel has a Triton fall-back path. For such kernels, the check can be disabled:
from kernels import get_kernel, has_kernel
has_kernel("kernels-community/flash-attn3", version=1, check_arch=False)
flash_attn3 = get_kernel("kernels-community/flash-attn3", version=1, check_arch=False)Kernels can now depend on other kernels from the Hub. Dependencies are declared in the general section of build.toml. For example:
[general]
# ...
kernel-depends = [
{ repo-id = "kernels-community/activation", version = 1 },
{ repo-id = "kernels-community/einops", version = 1 },
]The kernel dependencies can than be retrieved when the kernel is imported using the new get_kernel_dep function:
import kernels
activation = kernels.get_kernel_dep("kernels-community/activation")
einops = kernels.get_kernel_dep("kernels-community/einops")Warning: using kernel dependencies breaks support for kernels<0.17. For this reason, we recommend you to only start using kernel dependencies when version 0.17.0 has been out for a while. The kernels minimum version metadata described below will be ported to kernels 0.16 and will help with this and future migrations.
kernels minimum versionTo introduce new features more gracefully in the future, using newer kernels features (such as kernel dependencies) will write the minimum required version to the kernel metadata. The kernels client will use this to verify that it is compatible with the kernel and, if not, suggest what version of kernels to install.
use_kernel_forward_from_hub now accepts an optional condition argument. This condition is applied to the layer when kernelize is called. This is useful when a layer supports multiple configurations, but kernelization is only supported for one configuration:
@use_kernel_forward_from_hub(
"SwiGLUMLP",
condition=lambda module: module.config.hidden_act == "silu",
)
class MyMLP(nn.Module):
...kernel-builder now has experimental support for building TPU torch_tpu/Pallas kernels. backends = ["tpu"] produces a torch-tpu noarch variant.
The XPU/SYCL build supports Intel Crescent Island.
The Helion DSL is now supported through the new helion Python dependency (python-depends = ["helion"]). See the blog post Helion x 🤗 HF Kernels: Building and Shipping Out-of-the-box Performant Kernels for more information.
kernels only loads kernels from a curated set of trusted publishers. Until now, the only way to load a kernel from another publisher was to opt in with trust_remote_code=True, which disables the check entirely. You can now pass a list of repository IDs instead, so that only the given repositories are allowed:
get_kernel(
"some-other-org/my-kernel",
version=1,
trust_remote_code=["some-other-org/my-kernel"],
)kernels infoThe new kernels info subcommand describes a kernel from a Hub repo ID or a local path, printing its name, version, license, upstream/source repositories, Python dependencies, and supported backends:
$ kernels info kernels-community/activation
Repository: kernels-community/activation
Revision: v1
Name: activation
Version: 1
License: Apache-2.0
Upstream: -
Source: -
Python dependencies: -
Backends: cuda, metal
Kernel metadata now records the provenance of a build in a provenance object. The provenance information contains the Git commit of the kernel and the builder and whether both were in a clean or dirty state.
kernels 0.15.1 added experimental support for the Torch stable ABI. The stable ABI was enabled by setting torch.stable-abi to the stable ABI version that the kernel should be compiled with. Unfortunately, the stable ABI is not complete enough yet to be supported on all backends. For this reason, we have decided to change the torch.stable-abi option to take a table with versions per-backend:
[torch.stable-abi]
cuda = "2.11"
rocm = "2.9"Backends that are not listed are compiled without the stable ABI. This makes it possible to opt-in to the stable ABI per backend and use different stable ABI versions per backend.
This release adds support for Torch 2.13 and 2.14 and removes Torch 2.11 and 2.12.
stable-abi is now a table rather than a single Torch version. build.toml files can be migrated to edition 5 with kernel-builder update-build.check_arch=False.get_kernel instead of importing the kernel directly.kernels-data is not a separate Python package anymore, but is directly integrated into the kernels package.
cxx-flags for Torch and tvm-ffi sources by @danieldk in #680HIP_ARCHITECTURES gets properly cleared by @danieldk in #684kernel-depends option to general options in build.toml by @danieldk in #749GitHash by GitStatus by @danieldk in #765fetchFromHuggingFace: add by @danieldk in #766build.toml parsing to kernels-data by @danieldk in #768kernels-data by @danieldk in #771{rocm,xpu}Packages: read version manifests from directory by @danieldk in #797use_kernel_forward_from_hub by @danieldk in #796user_agent in layers by @sayakpaul in #805kernels-data Python binding into kernels by @danieldk in #818kernels-data to kernels-common by @danieldk in #821Full Changelog: v0.16.1...v0.17.0
This release adds a version check to require the user to upgrade when a kernel does not support this kernels version.
This release adds a version check to require the user to upgrade when a kernel does not support this kernels version.
Nothing published for this version
use_kernel_func_from_hub , FuncRepository , LocalFuncRepository , and LockedFuncRepository are now deprecated.
This release adds preliminary support for signing kernels. Kernel signing is currently in an experimental phase and the details may still change. For this reason, signatures are currently not yet validated when downloading a kernel. Kernel verification consists of two parts:
metadata.json. During verification, all files must be present and have the correct hash. Verification also fails when there are files that are not specified in the digest.metadata.json.sigstore. The signature is made using sigstore, which signs artifacts using ephemeral signing keys, reducing the impact of key theft. During verification, the signature is used to validate that metadata.json was not tampered with and that the signature was made from a trusted repository and workflow.Since the kernel is not verified yet during retrieval during the experimental stage, we provide the kernels verify-signature command-line utility that can be used to verify a kernel on the Hub. The kernel and kernel version to verify should be provided as arguments:
$ kernels verify-signature kernels-community/flash-attn4 0
✅ torch-cuda: kernel metadata is correctly signedAll kernels-community kernels are signed. If you want to experiment with kernel signing yourself, there are two changes you need to make:
nix flake update to get the latest version of kernel-builder. The latest version embeds the kernel digest in the metadata.cosign to sign the kernel. This cannot be done as part of the build itself, since signing using ephemeral kernels requires internet access and the kernel build sandbox does not provide internet access. You can use the kernels-community workflow as an example of how to set up metadata signing.kernel-builder now supports the cpu-kernels skill for writing, optimizing, and benchmarking C++ kernels using AVX2/AVX512. For example, to add the skill to Claude, use:
$ kernel-builder skills add --skill cpu-kernels --claudeuse_kernel_func_from_hub, FuncRepository, LocalFuncRepository, and LockedFuncRepository are now deprecated.
To make a function extensible by a layer, you can now use the same decorator as for layers (use_kernel_forward_from_hub). This makes it clearer that the function is actually replaced by a layer. We have also added the use_kernelized_func decorator to attach such a function to the layer wherein it is used to make it discoverable by kernelize. Here is a full example:
# Make silu_and_mul replaceable with a kernel layer registered as `silu_and_mul`.
@use_kernel_forward_from_hub("silu_and_mul")
def silu_and_mul(x: torch.Tensor) -> torch.Tensor:
d = x.shape[-1] // 2
return F.silu(x[..., :d]) * x[..., d:]
# Attach the function to the layer where it is used to make it discoverable by `kernelize`.
@use_kernelized_func(silu_and_mul)
class FeedForward(nn.Module):
def __init__(self, in_features: int, out_features: int):
self.linear = nn.Linear(in_features, out_features)
def forward(self, x: torch.Tensor) -> torch.Tensor:
return silu_and_mul(self.linear(x))The FuncRepository, LocalFuncRepository, and LockedFuncRepository classes will not be replaced. They allowed using an arbitrary function from a kernel as a layer. However, this was easily misused and did not have a clean way of marking such a function as supporting torch.compile or backwards passes. Going forward, they should be made available as regular kernel layers that can be used with LayerRepository and its local/locked versions.
For more information, see the layer documentation.
Kernels support a small set of curated Python dependencies, such as einops, nvidia-cute-dsl, and apache-tvm-ffi. These dependencies are now also provided as extras of the kernels package, curated for CUDA and curated-xpu for XPU:
# CUDA
$ pip install 'kernels[curated]'
# XPU
$ pip install 'kernels[curated-xpu]'This can be used to install all dependencies that a kernel might use.
The documentation now provides an overview of the kernel-builder architecture.
FuncRepository] Add ability of detecting flags by @vasqu in #607hash subcommand and hook up in Nix by @danieldk in #618--filter-unsigned by @danieldk in #653OIDCSourceRepositoryURI by @danieldk in #652pyext in torch-noarch by @danieldk in #658kernel repo publishing rights by @sayakpaul in #662ptxas version than CUDA version by @danieldk in #671Full Changelog: v0.15.1...v0.16.0
This release adds support for can_torch_compile / can_backward to FuncRepository .
This release adds support for can_torch_compile/can_backward to FuncRepository.
As announced by deprecation warnings in previous releases, specifying the kernel version is now required when loading a kernel. E.g.
As announced by deprecation warnings in previous releases, specifying the kernel version is now required when loading a kernel. E.g.
# Not valid anymore!
activation = kernels.get_kernel("kernels-community/activation")is now invalid, instead use:
activation = kernels.get_kernel("kernels-community/activation", version=1)The Hub page for a kernel shows the latest available kernel version. kernels will also warn if the loaded kernel is not the latest version. Full specification of the version helps avoiding breaking existing code as a result of kernel API changes. When the API of a kernel changes, the kernel author must bump up the API version so that downstream code that hasn't been updated for the API change yet, can continue to use the previous version.
This release adds support for the Torch stable ABI. When a kernel uses the Torch stable API and sets the the ABI version in build.toml, the kernel will be built to be compatible with that Torch version and later. For instance, the targeted Torch version can be set to 2.10 by setting stable-abi in build.toml:
[torch]
stable-abi = "2.10"Using the stable ABI has large benefits:
Functions like get_kernel that normally use rely on network access now work with HF_HUB_OFFLINE=1. The kernel will be loaded if it was downloaded before, otherwise an exception will be raised. Using HF_HUB_OFFLINE disables trusted publisher verification (since this requires internet access).
kernel-builder now offers a skill for writing Intel XPU kernels contributed by @danielfleischer. For instance, to add the XPU kernels skill for Claude, use:
$ kernel-builder skills add --claude --skill xpu-kernelsUp till this release, kernel-builder has always linked libstdc++ statically. However, this lead to issues for some kernels where both the statically linked instance and the dynamically linked instance would try to initialize the same global memory, leading to segfaults and other issues. We didn't encounter this behavior before because most kernels only use C++ code for simple wrapping of the actual compute functions. However, we have encountered some kernels using facilities like C++ std::regex, which triggers global locale initialization. To resolve these issues, we switched to dynamic linking of libstdc++.
To enable dynamic linking while still being fully compatile with manylinux_2_28, we rewrap the EL8 gcc toolchain that is used by manylinux_2_28 using Nix and expose it as a stdenv. This allows us to build kernels with this toolchain. For more technical details, see: https://huggingface.co/docs/kernels/builder/design-nix-builder#manylinux228-compatibility
We now have a page that describes how to set up a kernel development environment in your IDE. Currently Visual Studio Code is covered, but we plan to add additional IDEs and editors in the future.
--major option by @danieldk in #550Device by @danieldk in #571LocalLayerRepository.__str__ by @danieldk in #597kernel-builder check-abi by @danieldk in #596Full Changelog: v0.14.1...v0.15.1
This is a bugfix release to use the Hub API to check that a publisher is trusted.
This is a bugfix release to use the Hub API to check that a publisher is trusted.
deprecation for version and revision check. by @sayakpaul in #450
Kernels are now a separate repository type on the Hub. This brings many usability improvements to kernels. For example, you can view all kernels that are hosted on the Hub on the kernel overview page:
https://huggingface.co/kernels
This page allows you to filter kernels by supported backends (CUDA, XPU, Metal, etc.) and specific accelerators such as NVIDIA H100 or Apple M5 Max. The page for a kernel will also show the supported accelerators, operating systems, architectures and Torch versions. For instance, the flash-attn2 page shows all supported hardware and architectures:
https://huggingface.co/kernels/kernels-community/flash-attn2
Starting with kernels 0.14, we only support the new kernel repository type. If you would like to upload kernels to the Hub, you can request support for kernel repositories for your user or organization under Settings -> Account.
Kernels have had support for optional metadata to state dependencies, etc. However, we have made including metadata mandatory. This allows users of kernels to query metadata such as the kernel's license and name. But it also made it possible to load kernels by their unique identifier that is also used in their Torch ops name. This makes it easier to debug kernels, since the dynamically loaded name corresponds to the operator name.
We have also added support for querying metadata of kernels that have been loaded:
>>> for kernel in kernels.get_loaded_kernels():
... metadata = kernel.metadata
... repo_info = kernel.repo_info
... print(metadata.id, repo_info.repo_id, metadata.backend.backend_type)
_relu_metal_c835f43 kernels-community/relu metalTo improve security and restrict loading of arbitrary code, kernels will by default only load kernels from trusted publishers. To load other kernels, use the trust_remote_code option:
get_kernel("some-other-org/my-kernel", version=1, trust_remote_code=True)This release adds support for Torch 2.12, currently based on RC9. The main branch will be updated with the final release when it is available, but there are typically no ABI changes in (late) release candidates, so building kernels with RC9 should also work on the final release.
get_loaded_kernels() by @cbensimon in #428result, build, or target dir by @danieldk in #464Metadata by @danieldk in #481build.toml when specified by @danieldk in #485None by @danieldk in #493make pin-actions target to pin all GitHub actions by @danieldk in #498Metadata from kernels-data by @danieldk in #499Full Changelog: v0.13.0...v0.14.0
* feat: add test and docs for get_loaded_kernels.
Fix failing tests (#497)
* feat: add test and docs for get_loaded_kernels.
* fix existing failing tests
* fix 2
* style
---------
Co-authored-by: Daniël de Kok <me@danieldk.eu>
Nothing published for this version
kernels 0.13.0 is a feature-packed release with among other things an improved CLI for building kernels ( kernel-builder ), Torch 2.11 support, and a
kernels 0.13.0 is a feature-packed release with among other things an improved CLI for building kernels (kernel-builder), Torch 2.11 support, and a tech-preview of TVM FFI support.
kernel-builder CLI overhaulThe build2cmake command has been renamed to kernel-builder. This new tool can be used to develop, build, and upload kernels without directly using Nix.
These are the main subcommands for the new kernel-builder CLI:
kernel-builder init: scaffold a new kernel, including tests and benchmarks.kernel-builder build: build a kernel.kernel-builder build-and-copy: build a kernel and copy artifacts to the build directory.kernel-builder build-and-upload: build a kernel and upload it to the Hub.kernel-builder create-pyproject: create Python project such as pyproject.toml to develop kernels in IDEs and editors.kernel-builder devshell / kernel-builder testshell — drop into a development or test shell for a kernel.kernel-builder upload: upload a built kernel to the Hugging Face Hub.kernel-builder list-variants — list all supported build variants for a kernel.The build, devshell, and testshell subcommands accept a --variant flag to select a specific build variant. All subcommands accept a directory argument instead of requiring a specific working directory.
An installation script is also provided to help new users get a working kernel-builder environment set up quickly, including Nix, the binary cache, and the required trusted-user configuration. Go to the following page for information on how to get started:
https://huggingface.co/docs/kernels/main/en/builder/writing-kernels#quick-install
kernel-builder now supports Torch 2.11. Torch 2.9 support has been removed in accordance with our policy of supporting the two latest PyTorch versions.
kernels 0.13 adds support for TVM FFI kernels. TVM FFI aims to be a single ABI for multiple frameworks, such as Torch, JAX, NumPy, and CuPy. TVM FFI support is a tech preview. For instance, we might still make changes to the build.toml options for TVM FFI, change the kernel source layout, or change the provided helper functions.
The kernels examples directory provides ReLU and CUTLASS example kernels that use TVM FFI.
kernel-builder now supports card filling. If the kernel source repository contains a CARD.md template, building a kernel will fill the template with details about the kernel. When a kernel is uploaded (with kernel-builder upload or kernel-builder build-and-upload), the card will be uploaded as the README.md of the Hub repository. The default card template can be generated with kernels init.
kernels skillsWe added a new CLI command for installing an agent-compatible skill. Use kernels skills add to install the skills for AI coding assistants like Claude, Codex, and OpenCode. For now, only the cuda-kernels skill is supported. Skill files are downloaded from the huggingface/kernels directory in this repository. ROCm kernel skills are on the way.
Kernels can now be overridden locally without changing any get_kernel call sites. Set the LOCAL_KERNELS environment variable to a colon-separated list of org/repo=local_path pairs:
LOCAL_KERNELS=kernels-community/activation=/path/to/local/activation
This is useful for testing kernel changes locally before uploading them to the Hub.
This is useful when running some operations on CPU while the rest of the model runs on a GPU.
Large kernel uploads are now automatically split across multiple commits to stay within Hub limits, rather than failing or requiring manual intervention for kernels with many files.
metadata.json is correctly added to the output of Windows builds by @danieldk in #242torchVersions argument of genKernelFlakeOutputs by @danieldk in #246render_binding and render_extensions by @danieldk in #248setup.py and move writing to common module by @danieldk in #250common by @danieldk in #249render_deps function by @danieldk in #251cli by @danieldk in #269get_kernel: support specifying the backend by @danieldk in #268kernels skills add to the cli by @burtenshaw in #278ci-test package by @danieldk in #281repo_id in the card usage. by @sayakpaul in #284python_dependencies.json path by @danieldk in #310build2cmake -> kernel-builder, move nix bits to nix-builder by @danieldk in #340kernel-builder generate -> kernel-builder create-pyproject by @danieldk in #344pyproject::compat -> pyproject::common by @danieldk in #366completions subcommand for generating shell completions by @danieldk in #388main by @sayakpaul in #389devshell/testshell by @danieldk in #394metadata.json and build.toml datastructures to separate crate by @danieldk in #395Note truncated.
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Your coding agent can read these notes before it upgrades. Set up the MCP server →