NewYour coding agent can read the release notes before it upgrades.Set up the MCP server →
Go modules · #563 by repository stars
Last release 3 days ago
06 Oct 2026
Ships on a steady schedule
a new release about every 1 weeks
Nearly every release is documented
notes for 11 of 11 stable releases
Nothing withdrawn
no release was ever pulled
7 years old
456 releases · first in 2020
One column per quarter.
- Added startupProbe to NVIDIA driver container to allow RollingUpgrades to progress to other nodes only after driver modules are successfully loaded
Added startupProbe to NVIDIA driver container to allow RollingUpgrades to progress to other nodes only after driver modules are successfully loaded on current one.
Added support for driver.rollingUpdate.maxUnavailable parameter to specify maximum nodes for simultaneous driver upgrades. Default is 1.
NVIDIA driver container will auto-disable itself on the node with pre-installed drivers by applying label nvidia.com/gpu.deploy.driver=pre-installed . This is useful for heterogeneous clusters where only some GPU nodes have pre-installed drivers(e.g. DGX OS).
Apply tolerations to cuda-validator and device-plugin-validator Pods based on deamonsets.tolerations in ClusterPolicy . For more info refer here .
Fixed an issue causing cuda-validator Pod to fail when accept-nvidia-visible-devices-envvar-when-unprivileged = false is set with NVIDIA Container Toolkit. For more info refer here .
Fixed an issue which caused recursive mounts under /run/nvidia/driver when both driver.rdma.enabled and driver.rdma.useHostMofed are set to true . This caused other GPU Pods to fail to start.
- The gpu-operator:v1.11.0 and gpu-operator:v1.11.0-ubi8 images have been released with the following known HIGH Vulnerability CVEs. These are from th…
Support for NVIDIA Data Center GPU Driver version 515.48.07 .
Support for NVIDIA AI Enterprise 2.1.
Support for NVIDIA Virtual Compute Server 14.1 (vGPU).
Support for Ubuntu 22.04 LTS.
Support for secure boot with GPU Driver version 515 and Ubuntu Server 20.04 LTS and 22.04 LTS.
Support for Kubernetes 1.24.
Support for Time-Slicing GPUs in Kubernetes .
Support for Red Hat OpenShift on AWS, Azure and GCP instances. Refer to the Platform Support Matrix for the supported instances.
Support for Red Hat Openshift 4.10 on AWS EC2 G5g instances(ARM).
Support for Kubernetes 1.24 on AWS EC2 G5g instances(ARM).
Support for use with the NVIDIA Network Operator 1.2.
[Technical Preview] - Support for KubeVirt and Red Hat OpenShift Virtualization with GPU Passthrough and NVIDIA vGPU based products .
[Technical Preview] - Kubernetes on ARM with Server Base System Architecture (SBSA).
GPUDirect RDMA is now supported with CentOS using MOFED installed on the node.
The NVIDIA vGPU Manager can now be upgraded to a newer branch while using an older, compatible guest driver.
DGX A100 and non-DGX servers can now be used within the same cluster.
Improved user interface while deploying a ClusterPolicy instance(CR) for the GPU Operator through Red Hat OpenShift Console.
Improved the container-toolkit to handle v1 containerd configurations.
Fix for incorrect reporting of DCGM_FI_DEV_FB_USED where reserved memory is reported as used memory. For more details refer to GitHub issue .
Fixed nvidia-peermem sidecar container to correctly load the nvidia-peermem module when MOFED is directly installed on the node.
Fixed duplicate mounts of /run/mellanox/drivers within the driver container which caused driver cleanup or re-install to fail.
Fixed uncordoning of the node with k8s-driver-manager whenever ENABLE_AUTO_DRAIN env is disabled.
Fixed readiness check for MOFED driver installation by the NVIDIA Network Operator. This will avoid the GPU driver containers to be in CrashLoopBackOff while waiting for MOFED drivers to be ready.
All worker nodes within the Kubernetes cluster must use the same operating system version.
The NVIDIA GPU Operator can only be used to deploy a single NVIDIA GPU Driver type and version. The NVIDIA vGPU and Data Center GPU Driver cannot be used within the same cluster.
See the limitations sections for the [Technical Preview] of GPU Operator support for KubeVirt.
The clusterpolicies.nvidia.com CRD has to be manually deleted after the GPU Operator is uninstalled using Helm.
nouveau driver has to be blacklisted when using the NVIDIA vGPU. Otherwise the driver will fail to initialize the GPU with the error Failed to enable MSI-X in the system journal logs and all GPU Operator pods will be stuck in init state.
The gpu-operator:v1.11.0 and gpu-operator:v1.11.0-ubi8 images have been released with the following known HIGH Vulnerability CVEs. These are from the base images and are not in libraries used by GPU Operator:
xz-libs - CVE-2022-1271
- The gpu-operator:v1.10.1 and gpu-operator:v1.10.1-ubi8 images have been released with the following known HIGH Vulnerability CVEs. These are from th…
Validated secure boot with signed NVIDIA Data Center Driver R510.
Validated cgroup v2 with Ubuntu Server 20.04 LTS.
Fixed an issue when GPU Operator was installed and MIG was already enabled on a GPU. The GPU Operator will now install successfully and MIG can either be disabled via the label nvidia.com/mig.config=all-disabled or configured with the required MIG profiles.
The gpu-operator:v1.10.1 and gpu-operator:v1.10.1-ubi8 images have been released with the following known HIGH Vulnerability CVEs. These are from the base images and are not in libraries used by GPU Operator:
openssl-libs - CVE-2022-0778
zlib - CVE-2018-25032
gzip - CVE-2022-1271
- The gpu-operator:v1.10.0 and gpu-operator:v1.10.0-ubi8 images have been released with the following known HIGH Vulnerability CVEs. These are from th…
Support for NVIDIA Data Center GPU Driver version 510.47.03 .
Support NVIDIA A2, A100X and A30X
Support for A100X and A30X on the DPU’s Arm processor.
Support for secure boot with Ubuntu Server 20.04 and NVIDIA Data Center GPU Driver version R470.
Support for Red Hat OpenShift 4.10.
Support for GPUDirect RDMA with Red Hat OpenShift.
Support for NVIDIA AI Enterprise 2.0.
Support for NVIDIA Virtual Compute Server 14 (vGPU).
Enabling/Disabling of GPU System Processor (GSP) Mode through NVIDIA driver module parameters.
Ability to avoid deploying GPU Operator Operands on certain worker nodes through labels. Useful for running VMs with GPUs using KubeVirt.
Increased lease duration of GPU Operator to 60s to avoid restarts during etcd defrag. More details here .
Avoid spurious alerts generated of type GPUOperatorOpenshiftDriverToolkitEnabledNfdTooOld on RedHat OpenShift when there are no GPU nodes in the cluster.
Avoid uncordoning nodes during driver pod startup when ENABLE_AUTO_DRAIN is set to false .
Collection of GPU metrics in MIG mode is now supported with 470+ drivers.
Fabric Manager (required for NVSwitch based systems) with CentOS 7 is now supported.
Upgrading to a new NVIDIA AI Enterprise major branch:
Upgrading the vGPU host driver to a newer major branch than the vGPU guest driver will result in GPU driver pod transitioning to a failed state. This happens for instance when the Host is upgraded to vGPU version 14.x while the Kubernetes nodes are still running with vGPU version 13.x.
To overcome this situation, before upgrading the host driver to the new vGPU branch, apply the following steps:
kubectl edit clusterpolicy
modify the policy and set the environment variable DISABLE_VGPU_VERSION_CHECK to true as shown below:
driver : env : - name : DISABLE_VGPU_VERSION_CHECK value : "true"
write and quit the clusterpolicy edit
The gpu-operator:v1.10.0 and gpu-operator:v1.10.0-ubi8 images have been released with the following known HIGH Vulnerability CVEs. These are from the base images and are not in libraries used by GPU Operator:
openssl-libs - CVE-2022-0778
- Improved logic in the driver container for waiting on MOFED driver readiness. This ensures that nvidia-peermem is built and installed correctly.
Improved logic in the driver container for waiting on MOFED driver readiness. This ensures that nvidia-peermem is built and installed correctly.
Allow driver container to fallback to using cluster entitlements on Red Hat OpenShift on build failures. This issue exposed itself when using GPU Operator with some Red Hat OpenShift 4.8.z versions and Red Hat OpenShift 4.9.8. GPU Operator 1.9+ with Red Hat OpenShift 4.9.9+ doesn’t require entitlements.
Fixed an issue when DCGM-Exporter didn’t work correctly with using the separate DCGM host engine that is part of the standalone DCGM pod. Fixed the issue and changed the default behavior to use the DCGM Host engine that is embedded in DCGM-Exporter. The standalone DCGM pod will not be launched by default but can be enabled for use with DGX A100.
Update to latest Go vendor packages to avoid any CVE’s.
Fixed an issue to allow GPU Operator to work with CRI-O runtime on Kubernetes.
Mount correct source path for Mellanox OFED 5.x drivers for enabling GPUDirect RDMA.
- Automatic detection of default runtime used in the cluster. Deprecate the operator.defaultRuntime parameter.
Support for NVIDIA Data Center GPU Driver version 470.82.01 .
Support for DGX A100 with DGX OS 5.1+.
Support for preinstalled GPU Driver with MIG Manager.
Removed dependency to maintain active Red Hat OpenShift entitlements to build the GPU Driver. Introduce entitlement free driver builds starting with Red Hat OpenShift 4.9.9.
Support for GPUDirect RDMA with preinstalled Mellanox OFED drivers.
Support for GPU Operator and operands upgrades using Red Hat OpenShift Lifecycle Manager (OLM).
Support for NVIDIA Virtual Compute Server 13.1 (vGPU).
Automatic detection of default runtime used in the cluster. Deprecate the operator.defaultRuntime parameter.
GPU Operator and its operands are installed into a single user specified namespace.
A loaded Nouveau driver is automatically detected and unloaded as part of the GPU Operator install.
Added an option to mount a ConfigMap of self-signed certificates into the driver container. Enables SSL connections to private package repositories.
Fixed an issue when DCGM Exporter was in CrashLoopBackOff as it could not connect to the DCGM port on the same node.
GPUDirect RDMA is only supported with R470 drivers on Ubuntu 20.04 LTS and is not supported on other distributions (e.g. CoreOS, CentOS etc.)
The GPU Operator supports GPUDirect RDMA only in conjunction with the Network Operator. The Mellanox OFED drivers can be installed by the Network Operator or pre-installed on the host.
Upgrades from v1.8.x to v1.9.x are not supported due to GPU Operator 1.9 installing the GPU Operator and its operands into a single namespace. Previous GPU Operator versions installed them into different namespaces. Upgrading to GPU Operator 1.9 requires uninstalling pre 1.9 GPU Operator versions prior to installing GPU Operator 1.9
Collection of GPU metrics in MIG mode is not supported with 470+ drivers.
The GPU Operator requires all MIG related configurations to be executed by MIG Manager. Enabling/Disabling MIG and other MIG related configurations directly on the host is discouraged.
Fabric Manager (required for NVSwitch based systems) with CentOS 7 is not supported.
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Your coding agent can read these notes before it upgrades. Set up the MCP server →