NewYour coding agent can read the release notes before it upgrades.Set up the MCP server →
Go modules · #563 by repository stars
Last release 3 days ago
01 Oct 2026
Ships on a steady schedule
a new release about every 1 weeks
Nearly every release is documented
notes for 11 of 11 stable releases
Nothing withdrawn
no release was ever pulled
7 years old
452 releases · first in 2020
One column per quarter.
- Fixed an issue where Driver Daemonset was spuriously updated on RedHat OpenShift causing repeated restarts in Proxy environments.
Fixed an issue where Driver Daemonset was spuriously updated on RedHat OpenShift causing repeated restarts in Proxy environments.
The MIG Manager version was bumped to v0.1.3 to fix an issue when checking whether a GPU was in MIG mode or not. Previously, it would always check for MIG mode directly over the PCIe bus instead of using NVML. Now it checks with NVML when it can, only falling back to the PCIe bus when NVML is not available. Please refer to the Release notes for a complete list of fixed issues.
Container Toolkit bumped to version v1.7.1 to fix an issue when using A100 80GB.
Added support for user-defined MIG partition configuration via a ConfigMap .
Nothing published for this version
- Fixed an issue with using the NVIDIA License System in NVIDIA AI Enterprise deployments.
Fixed an issue with using the NVIDIA License System in NVIDIA AI Enterprise deployments.
- Support for NVIDIA Data Center GPU Driver version 470.57.02 .
Support for NVIDIA Data Center GPU Driver version 470.57.02 .
Added support for NVSwitch systems such as HGX A100. The driver container detects the presence of NVSwitches in the system and automatically deploys the Fabric Manager for setting up the NVSwitch fabric.
The driver container now builds and loads the nvidia-peermem kernel module when GPUDirect RDMA is enabled and Mellanox devices are present in the system. This allows the GPU Operator to complement the NVIDIA Network Operator to enable GPUDirect RDMA in the Kubernetes cluster. Refer to the RDMA documentation on getting started.
Note
This feature is available only when used with R470 drivers on Ubuntu 20.04 LTS.
Added support for upgrades of the GPU Operator components. A new k8s-driver-manager component handles upgrades of the NVIDIA drivers on nodes in the cluster.
NVIDIA DCGM is now deployed as a component of the GPU Operator. The standalone DCGM container allows multiple clients such as DCGM-Exporter and NVSM to be deployed and connect to the existing DCGM container.
Added a nodeStatusExporter component that exports operator and node metrics in a Prometheus format. The component provides information on the status of the operator (e.g. reconciliation status, number of GPU enabled nodes).
Reduced the size of the ClusterPolicy CRD by removing duplicates and redundant fields.
The GPU Operator now supports detection of the virtual PCIe topology of the system and makes the topology available to vGPU drivers via a configuration file. The driver container starts the nvidia-topologyd daemon in vGPU configurations.
Added support for specifying the RuntimeClass variable via Helm.
Added nvidia-container-toolkit images to support CentOS 7 and CentOS 8.
nvidia-container-toolkit now supports configuring containerd correctly for RKE2.
Added new debug options (logging, verbosity levels) for nvidia-container-toolkit
The driver container now loads ipmi_devintf by default. This allows tools such as ipmitool that rely on ipmi char devices to be created and available.
GPUDirect RDMA is only supported with R470 drivers on Ubuntu 20.04 LTS and is not supported on other distributions (e.g. CoreOS, CentOS etc.)
The operator supports building and loading of nvidia-peermem only in conjunction with the Network Operator. Use with pre-installed MOFED drivers on the host is not supported. This capability will be added in a future release.
Support for DGX A100 with GPU Operator 1.8 will be available in an upcoming patch release.
This version of GPU Operator does not work well on RedHat OpenShift when a cluster-wide proxy is configured and causes constant restarts of driver container. This will be fixed in an upcoming patch release v1.8.2 .
Nothing published for this version
- NFD version bumped to v0.8.2 to support correct kernel version labeling on Anthos nodes. See NFD issue for more details.
NFD version bumped to v0.8.2 to support correct kernel version labeling on Anthos nodes. See NFD issue for more details.
Nothing published for this version
Nothing published for this version
- Support for NVIDIA Data Center GPU Driver version 460.73.01 .
Support for NVIDIA Data Center GPU Driver version 460.73.01 .
Added support for automatic configuration of MIG geometry on NVIDIA Ampere products (e.g. A100) using the k8s-mig-manager .
GPU Operator can now be deployed on systems with pre-installed NVIDIA drivers and the NVIDIA Container Toolkit.
DCGM-Exporter now supports telemetry for MIG devices on supported Ampere products (e.g. A100).
Added support for a new nvidia RuntimeClass with containerd .
The Operator now supports PodSecurityPolicies when enabled in the cluster.
Changed the label selector used by the DaemonSets of the different states of the GPU Operator. Instead of having a global label nvidia.com/gpu.present=true , each DaemonSet now has its own label, nvidia.com/gpu.deploy.<state>=true . This new behavior allows a finer grain of control over the components deployed on each of the GPU nodes.
Migrated to using the latest operator-sdk for building the GPU Operator.
The operator components are deployed with node-critical PriorityClass to minimize the possibility of eviction.
Added a spec for the initContainer image, to allow flexibility to change the base images as required.
Added the ability to configure the MIG strategy to be applied by the Operator.
The driver container now auto-detects OpenShift/RHEL versions to better handle node/cluster upgrades.
Validations of the container-toolkit and device-plugin installations are done on all GPU nodes in the cluster.
Added an option to skip plugin validation workload pod during the Operator deployment.
The gpu-operator-resources namespace is now created by the Operator so that they can be used by both Helm and OpenShift installations.
DCGM does not support profiling metrics on RTX 6000 and RTX 8000. Support will be added in a future release of DCGM Exporter.
After uninstall of GPU Operator, NVIDIA driver modules might still be loaded. Either reboot the node or forcefully remove them using sudo rmmod nvidia nvidia_modeset nvidia_uvm command before re-installing GPU Operator.
When MIG strategy of mixed is configured, device-plugin-validation may stay in Pending state due to incorrect GPU resource request type. User would need to modify the pod spec to apply correct resource type to match the MIG devices configured in the cluster.
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Your coding agent can read these notes before it upgrades. Set up the MCP server →