1. Fundamentals: GPUs in Kubernetes¶
1.1 How Kubernetes sees hardware¶
Kubernetes natively schedules only CPU, memory, ephemeral storage, and hugepages. It knows nothing about GPUs. GPUs reach Kubernetes through two mechanisms:
| Mechanism | Description | Status |
|---|---|---|
| Device Plugin framework | A DaemonSet registers an extended resource (for example nvidia.com/gpu) with the kubelet over gRPC. The kubelet advertises an integer count, and the device plugin decides which devices a container gets. |
Stable, the most widely used |
| Dynamic Resource Allocation (DRA) | A richer API with ResourceClaim, DeviceClass, and ResourceSlice. It supports attribute-based selection, sharing, and partitioning. |
GA in Kubernetes 1.34 (resource.k8s.io/v1) |
1.2 What a container needs to use a GPU¶
For a containerized process to run CUDA code, all of the following must be in place:
- The NVIDIA kernel driver (
nvidia.ko,nvidia-uvm.ko,nvidia-modeset.ko, and optionallynvidia-peermem.ko) loaded on the host. - Device nodes (
/dev/nvidia0,/dev/nvidiactl,/dev/nvidia-uvm, and so on) exposed inside the container. - User-space driver libraries (
libcuda.so,libnvidia-ml.so, and so on) that match the kernel driver version, mounted into the container. - A container runtime hook or CDI spec (NVIDIA Container Toolkit) that injects items 2 and 3.
- A scheduler-visible resource (from the device plugin or DRA) so pods land on GPU nodes.
- The CUDA runtime and your application inside the container image (for example
nvidia/cuda:12.x-runtime).
Key idea: The container image ships the CUDA toolkit/runtime. The host provides the driver. A newer driver supports older CUDA runtimes (backward compatibility). Forward compatibility, where an older driver runs a newer CUDA, is possible only on datacenter GPUs with the
cuda-compatpackage.
1.3 The manual way (and why it hurts)¶
Without the operator, you would have to do the following on every GPU node:
- Install the driver package that matches the kernel, and rebuild it on every kernel update.
- Install and configure the NVIDIA Container Toolkit for containerd or CRI-O.
- Deploy the device plugin DaemonSet.
- Deploy DCGM and the DCGM exporter for metrics.
- Deploy Node Feature Discovery and GPU Feature Discovery for labels.
- Configure MIG by hand.
- Coordinate driver upgrades with node drains.
The GPU Operator automates all of this.