6. Installation Scenarios (Intermediate)¶
6.1 A production-style values.yaml¶
# values-prod.yaml
operator:
defaultRuntime: containerd
nfd:
enabled: true
driver:
enabled: true
version: "<driver-version>" # pin it, e.g. a production branch (R550 / R570 / R580)
kernelModuleType: auto # auto | open | proprietary
rdma:
enabled: false # true for GPUDirect RDMA (with Network Operator)
upgradePolicy:
autoUpgrade: true
maxParallelUpgrades: 1
maxUnavailable: 25%
drain:
enable: true
force: false
deleteEmptyDir: true
timeoutSeconds: 300
gpuPodDeletion:
force: false
deleteEmptyDir: true
timeoutSeconds: 300
waitForCompletion:
timeoutSeconds: 0
podSelector: "" # e.g. "app=training" to wait for jobs to finish
toolkit:
enabled: true
cdi:
enabled: true
default: false
devicePlugin:
enabled: true
config:
name: "" # ConfigMap for time-slicing/MPS (Section 9)
default: ""
mig:
strategy: single # single | mixed
migManager:
enabled: true
dcgm:
enabled: true
dcgmExporter:
enabled: true
serviceMonitor:
enabled: true # requires Prometheus Operator CRDs
interval: 15s
gfd:
enabled: true
validator:
plugin:
env:
- name: WITH_WORKLOAD
value: "false"
# Keep operator components off non-GPU nodes and tolerate GPU taints
daemonsets:
tolerations:
- key: nvidia.com/gpu
operator: Exists
effect: NoSchedule
priorityClassName: system-node-critical
helm upgrade --install gpu-operator nvidia/gpu-operator \
-n gpu-operator --version=<chart-version> -f values-prod.yaml --wait
6.2 Drivers already installed on the host¶
If the Container Toolkit is also pre-installed on the host, add --set toolkit.enabled=false.
6.3 Precompiled drivers (faster, Secure Boot friendly)¶
Precompiled images exist only for specific kernel flavors (mainly Ubuntu). They skip on-node compilation, so nodes come up in seconds rather than minutes.
6.4 Per-node-pool drivers with the NVIDIADriver CRD¶
You can run different driver versions or types on different node pools, for example a stable branch on inference nodes and a new branch on training nodes.
helm install gpu-operator nvidia/gpu-operator -n gpu-operator \
--set driver.nvidiaDriverCRD.enabled=true \
--set driver.nvidiaDriverCRD.deployDefaultCR=false
apiVersion: nvidia.com/v1alpha1
kind: NVIDIADriver
metadata:
name: h100-training
spec:
driverType: gpu
kernelModuleType: open
version: "<driver-version>"
repository: nvcr.io/nvidia
image: driver
nodeSelector:
nvidia.com/gpu.product: NVIDIA-H100-80GB-HBM3
---
apiVersion: nvidia.com/v1alpha1
kind: NVIDIADriver
metadata:
name: l4-inference
spec:
driverType: gpu
version: "<other-driver-version>"
repository: nvcr.io/nvidia
image: driver
nodeSelector:
nvidia.com/gpu.product: NVIDIA-L4
The node selectors of different
NVIDIADriverCRs must not overlap.
6.5 Containerd variants (k3s, RKE2, MicroK8s)¶
These distributions keep containerd config in non-standard paths:
toolkit:
env:
- name: CONTAINERD_CONFIG
value: /var/lib/rancher/rke2/agent/etc/containerd/config.toml.tmpl # RKE2
- name: CONTAINERD_SOCKET
value: /run/k3s/containerd/containerd.sock
- name: CONTAINERD_RUNTIME_CLASS
value: nvidia
- name: CONTAINERD_SET_AS_DEFAULT
value: "true"
For k3s the config is /var/lib/rancher/k3s/agent/etc/containerd/config.toml.tmpl. Recent operator versions also support containerd drop-in config files.
6.6 Managed clouds¶
| Platform | Notes |
|---|---|
| AWS EKS | EKS GPU AMIs (AL2023 NVIDIA, Bottlerocket NVIDIA) ship with drivers, so use driver.enabled=false and often toolkit.enabled=false. With a plain Ubuntu AMI, use the full operator |
| GKE | GKE can install drivers itself (gpu-driver-version=latest on the node pool). If you use the operator, set driver.enabled=false and toolkit.enabled=false and use GKE's device plugin, or create node pools with gpu-driver-version=disabled and Ubuntu images and let the operator manage everything. Don't mix the two approaches |
| AKS | Create GPU node pools with --skip-gpu-driver-install (or the equivalent option in current AKS), then install the full operator |
| OpenShift | Install NFD and the GPU Operator from OperatorHub. Drivers are built against RHCOS via the Driver Toolkit, and entitlement is not needed on modern OCP |
6.7 Air-gapped / disconnected clusters¶
- Mirror all images (operator, driver for each OS and kernel, toolkit, device plugin, DCGM, exporter, validator, NFD, MIG manager) to an internal registry.
- Set
--set operator.repository=...,driver.repository=..., and so on, or use a global registry override. - For on-node driver compilation, provide a local package repository via a ConfigMap (
driver.repoConfig.configMapName), or use precompiled drivers. - Provide custom CA certificates via
driver.certConfig.nameif your mirror uses a private CA.