Skip to content

22. Cheat Sheet and Checklists

22.1 Commands

# Install
helm install gpu-operator nvidia/gpu-operator -n gpu-operator --create-namespace

# Status
kubectl get clusterpolicy cluster-policy -o jsonpath='{.status.state}'
kubectl get nodes -L nvidia.com/gpu.product,nvidia.com/gpu.count,nvidia.com/mig.config.state

# GPU info
kubectl exec -n gpu-operator ds/nvidia-driver-daemonset -- nvidia-smi
kubectl exec -n gpu-operator ds/nvidia-driver-daemonset -- nvidia-smi topo -m
kubectl exec -n gpu-operator ds/nvidia-driver-daemonset -- nvidia-smi mig -lgip

# Time-slicing / MPS on
kubectl patch clusterpolicies.nvidia.com/cluster-policy --type merge \
  -p '{"spec":{"devicePlugin":{"config":{"name":"<cm>","default":"any"}}}}'

# MIG layout
kubectl label node <n> nvidia.com/mig.config=all-1g.10gb --overwrite

# Per-node sharing profile
kubectl label node <n> nvidia.com/device-plugin.config=<profile> --overwrite

# Metrics
kubectl port-forward -n gpu-operator svc/nvidia-dcgm-exporter 9400:9400
curl -s localhost:9400/metrics | grep DCGM_FI_DEV_GPU_UTIL

# Health
kubectl exec -n gpu-operator ds/nvidia-dcgm -- dcgmi diag -r 1

22.2 Production readiness checklist

Setup - [ ] Operator chart and driver version pinned in Git - [ ] Driver on a production branch. Open kernel modules on Hopper and newer - [ ] CDI enabled - [ ] gpu-operator namespace privileged. Other namespaces restricted - [ ] GPU nodes tainted. ExtendedResourceToleration enabled - [ ] Images mirrored (if air-gapped or rate-limited)

Observability - [ ] DCGM exporter with profiling metrics, a ServiceMonitor, and a Grafana dashboard - [ ] Alerts: XID, DBE ECC, temperature, operator not ready, idle allocated GPUs - [ ] Per-namespace GPU-hours (chargeback/showback)

Efficiency - [ ] Sharing strategy chosen per node pool (MIG/MPS/time-slicing/none) - [ ] Kueue (or KAI/Volcano) quotas with borrowing and preemption - [ ] Bin-packing scheduler profile for GPUs - [ ] Node autoscaling (scale to zero where possible), and KEDA for inference - [ ] Idle culling for notebooks. TTLs on Jobs

Performance - [ ] /dev/shm sized for training pods - [ ] CPU manager static plus Topology Manager for latency- or throughput-critical nodes - [ ] GPUDirect RDMA verified with nccl-tests (multi-node) - [ ] Mixed precision/FP8 in training. Optimized serving engines for inference - [ ] Data pipeline verified to keep SM_ACTIVE high (Nsight Systems) - [ ] Images and models cached on nodes

Reliability - [ ] Driver upgrade policy with drain and maxParallelUpgrades - [ ] Canary node pool for driver and operator upgrades - [ ] Periodic DCGM diagnostics. Automated cordoning of unhealthy GPUs - [ ] Training jobs checkpoint regularly (preemption- and failure-safe)