| No operator pods on a GPU node |
NFD didn't label the node (pci-10de.present) |
Check NFD worker pods. Check for duplicate NFD installs |
Driver pod CrashLoopBackOff, "Could not resolve Linux kernel version" |
Kernel headers unavailable for this kernel |
Update the OS repos, use a supported kernel, or use precompiled drivers |
| Driver fails, "Key was rejected by service" |
Secure Boot blocks unsigned modules |
Disable Secure Boot or use signed/precompiled drivers |
| Driver fails, "nouveau" in use |
Nouveau loaded |
Blacklist nouveau and reboot |
| Driver pod hangs, "Unable to load nvidia module: already in use" |
Host driver is pre-installed |
driver.enabled=false |
Toolkit OK but pods fail with could not select device driver "" with capabilities: [[gpu]] |
Runtime not configured, or the wrong containerd config path |
Check toolkit.env CONTAINERD_CONFIG/SOCKET for your distro. Restart containerd |
nvidia.com/gpu: 0 allocatable |
Device plugin not registered, all GPUs unhealthy, or MIG mixed strategy hiding GPUs |
Device plugin logs. nvidia-smi. Check MIG strategy |
Pod Pending: Insufficient nvidia.com/gpu |
Real shortage, or fragmentation |
Bin-packing, autoscaling, check taints and tolerations |
TopologyAffinityError |
Topology Manager can't align NUMA |
Use restricted/best-effort, or fewer GPUs per pod |
CUDA error 802 system not yet initialized |
Fabric Manager not running (NVSwitch systems) |
Check the fabric manager in driver pod logs. Match versions |
| CUDA error 804 / "forward compatibility" |
Container CUDA newer than the driver |
Upgrade the driver or use an older CUDA image |
Failed to initialize NVML: Unknown Error after a while |
cgroup driver/systemd reload removed device access (cgroup v2 + runc) |
Use CDI, update the toolkit, or use the systemd cgroup driver consistently |
mig.config.state=failed |
GPU in use, or MIG enable needs a reset |
Drain workloads, check MIG manager logs, reboot if needed |
DCGM exporter shows no PROF_ metrics |
Profiling metrics not in CSV, or not supported (some GPUs, vGPU) |
Custom CSV (Section 11.3) |
| Two pods see the same GPU unexpectedly |
Time-slicing on, or NVIDIA_VISIBLE_DEVICES=all in the image |
Check sharing config. Harden env var handling |
Bus error in PyTorch DataLoader |
/dev/shm too small |
Memory-backed emptyDir (Section 13.2) |
NCCL slow, logs show NET/Socket |
RDMA not used |
Network Operator, IPC_LOCK, NCCL_IB_HCA, GDR settings |