11. Monitoring and Observability (DCGM)¶
You can't optimize what you don't measure. The DCGM exporter is the foundation.
11.1 Wire it into Prometheus¶
With kube-prometheus-stack installed:
helm upgrade gpu-operator nvidia/gpu-operator -n gpu-operator --reuse-values \
--set dcgmExporter.serviceMonitor.enabled=true
If Prometheus selects ServiceMonitors by label, add dcgmExporter.serviceMonitor.additionalLabels.release=<prometheus-release>.
Import the official NVIDIA DCGM Exporter Grafana dashboard (Grafana ID 12239).
11.2 The metrics that actually matter¶
| Metric | What it tells you | Watch for |
|---|---|---|
DCGM_FI_DEV_GPU_UTIL |
% of time any kernel was running | A coarse metric. 100% can still mean poor use of the hardware |
DCGM_FI_PROF_GR_ENGINE_ACTIVE |
Graphics/compute engine active ratio | A more accurate version of "busy" |
DCGM_FI_PROF_SM_ACTIVE |
Fraction of SMs with at least one warp resident | Low = kernels too small or poor parallelism |
DCGM_FI_PROF_SM_OCCUPANCY |
Warps resident relative to the maximum | Low = register or shared memory pressure, small blocks |
DCGM_FI_PROF_PIPE_TENSOR_ACTIVE |
Tensor Core usage | Near 0 on DL training = not using mixed precision |
DCGM_FI_PROF_DRAM_ACTIVE |
Memory bandwidth usage | High = memory bound |
DCGM_FI_DEV_FB_USED / FB_FREE |
VRAM usage | Over-provisioned GPUs, OOM risk |
DCGM_FI_PROF_PCIE_TX/RX_BYTES |
Host-device transfer | High = data loading bottleneck |
DCGM_FI_PROF_NVLINK_TX/RX_BYTES |
GPU-to-GPU traffic | Collective communication health |
DCGM_FI_DEV_POWER_USAGE |
Watts | Power capping, efficiency |
DCGM_FI_DEV_GPU_TEMP, DCGM_FI_DEV_MEMORY_TEMP |
Temperature | Thermal throttling |
DCGM_FI_DEV_CLOCK_THROTTLE_REASONS (or CLOCKS_EVENT_REASONS) |
Why clocks dropped | Power, thermal, or sync boost |
DCGM_FI_DEV_XID_ERRORS |
Last XID error code | Hardware or driver faults |
DCGM_FI_DEV_ECC_DBE_VOL_TOTAL |
Double-bit ECC errors | Failing memory, so retire the GPU |
DCGM_FI_DEV_ROW_REMAP_FAILURE |
Row remapping failed | Needs RMA |
GPU_UTILlies. A tiny kernel running nonstop shows 100% utilization while using 1 of 132 SMs. Use SM_ACTIVE, SM_OCCUPANCY, and PIPE_TENSOR_ACTIVE to judge real efficiency.
11.3 Enabling profiling metrics¶
The DCGM_FI_PROF_* metrics may not be in the default CSV. Provide a custom metrics file:
apiVersion: v1
kind: ConfigMap
metadata:
name: dcgm-metrics
namespace: gpu-operator
data:
dcgm-metrics.csv: |
DCGM_FI_DEV_GPU_UTIL, gauge, GPU utilization (%)
DCGM_FI_DEV_FB_USED, gauge, Framebuffer used (MiB)
DCGM_FI_DEV_FB_FREE, gauge, Framebuffer free (MiB)
DCGM_FI_DEV_POWER_USAGE, gauge, Power (W)
DCGM_FI_DEV_GPU_TEMP, gauge, GPU temperature (C)
DCGM_FI_DEV_SM_CLOCK, gauge, SM clock (MHz)
DCGM_FI_DEV_XID_ERRORS, gauge, Last XID error
DCGM_FI_DEV_ECC_DBE_VOL_TOTAL, counter, Volatile DBE errors
DCGM_FI_PROF_GR_ENGINE_ACTIVE, gauge, Graphics engine active ratio
DCGM_FI_PROF_SM_ACTIVE, gauge, SM active ratio
DCGM_FI_PROF_SM_OCCUPANCY, gauge, SM occupancy ratio
DCGM_FI_PROF_PIPE_TENSOR_ACTIVE,gauge, Tensor pipe active ratio
DCGM_FI_PROF_DRAM_ACTIVE, gauge, DRAM active ratio
DCGM_FI_PROF_PCIE_TX_BYTES, gauge, PCIe TX bytes/s
DCGM_FI_PROF_PCIE_RX_BYTES, gauge, PCIe RX bytes/s
DCGM_FI_PROF_NVLINK_TX_BYTES, gauge, NVLink TX bytes/s
DCGM_FI_PROF_NVLINK_RX_BYTES, gauge, NVLink RX bytes/s
kubectl apply -f dcgm-metrics.yaml
helm upgrade gpu-operator nvidia/gpu-operator -n gpu-operator --reuse-values \
--set dcgmExporter.config.name=dcgm-metrics
11.4 Per-pod attribution¶
The DCGM exporter maps GPUs to pods via the kubelet pod-resources API, adding pod, namespace, and container labels. This enables per-team dashboards and chargeback. For MIG, metrics carry GPU_I_PROFILE and GPU_I_ID labels.
11.5 Useful PromQL¶
# Average real SM activity per namespace
avg by (namespace) (DCGM_FI_PROF_SM_ACTIVE{namespace!=""})
# Idle allocated GPUs (allocated to a pod but <5% busy for 1h) -> waste
avg_over_time(DCGM_FI_DEV_GPU_UTIL{pod!=""}[1h]) < 5
# Memory headroom per GPU (%)
100 * DCGM_FI_DEV_FB_USED / (DCGM_FI_DEV_FB_USED + DCGM_FI_DEV_FB_FREE)
# Tensor Core usage on training jobs
avg by (pod) (DCGM_FI_PROF_PIPE_TENSOR_ACTIVE{namespace="training"})
# Cluster-wide allocation ratio (needs kube-state-metrics)
sum(kube_pod_container_resource_requests{resource="nvidia_com_gpu"})
/ sum(kube_node_status_allocatable{resource="nvidia_com_gpu"})
11.6 Alerts to have¶
groups:
- name: gpu.rules
rules:
- alert: GPUXidError
expr: DCGM_FI_DEV_XID_ERRORS > 0
for: 1m
labels: {severity: critical}
annotations:
summary: "XID {{ $value }} on {{ $labels.Hostname }} GPU {{ $labels.gpu }}"
- alert: GPUDoubleBitECC
expr: increase(DCGM_FI_DEV_ECC_DBE_VOL_TOTAL[10m]) > 0
labels: {severity: critical}
- alert: GPUHighTemp
expr: DCGM_FI_DEV_GPU_TEMP > 85
for: 5m
labels: {severity: warning}
- alert: GPUAllocatedButIdle
expr: avg_over_time(DCGM_FI_DEV_GPU_UTIL{pod!=""}[2h]) < 5
labels: {severity: info}
annotations:
summary: "Pod {{ $labels.namespace }}/{{ $labels.pod }} holds a GPU but is idle"
- alert: GPUOperatorNotReady
expr: gpu_operator_reconciliation_status == 0
for: 15m
labels: {severity: warning}