Skip to content

10. Dynamic Resource Allocation (DRA) for GPUs

DRA (resource.k8s.io/v1, GA in Kubernetes 1.34) replaces "count of opaque devices" with claims over devices that have attributes. NVIDIA provides the NVIDIA DRA Driver for GPUs (k8s-dra-driver-gpu), which the GPU Operator can deploy alongside or instead of the device plugin.

10.1 What DRA adds

  • Select GPUs by attributes using CEL, for example memory ≥ 40Gi, a specific architecture, or a driver version.
  • Shared claims: Multiple pods or containers can reference the same ResourceClaim, so they share one GPU explicitly.
  • Dynamic MIG and partitioning without pre-labeling nodes (driver feature, check maturity).
  • ComputeDomains for multi-node NVLink (GB200 NVL72): IMEX channels are managed for you.
  • Admin access and device health status reporting.

10.2 Example

apiVersion: resource.k8s.io/v1
kind: ResourceClaimTemplate
metadata:
  name: big-gpu
spec:
  spec:
    devices:
      requests:
        - name: gpu
          exactly:
            deviceClassName: gpu.nvidia.com
            selectors:
              - cel:
                  expression: |
                    device.capacity['gpu.nvidia.com'].memory.compareTo(quantity('40Gi')) >= 0
---
apiVersion: v1
kind: Pod
metadata:
  name: dra-gpu-pod
spec:
  resourceClaims:
    - name: gpu
      resourceClaimTemplateName: big-gpu
  containers:
    - name: app
      image: nvcr.io/nvidia/cuda:12.6.0-base-ubuntu22.04
      command: ["nvidia-smi", "-L"]
      resources:
        claims:
          - name: gpu

10.3 Two containers sharing one GPU

apiVersion: resource.k8s.io/v1
kind: ResourceClaim
metadata:
  name: shared-gpu
spec:
  devices:
    requests:
      - name: gpu
        exactly:
          deviceClassName: gpu.nvidia.com
---
# Pods A and B both reference resourceClaimName: shared-gpu
spec:
  resourceClaims:
    - name: gpu
      resourceClaimName: shared-gpu

Device plugin or DRA? Use the device plugin today for simple, stable fleets. Adopt DRA when you need attribute-based selection, explicit sharing, or GB200 ComputeDomains. Don't advertise the same GPUs through both on the same node.