NVIDIA GPU Operator on Kubernetes: The Complete Guide¶
From the basics to advanced setup, performance tuning, and getting the most out of your GPUs
Version note: The GPU Operator, its components, and Kubernetes change quickly. The commands and values here follow current GPU Operator conventions (v24.x/v25.x). Before you run anything in production, check the exact chart version, the driver branch, and the support matrix in the official docs.
Contents¶
- Fundamentals: GPUs in Kubernetes
- What the GPU Operator Is and Why You Need It
- Architecture and Components
- Prerequisites and Planning
- Installation (Basic)
- Installation Scenarios (Intermediate)
- Verifying the Installation
- Scheduling GPU Workloads
- GPU Sharing: Time-Slicing, MPS, MIG, vGPU
- Dynamic Resource Allocation (DRA) for GPUs
- Monitoring and Observability (DCGM)
- Performance Optimization: Node and Hardware Level
- Performance Optimization: Kubernetes Level
- Performance Optimization: Networking and Multi-Node (RDMA, NCCL)
- Performance Optimization: Storage and Data Pipelines
- Performance Optimization: Application Level
- Using GPUs Efficiently: Utilization, Quotas, Autoscaling, Cost
- Day-2 Operations: Upgrades, Health, Reliability
- Security Considerations
- Troubleshooting Playbook
- Reference Architectures
- Cheat Sheet and Checklists
- Further Reading