16. Performance Optimization: Application Level¶
The biggest gains usually come from the workload itself.
16.1 Training¶
| Technique | Gain | Notes |
|---|---|---|
| Mixed precision (BF16/FP16) | 2–3× vs FP32 | Uses Tensor Cores. Check PIPE_TENSOR_ACTIVE |
| FP8 (Hopper/Blackwell) | Up to ~1.5–2× over BF16 | Transformer Engine. FP4/MXFP formats on Blackwell |
torch.compile |
1.2–2× | Kernel fusion, fewer launches |
| FlashAttention / SDPA | Large for attention | Memory and speed |
| Larger batch size to fill VRAM | Higher utilization | Use gradient accumulation if memory bound |
| Activation checkpointing | Fits bigger models/batches | Trades compute for memory |
| FSDP / ZeRO / tensor/pipeline parallel | Scale beyond one GPU's memory | Megatron-Core, DeepSpeed, NeMo |
| CUDA Graphs | Reduces CPU launch overhead | Good for small, static-shaped steps |
| Overlap comm/compute | Hides NCCL time | Bucketed DDP, FSDP prefetch |
Fused optimizers (fused=True AdamW, Apex) |
5–15% | Fewer kernels |
16.2 Inference¶
| Technique | Notes |
|---|---|
| Optimized serving engines | vLLM, SGLang, TensorRT-LLM, NVIDIA Dynamo, Triton Inference Server, NVIDIA NIM |
| Continuous/in-flight batching | Large throughput gain over static batching |
| PagedAttention / KV-cache management | Higher concurrency per GPU |
| Quantization (FP8, INT8, INT4 AWQ/GPTQ, NVFP4) | Smaller footprint, higher throughput, often fits on a cheaper GPU or MIG slice |
| Speculative decoding | Lower latency |
| Disaggregated prefill/decode | Dynamo, llm-d. Scales prefill and decode pools separately |
| KV-cache-aware routing | Gateway API Inference Extension, Dynamo router. Reuses the prefix cache |
| TensorRT / ONNX Runtime with TRT EP | For CV and classic DL models |
| Dynamic batching in Triton | max_queue_delay_microseconds tuned to the latency SLO |
| Right-size the GPU | A 7B INT4 model doesn't need an H100. L4, a MIG slice, or time-slicing may do |
16.3 Profiling tools¶
- Nsight Systems (
nsys): timeline of CPU, GPU, NCCL, and data loading. Finds idle gaps. - Nsight Compute (
ncu): per-kernel analysis. - PyTorch Profiler with TensorBoard.
- DCGM profiling metrics for continuous fleet-level monitoring (Section 11).
Run Nsight Systems in a pod:
nsys profile -o /results/profile --trace=cuda,nvtx,osrt,cudnn,cublas \
--duration=60 python train.py
(Profiling counters may require NVreg_RestrictProfilingToAdminUsers=0 in the driver module parameters, set through the driver's kernelModuleConfig ConfigMap, or the SYS_ADMIN capability.)