GPU cluster burn-in testing is the acceptance process used to expose hardware, firmware, topology, cooling, power, memory, PCIe, NVLink, network, and performance problems before a cluster is trusted with long-running AI jobs. A node that passes nvidia-smi can still fail under sustained matrix load, memory traffic, full rack power, all-to-all collectives, or checkpoint I/O.
Within NVIDIA AI Infrastructure, burn-in is the bridge between installation and production scheduling. NVIDIA DCGM provides a strong GPU-level diagnostic base, while cluster acceptance should add fabric, storage, collective, and sustained multi-node workload tests.
Current DCGM diagnostics include software/deployment checks, memory tests, PCIe, sustained diagnostic compute, memory bandwidth, targeted stress, targeted power, NVLink bandwidth, memtest, and other plugins depending on GPU/product support.
Begin with software and inventory validation
Confirm GPU count/model/UUID, driver, CUDA compatibility, Fabric Manager where required, device-node access, persistence/compute mode, firmware, NVLink/NVSwitch state, and management tooling.
DCGM level-1/software checks can expose missing libraries, inaccessible device nodes, conflicting environments, memory-health state, and other deployment issues before heavy stress begins.
Burn-in results are meaningless if half the node is unavailable because of a cgroup or driver mismatch rather than hardware failure.
Run DCGM diagnostics at a level appropriate to the acceptance window
DCGM higher-level diagnostic suites exercise memory, PCIe, compute, and other components and can consume substantial GPU, CPU, memory, power, and fabric resources.
Coordinate tests so no production workload shares the node and ensure rack power/cooling is ready for all nodes to draw high power simultaneously.
Record the exact DCGM version, diagnostic suite/tests, parameters, duration, and resulting failure codes for reproducibility.
Sustained compute catches marginal behavior short tests miss
DCGM’s diagnostic plugin performs matrix work and framebuffer verification while monitoring XIDs, ECC events, thermal excursions, and performance consistency.
Longer targeted-stress runs can expose thermal saturation, power delivery weakness, clock throttling, or intermittent GPU errors that a one-minute smoke test misses.
Burn-in duration should reflect deployment risk and job length; a cluster intended for multi-day training should not be accepted solely from a few seconds of compute.
Memory tests should exercise capacity and integrity
Use DCGM memory/memtest/diagnostic coverage to identify framebuffer allocation/read/write errors, ECC conditions, row-remap/page-retirement issues, and memory data-pattern failures.
Test near the expected usable memory footprint without overcommitting the host into unrelated swapping or OOM behavior.
GPU ECC Error Monitoring explains how to interpret memory-health signals that appear during burn-in.
PCIe and NVLink tests should match the actual topology
DCGM PCIe and NVLink-related tests can identify low bandwidth, excessive replay, broken peer paths, or link-state problems.
Compare measured GPU-to-host and GPU-to-GPU connectivity with nvidia-smi topo -m and the rack/server design.
A GPU that performs correctly alone can still be unacceptable if one NVLink is down or one PCIe link trains at a lower generation/width than the rest of the node.
Collective communication should be tested across the intended scale
Use NCCL collectives or an equivalent supported workload across one node, one rack, and multiple racks to validate GPU/NIC/RDMA/fabric paths.
All-reduce, all-gather, and reduce-scatter expose different traffic patterns and can identify weak links that point-to-point tests miss.
Track per-rank bandwidth and outliers; one slow host can reduce effective performance of the whole synchronous job.
Storage burn-in should run concurrently with GPU stress
AI clusters stress storage during dataset reads and checkpoint writes, often at the same time that GPUs and the network are busy.
Run representative read/write/checkpoint patterns while monitoring GPU duty cycle and fabric contention.
AI Storage Throughput covers how to size and interpret that path; burn-in confirms the installed cluster reaches the expected envelope.
Rack-level power and cooling are part of the test
Run enough nodes simultaneously to approach realistic rack power and thermal load.
Monitor inlet/outlet temperatures, GPU temperature, clock/power throttling, PSU/PDU state, fan behavior, and facility alarms.
A node that passes in isolation may throttle once a full rack exhausts hot air into the neighboring chassis or hits a shared power limit.
Outlier analysis should compare nominally identical GPUs
Current DCGM diagnostics can compare performance metrics across tested GPUs and flag large deviations depending on parameters.
Build your own baseline as well: GPU compute, memory bandwidth, PCIe, NVLink, NCCL, storage, temperature, and power should fall inside a narrow enough distribution for the hardware class.
Investigate persistent outliers even when they do not fail a vendor threshold; distributed workloads are often limited by the slowest participant.
Failures should be classified into retry, reset, isolate, or replace paths
Not every diagnostic failure proves a dead GPU. Software configuration, conflicting process, bad driver, transient fabric issue, pending row remap, thermal condition, or true hardware fault can produce different actions.
DCGM assigns diagnostic categories/severities that help automate triage, but operations should maintain a hardware-service policy tied to repeated evidence.
Keep failed nodes out of the production scheduler until the root cause is understood and a clean re-test passes.
Burn-in is successful when production starts with a measured baseline
The output should be more than pass/fail: store node/GPU inventory, topology, diagnostic results, NCCL/fabric throughput, storage performance, temperatures, power, ECC state, software versions, and test duration.
That baseline helps later incident response distinguish “this node was always 10% slower” from a real degradation that appeared after deployment.
Node acceptance should include firmware consistency across GPU, NIC/DPU, BIOS, BMC, NVSwitch, and storage components. Small version differences can produce performance variance or feature fallbacks that only appear under scale. Capture the complete firmware/software bill of materials with the burn-in result so later drift can be compared against the known-good baseline.
Fabric burn-in should include congestion, not only unloaded point-to-point bandwidth. Run simultaneous collective traffic across many nodes and observe switch port utilization, ECN/PFC or InfiniBand congestion counters, retransmissions, and tail latency. A fabric can meet single-flow bandwidth while collapsing when many ranks communicate at once.
Power testing should exercise transition behavior as well as sustained draw. Rapid changes between idle and full matrix/power load can expose board or power-delivery instability that steady-state tests miss. DCGM includes targeted power and pulse-style diagnostics for supported products, and rack telemetry should be observed at the same time.
Thermal soak should continue long enough for the chassis and rack to reach steady temperature. Five minutes of high power may not expose a cooling design that degrades after twenty or thirty minutes. Record GPU temperatures, clocks, power caps, fan speeds, inlet temperatures, and any thermal clock events throughout the run.
Cluster burn-in should validate scheduling labels and health automation too. Mark a deliberately failed node unhealthy and confirm the scheduler/health controller keeps new jobs away from it. The hardware test is incomplete if the operations plane cannot convert a diagnostic result into safe cluster behavior.
Network and storage cable/port mapping should be reconciled with inventory. A node can pass bandwidth while connected to the wrong switch rail or storage network, creating hidden oversubscription later when topology-aware jobs assume the documented rack design. Burn-in should compare actual topology to the intended rack/rail map.
Repeatability matters. Run a subset of diagnostics again after firmware upgrades, rack moves, cable replacement, or hardware repair and compare to the original baseline. A cluster should have an acceptance envelope, not one historical pass/fail event that is never revisited after material change.
Burn-in data should be searchable by node, GPU UUID, serial number, rack, and date. When a production job later reports one slow rank or ECC fault, engineers can determine whether that same GPU showed marginal bandwidth, high temperature, or transient errors during commissioning. This turns burn-in into long-term reliability evidence rather than disposable installation output.
Acceptance should include repeated cold boots and re-enumeration. Some PCIe, NVLink, NIC, or firmware issues appear only after power cycles, not during a continuous uptime test. Verify device count, PCIe link width/speed, GPU UUID mapping, NVSwitch/Fabric Manager state, and network interfaces after multiple restart sequences.
Job-level burn-in should use a representative framework stack in addition to synthetic diagnostics. Launch a small distributed training or inference workload using the same container image, CUDA/NCCL/framework versions, scheduler, and storage path production will use. This catches integration problems that DCGM intentionally does not test, such as framework/library incompatibility or launcher misconfiguration.
Failure injection improves confidence. Disable one fabric link, cordon one node, introduce a controlled storage slowdown, or remove one GPU from the scheduler and observe whether jobs, monitoring, and alerting behave as expected. Burn-in should validate the operational system around the hardware, not only the silicon.
Acceptance criteria should be published before the run begins: allowed ECC/XID state, minimum memory/PCIe/NVLink bandwidth, acceptable temperature and clock behavior, collective bandwidth range, storage thresholds, and maximum outlier deviation. Predefined criteria reduce the temptation to accept marginal nodes simply because deployment deadlines are approaching and make vendor/RMA conversations much more objective.
Cluster acceptance should also test observability continuity during stress. Metrics, logs, BMC telemetry, DCGM exporter, scheduler health, and alert delivery should remain available while GPUs, NICs, storage, and CPUs are saturated. A cluster that performs well but loses monitoring at peak load is not ready for production because the first real incident will remove the evidence needed to diagnose it.
Keep the burn-in profile as a comparison point for later incidents. Temperatures, clocks, power, memory errors, interconnect behavior, and representative throughput captured before production can reveal slow degradation that would otherwise be dismissed as normal workload variation.