NVIDIA Data Center GPU Manager (DCGM) provides telemetry, health monitoring, diagnostics, topology, profiling, job statistics, and programmatic APIs for NVIDIA data-center accelerators. DCGM Exporter exposes selected DCGM fields to Prometheus, making it the common monitoring path for GPU Kubernetes and bare-metal environments. Current DCGM documentation distinguishes passive health watches from active diagnostics: health interprets retained telemetry while workloads keep running, whereas diagnostics actively exercise GPU, interconnect, memory, and related components.
Within NVIDIA AI Infrastructure, DCGM should be the GPU hardware/telemetry layer that complements application metrics. GPU ECC Error Monitoring provides the deeper memory-error context.
Use DCGM field watches as the telemetry foundation
DCGM collects fields for utilization, clocks, temperature, power, memory, ECC, XID errors, PCIe, NVLink/NVSwitch, ConnectX and other supported hardware subsystems.
Choose a sampling interval appropriate to the failure or performance signal.
One-second utilization may be useful for scheduler efficiency, while slower thermal/power trends can use longer intervals.
DCGM Exporter turns fields into Prometheus metrics
dcgm-exporter selects fields from its collector configuration and serves them in Prometheus format on the /metrics endpoint.
Use the default collector as a starting point and add fields only when dashboards or alerts have a clear operational purpose.
High-cardinality labels and unnecessary high-frequency metrics increase monitoring cost without improving diagnosis.
Passive health checks can run with production workloads
DCGM health monitoring watches retained telemetry and reports warnings/failures for enabled subsystems without launching a dedicated workload.
Current health watch coverage includes PCIe, memory, InfoROM, driver, thermal, power, NVLink, NVSwitch, IMEX, ConnectX, and other supported entities.
A PASS means no enabled health rule found an incident; it does not prove every component was watched or stress-tested.
Active diagnostics answer a different question
dcgmi diag runs tests that exercise deployment, hardware, software integration, memory, compute, or interconnect according to the selected diagnostic suite.
These tests can consume substantial GPU/CPU/memory/power/fabric resources.
Run quick readiness checks before jobs and deeper tests after failures or during maintenance windows when GPUs can be taken out of normal service.
Monitor XID errors as events, not only counters
XID errors can indicate driver, GPU memory, PCIe, application, or hardware conditions with very different severity.
DCGM Exporter can expose XID count/total metrics and labels for observed XID values.
Alert on relevant XIDs and correlate with job ID, GPU UUID, ECC, temperature, NVLink and application crash evidence before deciding whether to drain or RMA a node.
Track ECC and memory pressure together
Correctable ECC events, uncorrectable errors, retired pages, memory usage, and OOM/application failures tell a more complete story than one ECC number.
A rising corrected-error trend may justify maintenance before it becomes an uncorrectable fault.
Use GPU Cluster Burn-In Testing after hardware replacement or repeated errors to validate the node under controlled load.
NVLink and P2P status belong on the same dashboard as GPU health
GPU compute can appear healthy while one NVLink path is degraded or disabled.
DCGM exposes NVLink status/error counters and current exporter metrics can expose P2P status relationships.
Correlate collective-performance regressions with interconnect health before changing NCCL parameters.
Job-level attribution makes utilization actionable
DCGM supports process/job statistics and workload labeling patterns so GPU metrics can be tied to scheduler allocations, Kubernetes pods, or applications.
Dashboards should show utilization, memory, power, errors, and interconnect health by node and by workload.
This separates an underutilized GPU caused by the model from an idle GPU caused by data/network starvation.
MIG requires entity-aware monitoring
On Multi-Instance GPU systems, metrics may apply to physical GPUs, GPU instances, or compute instances depending on field support.
Do not aggregate every metric only at the physical GPU level if the scheduler allocates MIG slices independently.
Use entity labels and inventory so an alert identifies the affected slice/workload without confusing it with another tenant on the same card.
Prometheus alerts should encode operational actions
Create alerts for thermal/power throttling, sustained low utilization where jobs should be busy, XID events, uncorrectable ECC, degraded NVLink, health FAIL/WARN, and exporter/host-engine absence.
Each alert should state whether to observe, drain, reset, run diagnostics, or escalate hardware replacement.
A wall of GPU metrics is not monitoring until responders know what state transition each threshold requires.
DCGM monitoring succeeds when telemetry, health, diagnostics, and workload context are connected
The mature cluster exports a curated metric set, watches passive health, uses active diagnostics in controlled windows, attributes telemetry to jobs, handles MIG entities correctly, and correlates GPU, PCIe, NVLink, ECC, power and application evidence.
DCGM turns accelerator health into an operational signal when dashboards can answer both “which GPU is unhealthy?” and “which workloads were affected?”
DCGM can run with an embedded host engine inside tools such as dcgm-exporter or with a separately managed host engine. Standardize one operating model per environment so multiple processes do not create confusing ownership or sampling behavior. Kubernetes deployments commonly use the GPU Operator/Exporter integration, while bare metal may run hostengine as a system service.
Metric selection should distinguish counters from gauges and events. Power, temperature, utilization, memory use, and clocks are current-state or sampled values; XID and some error metrics are event/counter oriented. Alert expressions should match semantics so a historical one-time XID does not look like a continuously increasing active fault forever.
Health watches have retention windows. If the monitoring system polls less often than DCGM retains the relevant samples, an incident can disappear before the external alert sees it. Choose update interval and max-keep-age long enough for the collector and incident-response cadence, especially on batch clusters where jobs finish overnight.
DCGM groups can scope monitoring and diagnostics to the GPUs allocated to one job or node role. Create transient job groups or scheduler integration when investigating a specific allocation rather than running disruptive diagnostics across every GPU in the chassis.
Power and clock-throttle reasons should be visible next to utilization. A GPU at 70% utilization with power/thermal throttling has a different optimization path from one at 70% because the application is input-bound. Current exporter-owned metrics can expose clock-event reasons that help distinguish idle, configured clocks, thermal slowdown, power cap, or hardware braking conditions.
PCIe telemetry belongs with NVLink because external communication can bottleneck on either. Monitor link generation/width where available, replay/error conditions, and topology health. A training job spanning nodes may use NVLink locally and GPUDirect RDMA through PCIe/NIC externally; one degraded PCIe path can appear as an NCCL problem.
Exporter absence should be a first-class alert. If dcgm-exporter, hostengine, driver, or permissions fail, dashboards may simply show no data. Alert on scrape failure and expected GPU count mismatch so loss of observability is not mistaken for a healthy quiet node.
Kubernetes labels can connect GPU metrics to pods/namespaces when configured appropriately. Preserve workload attribution without exploding cardinality; avoid labels that include ephemeral request IDs or unbounded user strings. The monitoring system should answer which job used the GPU while staying operational at cluster scale.
Diagnostics should be integrated with node-drain workflows. When DCGM health or repeated XIDs cross a threshold, drain the node, stop workloads, run the appropriate diagnostic suite, collect results, then decide reset, driver remediation, cable/interconnect work, or hardware replacement. Automating this sequence reduces repeated job failures on a known-bad node.
RMA decisions should preserve diagnostic evidence. Record GPU UUID/serial, driver/DCGM version, XID/ECC history, health results, active diagnostic output, temperatures/power, NVLink/PCIe state, and workload symptoms. DCGM is most valuable when operations can move from a vague ‘GPU flaky’ report to repeatable hardware evidence.
DCGM field identifiers and exporter names should be versioned in dashboards. New exporter releases can add or deprecate fields, labels, or derived metric families. Pin dashboards/alert rules to tested exporter/DCGM versions and validate the effective /metrics output after upgrades rather than assuming every configured field is emitted on every GPU generation.
Health WARN and FAIL should map to scheduler actions by category. A transient performance-threshold warning might justify observation, while memory, NVLink, PCIe, or driver failures may require draining. Use the health error category/severity labels in current exporter metrics so automation does not treat every non-PASS result identically.
DCGM diagnostics should be sequenced from least disruptive to most disruptive. Start with passive health and quick diagnostics, then move to medium or targeted tests after a job failure, reserving full stress validation for maintenance. This reduces unnecessary GPU downtime while still escalating when evidence points to hardware.
Monitoring should compare expected and discovered GPU/MIG inventory after reboot or maintenance. A missing GPU, disabled NVLink, changed MIG layout, or driver attach failure can reduce capacity without triggering an application-specific alert. Inventory drift should block scheduling until the node matches its approved hardware profile.
DCGM dashboards should separate fleet health from performance efficiency. Fleet health answers whether hardware is reliable; efficiency answers whether expensive GPUs are being used well. Keep XID/ECC/thermal/NVLink alerts distinct from low-utilization or memory-underuse optimization alerts so operations and ML performance teams receive the right work.
Node qualification after repair should combine DCGM diagnostics with a known workload or burn-in. Passing hardware diagnostics proves specific components under test, while an application-level soak can reveal intermittent PCIe, fabric, or thermal behavior that only appears under sustained distributed load.