Skip to main content

GPU Observability & Reliability

A large training job runs overnight, then crashes on an XID error from one bad GPU. Without monitoring, your team restarts on the same node and loses another day. With proper observability, the failing GPU is flagged and drained quickly.

One source of truth

DCGM telemetry into Prometheus and Grafana, so per-GPU health, utilization and errors live in one place your on-call can trust.

Catch it before the crash

XID and ECC trends flag a degrading GPU while the job is still healthy, so hardware is drained on your schedule, not the job's.

Recover automatically

Fault detection wired to cordon, GPU reset, DCGM diagnostics and escalation, so jobs restart on healthy hardware.

What we do

See every GPU, catch every fault.

From the exporter on each node to the alert that pages your on-call, we build the reliability layer around your cluster.

01
DCGM metrics stack
DCGM Exporter deployment, custom field groups per workload type, and collection intervals tuned for training versus inference.
02
Prometheus integration
ServiceMonitor and PodMonitor setup, recording rules for cluster aggregations, and remote write to Thanos or Cortex for large clusters.
03
Grafana dashboards
Cluster overview, per-node GPU detail, job performance correlation, hardware health trends and capacity planning.
04
XID & health detection
Real-time XID monitoring from kernel logs and DCGM with severity classification and automated node drain for critical codes, plus ECC trend tracking, thermal throttling and PCIe link degradation alerts.
05
Automated recovery
Detect fault, cordon node, attempt GPU reset, run DCGM diagnostics, then uncordon or escalate, with Node Problem Detector integration and tuned alerting rules.
How we work

Scope, design, build, hand off.

01

Assess

Audit current monitoring gaps. Most clusters have zero GPU observability.

02

Deploy

DCGM Exporter, Prometheus, Grafana, alerting and recovery automation.

03

Tune

Adjust thresholds and collection intervals for your SLOs.

04

Transfer

Dashboards, runbooks, alert playbooks and on-call procedures.

Our engineers contribute upstream to the projects this layer runs on: KAI Scheduler, Network Operator, DOCA driver build, ipoib-cni. See the contributions →
Technologies we work with
DCGMDCGM ExporterPrometheusGrafanaAlertmanagerThanosNode Problem DetectorGPU Operatornvidia-smiKubernetesSlurm
FAQ
What is DCGM?
NVIDIA DCGM (Data Center GPU Manager) is the official toolkit for GPU telemetry, diagnostics and policy management. It exposes per-GPU utilization, memory, temperature, power, ECC errors, XID events and PCIe metrics through a Prometheus exporter. It is the reliable source of truth for GPU health.
Which GPU metrics matter for reliability?
SM utilization, memory bandwidth utilization, XID errors, ECC double-bit and single-bit counts, power draw, thermal throttling events, PCIe replay counts and NVLink error counters. For training, add NCCL timeouts and AllReduce duration. These catch the majority of hardware and driver issues before jobs crash.
What is an XID error?
XID errors are NVIDIA driver events reported via the kernel log when something goes wrong: ECC failures, a GPU falling off the bus, hardware errors or timeouts. Some are transient, others (like XID 79) are fatal and require node replacement. Mature monitoring alerts on these and automates node draining for critical codes.
How fast can GPU observability detect failures?
With DCGM scraping at short intervals and proper alerts, most failures (thermal throttling, ECC storms, XID events, PCIe link downgrade) are detected within about a minute. Fail-fast controllers can drain the affected pod automatically, so jobs restart on healthy hardware.
Can you integrate with our existing monitoring stack?
Yes. DCGM exports Prometheus metrics, so it drops into any stack built on Prometheus, Grafana, VictoriaMetrics, Mimir, Datadog or Grafana Cloud. We also integrate with PagerDuty, Opsgenie, Loki and Elastic, and build GPU-specific Grafana dashboards on top of your existing setup.
Tired of GPU failures going undetected?

We build monitoring that catches GPU issues before they crash your training jobs.