NVIDIA GPU Operator & DCGM Observability

Live NVIDIA GPU Operator and DCGM exporter observability. Covers all 23 active DCGM metric families and GPU Operator health metrics; every panel has a Grafana hover-info description.

NVIDIA GPU Operator & DCGM Observability screenshot 1
NVIDIA GPU Operator & DCGM Observability screenshot 2

OKE NVIDIA GPU Operator & DCGM Dashboard

Classic Grafana dashboard for NVIDIA GPU Operator and DCGM Exporter telemetry on OKE or Kubernetes.

Import

  1. In Grafana, choose Dashboards → New → Import and select dashboard.json.
  2. Bind the Prometheus datasource when prompted.
  3. Select a GPU node, GPU, or workload using the dashboard filters.

Requirements

  • Prometheus must scrape the DCGM exporter and GPU Operator.
  • The default scrape jobs are nvidia-dcgm-exporter and gpu-operator. If yours differ, edit the job selectors after import.
  • DCGM workload attribution requires the labels Hostname, UUID, gpu, exported_namespace, exported_pod, and exported_container.

The dashboard includes the DCGM metric families emitted by the reference GPU Operator configuration. Metrics unsupported by your GPU or collector version may show No data.

For Grafana.com sharing, upload dashboard.json as a Classic JSON dashboard. Do not upload dashboard.v2-spec.original.json.

Revisions
RevisionDescriptionCreated

Get this dashboard

Import the dashboard template

or

Download JSON

Datasource
Dependencies