Advanced GPU Observability

Advanced GPU Observability for AKS using Inspektor Gadget eBPF CUDA metrics (enabled via the Inspektor Gadget cluster extension and scraped by Azure Managed Prometheus) and Pyroscope profiles. Tracks per-pod CUDA memory usage, top processes, allocation/free rates, and CUDA memory profiles.

Advanced GPU Observability screenshot 1
The Advanced GPU Observability dashboard uses the geneva-datasource and grafana-azureprometheus-datasource data sources to create a Grafana dashboard with the flamegraph, text and timeseries panels.
Revisions
RevisionDescriptionCreated

Get this dashboard

Import the dashboard template

or

Download JSON

Datasource
Dependencies