SAS Viya Operations Dashboard

A SAS Viya dashboard that provides an at-a-glance view with minimal filtering. Built by curating panels from multiple dashboards.

SAS Viya Operations Dashboard screenshot 1

SAS Viya Operations Dashboard

A collection of Grafana panels for daily operations of SAS Viya.

The dashboard helps administrators identify abnormal conditions in a few minutes during routine health checks.

Import

Import the dashboard JSON file from the Grafana Dashboard Import page.

Features

  • Focused on daily operational checks.
  • Designed to reduce navigation and filtering.
  • Includes commonly used metrics for CPU, memory, storage, Kubernetes, and PostgreSQL.
  • Excludes containers that consistently show high utilization and are less useful for daily health checks (for example, pgBackRest).

Design Principles

This dashboard is not intended for deep investigation.

Instead, it is designed to answer a simple question:

Is everything healthy today?

The panels were modified from existing dashboards to reduce the number of operations required during routine monitoring.

Customization

This dashboard is intended as a starting point for SAS Viya operations.

If you find new metrics that should be monitored in daily operations, feel free to add them and customize the dashboard for your environment.

Notes

Some containers regularly appear at the top of utilization rankings even when the environment is healthy.

Use the Hidden dashboard variable to exclude such containers from charts and rankings.

Example:

pgbackrest|rabbitmq|sas-opendistro|sas-programming-environment|twistlock-defender|fluent-bit

Acknowledgements

This dashboard includes panels derived from the following sources.

SAS Viya Monitoring for Kubernetes

Original DashboardOriginal PanelNotes
Kubernetes HeadroomResource by NodeModified
Perf / Node UtilizationCPU UtilizationModified
CPU Saturation (load per CPU)Modified
Memory UtilizationModified
Memory Saturation (Major Page Faults)Modified
Disk IO UtilizationModified
Net Utilization (Bytes Receive/Transmit)Modified
Disk Capacity UsedModified
Disk IO SaturationModified
Perf / Container Utilization% CPU use / limModified
% MEM use / limModified
Kubernetes / Persistent VolumesVolume Space UsageModified
SAS Java ServicesJVM HeapModified
PostgreSQLConnectionsModified
Replication Lat Time (replica Only)Modified

Grafana Labs

Original DashboardOriginal PanelNotesID
OOM and RestartsrestartModified16718
OOM Killed processes pr node deltaModified16718
Spring Boot HTTPDuration of SUCCESSFUL requestsModified20820
Mean response timeModified20820

The original dashboards were used as references and were modified for daily operational monitoring.

Revisions
RevisionDescriptionCreated

Get this dashboard

Import the dashboard template

or

Download JSON

Datasource
Dependencies