Webinar

How Pega Monitors 150,000+ Pods and Accelerates Root Cause Analysis with Grafana Cloud

You are registered for this webinar Thanks for registering
You'll receive an email confirmation, and a reminder on the day of the event. You'll receive an email when the on-demand video is available.
From blueprint to production: how Pega Cloud operates thousands of AI-generated enterprise applications with Grafana

Pega, the enterprise low-code platform behind mission-critical case-management applications for telecom, healthcare, insurance, and government customers, wants to catch problems before its clients ever notice them — and it’s chasing that goal across a serious footprint: 1,600+ Grafana users, 10,000+ monitored clusters, 150,000+ pods, and roughly 120 terabytes of telemetry ingested into Grafana Cloud every month.

The challenge

Early on, Pega’s own architecture diagrams had a gap. Pega’s simplified cloud architecture included web nodes, batch nodes, decisioning services, Postgres, Elasticsearch, Kafka, and Cassandra. But what was missing was true monitoring and observability.

As Pega scaled to thousands of Kubernetes clusters running case-management workloads for its clients, that gap became a real operational risk. A single underlying issue – a slow database query, a failing dependency – could ripple upward and get reported by dozens of different teams as dozens of different problems. At that scale, high cardinality also meant a fast-growing bill, not too mention, once dashboarding access opened up company-wide, dashboard sprawl followed. They needed a place to centralize their observability.

The solution

Pega now feeds its entire stack – HTTP and Tomcat response times, database queries, Kafka and Cassandra throughput, Kubernetes cluster health, and load-balancer traffic – into Grafana Cloud so SRE, operations, and client-facing field teams (1,600 of them) see the same picture of any of Pega’s 10,000+ monitored clusters, instead of stitching it together themselves.

To keep cost and noise under control at that scale, Pega adopted Adaptive Metrics and Adaptive Logs. On governance, Pega now separates the dashboards anyone can build from the smaller set the team uses for root-cause analysis and SLA reporting – those are locked down – and applies access controls to logs, since they can contain customer PII and Pega operates in heavily regulated environments, including the U.S. government’s FedRAMP program.

The impact

The result is a platform that’s shifting from reactive troubleshooting to proactive detection at enterprise scale.

  • Gives 1,600+ users across SRE, operations, and client-facing field teams one shared view of platform health
  • Monitors 10,000+ clusters and 150,000+ pods from a single Grafana Cloud stack
  • Ingests roughly 120 terabytes of telemetry monthly, with 130 million+ active metric series
  • Cut costs with Adaptive Metrics and Adaptive Logs by trimming unused cardinality
  • Replaced ungoverned dashboard sprawl (2,000+ dashboards) with a curated set for RCA and SLA reporting
  • Detects problems proactively before a client reports them for a significant percentage of issues.

Looking ahead

Pega’s next projects include monitoring its growing fleet of AI agents for performance, misuse, and token consumption, further improving its signal-to-noise ratio, and automating SLA reports for clients directly from its observability data.

More great videos and webinars