Webinar

How PointsBet Put AI at the Center of Incident Response with Grafana Cloud

You are registered for this webinar Thanks for registering
You'll receive an email confirmation, and a reminder on the day of the event. You'll receive an email when the on-demand video is available.
Inside PointsBet's move from vendor lock-in to open, AI-driven observability

Company: PointsBet
Industry: Media & Entertainment

PointsBet operates a wagering platform across Australia and Ontario, Canada – sportsbook, racing, and online casino products in an industry where every second of downtime carries real revenue and reputational risk. Staff Site Reliability Engineer Adam Dagdelen’s team rebuilt PointsBet’s observability practice around Grafana Cloud and OpenTelemetry to break free of vendor lock-in, then layered in Grafana Assistant and Assistant Investigations to cut the time it takes to find a root cause during an incident.

The challenge

When an alert fired, PointsBet’s on-call engineers faced a wall of disconnected tools before they’d even started troubleshooting: Jira for incident management, PagerDuty for paging, Application Insights for application telemetry, Datadog for infrastructure, and a separate Azure Managed Grafana instance for dashboards.

Underneath the sprawl sat a harder problem: real vendor lock-in. PointsBet had built the Application Insights SDK into every .NET service it ran, so any future change in tooling meant re-instrumenting the entire estate from scratch. The team also had genuine blind spots – no real user monitoring, no continuous profiling, no SLOs – and no way to see which team, service, or signal was actually driving the observability bill.

The solution

PointsBet standardized on OpenTelemetry before picking a vendor, decoupling how it instrumented code from where that telemetry landed. Grafana Cloud won out for being open by default (OpenTelemetry- and Prometheus-native and Adaptive Telemetry made cost transparent_. The team migrated via dual export to both Application Insights and Grafana Cloud at once, using an OpenTelemetry defaults package, consistent deploy-time labeling, and a Collector scaled through GitOps ownership between infrastructure and delivery teams.

Grafana Assistant also turns plain-language questions into PromQL, removing one of the biggest hidden costs of any migration: relearning an entire query language. Assistant Investigations now runs automatically against P1 alerts, building a root-cause hypothesis from application telemetry, infrastructure telemetry, and recent deploys before an engineer opens a dashboard.

The impact

PointsBet shifted from firefighting across four disconnected tools to running one platform that can answer what broke, why, and what it cost.

  • Estimated 35%+ reduction in observability spend, by using Adaptive Telemetry
  • Every delivery team migrated to new tooling without ever losing its old safety net mid-incident
  • Grafana Assistant turns plain language into PromQL, removing the query-language learning curve for delivery teams
  • Assistant Investigations run first-pass root cause analysis on P1 alerts, forming a hypothesis before a human opens a dashboard
  • Developers now instrument what they’re unsure about, check usage after a week, and drop what isn’t earning its keep
  • Full telemetry and cost attribution by team, service, and signal, replacing what Dagdelen called “obscure costing”
  • One platform, one bill, and one place to answer what broke, why, and what it cost

“What we got was a mindset shift with the feature of Adaptive Telemetry. Being able to provide us that transparent feedback…and either take those cost savings or reinvest them in areas that we do need.”\

Adam Dagdelen, Staff Site Reliability Engineer, PointsBet

Looking ahead

PointsBet is extending the same paved road to more signals: real user monitoring, database observability, Grafana IRM in place of PagerDuty, and business metrics and SLOs. Dagdelen expects triage itself to keep changing shape, from reading a dashboard to having a conversation with an agent, with dashboards remaining in place as evidence rather than the primary tool.

More great videos and webinars