Instrument and observe production-ready agentic apps with Grafana Cloud, AWS, and Claude
GenAI applications do far more than chat: They reason, call tools, retrieve knowledge, and take actions inside business-critical workflows. This all creates new operational challenges in understanding whether your agents are behaving correctly, safely, efficiently, and cost effectively in production.
In this workshop, you'll see how to build an agent using Amazon Bedrock and Amazon Bedrock AgentCore, which are AWS services designed to help teams build, deploy, and operate agents securely at scale. You will also learn from Anthropic why to reach for the Claude Agent SDK when you're building agents, where it fits across the agent lifecycle from prototype to production, and how the instrumentation it exposes connects directly to the observability and ROI questions the rest of the workshop takes on. From there, you’ll use Grafana Cloud's Agent Observability to monitor a real multi-agent AI application and add the observability, evaluation, alerting, and guardrail layers needed to run the agent with confidence.
Through a series of guided labs, you’ll inspect conversations, prompts, tool calls, and agent execution to understand why an agent behaved the way it did. You'll create evaluators to automatically detect prompt injection, toxicity, and bias, configure evaluation rules, create alerts when AI behavior degrades, and finally deploy guards that block malicious requests before they ever reach the model.
By the end of the session, you’ll understand how to move beyond traditional application monitoring and build a complete operational layer for AI systems. You’ll leave with practical patterns for optimizing cost and performance, debugging agent behavior, protecting users, and continuously improving AI applications in production.
Whether you're building copilots, autonomous agents, or customer-facing AI applications, you'll leave with practical techniques you can immediately use to observe, evaluate, govern, and improve AI systems running in production.
Speakers

Ryan Davies
Member of Technical Staff, Anthropic

Saurabh Shanbhag
Senior Partner Solutions Architect, AWS

Devin Cheevers
Director of Product, Grafana Labs

Patrick Easters
Staff Field Engineer, Grafana Labs

Matt Simonsen
Senior Solutions Engineer, Grafana Labs
What you'll accomplish
See how to build or connect an AI agent using Amazon Bedrock and Amazon Bedrock AgentCore
Understand where the Claude Agent SDK fits in the agent lifecycle and how it ties into observability and ROI, presented by Anthropic
Learn how to instrument AI agents and capture conversations, prompts, tool calls, and traces
Debug agent behavior using end-to-end conversation and execution visibility
Build automated evaluations for quality, safety, and prompt injection detection
Configure sampling strategies, alerts, and quality monitoring for production workloads
Deploy guards that block attacks and redact sensitive information before LLM execution
Apply proven patterns for operating AI applications in production with Grafana Cloud