Grafana Cloud

Agent Observability pricing

Grafana Agent Observability bills on usage across three separate dimensions: the generations your agents produce, the tokens consumed by LLM-based evaluations and guards, and the standard telemetry your SDKs emit.

For current rates and included usage, refer to the pricing page.

How Agent Observability is billed

Three meters run independently:

  • Generations for the AI model calls your monitored agents and applications make.
  • Evaluation and guard tokens for evaluators that call an LLM.
  • Telemetry for the OpenTelemetry metrics and traces the SDKs emit, at standard Grafana Cloud rates.

Generation usage is billed separately from token usage. OpenTelemetry metrics and traces aren’t included in generation pricing.

Definitions

Generation: One call to an AI model made by your monitored agent or application. Each time your application calls an LLM provider and the SDK exports that interaction, it’s counted as a generation. One agent task can produce several generations if it makes several model calls.

Evaluation: An automatic scoring of your agent’s production traffic against criteria you define, using one or more evaluators. For details, refer to Set up online evaluation.

Guard: A synchronous check that runs before an LLM call to block harmful prompts, redact sensitive data, or filter dangerous tool calls. For details, refer to Set up guards.

LLM-based evaluator: An evaluation or guard that calls an LLM to judge or guard an interaction, such as an LLM judge evaluation.

Deterministic evaluator: An evaluation or guard that doesn’t call an LLM, such as regular expression, JSON schema, heuristic, redact, and tool filter rules.

System initiated token pool: The organization wide token allowance that LLM-based evaluation and guard usage draws from. The pool is shared with other Grafana Cloud AI features. System initiated Grafana Assistant usage and automatically triggered Assistant Investigations draw from the same allowance.

How usage is calculated

Generation billing covers ingesting and storing generation telemetry. It doesn’t cover the cost of running your model, which your LLM provider bills separately.

Free and Pro plans include 30,000 generations each billing month. On the Free plan this is a hard limit, and you can’t purchase additional generations. On the Pro plan, usage above the included amount is charged at the rate on the pricing page. Contracted plans use a custom generation rate based on the annual commitment.

LLM judge evaluations and LLM-based guards consume tokens because Grafana runs an LLM to perform them, and that usage counts toward the system initiated token pool for your organization. Grafana provides the LLM used for evaluations and guards; you can’t supply your own. Deterministic evaluators and guards don’t consume tokens unless they call an LLM-backed evaluator.

Free and Pro plans include 25 million system initiated tokens per organization each billing month. Contracted accounts include a larger allowance; refer to your contract for the terms that apply to your organization. Only token usage above the allowance is charged. Allowances reset each billing month.

Telemetry follows standard Grafana Cloud pricing and limits for metrics and traces. Agent Observability doesn’t add a separate usage cap for it.

Special considerations

When billing starts

Pricing and metering for generations and for evaluation and guard token usage start on October 1, 2026. Usage before that date isn’t billed retroactively.

If you used LLM-based evaluations and guards before October 1, the tokens they consume weren’t previously metered. From October 1, they draw from your organization’s system initiated token pool, and usage above the allowance is charged.

Telemetry billing follows standard Grafana Cloud pricing and is already active.

Contracted accounts

Contracted plans use a custom generation rate based on the annual commitment, and volume discounts can apply to generation usage. Discounts don’t apply to token usage for evaluations and guards. Refer to your contract or contact your Grafana Labs account team for the terms that apply to your organization.

View your usage

Agent Observability analytics views show generation volume, latency, errors, token usage, and cost patterns, which helps you identify expensive workflows.

For usage alongside the rest of your Grafana Cloud spend, go to Cost Management and Billing > Usage in your stack.