Documentation for automated readers
A curated documentation index is available at: https://grafana.com/llms.txt
A complete documentation index is available at: https://grafana.com/llms-full.txt
These indexes can help with page discovery before fetching individual documents.
This page is also available in Markdown, which may be easier for automated readers and AI tools to parse than HTML. The Markdown version is available at https://grafana.com/docs/grafana-cloud/observe-and-act/agent-observability/guides/cost-optimization.md, or by sending Accept: text/markdown to https://grafana.com/docs/grafana-cloud/observe-and-act/agent-observability/guides/cost-optimization/. For broader documentation discovery, the curated index is available at https://grafana.com/llms.txt and the complete index is available at https://grafana.com/llms-full.txt.
Optimize cost and performance
Agent Observability captures detailed token usage, cost, and cache data for every generation. Use this data to find optimization opportunities and reduce your AI spending.
Analyze token usage
Open Analytics and review the tokens and cost dashboard. Look for:
- High-token agents: agents that consistently use more tokens than expected may have verbose prompts or unnecessary context.
- Model selection: some calls may use expensive models where a cheaper alternative would produce equivalent results.
- Reasoning tokens: if reasoning tokens are high, check whether thinking mode is enabled unnecessarily.
Improve cache efficiency
Prompt caching reduces costs by reusing previously processed prompt prefixes. Agent Observability tracks cache read and cache write tokens for each generation.
To improve cache rates:
- Keep the system prompt stable. Changing the system prompt frequently invalidates the cache.
- Place static content (system prompt, tool definitions) at the beginning of the message sequence.
- Minimize dynamic content that changes between requests.
Monitor cache efficiency in the analytics dashboard. A healthy cache read ratio means you’re reusing processed prompt prefixes effectively.
Note
If you query the token metric yourself rather than reading the dashboard, check the
gen_ai_token_semanticslabel first. Where it readsinclusive, theinputseries already contains the cache tokens, so addingcache_readto it counts them twice, and onlyinput - cache_read - cache_write - cache_creationpays the full prompt rate.Where the label is absent, check the provider before you add the buckets up. An absent label means nothing was declared, not that the buckets are separate: OpenAI and Gemini report an inclusive
inputwhether or not the label is there. The dashboard handles both; a hand-written query has to choose.
Reduce tool call overhead
The tools analytics page shows tool call frequency and duration. Look for:
- Unused tools: tools that are defined but never called add prompt tokens without value. Consider removing them.
- Slow tools: tools with high execution time may benefit from timeouts or caching.
- Excessive tool calls: agents that call too many tools per turn may need prompt adjustments.
Attribute cost to a team, project, or developer
For SDK instrumentation, set client tags to split cost by team, project, or environment. Use the AGENTO11Y_TAGS environment variable or the tags field in the client configuration:
export AGENTO11Y_TAGS=team=ai,env=productionClient tags reach both the generation export and the OpenTelemetry metrics, so each key becomes a Prometheus label named agento11y_tag_<KEY>. Turn on Break down in the tokens and estimated cost charts and select that label to compare spend across tag values.
Tags that you pass on a single generation stay on that generation and never become metric labels, so they can’t drive this breakdown. For the behavior of each mechanism, refer to Client tags.
Coding-agent integrations can resolve user, repository, and branch labels automatically. For setup instructions, refer to Attribute cost to a user, repository, or branch.
Compare agent versions
When you change prompts or tools, use the agent catalog to compare the new version’s cost and performance against the previous version. Check whether the change improved quality without significantly increasing cost, or vice versa.
Next steps
Was this page helpful?
Related resources from Grafana Labs


