Documentation for automated readers
A curated documentation index is available at: https://grafana.com/llms.txt
A complete documentation index is available at: https://grafana.com/llms-full.txt
These indexes can help with page discovery before fetching individual documents.
This page is also available in Markdown, which may be easier for automated readers and AI tools to parse than HTML. The Markdown version is available at https://grafana.com/docs/grafana-cloud/observe-and-act/agent-observability/configure/model-prices.md, or by sending Accept: text/markdown to https://grafana.com/docs/grafana-cloud/observe-and-act/agent-observability/configure/model-prices/. For broader documentation discovery, the curated index is available at https://grafana.com/llms.txt and the complete index is available at https://grafana.com/llms-full.txt.
Configure your model prices
Agent Observability estimates the cost of each generation from its token usage and a price for the model that served it. By default, it uses public list prices from the model card catalog. If you pay a different price, set your own price for that model. Examples are a negotiated rate, a committed-use discount, a reseller agreement, or a model the catalog doesn’t carry. Agent Observability uses your price for the generations it receives from then on.
How your prices apply
Your price for a provider and model replaces the catalog price for that pair:
- It replaces the catalog price completely. Agent Observability doesn’t fill in missing token types from the catalog. Set a price for every token type your contract covers. A token type without a price isn’t charged.
- It applies from the moment you set it. Generations that Agent Observability already received keep the cost they were given.
- A change adds a new price. To change a price, set a new one. The earlier price still covers the period when it applied.
- Prices are in USD per million tokens, the way providers publish them. You can also set a per-call fee, and long-context prices. When a call’s input, including cached tokens, is above the long-context threshold, every token in that call is charged at the long-context prices. A token type without a long-context price keeps its normal price.
Where you see your prices
Cost figures in Analytics, Agents, Conversations, and Needs attention use your prices. In a conversation, each call is priced at the price in force when it ran. The cost tooltip says whether a figure comes from your prices or from public list prices.
Some figures still use public list prices:
- Filtering or breaking down cost by a label other than provider, model, or agent calculates cost from token counts at list prices. The page shows a Cost shown at list prices notice when this happens.
- LLM judge cost estimates use catalog prices, because Agent Observability runs judges with its own provider credentials.
- Experiment costs use catalog prices.
Who can see your prices
Your prices aren’t secret from people who can view Agent Observability data. Anyone with the grafana-agento11y-app.data:read permission sees costs calculated with your prices, and can work out a price by dividing a cost by its tokens. The conversation view also reads your prices to price each call. Only users with the grafana-agento11y-app.pricing:write permission can set, list, or delete them.
Before you begin
Before you set your prices, make sure you have the following:
- gcx 1.4.0 or later, logged in to your Grafana Cloud stack.
- The
grafana-agento11y-app.pricing:writepermission. The Agent Observability Admin role includes it, and Grafana Admins have that role by default. For more information, refer to Configure RBAC.
Set a price for a model
Use the provider and model names exactly as Agent Observability shows them, for example in the provider and model filters on Analytics.
gcx agento11y model-rates create --provider <PROVIDER> --model <MODEL> \
--price-input <INPUT_PRICE> --price-output <OUTPUT_PRICE>Replace the following:
- PROVIDER: the provider, for example
openai. - MODEL: the model, for example
gpt-5. - INPUT_PRICE: what you pay in USD per million input tokens.
- OUTPUT_PRICE: what you pay in USD per million output tokens.
If your contract has other charges, add --price-cache-read, --price-cache-write, or --price-request. For a long-context tier, add --long-context-threshold and the --long-context-price-* flags.
List your prices
To review the prices you set and when each took effect, run:
gcx agento11y model-rates listDelete a price
To delete a price, find the time it took effect in the list output, then run the following command. gcx asks you to confirm before it deletes the price.
gcx agento11y model-rates delete --provider <PROVIDER> --model <MODEL> \
--effective-from <EFFECTIVE_FROM>Replace EFFECTIVE_FROM with the time from the list output.
After you delete the price in force, Agent Observability uses your previous price for that model, or the catalog price if there’s none. Dashboards keep the costs they already calculated. A conversation prices each call from your remaining prices, so the calls the deleted price covered show the price that applies without it.
Next steps
Was this page helpful?
Related resources from Grafana Labs


