Grafana Cloud

Configure online evaluation

Online evaluation continuously scores live generation traffic. Evaluators define the scoring logic, and evaluations define which traffic to score.

Create evaluators

Use Agent Observability or the evaluation API to create evaluators. Four evaluator types are available:

LLM judge

An LLM judge uses an LLM to score generations based on criteria you define in a prompt template.

Key settings:

  • provider and model: the LLM to use for judging.
  • system_prompt and user_prompt: prompt templates with variables.
  • max_tokens and temperature: generation controls.
  • reasoning_effort: optional none, minimal, low, medium, high, or xhigh. none disables reasoning; omission leaves it unspecified. Provider APIs validate model support.
  • timeout_ms: deadline for each judge attempt, in milliseconds. The default is 30000.

Use OpenRouter

Enable the OpenRouter judge provider and provide an API key to run LLM judges through OpenRouter:

Bash
export SIGIL_EVAL_OPENROUTER_ENABLED=true
export SIGIL_EVAL_OPENROUTER_API_KEY="<OPENROUTER_API_KEY>"

You can also use the conventional OPENROUTER_API_KEY variable. The provider uses https://openrouter.ai/api/v1 by default and appears as OpenRouter in the evaluator provider list. Its model list includes text models that support structured output, which LLM judges require.

Use SIGIL_EVAL_OPENROUTER_ALLOWED_MODELS to restrict the available model slugs. Separate multiple values with commas, for example openai/gpt-5,anthropic/claude-sonnet-4.5.

To use OpenRouter when an evaluator does not override its judge target, set SIGIL_EVAL_DEFAULT_JUDGE_MODEL to an OpenRouter target such as openrouter/anthropic/claude-sonnet-4.5.

For optional OpenRouter app attribution, set SIGIL_EVAL_OPENROUTER_HTTP_REFERER to your app URL. You can also set SIGIL_EVAL_OPENROUTER_APP_TITLE to choose its display name. Set SIGIL_EVAL_OPENROUTER_BASE_URL only when you route requests through a compatible proxy.

Use TypeSafe Jev for typed decisions

TypeSafe Jev is a decision model rather than a text generator. Use it when your evaluator needs one bounded answer that your rule can act on directly.

Enable the provider in every querier and evaluation worker:

text
SIGIL_EVAL_TYPESAFE_ENABLED=true
SIGIL_EVAL_TYPESAFE_API_KEY=<TYPESAFE_API_KEY>

Both settings are required, and an administrator must enable the sigil.eval-typesafe feature flag for the tenant. If deployment configuration or the tenant flag is missing, Agent Observability omits the TypeSafe Jev provider, its models, and the built-in Jev trace-triage template from the UI. Turning the tenant flag off also stops existing Jev evaluation and guard calls without disabling other evaluators.

Select TypeSafe Jev as the provider and jev-latest as the model. Configure one of these output contracts:

  • Boolean, which Agent Observability maps to a Jev Noul question. A probability of 0.5 or greater becomes true.
  • String with 2-255 allowed values, which maps to a Jev Choice question.
  • Number with integer minimum and maximum values spanning 2-10 levels, which maps to a Jev Score question. In Output description, provide one concrete label for every value, separated by semicolons or new lines, for example 1: harmless; 2: needs review; 3: urgent. Agent Observability uses the values to order the rubric and sends only the descriptive labels to Jev.

Jev returns a typed value and probabilities instead of generated rationale. Agent Observability creates the displayed explanation from that answer, shows the probability distribution and concrete Jev model version in test and score details, and still derives pass or fail from your configured pass condition.

Agent Observability sends one question per Jev evaluation, so Jev’s 32k-token limit for state plus the longest question is the effective context window. The separate 64k limit applies to state plus all questions and does not increase this one-question request. Because TypeSafe does not publish a preflight tokenizer, Agent Observability applies a conservative request budget before the call. If a rendered trace is too large, it preserves the beginning and end, inserts an explicit omission marker, and records the original and submitted sizes. Result details and copied summaries show Input compacted whenever this happens. Shorten the prompt or evaluate a narrower scope if the omitted middle contains evidence your decision needs.

Max-token and temperature settings do not apply to System One, so Agent Observability hides those controls when you select TypeSafe Jev.

Use jev-latest while you explore. Pin a versioned model ID after you calibrate confidence thresholds or routing policy because the alias advances when TypeSafe releases a new stable model. If you set SIGIL_EVAL_TYPESAFE_ALLOWED_MODELS, add the pinned ID to that list. Jev performs best on English; validate it against representative examples before you rely on a non-English workload.

The rendered evaluator content is sent to TypeSafe as Jev state. Apply the same privacy and data-handling review that you use for any external judge provider before you send production traces.

JSON schema

Validates that the agent response matches a JSON schema. Returns true or false.

Regex

Checks the agent response against one or more regex patterns. Use reject: true to invert the match.

Heuristic

Applies a rule tree with AND/OR logic. Supported checks: not_empty, contains, not_contains, min_length, max_length.

Choose the content to check

JSON schema, regex, and heuristic evaluators evaluate the agent response by default. You can choose another field instead:

TargetDescription
responseAssistant response text (default).
inputUser input text.
system_promptThe system prompt.

Under Content to check, select the Generation field for a JSON Schema or Regex evaluator. Heuristic evaluators select a field for each condition. In evaluator config JSON, continue to use the target field shown above.

Your application can send the system prompt in either of two ways, depending on the SDK and framework you use: in the generation’s own system prompt field, or as a message with the system role inside the chat history. The system_prompt target and the {{system_prompt}} template variable accept both. They read the system prompt field first, and the text of the system messages when that field is empty.

Use this for lightweight detection of injected content in generation input without requiring an LLM judge.

Configure pass verdicts

Boolean outputs from heuristic, regex, JSON schema, and LLM judge evaluators record a pass or fail verdict only when you explicitly configure a pass_value on the output key. When pass_value is omitted, the score is recorded without a verdict.

Use template variables

LLM judge prompts support three explicit variable scopes. The prompt editor puts the scope that matches the rule trigger first:

ScopeExample variablesContent
Generation{{generation.latest_user_message}}, {{generation.agent_response}}, {{generation.agent_thinking}}, {{generation.agent_sequence}}, {{generation.system_prompt}}, {{generation.tool_calls}}, {{generation.tool_results}}, {{generation.tools}}, {{generation.stop_reason}}, {{generation.call_error}}Only the model call selected by the trigger.
Turn{{turn.transcript}}, {{turn.latest_user_message}}, {{turn.agent_response}}, {{turn.agent_sequence}}, {{turn.tool_calls}}, {{turn.tool_results}}, {{turn.tools}}The user request, intermediate tool steps, and final response. Available for the User-visible response trigger.
Conversation{{conversation.transcript}}, {{conversation.latest_user_message}}, {{conversation.agent_response}}, {{conversation.agent_sequence}}, {{conversation.tool_calls}}, {{conversation.tool_results}}, {{conversation.tools}}All captured turns. Available for user-visible-response and conversation-idle evaluation.

Use {{turn.tool_calls}} instead of {{generation.tool_calls}} when evaluating a final response: the final generation usually contains the response, while the tool calls occurred in earlier generations in that turn.

Existing unscoped variables such as {{agent_response}} and {{tool_calls}}, underscore turn variables such as {{turn_tool_calls}}, and assistant_* aliases remain supported. Prefer scoped variables in new prompts because their meaning does not change with the trigger.

Configure evaluation traffic

An evaluation connects one or more evaluators to generation or conversation traffic. Configure:

  • Trigger: when evaluation runs:
    • User-visible response (user_visible_turn): when the agent produces final text for a user request. Use turn.* variables to include its intermediate tool steps.
    • Every agent generation (all_assistant_generations): after any agent model call.
    • Tool call requested (tool_call_steps): when a generation calls tools.
    • Generation failed (errored_generations): when a generation ends with an error.
    • Conversation idle (conversation): after the configured idle window.
  • Matching conditions: additional criteria to narrow the selection (agent name, model, mode, tags, errors, and conversation.tool_used for targeting conversations where a tool was called). For that tool filter, Any matches when one pattern hits a tool used in the conversation, and All matches when every pattern does.
  • Sampling rate: percentage of matching generations to evaluate.
  • Evaluators: the evaluators to run.

The legacy turn selector remains accepted for existing API clients and stored rules, but it is not a separate authoring choice because it fires at the same event as user_visible_turn.

Create an evaluation

Open Evaluations → New evaluation. Choose how to start:

  • Build with Assistant — describe the quality you care about; Grafana Assistant creates the evaluator and evaluation.
  • See evaluator library — open the evaluator library to review what your workspace already has.
  • Create new evaluation — fill the rule and judge by hand in the shared builder.

Scoring starts on new traffic after you create the evaluation; historical generations are not backfilled.

Use score conditions to trigger actions

An evaluation action adds matching conversations to a saved collection. Use it to send failures, examples that need review, or scores that meet a threshold to a review queue.

Choose one of these triggers when you add an action:

  • When passes: runs when every evaluator produces a passing verdict.
  • When fails: runs when every evaluator produces a failing verdict.
  • When output value: runs when one evaluator output matches a value that you choose.

For an output-value trigger, select the evaluator and output key, then choose a comparison and value. Numeric outputs support equality and greater-than or less-than comparisons. String outputs support equality and contains. Boolean outputs support equality. Use this trigger for a specific score, such as adding conversations whose toxicity score is greater than a threshold, even when the evaluator does not produce a pass/fail verdict.

Preview an evaluation before you save it

When you create or configure an evaluation, the rule builder shows a Playground beside the form so you can check your rule before you save it. The playground has two tabs.

What gets evaluated

This tab is a volume simulation: it applies your current criteria and sample rate to recent traffic (for example, the last 6 hours). It is not a forecast of future volume. On high-volume windows, the preview is based on a sample of that traffic. In that case, projected generation volume is labelled as an estimate, and actual volume may differ if the traffic mix changed. Matching traffic is grouped into conversations. Open a conversation to see each generation and whether it would have been evaluated, matched but sampled out, or did not match. Select an agent, model, or tag on a generation to add it to your match criteria. When the rule includes LLM judges, the playground also shows an approximate projected cost per day based on sample-rendered judge prompts, model pricing, and projected evaluation volume; real spend varies with transcript length and judge completions.

When the selector is conversation, the playground shows conversation-level simulation only: each card reports whether the conversation would be evaluated (one evaluation per matching conversation after the idle window) or was sampled out, and links out to the full conversation instead of a per-generation drill-down. For a capped scan, the playground reports conversations selected for evaluation after sampling but does not project hourly or daily unique-conversation volume.

Use this tab to tune your selector, match filters, and sampling rate without waiting for scores to arrive asynchronously.

How it’s evaluated

This tab shows what the LLM judge receives for a selected generation — or a selected conversation when the rule selector is conversation. Choose which attached evaluator to test, pick a Source (Matched traffic from the volume preview by default, or Saved conversations), then pick a generation or conversation. Inspect the rendered judge prompt with template variables filled in and highlighted. Run evaluator evaluates the selection with that evaluator’s work-in-progress config so you can check pass/fail and rationale before saving.

Gate LLM judges with deterministic evaluators

When a rule contains more than one evaluator, choose Sequential gate to run them in the displayed order. Each evaluator must pass before the next evaluator runs. Use a deterministic evaluator first to filter out generations that don’t need an LLM judge.

Configure a pass condition for every output of each non-final evaluator. If a gate doesn’t pass, later evaluators don’t run. If a gate has an execution error, the evaluation retries according to its normal retry policy.

Choose Parallel when you want every evaluator to score the same matching traffic independently.

Evaluate conversations

Conversation-scope rules evaluate the complete conversation instead of one generation. Set the selector to conversation and configure min_idle_seconds so Agent Observability waits for the interaction to settle before it scores the transcript.

Conversation-scope LLM judge prompts can use {{conversation_transcript}} to include the full interaction. Generation-specific variables, such as {{agent_response}}, aren’t populated for conversation-scope evaluations.

Next steps