%3Aquality(90)%2F&w=3840&q=75)
Tempo 3.1 release: new features for Kafka, TraceQL metrics updates, trace redaction, and more
Building on the major release of Tempo 3.0, Tempo 3.1 is here, delivering community-contributed Kafka client improvements, query-based trace redaction, sampling-aware TraceQL metrics, and more.
Together, the updates in 3.1 make it easier to operate Tempo, get accurate insights from your trace data, and investigate issues more efficiently. You can continue reading and check out the video below to learn more about the latest features. The Tempo 3.1 release notes and changelog provide more in-depth details and include all changes in this release.

Kafka client improvements from the community
Tempo 3.1 adds several community-contributed improvements to Kafka-based ingestion, with new options to secure connections, reduce data transfer costs, and support more Kafka-compatible backends.
TLS and more SASL mechanisms
Tempo 3.0 moved microservices-mode ingestion onto Kafka, and the client it shipped with could authenticate, but only with SASL PLAIN and only over an unencrypted connection. It had no TLS settings, so a broker that requires TLS would refuse the connection. Tempo 3.1 adds TLS and four more mechanisms: SCRAM-SHA-256, SCRAM-SHA-512, OAUTHBEARER, and AWS_MSK_IAM.
ingest:
kafka:
address: kafka.example.com:9093
topic: tempo-traces
sasl_mechanism: SCRAM-SHA-512
sasl_username: ${KAFKA_USERNAME}
sasl_password: ${KAFKA_PASSWORD}
tls_enabled: true
tls_ca_path: /etc/tempo/kafka-ca.pem
tls_cert_path: /etc/tempo/kafka-client.crt # optional, for mTLS
tls_key_path: /etc/tempo/kafka-client.key # optional, for mTLSPass -config.expand-env=true to expand the environment variables.
Thanks to @heytrav for this contribution. To learn more, see the PR and the Configure authentication and TLS docs.
Rack-aware fetching
Tempo always fetched from the partition leader, wherever it happened to live. When the leader sits in another availability zone, every trace you read crosses a zone boundary, and your cloud provider bills you for the transfer.
The new client_rack option helps reduce those costs by enabling rack-aware fetching (KIP-392), so consumers can read from a replica in their own zone:
ingest:
kafka:
client_rack: us-east-1aThis is a read-side setting, so it applies to block-builders, live-stores, and metrics-generators, not to the distributor writing records into Kafka. Your brokers need rack IDs configured for it to take effect.
Thanks to @KyriosGN0 for this contribution. To learn more, see the PR and the ingest configuration docs.
Configurable producer compression
Tempo's distributor compresses every batch it writes to Kafka, and until 3.1, the codec wasn't configurable. That's fine unless your backend accepts only one codec: for example, Azure Event Hubs supports the Kafka protocol but accepts only gzip, which ruled it out as a Tempo backend.
The new producer_compression option lets you pick from none, gzip, snappy, lz4, or zstd:
ingest:
kafka:
producer_compression: gzipTogether with the TLS support above, this option makes Event Hubs a usable Kafka backend for Tempo. If you leave producer_compression unset, the Kafka client uses snappy by default. If your backend doesn’t support snappy, set producer_compression to a supported algorithm, or set it to none to disable compression.
Thanks to @fleighton for this contribution. To learn more, see the PR and the ingest configuration docs.
Protect sensitive data: redact traces with a TraceQL query
In Tempo 3.0, we added trace redaction, allowing you to permanently remove sensitive data, such as an email address, an auth token, or an account number that ended up in a span attribute, from your trace data without waiting for retention to expire. But tempo-cli redact only accepted trace IDs, so you had to enumerate every affected trace. Sensitive data usually lands in whatever traffic hits a given code path, which can be more traces than you can practically list. Any trace you miss is data still sitting in object storage, and potentially queryable.
Now, with Tempo 3.1, you can redact using a TraceQL query instead:
tempo-cli redact \
--tenant=<TENANT_ID> \
--query '{span.attribute = "<leaked PII>"}' \
--dry-run \
<SCHEDULER_ADDRESS>:<GRPC_PORT>Start with --dry-run, as shown above. Tempo evaluates the query and counts what it would remove without modifying any blocks. The command prints the batch ID and the number of jobs created. Job counts and how many traces were matched or removed are per-tenant metrics that can be seen on the Redaction row of the Backend Work dashboard, which ships with Tempo's monitoring mixin. Once the count matches what you expect, run the same command without --dry-run to rewrite the blocks. This cannot be undone.
Because a wrong query permanently deletes data you didn't mean to remove, we're deliberately keeping the accepted syntax small for now: a single spanset filter with equality comparisons against resource.* and span.* attributes, combined using && and ||. Anything outside that is rejected when you submit the job. We are continuing to improve this functionality.
If you know the time range where data needs to be redacted, you can use it to run the job faster and more efficiently. This is important in high-volume installs because Tempo holds compaction off for a tenant while a redaction is applying. Use --start and --end, which accept now, a relative offset like now-7d, or an RFC3339 timestamp. Compaction catches up between runs.
Note: Only use --start and --end once every scheduler and worker in your cell is running at least Tempo 3.1. An older worker ignores the window and removes every query match in each block it is given, regardless of timestamp, with no error and no way to recover the data.
To learn more, see the query selector PR, the time range PR, and the Redact traces docs for the full query syntax and constraints.
Query trace metrics with greater accuracy, flexibility, and speed: updates to TraceQL metrics
In Tempo 3.0, TraceQL metrics became generally available, letting you query ad-hoc metrics directly from trace data. This makes it easier to answer questions about performance, error rates, and service behavior across distributed systems.
With the 3.1 release, we’re rolling out several updates that make TraceQL metrics more flexible and efficient, and more accurate when working with sampled trace data.
Sampling-aware metrics queries
Let’s say you sample 50% of your traces before they reach Tempo. If you ran a rate() query against a service, you would get half the traffic it actually served, because Tempo counts only the spans that made it through sampling.
This is especially useful with Adaptive Traces and other tail sampling setups, where sampling rates can vary across services. With sampling-aware queries, TraceQL metrics can account for those differences and more accurately reflect the underlying traffic.
With Tempo 3.1, you can correct for sampling at query time with the new experimental with(extrapolate=true) hint:
{ } | rate() with(extrapolate=true)%3Aquality(90)%2F&w=3840&q=75)
The same query on 50%-sampled traces, without the hint on the left and with it on the right. The two graphs have the same shape, but the y-axis tops out at 4 on the left and 8 on the right.
Samplers that implement OpenTelemetry's probability sampling specification, which is still in development and not supported in every language yet, stamp the rate they sampled at onto the span, in a field called tracestate. The hint reads that rate back and scales the span accordingly. At 50% sampling, each stored span counts as two, so a query that matches 1,000 spans reports 2,000. Spans that arrive with no sampling rate recorded count as one, which means partial adoption is safe: if only two of your services sample, the other services' numbers don't change.
Nothing new is written to your blocks. tracestate is already stored with every span, and queries that don't use the hint don't read it. If you already correct for sampling in the metrics-generator by setting enable_tracestate_span_multiplier, which we added in Tempo 3.0, the query path reads the rate the same way, so an ad-hoc query and your pre-aggregated tempo_spanmetrics_* series agree.
Extrapolation applies to rate, count_over_time, sum_over_time, avg_over_time, histogram_over_time, quantile_over_time, and compare. It doesn't apply to min_over_time or max_over_time, because sampling doesn't change the smallest or largest value you actually observed.
Note: This hint is experimental and requires vParquet4 blocks or later.
To learn more, see the PR and the TraceQL metrics functions docs.
Combine metrics queries with arithmetic
Many of the metrics you want to derive from traces require combining two metrics queries rather than running just one. A common example is error rate. Traditionally, TraceQL could measure errors and total traffic separately but couldn't divide one by the other, so you’d have to run two queries and divide them with a math expression in Grafana.
With Tempo 3.1, you can write the division in TraceQL itself. The operators +, -, *, and / work between two metrics queries, so an error rate is one query:
({ status = error } | rate() by (resource.service.name)) / ({ } | rate() by (resource.service.name))%3Aquality(90)%2F&w=3840&q=75)
The same division with by (resource.service.name) on both sides, so each service gets its own error rate. By contrast, with the query ({status=error} | rate() by (resource.service.name)) / ({} | rate()), only the numerator is grouped by service, while the denominator is calculated across all data.
Each sub-query needs its own parentheses. A scalar can go on either side of the operator:
# Calculate error percent in range of 0 .. 100
100 * ({status=error} | rate()) / ({ } | rate())
# Calculate p95 in milliseconds
({} | quantile_over_time(duration, 0.95)) * 1000If you group your queries, two series combine only when their label sets match exactly, and a side with no labels is broadcast across every series on the other side.
You can combine with(extrapolate=true) and arithmetic. One hint at the end of the expression applies to both sub-queries:
({ status = error } | rate()) / ({ } | rate()) with(extrapolate=true)
The hint has to go at the end. Putting with(...) inside a sub-query is a syntax error.
To learn more, see the arithmetic PR, the scalars PR, and the arithmetic expressions docs, and the July Community Call.
Faster metrics queries by default
Metrics queries are also faster: we've seen simple ones run close to twice as fast. The span-only fetch path we introduced as experimental in Tempo 3.0 is now on by default for vParquet5 blocks, which is what Tempo 3.1 writes.
Queries that don't need the full trace structure now process individual spans instead of whole traces, which reduces latency and memory use. In one warmed-up test, { } | rate() with span-only fetch enabled completed in slightly over half the time it took with span-only fetch disabled. You don't have to change your queries to use it, and you can opt out per tenant or per query if needed.
To learn more, see the PR and the faster read path docs.
Enhancements to the metrics-generator
Metrics-generator is an optional Tempo component that derives metrics from ingested spans, giving you both RED (Rate/Error/Duration) metrics and service graphs, which show the relationships between your services. Here we cover a few of the improvements to metrics-generator included in Tempo 3.1, but please refer to the changelog and release notes for the complete list.
See why your service map is missing edges
Tempo builds a service graph edge by pairing the client and server spans for a connection. When one side never arrives, the edge expires and leaves a hole in your map, and you can't tell a connection that doesn't exist from one Tempo couldn't match.
In Tempo 3.1, we added an unmatched_span_kind label to the expired edges counter, so you can see which side went missing. A lot of expired edges labeled SPAN_KIND_SERVER, for example, point to server-to-server instrumentation that will never resolve into a node.
To learn more, see the PR, the service graphs docs, and the expired edges troubleshooting docs.
Recognize db.system.name from updated OTel semantic conventions
OpenTelemetry semantic conventions v1.30.0 renamed the db.system attribute to db.system.name. Instrumentation that emits only the new name stopped being recognized as a database call, so those nodes disappeared from your service map after an SDK upgrade, with no error to tell you.
Tempo 3.1 accepts both attributes for identifying database requests and for naming virtual nodes. When a span carries both, db.system takes precedence, so upgrading your SDKs won't rename the nodes your dashboards and alerts already point at.
Thanks to @iamrajiv for this contribution. To learn more, see the PR and the database name attributes docs.
Cut per-span metrics-generator costs
The generator processes every span you ingest, so per-span work sets the floor on what it costs to run. Building label sets accounted for a significant share of that work.
The span-metrics and service-graphs processors now borrow pooled label buffers from the registry instead of allocating new ones each time, and our steady-state service-graphs benchmarks report zero allocations per operation. Metric names, labels, and values are unchanged, so you won't see a difference in your dashboards, just generator CPU and memory you no longer have to provision.
To learn more, see the span metrics PR and the service graph PR.
Make large traces easier to read with span pruning
A single user request can produce many spans. If your application runs 100 nearly identical database queries behind that request, each one becomes its own span, and you end up with 1 root span and 100 database spans that all look the same. A trace like that is harder to read, and the more queries the request makes, the worse it gets.
With Tempo 3.1, the trace-by-ID v2 endpoint collapses each group of similar leaf spans into a single summary span, which carries aggregation.span_count and the minimum, maximum, average, and total duration of the spans it replaced. In the example above, you get back 2 spans instead of 101, in a correspondingly smaller response. Spans only group when their name, kind, status, and parent name match, so an error is never folded in with successes. This happens at read time, so nothing leaves storage.
Ask for it per query:
GET /api/v2/traces/<traceID>?span_pruning=trueSet span_pruning_enabled_by_default to true in your Tempo configuration, and you won’t need to pass span_pruning on each request:
query_frontend:
trace_by_id:
span_pruning_enabled_by_default: trueThat also covers the Grafana UI and the Tempo MCP server's get-trace tool, neither of which sends the parameter yet. An explicit span_pruning in a request still takes precedence, so a caller that needs the complete trace can always ask for it. You can also set that default per tenant.
We also added TraceQL filter support to the trace-by-ID v2 endpoint. It takes a single spanset filter on the q parameter, so you can fetch one trace and get back only the spans that match instead of all of them. Those spans come back on their own by default; add keep_hierarchy=true, and you get each match's ancestor path to the root too.
GET /api/v2/traces/<traceID>?q={ span.http.status_code = 500 }&keep_hierarchy=trueThe gcx CLI is being updated to expose both on gcx traces get: --prune for pruning and --filter for the query.
Note: Span pruning is experimental. The span_pruning parameters, their behavior, and the response format may change in future releases.
To learn more, see the trace pruning PR, the filter PR, the Query V2 API docs for the request parameters, and the query-frontend configuration docs.
Make trace comparisons faster and more efficient for AI
Sometimes you want to compare two traces: an endpoint got slower after a deploy, and you have one trace from before and one from after, or the same request succeeds for one caller and fails for another. Either way, comparing them by hand means scrolling through two traces looking for the span that changed, appeared, or disappeared.
With Tempo 3.1, you can use traces diff to send two trace IDs to the query-frontend and get back what changed, instead of pulling both traces and comparing them yourself. This creates a much smaller context window for an LLM, too, which makes AI-driven investigations faster and cheaper:
POST /api/v2/traces/diff{
"base": { "traceId": "<BASE_TRACE_ID>" },
"compare": { "traceId": "<COMPARE_TRACE_ID>" }
}
By default, you get a span-level patch listing every span that was added, removed, or modified. Partial traces are rejected, since spans missing from an incomplete trace would look like spans that were removed.
You can also ask for a summary instead of the patch. It reports whether trace latency, errors, and span count went up or down, whether the structure changed, and which services were involved, with a millisecond delta for the ones that moved. It tells you which direction each number moved, not whether that's a regression. A trace that got faster might just be failing earlier.
The summary also catches drift that the patch filters out. When every span in a service is slightly slower, no single span clears the per-span tolerance, but the service's summed span duration does, so the service still shows up as changed.
There are multiple tools you can use to get the traces diff. The gcx CLI wraps the API in an experimental gcx traces diff, so you can compare two IDs from the terminal without fetching either trace yourself:
gcx traces diff --context prod -d <DATASOURCE_UID> <BASE_TRACE_ID> <COMPARE_TRACE_ID>tempo-cli experimental traces-diff diffs two local trace JSON files instead, which is handy when you've already exported them. It emits the patch by default, and you can get the summary with --format. The Tempo MCP server we added in 2.9 now includes a traces-diff tool that returns the summary first, so an AI assistant connected to Tempo gets a usable answer without always pulling back a large patch.
Note: This endpoint is experimental. The request and response formats may change in future releases.
To learn more, see the API PR, the CLI PR, the MCP PR, the traces diff API docs, the experimental traces diff Tempo CLI docs, the gcx traces diff reference, and the Compare traces Tempo MCP docs.
Improved query performance and lower memory usage: vParquet5 is now the default block format
Tempo stores traces in the Apache Parquet columnar block format, and continues to iterate on the layout for better performance and efficiency. We introduced vParquet5 as an opt-in format in Tempo 2.10, and in 3.1, it becomes the default.
The faster metrics read path mentioned earlier in this post is one payoff of the new default. It's implemented for vParquet5 only, and Tempo falls back to the standard path for blocks in older formats, so your metrics queries get faster as new blocks accumulate. Another payoff is that newly written blocks support the span:childCount intrinsic. We added it in 2.10, but only vParquet5 supports it, so previously you had to opt into vParquet5 to use it.
vParquet5 also offers more customization for the way the Parquet columns are configured. You can now have twice as many dedicated columns for strings, support for integers, and very long or high-cardinality strings (known as "blobs"), at each of resource, span, and event levels. The best way to take advantage of this is to use the new tempo-cli command suggest columns. It will analyze your data and tell you the optimal layout, and make it easy to copy and paste the output into your Tempo configuration.
To learn more about the features vParquet5 brings, see the docs for dedicated columns and the vParquet5 section of the Tempo 2.10 release blog post. You can also see the PR and the Apache Parquet block format docs.
How to learn more
To see the full list of improvements, bug fixes, and breaking changes in Tempo 3.1, please refer to the release notes and changelog. Before you upgrade, please review Upgrade your Tempo installation for the changes that need action.
If you are interested in hearing more about Tempo news, please join us on the Grafana Labs Community Slack channel #tempo, post a question in our community forums, or join our monthly Tempo community call. See you there!
Grafana Assistant is the easiest way to get started with metrics, logs, traces, dashboards, and more in Grafana Cloud. We have a generous forever-free tier and plans for every use case. Sign up for free now!
%3Aquality(90)%2F&w=3840&q=75)