<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Grafana Labs blog on Grafana Labs</title><link>https://grafana.com/blog/</link><description>Recent content in Grafana Labs blog on Grafana Labs</description><generator>Hugo -- gohugo.io</generator><language>en</language><atom:link href="/blog/index.xml" rel="self" type="application/rss+xml"/><item><title>'Grafana's Big Tent' podcast: Anthropic on agentic coding, observability, and the future of software engineering</title><link>https://grafana.com/blog/-grafana-s-big-tent-podcast-anthropic-on-agentic-coding-observability-and-the-future-of-software-engineering/</link><pubDate>Fri, 10 Jul 2026 18:34:08</pubDate><author>Grafana Labs Team</author><guid>https://grafana.com/blog/-grafana-s-big-tent-podcast-anthropic-on-agentic-coding-observability-and-the-future-of-software-engineering/</guid><description>&lt;p>In this episode of "Grafana's Big Tent" podcast, hosts Mat Ryer, Senior Director of AI at Grafana Labs, and Tom Wilkie, CTO at Grafana Labs, sit down with Eric Burns, Field Executive Architect at Anthropic, to talk about why Anthropic bet early on running across every major cloud, what it's like watching a technology go from "interesting" to "obviously inevitable" in real time, and how agentic coding tools have changed the day-to-day of building software at Grafana Labs.&lt;/p>&lt;p>You can watch the full episode in the YouTube video below, or listen on&lt;a href="https://open.spotify.com/show/3beQvS8to0rYs1gxOnPrfD"> &lt;/a>&lt;u>&lt;a href="https://open.spotify.com/show/3beQvS8to0rYs1gxOnPrfD">Spotify&lt;/a>&lt;/u> or&lt;a href="https://podcasts.apple.com/us/podcast/grafanas-big-tent/id1616725129"> &lt;/a>&lt;u>&lt;a href="https://podcasts.apple.com/us/podcast/grafanas-big-tent/id1616725129">Apple Podcasts&lt;/a>&lt;/u>.&lt;/p>&lt;p>&lt;em>Note: The following are highlights from episode 9, season 3 of "Grafana's Big Tent" podcast. The transcript below has been edited for length and clarity.&lt;/em>&lt;/p>&lt;h2>&lt;strong>Why Anthropic bet on being everywhere, not just one cloud&lt;/strong>&lt;/h2>&lt;p>&lt;strong>Tom Wilkie:&lt;/strong> We're a pretty big fan of Anthropic models. We do use models from other vendors. I've heard they exist. But I think one of the things that attracted us to the Anthropic models was the availability. We run across all the major cloud providers, across Microsoft and Amazon and Google and so on, and the fact that you can get your models in-region across all of them was a big win for us. How did that come about, or why is that Anthropic's strategy versus, I know Anthropic and OpenAI and Google have very different strategies for model availability.&lt;/p>&lt;p>&lt;strong>Eric Burns:&lt;/strong> Yeah. I've only been here coming up on two years, which makes me near old guard at Anthropic, but not there for the early days in the origin story. It strikes me as kind of a classic second-mover situation where there was an early breakaway leader and they had kind of a slightly more vertically integrated, or at least single-partner, strategy. When you're trying to get your feet under you as a startup, you try to figure out where your opportunities are and you maximize what you can do within the opportunity space.&lt;/p>&lt;p>So I did a startup before this. One of my favorite quotes from that time was something like, "Great architecture is all in the constraints." If you're a second-mover lab, one of your constraints is that you need to line up a hosting partner and you need to tell a good story for that. I think it's a very significant maturity transition for a company to be able to work with many different partners and many different platforms. It's kind of ripping the Band-Aid at a technical level that you can put in all of this work before the first dollar of partner revenue arrives.&lt;/p>&lt;p>That's a substantial opportunity cost, and it might fail, and you might be left without a good result. But I think one of the things that makes Anthropic is the decision to pay that cost early and to generalize and be able to work across clouds. You can see that now in being the first model provider to reach all three major hyperscaler platforms. I think one of our core competencies is making our models run on a diverse set of hardware.&lt;/p>&lt;h2>&lt;strong>Natural language as the new UI for dashboards&lt;/strong>&lt;/h2>&lt;p>&lt;strong>Eric:&lt;/strong> One of the most fun things about working at Anthropic is that there's just this continuous progression of goosebumps moments where you're like, "Oh my gosh, I can't believe computers can do this." And also, "I can't believe I'm getting to see this up close in the lab as it's materializing."&lt;/p>&lt;p>One of those moments that I'll never forget was the first time that I saw that we had a model that was able to basically start with natural language and create streaming UX, just build charts at the speed of thought, build very capable React single-page apps in a single prompt. I thought, "Oh my gosh, this is going to completely transform how we deliver user experiences." &lt;/p>&lt;p>This has progressed into this idea that there's a UI paradigm that you can actually build and it's a legitimate choice now. Which is, you start with natural language and end with streaming charting and BI and structured data presentation. I feel ridiculous saying this out loud on a Grafana podcast, but it just does seem like the wind is at the back of that particular model, that UX design model right now.&lt;/p>&lt;p>&lt;strong>Tom:&lt;/strong> One of the longest open issues on the Grafana repo was, how do I batch-edit things? Because from a UX perspective, actually doing good batch-editing UX is really hard, it turns out. I now just believe natural language is the best UX for doing certain tasks, like batch-editing dashboards.&lt;/p>&lt;p>&lt;strong>Eric:&lt;/strong> Yes, absolutely.&lt;/p>&lt;p>&lt;strong>Tom:&lt;/strong> In one of the early demos that Mat gave me for the Assistant, which I thought was very impressive because I understand how hard it was to achieve, was to make all the themes for the panels purple and change all the panels. Obviously, you show that to a salesperson and they're like, "Why do I want that?" But for an engineer, it's like, "Oh wow, that's really difficult to pull off."&lt;/p>&lt;p>And suddenly, there's a lot of pre-AI dashboard slop around, but now the AI is actually better at producing dashboards, we've found, than humans. It labels the axes. It puts reasonable names on them and titles.&lt;/p>&lt;h2>&lt;strong>AI adoption across Grafana Labs' engineering org&lt;/strong>&lt;/h2>&lt;p>&lt;strong>Tom:&lt;/strong> We obviously have all witnessed the revolution in software engineering. I don't think, at Grafana Labs, any of our engineers are writing code anymore. Our penetration for things like Cursor, Claude Code, Codex, and so on is basically now at 100% in our engineering teams. But I guess this is an observability podcast and we should probably talk about observability at some point. As you go and talk to customers and execs and the community, what are you seeing in terms of the agentic use cases in observability?&lt;/p>&lt;p>&lt;strong>Eric:&lt;/strong> It's a very complex set of actors that are gradually reaching consensus. If I go back to what's happening with coding agents, I've seen this real shift at the executive discussion level over the last six months especially. Six months ago it was: Are the models going to get good enough? Are the coding agents actually going to get there? Is this something that we should inflict on our team, or can we kind of hang out in shadow IT land and some people come in and use it?&lt;/p>&lt;p>Now it has definitively shifted... If you equip all of your frontline engineers with a fire hose of output, all of your other systems are going to crumble under the amount of stress that they're putting on them. Now it's moving to a second-order problem of building heavy automated test coverage, which of course coding agents are great for, obviously having evals around anything that is non-deterministic, and basically stacking the way up to being able to deploy coding agents with some confidence.&lt;/p>&lt;p>Then the second realization, as orgs sort of get the baseline of internalizing officially blessed coding agents, whatever flavor they pick, is this realization that integration just got really, really cheap. Being able to take one internal system and string it to another one became a prompt and a one-shot and some smoke testing, as opposed to, if you're outsourcing it, massive spec document, throw it over the wall, several iterations, months later, you get this dashboard that has been scoped down to the shadow of what you were hoping it would be, and you're like, "I guess I'll do another turn."&lt;/p>&lt;p>The ability to wire all this stuff up just based on immediate need, and then just sort of chat with your data or build a dashboard based on some wild idea, I think this is putting even more pressure on strong, opinionated UI systems for delivering this stuff, and critically, ways to persist it. One of the uncomfortable flip sides of all this prolific output from product managers and people—who didn't think of themselves as coders, but they can write the requirements and they can vibe-code their way to stuff—is distribution and hosting is actually really hard.&lt;/p>&lt;p>If you've got these deeply rooted systems of record where all of your tribal knowledge and all of your state info and your telemetry lives, and you've got kind of a consistent way, for example Grafana, to deliver this to users, the integration is where the magic is happening. It is not just ripping the whole thing off and saying, "Well, anybody can have a database, and we can vibe-code all of the integration layers in between, so why would we need a rendering system?" People still need a way to onboard onto using a platform. It's like back to that old "Who Moved My Cheese?" [&lt;u>&lt;a href="https://en.wikipedia.org/wiki/Who_Moved_My_Cheese%3F">story&lt;/a>&lt;/u>]. If the UI is continuously shifting, you can't enable people on it. You can't train people, and you can't really document it.&lt;/p>&lt;h2>&lt;strong>Zigbee meshes and Home Assistant&lt;/strong>&lt;/h2>&lt;p>&lt;strong>Tom:&lt;/strong> I don't get to code that much. I've got a very large engineering team. I'm mostly a manager now, and I have a backlog of coding projects that I've wanted to do. I'm slowly ticking off that backlog.&lt;/p>&lt;p>This is the intersection of AI and observability, because I do a lot of home automation. Everything in my house is fully automated, and I really want to store a lot more telemetry about my Zigbee mesh, for instance. Sometimes you press a button and it doesn't quite work, or doesn't work as quickly as you would hope, and I want to know why. What do I need to optimize and fix?&lt;/p>&lt;p>I just opened this PR yesterday, actually, and it was mostly Claude Code. It's been on my backlog to go and instrument Zigbee2MQTT for a very long time. It's not hard. It's 500 lines of Node. I am not a JavaScript or TypeScript engineer. It would have taken me days, if not weeks, to have done that, because I would have had to learn a ton. But I can read it and it seems reasonable. It does the right thing. It's got 100% test coverage. The satisfaction of being able to tick off a bunch of personal projects has just been absolutely huge.&lt;/p>&lt;p>&lt;strong>Eric:&lt;/strong> There are two really interesting principles—not principles, attributes of coding agents for certain personalities. One of them is, I've also gone down the rabbit hole of Home Assistant and linking everything to everything and automating it all. I use Grafana and my Home Assistant to look at my InfluxDB, which is all my sensor outputs and so on.&lt;/p>&lt;p>&lt;strong>Tom:&lt;/strong> Terrific. A hot tip: if you use the Prometheus exporter in Home Assistant, I recently refactored all of that, again using Claude Code, two months ago, so it's now got loads more entities in Prometheus.&lt;/p>&lt;p>&lt;strong>Eric:&lt;/strong> Excellent. I know, I'm turning into a serious Home Assistant nerd, so perhaps there's a separate conversation.&lt;/p>&lt;h2>&lt;strong>The 'recovering engineer' getting the dopamine hit back&lt;/strong>&lt;/h2>&lt;p>&lt;strong>Eric:&lt;/strong> This zone of things that I can trust Claude to solve is now a fairly complex software project, but the only upside: there's no business value generated. I can click something on my phone that I used to click on a wall panel, right? Things like that. Or I can see high-fidelity data that I couldn't see before. The cost-benefit is getting completely transformed in terms of the effort that I imagine something is going to take. It has been in steep log decay to the point where basically nothing feels out of reach in this sprawling home integration project.&lt;/p>&lt;p>There were certain things where I was like, "OK, well, I'm just going to grit my teeth and reach in there and write the code, or I'm just not going to do it at all." So the first property is, many things that were total wastes of time got so cheap that just dashing off a Claude prompt and then checking in an hour later and saying, "Oh my goodness, I got this thing." I wasn't expecting that pure upside.&lt;/p>&lt;p>The second one I think is really profound for managers, and I experienced this. The last time I wrote production code was five years ago. There's a half-life of the quality of your dev situation where I can step away for a month and come back, and it would be eight hours before I was back in the flow. You pull down the latest, you've got some break, somebody forgot to check in this other thing, there was a framework migration, now you have to go read the docs for this framework, just on and on. The ability to ramp back up into the flow is almost always instantaneous now, because Claude will just bash through whatever nonsense is keeping you from being in your dev loop.&lt;/p>&lt;p>&lt;strong>Tom:&lt;/strong> We refer to managers in Grafana Labs as recovering engineers.&lt;/p>&lt;p>&lt;strong>Eric:&lt;/strong> Yeah, right. It's the encapsulation of exactly that idea. Suddenly that recovering engineer can get the dopamine hit of solving a problem that they didn't used to be able to.&lt;/p>&lt;h2>&lt;strong>Do we still need to read the code?&lt;/strong>&lt;/h2>&lt;p>&lt;strong>Mat Ryer:&lt;/strong> Do you think we will end up in a situation where we've stopped looking at the code, like Assembly? We don't really look at Assembly unless we need to. Most people can't, though they probably can now thanks to Claude, etc. Do you think we'll get to the point where the code is like, you look at it if you need to debug something, otherwise you're good. &lt;/p>&lt;p>&lt;strong>Eric: &lt;/strong>You know, if you talk to Boris [Cherny], the creator of Claude Code, his view is, "We're already there." About a year ago, I stopped dirtying my hands with reading the actual code. Now I operate at the pattern level.&lt;/p>&lt;p>Increasingly, lately, I've had this weird experience with Opus 4.6 where I'll think I see something smart and I'll interrupt it in a loop. Then it dawns on me that it's actually a step ahead of me, and I'm like, "Oh, I'm so sorry. I thought I understood that, but actually, I don't. You're already on the right track."&lt;/p>&lt;p>So there's a certain threshold where, you know, I fancy on a good day I'm a decent engineer and I kind of know what's going on. I had this very uncomfortable feeling of being in Claude's way as it was trying to solve the problem and benevolently deliver the thing that I was asking it for.&lt;/p>&lt;p>One tier of question is, should we go review the code? Another one entirely is, is a human, even a competent one that knows the code base, for some value of competence, adding value or reducing throughput? I think these are the questions that, back to the idea of software engineering collectively engaging with it, there's no one right answer. Again, it's a question of risk tolerance.&lt;/p>&lt;p>But I've definitely had this ratcheting sense of my value being pushed out of implementation in the same way that anybody that's ever written Assembly would feel that most likely a high-level language expressing programmer intent very well, and a strong performance and compilation stack, is going to outperform whatever any of us could do at the assembly level.&lt;/p>&lt;p>&lt;strong>Mat:&lt;/strong> Yeah, I see that future. I really do. It's closer than we think. Very exciting.&lt;/p>&lt;p>&lt;em>"Grafana's Big Tent" podcast wants to hear from you. If you have a great story to share, want to join the conversation, or have any feedback, please contact the Big Tent team at bigtent@grafana.com.&lt;/em>&lt;/p></description></item><item><title>Business intelligence plugins for Grafana: A support update</title><link>https://grafana.com/blog/business-intelligence-plugins-for-grafana-a-support-update/</link><pubDate>Thu, 09 Jul 2026 15:52:01</pubDate><author>Thanos Karachalios</author><guid>https://grafana.com/blog/business-intelligence-plugins-for-grafana-a-support-update/</guid><description>&lt;p>In January, we &lt;u>&lt;a href="https://grafana.com/blog/business-intelligence-plugins-for-grafana-whats-next/">announced that Grafana Labs&lt;/a>&lt;/u> had assumed maintenance of the business intelligence (BI) plugins created by Volkov Labs, and committed to a six-month maintenance period. &lt;/p>&lt;p>Today, we’re sharing an update: &lt;strong>we're extending our maintenance commitment through the end of 2026&lt;/strong>. As announced earlier this year, that commitment includes maintaining compatibility with recent Grafana releases while handling bug fixes, security updates, and community contributions on a best-effort basis.&lt;/p>&lt;p>We want to be transparent about why we’re extending the maintenance period, what we've learned over the past several months, and how we're thinking about support going forward.&lt;/p>&lt;h2>A quick look back—and where things stand&lt;/h2>&lt;p>In September 2025, our longtime partner Volkov Labs announced they had been acquired. In light of the news, we committed in January 2026 to taking over the maintenance and development of their popular &lt;u>&lt;a href="https://grafana.com/grafana/plugins/all-plugins/?search=Volkov&amp;pg=business-intelligence-plugins-for-grafana-whats-next&amp;plcmt=in-text">BI plugin suite&lt;/a>&lt;/u> to ensure continuity for the Grafana community. &lt;/p>&lt;p>Since then, we've made significant progress on the priorities we outlined in January:&lt;/p>&lt;ul>&lt;li>&lt;strong>Compatibility:&lt;/strong> The plugins have been brought up to date with &lt;u>&lt;a href="https://grafana.com/blog/grafana-13-release-all-the-latest-features/">Grafana 13&lt;/a>&lt;/u> and &lt;u>&lt;a href="https://grafana.com/blog/react-19-is-coming-to-grafana-what-plugin-developers-need-to-know/">React 19&lt;/a>&lt;/u>—the core engineering work for the maintenance window is largely complete.&lt;/li>&lt;li>&lt;strong>Bug fixes and community PRs:&lt;/strong> We've continued to triage issues and review community contributions on a best-effort basis.&lt;/li>&lt;li>&lt;strong>Security:&lt;/strong> We've continued to monitor and respond to first- and third-party vulnerabilities.&lt;/li>&lt;/ul>&lt;h2>What we've heard from you&lt;/h2>&lt;p>Over the past several weeks we've been meeting directly with customers and solutions architects who rely on these plugins. The goal was simple: understand &lt;em>how&lt;/em> and &lt;em>why&lt;/em> you use them before we make any decisions about their future.&lt;/p>&lt;p>What we've learned has been genuinely valuable. Many organizations use plugins like &lt;u>&lt;strong>&lt;a href="https://github.com/grafana/business-charts">Business Charts&lt;/a>&lt;/strong>&lt;/u>, &lt;u>&lt;strong>&lt;a href="https://github.com/grafana/business-variable">Business Variable&lt;/a>&lt;/strong>&lt;/u>, and &lt;u>&lt;strong>&lt;a href="https://grafana.com/grafana/plugins/marcusolsson-dynamictext-panel/">Business Text&lt;/a>&lt;/strong>&lt;/u>&lt;u>&lt;a href="https://grafana.com/grafana/plugins/marcusolsson-dynamictext-panel/"> &lt;/a>&lt;/u>to go above and beyond what core Grafana panels can do, building highly specialized, persona-specific visualizations tailored to industry verticals, proprietary data, and even custom physical environments (think dashboards designed to be legible on large displays from across a room). &lt;/p>&lt;p>For these users, the value isn't just in the plugins themselves, but in their ability to customize panels to meet their unique needs. &lt;/p>&lt;h2>Extending the support window&lt;/h2>&lt;p>As noted above, we're extending the current maintenance commitment through the end of 2026. This gives us time to keep listening to how you use these plugins, and to evaluate the right long-term support for each one. &lt;/p>&lt;p>We'll share more concrete plans as they take shape, and communicate any changes proactively. &lt;strong>No plugin will be removed or have support reduced without advance notice.&lt;/strong>&lt;/p>&lt;h2>A note on custom code and support scope&lt;/h2>&lt;p>Because several of these plugins let you write your own JavaScript, we want to set clear, honest expectations about what "supported" means.&lt;/p>&lt;p>Grafana Labs support covers the plugins themselves—their documented behavior, the stable APIs they expose, and bugs in the code we maintain. It does not extend to debugging user-provided code added to the editor of a panel plugin. User-provided code in the panel editors can break between Grafana versions, and we're not able to guarantee or troubleshoot it on your behalf.&lt;/p>&lt;p>As we firm up the roadmap, improving the documented extension points—so you can achieve the customization you need on a more stable foundation—is a key part of the conversation. We'll provide updates once those decisions are made.&lt;/p>&lt;h2>How to reach us and share feedback&lt;/h2>&lt;ul>&lt;li>&lt;strong>Grafana Cloud and Grafana Enterprise users:&lt;/strong> Please contact Grafana Labs support with any questions. &lt;/li>&lt;li>&lt;strong>Grafana OSS users:&lt;/strong> Please use the &lt;u>&lt;a href="https://community.grafana.com/">community forums&lt;/a>&lt;/u>, the plugin &lt;u>&lt;a href="https://github.com/grafana/business-charts">GitHub repositories&lt;/a>&lt;/u>, or the &lt;u>&lt;a href="https://grafana.slack.com/archives/C0Y4TLW74">Community Slack&lt;/a>&lt;/u> channels.&lt;/li>&lt;/ul>&lt;p>If these plugins are important to your workflows, we'd especially love to hear from you while the roadmap is still taking shape. Your input will directly influence where this goes next.&lt;/p></description></item><item><title>How to scale access control in Grafana Cloud</title><link>https://grafana.com/blog/how-to-scale-access-control-in-grafana-cloud/</link><pubDate>Tue, 07 Jul 2026 17:02:52</pubDate><author>Jake Batty</author><guid>https://grafana.com/blog/how-to-scale-access-control-in-grafana-cloud/</guid><description>&lt;p>One of the primary reasons organizations adopt Grafana Cloud is to create a single pane of glass across the data they collect from self-hosted systems, cloud providers, and third-party platforms. Bringing those signals together enables richer correlations, reduces tool sprawl, and makes it easier for teams to understand what's happening across their environment.&lt;/p>&lt;p>But as observability grows and becomes more centralized, access management becomes more important. When infrastructure metrics, application logs, business KPIs, and customer-specific data all live in the same platform, organizations need a scalable way to ensure the right people have access to the right resources without creating additional administrative overhead.&lt;/p>&lt;p>The good news is this can all be done directly in &lt;u>&lt;a href="https://grafana.com/products/cloud/">Grafana Cloud&lt;/a>&lt;/u>. To illustrate how this can work in practice, let's look at a fictional company modeled after a common scenario.&lt;/p>&lt;h2>Meet AcmeCloud (and the challenge of manual provisioning access) &lt;/h2>&lt;p>AcmeCloud, a platform-as-a-service provider, has set up a new Grafana Cloud instance to visualize infrastructure metrics, business metrics, sensitive application logs pertaining to each of their customers, and more—all in one place. Now they need to provision access across their organization.&lt;/p>&lt;p>They have hundreds of users that require access, and different roles have different needs:&lt;/p>&lt;ul>&lt;li>Admins need full control of data and users&lt;/li>&lt;li>SREs need to configure data sources, explore telemetry, and build dashboards&lt;/li>&lt;li>Contractors, working on behalf of specific AcmeCloud customers, need to log in but only to see the dashboards and data relevant to their tenant&lt;/li>&lt;/ul>&lt;p>At this scale, manual provisioning isn't an option. Neither is trying to keep users in sync between AcmeCloud's identity provider and Grafana Cloud by hand. It's simply too time-consuming and too prone to errors.&lt;/p>&lt;p>For organizations facing similar growth, this is often the point where access management starts to become a scalability challenge rather than an administrative task.&lt;/p>&lt;h2>Establishing a single source of truth with SSO and SCIM&lt;/h2>&lt;p>Instead of handling everything manually, AcmeCloud relies on SSO (Single Sign-On) for authentication and SCIM (System for Cross-domain Identity Management) for provisioning.&lt;/p>&lt;p>SSO handles how users logged in, while SCIM ensures the right users and groups are automatically available in Grafana Cloud. There are many options for authenticating users in Grafana Cloud, including native integrations with identity providers as well as support for generic authentication methods. To see the full list of supported methods and integrations, &lt;u>&lt;a href="https://grafana.com/docs/grafana/latest/setup-grafana/configure-access/configure-authentication/">read more here&lt;/a>&lt;/u>.&lt;/p>&lt;p>As users are added, removed, or updated in the identity provider, those changes are reflected in Grafana Cloud automatically. The same concept applies to groups, which are mapped directly to &lt;u>&lt;a href="https://grafana.com/docs/grafana/latest/administration/team-management/">Grafana teams&lt;/a>&lt;/u>. Groups within the identity provider typically organize users by job function, which correlate to a Grafana team in most cases, since users working in the same role typically have the same use cases and permission levels.&lt;/p>&lt;p>The result is a single source of truth for identity, with no need to manually reconcile users or worry about configuration drift.&lt;/p>&lt;p>For organizations implementing this approach, Grafana's SSO and SCIM integrations make it possible to manage users and groups from the identity provider rather than treating Grafana as a separate identity system.&lt;/p>&lt;p>If you are new to the concept of SCIM provisioning, check out &lt;u>&lt;a href="https://grafana.com/blog/introducing-scim-provisioning-in-grafana-enterprise-grade-user-management-made-simple/">this blog&lt;/a>&lt;/u> to learn more.&lt;/p>&lt;p>&lt;strong>Tip: &lt;/strong>Establish clear group naming conventions in your identity provider before enabling SCIM. It makes permission mapping significantly easier as your environment grows.&lt;/p>&lt;h2>Building a layered access model with RBAC&lt;/h2>&lt;p>With users now syncing into Grafana Cloud, the next step is defining what those users can actually do.&lt;/p>&lt;p>AcmeCloud opts to manage access primarily through role-based access control (RBAC).This helps to maintain scalability while providing greater flexibility and security through a combination of basic roles, team memberships, and resource-level permissions. Dive deeper on creating a RBAC strategy and configuration through the &lt;u>&lt;a href="https://grafana.com/docs/grafana/latest/administration/roles-and-permissions/access-control/plan-rbac-rollout-strategy/">guidance in our documentation&lt;/a>&lt;/u>.&lt;/p>&lt;p>Basic roles in Grafana set the foundation, with assignments mapped from the identity provider: &lt;/p>&lt;ul>&lt;li>Admins are assigned the Admin role&lt;/li>&lt;li>Internal AcmeCloud users from the SRE team receive the Editor role&lt;/li>&lt;li>Contractors receive no basic role at all &lt;/li>&lt;/ul>&lt;p>This allows the internal SRE team to access dashboards and data broadly. They can build dashboards and alerts, as well as query metrics, logs, and traces across all tenants. All that is missing now is the ability to configure or alter data sources.&lt;/p>&lt;p>Contractors can log in, but they're presented with an empty Grafana experience by default. For now, there are no available actions to take in Grafana, but they need the ability to view dashboards based on data solely containing their own tenant id.&lt;/p>&lt;p>To build on the permissions provided through basic roles, teams that were provisioned through SCIM are used to layer on additional access.&lt;/p>&lt;p>Users inherit permissions from both their basic role and any teams they belong to. This is an important distinction: permissions in Grafana are additive, not restrictive. A user's effective access is the combination of all assigned roles and team memberships.&lt;/p>&lt;p>For internal users from the SRE team, a shared team is granted additional capabilities, including the ability to configure and manage data sources. This extends their access beyond what the Editor role provides by default.&lt;/p>&lt;p>For organizations designing their own access model, a useful starting point is to assign the minimum role required and then use teams to grant additional capabilities as needed.&lt;/p>&lt;p>&lt;strong>Tip: &lt;/strong>Avoid using highly permissive basic roles as a shortcut. Teams are often a more scalable way to manage access as responsibilities evolve.&lt;/p>&lt;h2>Restricting dashboard visibility with teams and folders&lt;/h2>&lt;p>Contractors are grouped into tenant-specific teams, with no additional capabilities added at that level. Instead, access is controlled through folder permissions.&lt;/p>&lt;p>Each team is given view access to a specific dashboard folder, so contractors can only see only the dashboards relevant to their tenant and nothing else.&lt;/p>&lt;p>Folder permissions provided a simple and scalable way to partition dashboard visibility without creating separate Grafana instances for every customer.&lt;/p>&lt;p>For AcmeCloud, this approach balanced operational simplicity with tenant isolation. Organizations managing multiple customers, business units, or environments often find folder permissions to be one of the simplest ways to segment dashboard access.&lt;/p>&lt;p>&lt;strong>Tip:&lt;/strong> Design your folder structure with future growth in mind. Reorganizing hundreds of dashboards later can become a significant effort.&lt;/p>&lt;h2>Enforcing data-level isolation with LBAC&lt;/h2>&lt;p>Up to this point, AcmeCloud has been using RBAC to determine what users could access within Grafana Cloud. But dashboard visibility and permissions were only part of the problem. &lt;/p>&lt;p>AcmeCloud also needs to ensure that access to dashboards doesn't automatically mean access to all underlying data. This is where label-based access control (LBAC) came into play.&lt;/p>&lt;p>By using LBAC, data is scoped by the tenant. That way, even if a user has access to a dashboard, queries are restricted so they only return data associated with their assigned tenant. Read more &lt;u>&lt;a href="https://grafana.com/docs/grafana/latest/administration/data-source-management/teamlbac/">here&lt;/a>&lt;/u> for guidance on configuration and limitations.&lt;/p>&lt;p>This adds a final layer of protection, ensuring that customer data remains isolated-even within a shared Grafana instance.&lt;/p>&lt;p>For organizations operating multi-tenant environments, this distinction is important. Dashboard permissions control what users can see, but data-level controls determine what data they can access. Both are required to build a complete access strategy.&lt;/p>&lt;h2>Delivering the right experience for every user&lt;/h2>&lt;p>From the contractor's perspective, the experience is intentionally limited.&lt;/p>&lt;p>After logging in via SSO, they're placed into their tenant-specific team and land on a predefined home dashboard. They can only view the dashboards within their assigned folder and they can't explore beyond that scope—no visibility into other tenants, and no risk of accessing unrelated data.&lt;/p>&lt;p>What made this setup work isn't any single feature. Instead, it's how all these features are combined to address scalability, flexibility, and security demands.&lt;/p>&lt;p>SCIM handles scale and lifecycle management. RBAC—implemented through basic roles, teams, and folder permissions—defines who can access resources. LBAC then ensures data-level isolation within those resources.&lt;/p>&lt;p>Individually, each of these capabilities is straightforward. Together, they form a layered access model that allows organizations to onboard large numbers of users, both internal and external, without sacrificing control.&lt;/p>&lt;h2>Bringing it all together&lt;/h2>&lt;p>As Grafana usage grows, access management becomes just as important as observability itself. Designing around user types, scope, and data boundaries from the start makes it possible to scale confidently without losing track of who can see and do what.&lt;/p>&lt;p>The AcmeCloud example demonstrates a pattern that many organizations can adopt: centralize identity with SSO and SCIM, use roles and teams to define capabilities, control visibility through folders, and enforce data isolation with LBAC.&lt;/p>&lt;p>A single pane of glass doesn't have to mean broad access to everything. With a layered approach to access control, organizations can centralize observability while still maintaining the security and governance required as their environments grow.&lt;/p></description></item><item><title>Full-stack observability in Grafana Cloud: How to investigate issues across services and infrastructure</title><link>https://grafana.com/blog/full-stack-observability-in-grafana-cloud-how-to-investigate-issues-across-services-and-infrastructure/</link><pubDate>Tue, 30 Jun 2026 16:00:56</pubDate><author>Victor Padilla</author><guid>https://grafana.com/blog/full-stack-observability-in-grafana-cloud-how-to-investigate-issues-across-services-and-infrastructure/</guid><description>&lt;p>Many times, the hardest part of troubleshooting isn’t fixing the actual problem. It’s figuring out where to start. &lt;/p>&lt;p>As engineers, it’s easy to lose count of how many times we’ve opened logs, then 10 metrics tabs, and another 10 tabs with trace queries, only to end up back in the logs trying to find a root cause. Modern applications run across several layers of services and infrastructure, and understanding an issue often means connecting information scattered across different resources, teams, and observability signals.&lt;/p>&lt;p>&lt;a href="https://grafana.com/products/cloud/application-observability/">Grafana Cloud Application Observability&lt;/a> and &lt;a href="https://grafana.com/products/cloud/kubernetes/">Kubernetes Monitoring&lt;/a> bring that context together, providing a full-stack view across applications, infrastructure, and Kubernetes environments. Starting your investigation from a service, pod, node, namespace, or cluster, you can quickly jump to the logs, traces, and profiles that help explain what's happening, all within the workflows you already use.&lt;/p>&lt;p>In Grafana Cloud, this experience is powered by the &lt;a href="https://grafana.com/docs/grafana-cloud/knowledge-graph/">knowledge graph&lt;/a>, which automatically models your applications and infrastructure into a unified graph. This graph maps telemetry to each connected entity, including services, pods, nodes, clusters, databases, and cloud accounts. The resulting views help you visualize these relationships and observability data together in a single place, so you can correlate signals, understand dependencies, and move from symptom to root cause faster.&lt;/p>&lt;p>In this post, we'll walk through an example of how full-stack observability in Grafana Cloud helps you investigate issues across the application and infrastructure layers, and how you can customize knowledge graph configurations to fit your environment.&lt;/p>&lt;h2>From entity to insights: a workflow example &lt;/h2>&lt;p>Grafana Cloud includes multiple features and views for full-stack observability across your applications, infrastructure, and Kubernetes environments. This eliminates the need to manually write queries or jump between dashboards. By automatically bringing together signals and visualizing relationships between services and infrastructure, you can identify issues and find root causes faster. &lt;/p>&lt;p>Each of the following features is built for a different stage of investigation, with the entities, relationships, and insights behind them powered by the knowledge graph.  &lt;/p>&lt;ul>&lt;li>&lt;strong>RCA workbench&lt;/strong> helps you investigate incidents by bringing insights, dependencies, and telemetry together in a single timeline.&lt;/li>&lt;li>&lt;strong>Entity graph&lt;/strong> provides a visual representation of the relationships between services, infrastructure, and other components in your environment, making it easier to understand dependencies and identify potential root causes.&lt;/li>&lt;li>&lt;strong>Entity catalog&lt;/strong> acts as a central inventory of all services and infrastructure discovered by the knowledge graph, combining health status, insights, metrics, and metadata so you can quickly identify what needs attention.&lt;/li>&lt;/ul>&lt;p>From any of these features, you can launch directly into logs, traces, and profiles for a given entity. The embedded &lt;u>&lt;a href="https://grafana.com/docs/grafana-cloud/visualizations/simplified-exploration/">Grafana Drilldown&lt;/a>&lt;/u> tab opens automatically with filters derived from entity configurations (more on that below), making it easy to correlate errors detected from metrics with other telemetry signals.&lt;/p>&lt;p>To illustrate how this works in practice, let's use an example. It's late in the evening and you've just received an alert through &lt;u>&lt;a href="https://grafana.com/docs/grafana-cloud/alerting-and-irm/alerting/">Grafana Alerting&lt;/a>&lt;/u> in Slack that makes you break out into a cold sweat. Unsure where to start, you follow the provided link to RCA Workbench, so you can explore all potential causes for a particular issue correlated over time and dependency for the impacted service. There are some insights about your failing service, so you go and take a look at the telemetry your application is emitting.&lt;/p>&lt;p>You've identified the symptom, but what's actually failing? Is the issue with the Kubernetes pod? Another service? The database? Good news: you don’t need to exit the workbench to check the whole picture.&lt;/p>&lt;p>Instead of jumping between tools, you can explore the service's connected entities, including microservices and its frontend component, directly from the workbench. Several related services show activity, but one stands out: your PostgreSQL database appears to be failing.&lt;/p>&lt;p>A quick jump into the database’s logs reveals the root cause. The database has too many simultaneous connections and is refusing new ones, which is causing some related services to break. From there, you can begin to troubleshoot, whether that's increasing resources or horizontally scaling additional instances.&lt;/p>&lt;p>You can also create and share shortened URLs that bring teammates directly to the same view, making it easier to collaborate during investigations.&lt;/p>&lt;p>You might notice that when you open one of the Drilldown views, some filters are already applied to surface only the most relevant data. This is because the knowledge graph is configurable and can be tailored to fit your needs. Let’s take a closer look at how these configurations work and how you can customize them for your environment.&lt;/p>&lt;h2>Have it your way: customizing configurations&lt;/h2>&lt;p>As shown in the example above, Drilldown is a powerful tool that helps you understand your data without learning an entirely new query language. However, pinpointing the fields and labels that are most useful across your environments can be a challenge, as every system and team may follow different conventions for structuring and emitting telemetry.&lt;/p>&lt;p>There are default &lt;u>&lt;a href="https://grafana.com/docs/grafana-cloud/knowledge-graph/configure/telemetry-correlation/#default-configurations">configurations&lt;/a>&lt;/u> that control how the knowledge graph filters, narrows down, and correlates your observability data with the entities in your environment. These configurations cover common scenarios by mapping labels such as &lt;em>pod&lt;/em>, &lt;em>namespace&lt;/em>, and &lt;em>cluster&lt;/em>, along with standard OpenTelemetry fields like &lt;strong>service.name&lt;/strong> and &lt;strong>service.namespace&lt;/strong> for logs. &lt;/p>&lt;p>There are many potential setups: different labeling strategies, OpenTelemetry or non-OpenTelemetry, internal conventions, and more. This results in an almost endless number of possible scenarios.&lt;/p>&lt;p>Instead of trying to support every configuration out of the box, we empower users to customize their own experience within the knowledge graph.&lt;/p>&lt;h3>Creating and editing a configuration&lt;/h3>&lt;p>You can create configurations for specific environments, apply them only to certain entity types, or define matchers based on entity properties. You can even configure them to query any base data source you choose. &lt;/p>&lt;p>Take the following example (shown in the GIF below), which shows a new configuration being created to map the entity property &lt;code>deployment.environment&lt;/code> to the log label &lt;code>service_namespace&lt;/code>, and the entity property &lt;code>service&lt;/code> to the log label &lt;code>service_name&lt;/code>. Furthermore, filters ensure this configuration is applied only to entities whose deployment environment starts with &lt;code>prod&lt;/code>. This could represent a real scenario in which your production metrics use &lt;code>deployment_environment&lt;/code>, while your logs only include &lt;code>service_namespace&lt;/code>.&lt;/p>&lt;p>To learn more about creating and editing correlations, please check out &lt;u>&lt;a href="https://grafana.com/docs/grafana-cloud/knowledge-graph/configure/telemetry-correlation/#configure-telemetry-correlation">our docs&lt;/a>&lt;/u>. &lt;/p>&lt;h3>Resolving configuration conflicts&lt;/h3>&lt;p>In some cases, configurations may overlap or conflict due to matching scenarios. When this happens, the &lt;strong>priority order&lt;/strong> defined on the configuration page determines which configuration takes precedence.&lt;/p>&lt;p>Configurations are evaluated as an ordered list, so adjusting their priority allows you to control how conflicts are resolved.&lt;/p>&lt;h3>Handling unmatched configurations&lt;/h3>&lt;p>If no configuration matches (or if the mappings cannot be applied) you’ll be prompted with an additional screen that allows you to temporarily apply a configuration even if it doesn’t match automatically.&lt;/p>&lt;p>This ensures you can still explore telemetry signals without needing to immediately adjust your configuration.&lt;/p>&lt;h2>Best practices for configurations &lt;/h2>&lt;p>While the system is designed to be flexible, a few best practices can help ensure a smoother experience:&lt;/p>&lt;ul>&lt;li>&lt;strong>Create a sensible default configuration&lt;/strong> that matches most of your environments.&lt;/li>&lt;li>&lt;strong>Add more specific configurations&lt;/strong> for special cases, such as different teams using different namespaces or environment labels for logs or traces, to achieve more granular telemetry filtering.&lt;/li>&lt;li>&lt;strong>Place default configurations at the bottom&lt;/strong> of the priority list so more specific ones take precedence.&lt;/li>&lt;li>&lt;strong>Use consistent fields and labels across metrics, logs, traces, and profiles&lt;/strong> to make correlation easier.&lt;/li>&lt;li>&lt;strong>Add as many mappings as possible&lt;/strong> to narrow down searches. Mappings are optional, so if an entity property is missing, it simply won’t be applied as a filter.&lt;/li>&lt;/ul>&lt;p>You can also use the &lt;u>&lt;strong>&lt;a href="https://grafana.com/docs/grafana-cloud/as-code/infrastructure-as-code/terraform/">Grafana Terraform provider&lt;/a>&lt;/strong>&lt;/u> to automate the creation and management of configurations. To learn more, please check out our documentation for the &lt;u>&lt;a href="https://grafana.com/docs/grafana-cloud/knowledge-graph/configure/telemetry-correlation/">knowledge graph&lt;/a>&lt;/u> and &lt;u>&lt;a href="https://grafana.com/docs/grafana-cloud/as-code/infrastructure-as-code/terraform/terraform-knowledge-graph/log-configurations/">Terraform&lt;/a>&lt;/u>.&lt;/p>&lt;h2>How to learn more&lt;/h2>&lt;p>The knowledge graph in Grafana Cloud offers a powerful way to unify your observability signals and accelerate root cause analysis. To dive deeper into shaping your telemetry and optimizing your graph, explore the following resources:&lt;/p>&lt;ul>&lt;li>&lt;u>&lt;a href="https://grafana.com/docs/grafana-cloud/knowledge-graph/configure/telemetry-correlation/#configure-telemetry-correlation">Configure telemetry correlation&lt;/a>&lt;/u>: Learn how to define explicit mappings between entities and data sources in detail.&lt;/li>&lt;li>&lt;u>&lt;a href="https://grafana.com/docs/grafana-cloud/monitor-applications/application-observability/setup/instrumentation-quality">Instrumentation quality&lt;/a>&lt;/u>: Understand the baseline for shaping your traces and metrics to ensure your application data is correctly linked.&lt;/li>&lt;li>&lt;u>&lt;a href="https://grafana.com/docs/opentelemetry/ingest/">Sending OTLP data&lt;/a>&lt;/u>: Learn how to collect, process, and export telemetry data into the Grafana Cloud observability stack so you can check it directly on the entity views.&lt;/li>&lt;/ul></description></item><item><title>Grafana 13.1 release: observability as code updates, extending Grafana Assistant across more data sources, and more</title><link>https://grafana.com/blog/grafana-13-1-release-all-the-latest-features/</link><pubDate>Wed, 24 Jun 2026 21:11:19</pubDate><author>Grafana Labs Team</author><guid>https://grafana.com/blog/grafana-13-1-release-all-the-latest-features/</guid><description>&lt;p>Earlier this year, &lt;u>&lt;a href="https://grafana.com/blog/grafana-13-release-all-the-latest-features/">Grafana 13 laid the groundwork&lt;/a>&lt;/u> for making it easier and faster than ever to turn your data into actionable insights.  &lt;/p>&lt;p>With our latest minor release, Grafana 13.1, we're building on that foundation, expanding observability as code, bringing Grafana Assistant to more data sources, and streamlining the everyday workflows teams rely on to visualize, analyze, and act on their data. &lt;/p>&lt;p>&lt;/p>&lt;p>Below are just some of the highlights from Grafana 13.1. If you want to explore &lt;em>all&lt;/em> the latest updates, please refer to the &lt;u>&lt;a href="https://github.com/grafana/grafana/blob/main/CHANGELOG.md">changelog&lt;/a>&lt;/u> or our &lt;u>&lt;a href="https://grafana.com/docs/grafana/latest/whatsnew/whats-new-in-v13-1/">What’s New documentation&lt;/a>&lt;/u>. &lt;/p>&lt;h2>Managing dashboards as code: what's new in Git Sync&lt;/h2>&lt;p>&lt;u>&lt;a href="https://grafana.com/docs/grafana/latest/as-code/observability-as-code/git-sync/">Git Sync&lt;/a>&lt;/u>, a feature that brings native GitOps workflows into your Grafana instance, &lt;u>&lt;a href="https://grafana.com/blog/git-sync-grafana/">reached general availability&lt;/a>&lt;/u> with the release of Grafana 13. We added features to give you more flexibility and control when managing your dashboards as code, including GitHub App authentication and support for GitLab, BitBucket, and pure Git.&lt;/p>&lt;p>But we didn’t stop there. Grafana 13.1 brings four more enhancements to Git Sync that make it even easier to incorporate observability as code into your day-to-day workflows.&lt;/p>&lt;h3>Import dashboards straight into a provisioned folder&lt;br>&lt;/h3>&lt;p>&lt;em>Generally available in all editions of Grafana &lt;/em>&lt;/p>&lt;p>You can now &lt;u>&lt;a href="https://grafana.com/whats-new/2026-06-23-git-sync--import-dashboards-from-the-ui-to-simplify-adding-them-in-a-synced-folder/">import dashboard JSON&lt;/a>&lt;/u> straight into a Git Sync-provisioned folder, picking the file path, branch, commit message, and workflow as part of the import.&lt;/p>&lt;p>From a folder, hit &lt;strong>Import&lt;/strong> and Grafana walks you through a provisioned import flow: pick the file path, branch, commit message, and workflow, and the dashboard is committed back to your repository as part of the import. &lt;/p>&lt;p>Uniqueness is path-based, so two dashboards can share a title as long as they live at different paths in the repo, and a conflicting path stops the import before anything is overwritten.&lt;/p>&lt;h3>Sync dashboards at the root level&lt;/h3>&lt;p>&lt;em>Generally available in all editions of Grafana &lt;/em>&lt;/p>&lt;p>You can now &lt;u>&lt;a href="https://grafana.com/whats-new/2026-06-23-git-sync--dashboard-synchronisation-now-available-at-root-level/">sync dashboards at the root level&lt;/a>&lt;/u>, without a containing folder, so provisioned dashboards can live alongside your non-provisioned ones. This is useful when a repo represents your whole Grafana setup, or when forcing everything under one folder doesn’t align with how your team organizes dashboards.&lt;/p>&lt;p>Pick &lt;strong>Sync external storage directly at root level without a containing folder&lt;/strong> in the setup wizard and your provisioned dashboards land at the root, alongside everything else, instead of being scoped under a single folder.&lt;/p>&lt;h3>Make dashboard context visible by default &lt;/h3>&lt;p>&lt;em>Available in public preview in all editions of Grafana&lt;/em>&lt;/p>&lt;p>Git Sync-provisioned folders &lt;u>&lt;a href="https://grafana.com/whats-new/2026-06-23-git-sync--readmemd-files-added-to-a-folder-in-git-are-displayed-in-the-ui/">now render their &lt;/a>&lt;/u>&lt;code>README.md&lt;/code> inline by default, so the context for a folder travels with it.&lt;/p>&lt;p>Just drop a &lt;code>README.md&lt;/code> next to your dashboards in the repo and it shows up in Grafana, including links, ownership notes, runbooks, or whatever your team wants to see sitting alongside their dashboards.&lt;/p>&lt;h3>Sign commits automatically &lt;/h3>&lt;p>&lt;em>Generally available in in all editions of Grafana&lt;/em>&lt;/p>&lt;p>Git Sync can now &lt;u>&lt;a href="https://grafana.com/whats-new/2026-06-23-git-sync--verified-commits/">sign commits with GPG, SSH, or S/MIME keys&lt;/a>&lt;/u>, so your Git provider marks them as verified. This means teams with branch protection rules that require signed commits can now use Git Sync without friction. Until now, Git Sync could only create unsigned commits, which caused pushes to be rejected in those repositories.&lt;/p>&lt;p>To enable signing, configure a signing key on the repository, and Git Sync will automatically sign every commit it makes to that branch. If no signing key is configured, commits remain unsigned.&lt;/p>&lt;p>To learn more about Git Sync, check out our &lt;u>&lt;a href="https://grafana.com/docs/grafana/latest/as-code/observability-as-code/git-sync/">documentation&lt;/a>&lt;/u>. &lt;/p>&lt;h2>Extending the reach of Grafana Assistant&lt;/h2>&lt;p>We're continuing to &lt;u>&lt;a href="https://grafana.com/blog/grafana-assistant-everywhere/">expand where and how you can use Grafana Assistant&lt;/a>&lt;/u>, our AI-powered agent in Grafana Cloud. From connecting to more data sources across your stack to making Assistant easier to access in self-managed environments, these updates help bring AI-powered observability to wherever your data (and Grafana instance) lives. &lt;/p>&lt;h3>Using Assistant with additional data sources &lt;/h3>&lt;p>With Grafana 13.1, you can &lt;u>&lt;a href="https://grafana.com/whats-new/2026-05-30-query-snowflake--jira--dynatrace--and-five-more-directly-from-grafana-assistant/">use Assistant to directly query&lt;/a>&lt;/u> eight additional Grafana data sources: &lt;u>&lt;strong>&lt;a href="https://grafana.com/grafana/plugins/grafana-snowflake-datasource/">Snowflake&lt;/a>&lt;/strong>&lt;/u>,&lt;strong> &lt;/strong>&lt;u>&lt;strong>&lt;a href="https://grafana.com/grafana/plugins/grafana-oracle-datasource/">Oracle&lt;/a>&lt;/strong>&lt;/u>,&lt;strong> &lt;/strong>&lt;u>&lt;strong>&lt;a href="https://grafana.com/grafana/plugins/elasticsearch/">Elasticsearch&lt;/a>&lt;/strong>&lt;/u>,&lt;strong> &lt;/strong>&lt;u>&lt;strong>&lt;a href="https://grafana.com/grafana/plugins/grafana-dynatrace-datasource/">Dynatrace&lt;/a>&lt;/strong>&lt;/u>,&lt;strong> &lt;/strong>&lt;u>&lt;strong>&lt;a href="https://grafana.com/grafana/plugins/grafana-honeycomb-datasource/">Honeycomb&lt;/a>&lt;/strong>&lt;/u>,&lt;strong> &lt;/strong>&lt;u>&lt;strong>&lt;a href="https://grafana.com/grafana/plugins/alexanderzobnin-zabbix-app/">Zabbix&lt;/a>&lt;/strong>&lt;/u>, &lt;u>&lt;strong>&lt;a href="https://grafana.com/grafana/plugins/grafana-jira-datasource/">Jira&lt;/a>&lt;/strong>&lt;/u>, and &lt;u>&lt;strong>&lt;a href="https://grafana.com/grafana/plugins/grafana-mongodb-datasource/">MongoDB&lt;/a>&lt;/strong>&lt;/u> (shown below).&lt;/p>&lt;p>This makes it easier to ask a single question and get an answer that draws from across your observability stack, your databases, and your project-tracking tools—no context switching required. For example, an investigation that starts with an alert can pull in error rates from Dynatrace, query performance from Oracle and recent deployments from Jira, all in one conversation. &lt;/p>&lt;p>For each data source, Assistant queries your data using natural language, correlates signals across sources, and visualizes the results as Grafana dashboards.&lt;/p>&lt;h3>Assistant now pre-installed in Grafana Enterprise&lt;/h3>&lt;p>Grafana Assistant now &lt;u>&lt;a href="https://grafana.com/whats-new/2026-06-23-grafana-assistant-is-now-pre-installed-in-grafana-enterprise/">comes pre-installed in Grafana Enterprise&lt;/a>&lt;/u>, with no plugin installation required. If you're a Grafana Enterprise user, you can connect your Grafana Cloud account to start using Assistant right away. &lt;/p>&lt;p>If you're a Grafana OSS user, you can still get access to Assistant by installing the plugin from the &lt;u>&lt;a href="https://grafana.com/grafana/plugins/">Grafana plugin catalog&lt;/a>&lt;/u> and connecting your Grafana Cloud account.&lt;/p>&lt;h2>Faster, more flexible dashboarding&lt;/h2>&lt;p>Grafana 13.1 brings a batch of improvements that make building and exploring dashboards faster and more flexible. &lt;/p>&lt;h3>Section-level variables for rows and tabs&lt;/h3>&lt;p>&lt;em>Generally available in all editions of Grafana&lt;/em>&lt;/p>&lt;p>In Grafana 13, we introduced section-level variables, a feature that lets you apply variables to each row or tab in a dashboard, so you can reduce clutter and improve the overall organization of your dashboards. With the 13.1 release, these &lt;u>&lt;a href="https://grafana.com/whats-new/2026-06-11-section-level-variables-for-rows-and-tabs-now-generally-available/">variable types are now generally available&lt;/a>&lt;/u>.&lt;/p>&lt;p>Traditionally, dashboard variables have applied to the whole dashboard at once: if you changed an &lt;code>$instance&lt;/code> variable, for example, this would update every panel together. That was a big limitation when a single dashboard spanned more than one service, such an API gateway and a database. Teams would have to split services across separate dashboards just to give each its own filters, making it difficult to achieve a unified view of their data.&lt;/p>&lt;p>Section-level variables solve for this. Each row or tab can now carry its own independent variables, so an API gateway row can scope to one set of instances while a database row scopes to another, all in the same dashboard.&lt;/p>&lt;h3>A revamped query editor &lt;/h3>&lt;p>&lt;em>Available in public preview in all editions of Grafana&lt;/em>&lt;/p>&lt;p>The &lt;u>&lt;a href="https://grafana.com/blog/grafana-13-release-all-the-latest-features/#build-complex-queries-faster">improved query editor experience&lt;/a>&lt;/u> we introduced as private preview in Grafana 13, which makes complex panels easier to build and manage, is &lt;u>&lt;a href="https://grafana.com/whats-new/2026-06-19-revamped-query-editor--now-in-public-preview-with-multi-select-and-stacked-view/">now in public preview&lt;/a>&lt;/u>, with two new capabilities:&lt;/p>&lt;ol>&lt;li>&lt;strong>Multi-select with bulk action&lt;/strong>: You no longer have to manage queries, expressions, and transformations one at a time. Instead, you can click &lt;strong>Select… &lt;/strong>in the sidebar footer to enter multi-select mode. From there, you can check the specific items you want to work with. A new bulk actions bar also lets you delete, hide, or show several queries at once, switch the data source for multiple queries in a single step, or enable and disable transformations in bulk.&lt;/li>&lt;li>&lt;strong>Stacked view&lt;/strong>: When you want to see the whole pipeline at once, the stacked view lays out all of your queries, expressions, and transformations in a single scrollable list. &lt;/li>&lt;/ol>&lt;p>Overall, these two new features combined with incremental improvements to the original release make creating and editing complex queries more straight-forward and faster than ever before.&lt;/p>&lt;h3>More dashboarding and visualization updates&lt;/h3>&lt;p>In addition to the new query editor and section-level variables, other data visualization updates in Grafana 13.1 include:&lt;/p>&lt;ul>&lt;li>&lt;u>&lt;strong>&lt;a href="https://grafana.com/whats-new/2026-05-13-quick-filters-and-data-grouping-are-now-generally-available/">Quick filters and data grouping&lt;/a>&lt;/strong>&lt;/u>: The new &lt;strong>Filter and Group by&lt;/strong> dashboard control combines filtering and grouping in one place, so exploring data is faster and more intuitive.&lt;/li>&lt;li>&lt;u>&lt;strong>&lt;a href="https://grafana.com/whats-new/2026-05-08-faceted-filter-for-time-series-legends/">Series visibility through time series legends&lt;/a>&lt;/strong>&lt;/u>: The new&lt;strong> Series visibility &lt;/strong>filter in the time series visualization lets you narrow visible series interactively, by name, label, or both, without touching the underlying query. &lt;/li>&lt;li>&lt;u>&lt;strong>&lt;a href="https://grafana.com/whats-new/2026-05-09-flexible-grouping-rules-and-field-overrides-for-nested-tables/">Enhancements to nested tables&lt;/a>&lt;/strong>&lt;/u>: Table panels with nested rows are now much more configurable, with cell styling and improvements to aggregation.&lt;/li>&lt;li>&lt;u>&lt;strong>&lt;a href="https://grafana.com/whats-new/2026-05-26-copy-and-paste-panel-styles-are-now-generally-available/">Copy and paste panel styles&lt;/a>&lt;/strong>&lt;/u>: In a couple clicks, you can replicate colors, line styles, and more from one panel to another.&lt;/li>&lt;li>&lt;u>&lt;strong>&lt;a href="https://grafana.com/whats-new/2026-05-26-panel-styles-are-now-generally-available/">Panel style presets&lt;/a>&lt;/strong>&lt;/u>&lt;strong> &lt;/strong>(shown below): Apply curated colors, thresholds, and display options to time series, stat, gauge, bar gauge, and bar chart panels with a single click in the panel editor.&lt;/li>&lt;/ul>&lt;p>For more details, check out our &lt;u>&lt;a href="https://grafana.com/docs/grafana/latest/whatsnew/whats-new-in-v13-1/">What’s New&lt;/a>&lt;/u> and &lt;u>&lt;a href="https://grafana.com/docs/grafana/latest/visualizations/">data visualization docs&lt;/a>&lt;/u>. &lt;/p>&lt;h2>PDC support for more data sources&lt;/h2>&lt;p>&lt;em>Generally available in Grafana Cloud &lt;/em>&lt;/p>&lt;p>With &lt;u>&lt;a href="https://grafana.com/docs/grafana-cloud/connect-externally-hosted/private-data-source-connect/">Private Data Source Connect (PDC)&lt;/a>&lt;/u>, you can create a private, encrypted tunnel between your Grafana Cloud stack and data sources running inside private networks, VPCs, or on-premises environments. &lt;/p>&lt;p>With Grafana 13.1, we’ve added PDC support for three new data sources: &lt;u>&lt;strong>&lt;a href="https://grafana.com/grafana/plugins/grafana-mqtt-datasource/">MQTT&lt;/a>&lt;/strong>&lt;/u>, &lt;u>&lt;strong>&lt;a href="https://grafana.com/grafana/plugins/grafana-github-datasource/">GitHub&lt;/a>&lt;/strong>&lt;/u>, and&lt;strong> &lt;/strong>&lt;u>&lt;strong>&lt;a href="https://grafana.com/grafana/plugins/grafana-ibmdb2-datasource/">IBM Db2&lt;/a>&lt;/strong>&lt;/u>.&lt;/p>&lt;p>With this update, you can connect Grafana Cloud to MQTT brokers for real-time IoT and sensor data; GitHub Enterprise Server instances for source control and project metrics; and IBM Db2 databases running on-premises or in a private cloud. &lt;/p>&lt;p>To get started, deploy a PDC agent inside your private network, then configure your data source in Grafana Cloud using its internal DNS name. To learn more, refer to the &lt;a href="https://grafana.com/docs/learning-paths/private-data-source-connect/">Private Data Source Connect learning path&lt;/a>.&lt;/p>&lt;h2>Learn more about Grafana&lt;/h2>&lt;p>For an in-depth list of all the new features in Grafana, check out our &lt;u>&lt;a href="https://grafana.com/docs/grafana/latest/?pg=blog&amp;plcmt=body-txt">Grafana documentation&lt;/a>&lt;/u>, the &lt;u>&lt;a href="https://github.com/grafana/grafana/blob/main/CHANGELOG.md">Grafana changelog&lt;/a>&lt;/u>, or our &lt;u>&lt;a href="https://grafana.com/docs/grafana/latest/whatsnew/whats-new-in-v13-1/">What's New documentation&lt;/a>&lt;/u>.&lt;/p>&lt;h2>Join the Grafana Labs community&lt;/h2>&lt;p>We invite you to engage with the &lt;u>&lt;a href="https://community.grafana.com/?pg=blog&amp;plcmt=body-txt">Grafana Labs community forums&lt;/a>&lt;/u>. Share your experiences with the new features, discuss best practices, and explore creative ways to integrate these updates into your workflows. Your insights and use cases are invaluable in enriching the Grafana ecosystem.&lt;/p>&lt;h2>Upgrade to Grafana 13.1&lt;/h2>&lt;p>&lt;u>&lt;a href="https://grafana.com/grafana/download/13.1.0">Download Grafana 13.1&lt;/a>&lt;/u> today or experience all the latest features by signing up for Grafana Cloud, which offers an actually useful forever-free tier and plans for every use case. Sign up for a &lt;u>&lt;a href="https://grafana.com/auth/sign-up/create-user/">free Grafana Cloud account&lt;/a>&lt;/u> today. &lt;/p>&lt;p>Our &lt;u>&lt;a href="https://grafana.com/docs/grafana/latest/upgrade-guide/?pg=blog&amp;plcmt=body-txt">Grafana upgrade guide&lt;/a>&lt;/u> also provides step-by-step instructions for those looking to upgrade from an earlier version to ensure a smooth transition.&lt;/p>&lt;h2>Special thanks to our community&lt;/h2>&lt;p>We extend our heartfelt gratitude to the &lt;u>&lt;a href="https://grafana.com/blog/2023/12/12/the-story-of-grafana-documentary-the-community-behind-the-code/?pg=grafana-12-4-release-all-the-latest-features&amp;plcmt=in-text">Grafana community&lt;/a>&lt;/u>! &lt;/p>&lt;p>Your contributions, ranging from pull requests to valuable feedback, are crucial in continually enhancing Grafana. And your enthusiasm and dedication inspire us at Grafana Labs to persistently innovate and elevate the Grafana platform. &lt;/p>&lt;p>&lt;u>&lt;em>&lt;a href="https://grafana.com/products/cloud/?pg=blog&amp;plcmt=body-txt">Grafana Cloud&lt;/a>&lt;/em>&lt;/u>&lt;em> is the easiest way to get started with metrics, logs, traces, dashboards, and more. We have a generous forever-free tier and plans for every use case. &lt;/em>&lt;u>&lt;em>&lt;a href="https://grafana.com/auth/sign-up/create-user/?pg=blog&amp;plcmt=body-txt">Sign up for free now&lt;/a>&lt;/em>&lt;/u>&lt;em>! &lt;/em>&lt;/p></description></item><item><title>Post-incident review for TanStack npm supply chain ransom incident: No unauthorized access to customer production systems</title><link>https://grafana.com/blog/post-incident-review-for-tanstack-npm-supply-chain-ransom-incident/</link><pubDate>Tue, 23 Jun 2026 20:24:27</pubDate><author>Joe McManus</author><guid>https://grafana.com/blog/post-incident-review-for-tanstack-npm-supply-chain-ransom-incident/</guid><description>&lt;p>On May 27, we completed our internal investigation of the &lt;a href="https://grafana.com/blog/grafana-labs-security-update-latest-on-tanstack-npm-supply-chain-ransomware-incident/">recent TanStack supply chain ransom incident&lt;/a> and confirmed our initial findings: The incident was strictly limited to Grafana Labs' GitHub environment. There was no unauthorized access to customer production systems, and the Grafana Cloud platform was not affected. &lt;/p>&lt;p>For an additional, independent audit, we engaged Mandiant, a leader in cybersecurity and incident response. We provided them with API access to Grafana Labs' log environment to conduct queries across our systems for their investigation, which started on June 1. &lt;strong>Mandiant confirmed that there was “no evidence of code tampering or repository poisoning within public organizations or production repositories delivered to end users.” &lt;/strong>&lt;/p>&lt;p>Since we discovered the incident, the Grafana Labs security teams have been running two parallel workstreams: completing the investigation and hardening our security operations. We are publishing this blog in the spirit of transparency to share more details about our incident response and remediation efforts.&lt;/p>&lt;h2>Summary and impact &lt;/h2>&lt;p>If you’re looking for the short version instead of reading &lt;a href="https://grafana.com/blog/grafana-labs-security-update-latest-on-tanstack-npm-supply-chain-ransomware-incident/">our previous updates&lt;/a>, here is the TL;DR: The TanStack supply chain attack hit us on May 11 via the Mini Shai-Hulud campaign. At the time, we believed we had successfully rotated every credential involved in this incident. We missed one. I won’t blame this oversight on hubris; the data we had at the time simply led us to believe our rotation was exhaustive. We were mistaken.&lt;/p>&lt;p>A bad actor utilized that overlooked credential to clone our entire repository collection. They then reached out on May 16, demanding a ransom to prevent a code leak.&lt;/p>&lt;p>Since Grafana Labs is an open source company, you might wonder why this is a concern. While most of our source code is public, we do maintain private repos for things like internal tools and specific Grafana Cloud features. It was a heavy decision, but we stuck to our principles and the &lt;a href="https://www.fbi.gov/how-we-can-help-you/scams-and-safety/common-frauds-and-scams/ransomware">FBI’s documented guidance&lt;/a>: We did not pay. &lt;/p>&lt;p>We launched our mitigation efforts immediately, and we confirmed that there was no unauthorized access to customer production systems, and the Grafana Cloud platform was not affected. We also confirmed that while our codebase was downloaded, it was not altered. Our customers and open source users do not need to take any action.&lt;/p>&lt;h2>Grafana Labs’ response &lt;/h2>&lt;p>We were alerted to the incident on a Saturday, and teams across the entire company took action quickly and decisively. (Or to borrow a phrase from one my favorite rappers Big Daddy Kane, ain't no half-stepping at Grafana Labs.)&lt;/p>&lt;p>In response, Grafana Labs suspended all GitHub applications on May 17, initiated a global code freeze on May 18, and conducted a cross-platform audit of Vault, GitHub, Okta, Kubernetes, AWS, GCP, and host logs to verify that no production customer data was compromised.&lt;/p>&lt;p>In the weeks following, our engineering teams contributed to a comprehensive audit that included but was not limited to:&lt;/p>&lt;ul>&lt;li>Completing 1,500 security-focused PR reviews&lt;/li>&lt;li>Auditing 280 GitHub applications, stripping permissions and removing several&lt;/li>&lt;li>Scanning 1,200 repositories for any signs of tampering&lt;/li>&lt;li>Executing 2,300 PR reviews looking for unauthorized changes in a single critical repo&lt;/li>&lt;li>Finishing infrastructure audits and retiring legacy systems&lt;/li>&lt;li>Performing wide-ranging new access audits&lt;/li>&lt;/ul>&lt;p>It was a massive undertaking, but each team stepped up in an extraordinary way to do their part. Engineering, security, and cross-functional partners worked tirelessly to respond, demonstrating the collaboration and the shared commitment we have to our community and our customers that I have always valued here at Grafana Labs. &lt;/p>&lt;p>After the initial assessment, we found that in addition to source code, the downloaded content included GitHub repositories that some Grafana Labs teams use to collaborate on and store internal operational information and other details about our business. This includes, for example, business contact names and email addresses that would be exchanged in a professional setting and email addresses that were used in some past marketing campaigns. This was not information pulled from or processed through the use of production systems or the Grafana Cloud platform. &lt;/p>&lt;p>If you wish to know if email addresses with your domain were identified, please reach out to Grafana Labs support. &lt;/p>&lt;h2>Incident timeline&lt;/h2>&lt;p>All times are in UTC &lt;/p>&lt;ul>&lt;li>19:21 11 May - First malicious code executed on self-hosted runners by Shai Hulud threat actors, leaking credentials. Rotated credentials.&lt;/li>&lt;li>07:21 14 May - First malicious commit made by the threat actor using grafana-delivery-bot, leaked from Shai Hulud attackers.&lt;/li>&lt;li>13:28 14 May - Data exfiltration of repos begins.&lt;/li>&lt;li>20:57 15 May - Data extortion threat actor publishes their extortion demand.&lt;/li>&lt;li>08:30 16 May - Grafana Labs security team becomes aware of the claimed ransom and begins seeking confirmation.&lt;/li>&lt;li>17:39 16 May - Compromise confirmed; incident declared.&lt;/li>&lt;li>19:33 16 May - All known affected credentials and GitHub applications suspended/rotated. Suspension and rotation of all other GitHub applications and accessible credentials begins.&lt;/li>&lt;li>21:10 16 May - Suspension of all GitHub applications completed.&lt;/li>&lt;li>16:40 17 May - All code changes made by GitHub application accounts associated with the threat actor identified and reverted.&lt;/li>&lt;li>16:52 17 May - Root cause, attack chain of compromise identified.&lt;/li>&lt;li>17:21 17 May - DockerHub credentials determined not compromised.&lt;/li>&lt;li>17:51 17 May - All malicious workflow runs identified. Final list of affected secrets compiled and rotated. Rotation of all other ci/common secrets from affected repos continues. &lt;/li>&lt;li>23:23 17 May - Last of the potentially accessible credentials confirmed rotated or suspended.&lt;/li>&lt;li>03:08 18 May - Begin freeze of all non-critical code and deployment changes.&lt;/li>&lt;li>08:00 25 May - All-engineering security hardening week commences.&lt;/li>&lt;li>10:58 26 May - Commit review completed, service thawing begins. A repository needs to have been fully reviewed and transitioned to use a GitHub application token broker for short-term, finely-scoped credentials before being thawed. &lt;/li>&lt;li>10:54 27 May - Transition from repos directly pushing images to DockerHub to pushing to Google Cloud Artifact Registry occurs.&lt;/li>&lt;li>27 May - Internal investigation complete. No additional attack activity or compromised credentials were discovered. &lt;/li>&lt;li>08:00 2 June - All-engineering security hardening week concludes. &lt;/li>&lt;li>20:43 3 June - Review of repositories for data loss completed. &lt;/li>&lt;li>18 June - Mandiant investigation completed, corroborating internal investigation. &lt;/li>&lt;/ul>&lt;h2>What’s next &lt;/h2>&lt;p>The investigation is now closed, but our work to improve security operations at Grafana Labs will continue. Dostoyevsky once noted that "when reason fails, the devil helps!" I’m quoting “Crime and Punishment” to underscore our philosophy: We only wanted to implement changes that actually moved the needle on security. &lt;/p>&lt;p>We’ve spent the past month executing high-impact controls, including a token broker, fine-grained access controls, additional alerting, and static analysis. In addition, we have moved off of certain GitHub Actions and now use more tightly scoped actions with short-lived tokens. &lt;/p>&lt;p>We have also started the process of compartmentalizing our GitHub organizations and isolating all archived repos into a dedicated organization with actions disabled.&lt;/p>&lt;p>We will share an overview of our response efforts and the technical details of how we improved our security posture from our post-incident review in the coming weeks.&lt;/p></description></item><item><title>ObservabilityCON 2026 is coming to San Francisco!</title><link>https://grafana.com/blog/observabilitycon-2026-is-coming-to-san-francisco/</link><pubDate>Thu, 11 Jun 2026 21:50:09</pubDate><author>Grafana Labs Team</author><guid>https://grafana.com/blog/observabilitycon-2026-is-coming-to-san-francisco/</guid><description>&lt;p>ObservabilityCON 2026 is heading to San Francisco, and even as the fog rolls in over the Golden Gate Bridge, the future of observability will never be clearer. Join us at our &lt;a href="/events/observabilitycon/">flagship observability event&lt;/a> at Pier 27 from October 19–21, 2026, where the community will gather at the epicenter of tech to navigate the fast-moving AI era.&lt;/p>&lt;p>Whether you're building agentic workflows on your local machine or running them in production, you’ll get access to the latest innovations, thought leaders, and observability experts. This year's annual event will feature:&lt;/p>&lt;ul>&lt;li>An opening keynote exploring what's new—and what's next—in the observability space, including a look at the latest Grafana Cloud features and AI-powered solutions.&lt;/li>&lt;li>Live demos of the newest agentic tools and workflows in Grafana Cloud.&lt;/li>&lt;li>One-on-one time with engineers, peer success stories, and a clear view of where observability is heading.&lt;/li>&lt;li>Hands-on workshops and technical deep-dives related to AI observability, digital experience monitoring, and more. &lt;/li>&lt;li>Tons of networking opportunities to connect with hundreds of your closest observability friends.&lt;/li>&lt;/ul>&lt;p>Stay tuned for more details on the full &lt;a href="/events/observabilitycon/">ObservabilityCON 2026&lt;/a> agenda coming in July.&lt;/p>&lt;p>&lt;/p>&lt;h2>How to register&lt;/h2>&lt;p>Registration isn't open just yet, but you can &lt;a href="/events/observabilitycon/">sign up now for early access&lt;/a> to be first in line for tickets and discount pricing. Limited tickets will be available at 50% off, so you'll want to secure your spot before they're gone.&lt;/p>&lt;p>&lt;/p>&lt;h2>Speaking and sponsorship opportunities&lt;/h2>&lt;p>Have an observability story to tell? Whether your team reduced noise, cut costs, adopted OpenTelemetry, migrated to Grafana Cloud, built AI into your workflows, or overcame a major challenge, we'd love to put your story on the ObservabilityCON stage and have you share lessons learned and the big wins your team experienced. &lt;/p>&lt;p>The &lt;a href="https://pretalx.com/observabilitycon-2026/cfp">call for presentations (CFP)&lt;/a> for ObservabilityCON 2026 speakers is now open and you have until &lt;strong>July 1&lt;/strong> to submit your story for our San Francisco agenda. Need a little inspiration first? Check out these &lt;a href="https://grafana.com/videos/?language=en&amp;type=on-demand&amp;events=observabilitycon&amp;years=2025">on-demand talks &lt;/a>from last year's ObservabilityCON event.&lt;/p>&lt;p>We also have limited sponsorship opportunities available for ObservabilityCON 2026. Please email sponsors@grafana.com to learn more.&lt;/p>&lt;p>&lt;/p>&lt;h2>Can’t make it to San Francisco?&lt;/h2>&lt;p>Don’t worry: we're taking ObservabilityCON on a world tour.&lt;/p>&lt;p>&lt;a href="/events/observabilitycon-on-the-road/">ObservabilityCON on the Road&lt;/a> is coming to four cities this fall: São Paulo, London, Madrid, and Bengaluru. This one-day event is packed with technical deep dives, hands-on AI demos, and an Ask the Experts booth.&lt;/p>&lt;p>Here are the dates for this year's stops:&lt;/p>&lt;ul>&lt;li>&lt;a href="/events/observabilitycon-on-the-road/sao-paulo">São Paulo&lt;/a>: November 4 (presented in Portuguese)&lt;/li>&lt;li>&lt;a href="/events/observabilitycon-on-the-road/london">London&lt;/a>: November 5&lt;/li>&lt;li>&lt;a href="/events/observabilitycon-on-the-road/madrid">Madrid&lt;/a>: November 24 (presented in Spanish)&lt;/li>&lt;li>&lt;a href="/events/observabilitycon-on-the-road/bengaluru">Bengaluru&lt;/a>: December 8&lt;/li>&lt;/ul>&lt;p>Plus we will be adding more cities to the line-up soon! &lt;/p>&lt;p>Preview AI-powered solutions, deepen your observability expertise, drive bigger business impact, and connect with Grafana Labs experts face to face. &lt;a href="/events/observabilitycon-on-the-road/">Find your event&lt;/a> and sign up for early access.&lt;/p>&lt;p>&lt;em>The countdown to ObservabilityCON 2026 has officially begun! &lt;/em>&lt;em>&lt;a href="/events/observabilitycon/">Sign up today &lt;/a>&lt;/em>&lt;em>so we can keep you posted on all the latest updates. &lt;/em>&lt;/p></description></item><item><title>Automatically discover and remediate root causes with Grafana Assistant Investigations</title><link>https://grafana.com/blog/automatically-discover-and-remediate-root-causes-with-grafana-assistant-investigations/</link><pubDate>Tue, 09 Jun 2026 15:04:37</pubDate><author>Maurice Rochau</author><guid>https://grafana.com/blog/automatically-discover-and-remediate-root-causes-with-grafana-assistant-investigations/</guid><description>&lt;p>You can use &lt;u>&lt;a href="https://grafana.com/docs/grafana-cloud/machine-learning/assistant/guides/investigation/">Grafana Assistant Investigations&lt;/a>&lt;/u> to automatically discover incidents and help find root causes—and this AI-powered Grafana Cloud feature recently got a major upgrade to give you even more confidence in its findings. &lt;/p>&lt;p>You can read more about the behind-the-scenes effort in our new engineering blog &lt;em>&lt;a href="https://medium.com/grafana-labs/inside-the-harness-how-grafana-assistant-investigates-incidents-9a982b8ff01d">Unprompted&lt;/a>&lt;/em>, where we get into harness engineering, context compaction, benchmarking, and keeping agents alive and working well in long-running sessions. In this post, we'll focus on how you can get the most out of this product iteration in conjunction with our other features. &lt;/p>&lt;p>Keep reading to learn about all the ways Assistant Investigations, currently in public preview, can help you improve your incident response so you can run your own “human on the loop” auto-remediation workflows at scale in Grafana Cloud.&lt;/p>&lt;h2>Investigate anything with Assistant Investigations&lt;/h2>&lt;p>At its core, Assistant Investigations is a highly sensitive and tuned background agent that’s capable of investigating anything within the observability space or developer lifecycle. It’s your problem finder and validator.&lt;/p>&lt;ul>&lt;li>&lt;strong>Find instrumentation gaps in your setup:&lt;/strong> Task Assistant Investigations with looking at your metrics, logs, traces, profiles, services, and labels, and correlate that with your code to spot any improvements for your setup.&lt;/li>&lt;li>&lt;strong>Define degradation criteria and evaluate against them:&lt;/strong> Build a skill with Grafana Assistant that captures certain criteria you want to meet. Schedule AI-assisted investigations on top of them and get a report every day if you’re still within your operational parameters.&lt;/li>&lt;li>&lt;strong>Use profiles to raise PRs to make your software faster:&lt;/strong> Kick off an AI-assisted investigation with the purpose of looking at profiles and improving latency in your login and registration service. Based on a correlation of telemetry data, you’ll get improvements posted to GitHub either as issue, PR, PR draft, or branch.&lt;/li>&lt;li>&lt;strong>Analyze your user drop-off rate across services:&lt;/strong> As long as the data is available, Assistant Investigations can also look at business metrics for you. That way, you can correlate high latency with churn in your checkout process or timeouts on your payment provider with chargebacks. You connect the dots between business impact, value capture, and your telemetry.&lt;/li>&lt;/ul>&lt;h3>Customize Assistant Investigations to fit your specific needs&lt;/h3>&lt;p>Assistant Investigations also gives you full customization, just like &lt;u>&lt;a href="https://grafana.com/docs/grafana-cloud/machine-learning/assistant/get-started/">Grafana Assistant&lt;/a>&lt;/u>. And it follows our "big tent" philosophy, so you can use the tools and agents you prefer. For example, you can:&lt;/p>&lt;ul>&lt;li>Use &lt;u>&lt;a href="https://grafana.com/blog/add-skills-to-agents-use-assistant-skills-for-faster-answers-investigations/">skills&lt;/a>&lt;/u> with auto-approved tools in them that it will discover and use&lt;/li>&lt;li>Use &lt;u>&lt;a href="https://grafana.com/docs/grafana-cloud/machine-learning/assistant/configure/mcp-servers/">MCP server integrations&lt;/a>&lt;/u> to wire up your entire stack to Grafana Cloud and give Assistant Investigations access to all necessary systems&lt;/li>&lt;li>Connect to code via GitHub and GitLab; bring in business data from Snowflake or Salesforce; look at feature flags in LaunchDarkly; or manage CI/CD with Jenkins.&lt;/li>&lt;/ul>&lt;h2>Common use cases for Assistant Investigations&lt;/h2>&lt;p>Next, let's look at some of the use cases that can make Assistant Investigations so valuable to your observability practice. &lt;/p>&lt;h3>Multiplayer your problem&lt;/h3>&lt;p>With Assistant Investigations, there’s no need to tackle problems alone. When your on-call colleague kicks off an investigation for a problem and pages another team, the other team can easily jump into the conversation. They can retrieve the investigation, steer it, ask follow-up questions, or provide valuable context.&lt;/p>&lt;p>With the Assistant &lt;u>&lt;a href="https://grafana.com/docs/grafana-cloud/machine-learning/assistant/workspace/">workspace&lt;/a>&lt;/u> and its canvas, multiple users can put down relevant information so that whoever reviews the investigation during the incident or after it can see what’s going on.&lt;/p>&lt;h3>Kick start your response by connecting to alerts and incidents&lt;/h3>&lt;p>Incidents don't happen on your schedule, and every minute counts when you need to get your systems back online. With Assistant Investigations, you can easily integrate with Grafana Alerting and Grafana Cloud IRM to kick off an investigation when alerts fire or incidents are declared. &lt;/p>&lt;p>If you have GitHub, GitLab, or Cursor configured, you can also close the loop and raise PRs automatically.You can even configure Assistant to send Slack messages into a group channel to nudge colleagues to review the auto-fix.&lt;/p>&lt;p>With Alerting’s outgoing webhooks, you can go further down the customization route by sending a POST request to another agent that has access to Grafana Cloud via our &lt;a href="https://grafana.com/docs/grafana/latest/developer-resources/mcp/">hosted MCP server&lt;/a> or &lt;a href="https://grafana.com/blog/get-observability-in-the-terminal-for-you-and-your-agents-with-the-gcx-cli-tool/">gcx&lt;/a>. This allows you to kick off any agentic workflow to act on alerts how you need to.&lt;/p>&lt;p>And  the next time an alert fires and the investigation comes back with a false-positive, just ask Assistant to improve your alert.&lt;/p>&lt;h3>Catch up on what changed overnight&lt;/h3>&lt;p>Thanks to agentic coding, software moves faster than ever—and we’re here to move fast with you. With Assistant's &lt;u>&lt;a href="https://grafana.com/blog/spend-less-time-on-repetitive-tasks-with-the-new-automation-feature-in-grafana-assistant">automation capabilities&lt;/a>&lt;/u>, you can kick off structured investigations on a time basis and have the reports sent back to Slack to you personally or in a channel. &lt;/p>&lt;p>Don't let your users be the first one to tell you when something breaks. Even if you don't have the right alert set up, the next time your checkout service sees increased latency and churn, you’ll catch it faster. And if you set up a GitHub or GitLab integration, you can create the PR to fix it right from Slack&lt;/p>&lt;h2>One piece of the larger Assistant ecosystem&lt;/h2>&lt;p>Now let's look at a practical example that highlights how versatile the Assistant ecosystem is and where Assistant Investigations fits in. If you're not using Assistant already, you can read along to get a sense for what it can do. Or, if you're already using Assistant, you can follow these same steps in your environment.&lt;/p>&lt;p>To do so, we'll use the workspace feature, which lets you manage multiple Assistant conversations from a dedicated page.&lt;/p>&lt;ol>&lt;li>Open Assistant and select &lt;strong>Open in Workspace&lt;/strong> from the conversation menu.&lt;/li>&lt;li>Ask about a service of your choice. For the purposes of this example, I picked a demo payment service, but you can run this with any service you have enabled to support Assistant and Assistant Investigations.&lt;/li>&lt;li>Ask Assistant to explore your service and write a skill for it.&lt;/li>&lt;/ol>&lt;p>Congrats, you’ve customized Assistant with your first skill! Next time you talk about the payment service, the context from the skill gets pulled in.&lt;/p>&lt;p>Next, let's re-use that skill in Assistant Investigations.&lt;/p>&lt;ol>&lt;li>Go to &lt;strong>Assistant&lt;/strong> > &lt;strong>Settings&lt;/strong> > &lt;strong>Integrations&lt;/strong> >&lt;strong> IRM webhooks&lt;/strong> and configure it for the incidents and alerts you want Assistant Investigation to run on.&lt;/li>&lt;li>Go to &lt;strong>Assistant&lt;/strong> > &lt;strong>Settings&lt;/strong> > &lt;strong>Custom rules&lt;/strong> and add a rule scoped to Investigations that says to always search for skills when an AI-assisted investigation starts and to pick the service-relevant skills.&lt;/li>&lt;/ol>&lt;p>Whenever Assistant Investigation is triggered, you can be sure it follows your runbook now.&lt;/p>&lt;p>Note: You can set custom rules and skills to “Just me” or “Everybody.” With “Everybody,” everyone profits from your setup without having to set up anything themselves. Once configured, it applies to the whole stack.&lt;/p>&lt;p>Finally, let’s layer on some automations.&lt;/p>&lt;ol>&lt;li>Ask Assistant to create an automation for you that uses your skill and runs every morning at 8 a.m.&lt;/li>&lt;li>Tell it to send you a Slack DM.&lt;/li>&lt;/ol>&lt;p>From now on, every morning at 8 a.m., your skill is executed and all out-of-the-ordinary findings are reported.&lt;/p>&lt;p>Note: In my example, I didn’t connect Slack on purpose to highlight how Assistant actively helps you configure certain parts of it as they become necessary.&lt;/p>&lt;p>This is just one example of all the ways you can use Assistant and Assistant Investigations. You'll get the best outcomes by playing around and using the different features in conjunction with each other, so start testing it out today!&lt;/p></description></item><item><title>How to generate real-world load tests using Grafana Cloud k6 and production telemetry</title><link>https://grafana.com/blog/how-to-generate-real-world-load-tests-using-grafana-cloud-k6-and-production-telemetry/</link><pubDate>Wed, 03 Jun 2026 16:14:15</pubDate><author>Matt Wimpelberg</author><guid>https://grafana.com/blog/how-to-generate-real-world-load-tests-using-grafana-cloud-k6-and-production-telemetry/</guid><description>&lt;p>For many development teams, a load test starts with a set of assumptions. &lt;/p>&lt;p>You pick 100 virtual users because it sounds reasonable. You ramp for 30 seconds because that's what the tutorial showed. You set a 500ms threshold because it feels like a good target. The test passes, you ship the release, and production falls over at 6 p.m. on a Tuesday because your synthetic load never resembled how real users interact with your application. &lt;/p>&lt;p>The good news is that, if you're running &lt;u>&lt;a href="https://grafana.com/products/cloud/">Grafana Cloud&lt;/a>&lt;/u>, you already have the data you need to run realistic load tests. Your dashboards capture how your users behave, including request rates, latency distributions, and traffic patterns over time. The signal is already there; you just need to connect it to your test configuration.&lt;/p>&lt;p>Using production telemetry to shape your tests makes the results far more meaningful. Instead of validating against a hypothetical workload, you're testing against patterns your systems already experience, leading to more reliable baselines, better thresholds, and fewer surprises in production. &lt;/p>&lt;p>This post walks through how to do exactly that: pull real telemetry signals from Grafana Cloud and translate them into a &lt;u>&lt;a href="https://grafana.com/products/cloud/performance-load-testing-k6/">Grafana Cloud k6&lt;/a>&lt;/u> testing scenario that reflects production environments (rather than assumptions). &lt;/p>&lt;h2>VU count isn’t the only place to start&lt;/h2>&lt;p>&lt;u>&lt;a href="https://grafana.com/docs/k6/latest/reference/glossary/#virtual-user">Virtual user (VU)&lt;/a>&lt;/u> count is an important part of how Grafana Cloud k6, the fully managed performance testing platform powered by k6 OSS, runs tests. However, VU count doesn’t always need to be the first lever you pull. When your goal is to test how a service behaves under a known request rate, arrival-rate executors let you express that intent more directly.&lt;/p>&lt;p>What you may care about is the &lt;em>arrival&lt;/em>&lt;em>&lt;strong> &lt;/strong>&lt;/em>&lt;em>rate&lt;/em>, or how many requests per second your system handles. VU count is a byproduct of arrival rate and response time. If your p95 latency doubles, you need twice as many VUs to sustain the same throughput. Start from VUs and you're tuning the wrong variable.&lt;/p>&lt;p>This distinction often comes down to API testing vs. website testing. There is typically a difference between trying to simulate a number of "real users," including their actions and pauses and how those translate into backend capacity, vs. simulating a certain throughput on API endpoints. While the underlying test is ultimately similar, how you derive those traffic levels depends on the type of telemetry data you're working with. For instance, analyzing and simulating full user sessions would align more with &lt;u>&lt;a href="https://grafana.com/products/cloud/frontend-observability/">Grafana Cloud Frontend Observability&lt;/a>&lt;/u>, whereas the metrics described below are focused on deriving the traffic levels of your API endpoints.&lt;/p>&lt;p>k6 has two executor types that map directly to how real traffic works:&lt;/p>&lt;ul>&lt;li>&lt;code>constant-arrival-rate&lt;/code>:  a fixed number of iterations per second, regardless of how long each one takes. Use this when you want to hold a steady request rate and observe what happens to latency under that pressure.&lt;/li>&lt;li>&lt;code>ramping-arrival-rate&lt;/code>:  arrival rate changes over time according to stages you define. Use this when your traffic has a shape: a morning ramp, an evening peak, a midday lull.&lt;/li>&lt;/ul>&lt;p>Both of these need real numbers to be useful. Here's how to get them.&lt;/p>&lt;h2>&lt;strong>Step 1: Find your actual request rate&lt;/strong>&lt;/h2>&lt;p>Open Grafana Cloud and run this query against your &lt;a href="https://prometheus.io/">Prometheus&lt;/a> or &lt;u>&lt;a href="https://grafana.com/oss/mimir/">Mimir&lt;/a>&lt;/u> data source. Swap in your service name:&lt;/p>&lt;pre>&lt;code>sum(rate(http_server_requests_total{app="leaderboard-api"}[$__rate_interval]))&lt;/code>&lt;/pre>&lt;p>Look at this over a representative window. A full week is better than a single day. You're looking for two things:&lt;/p>&lt;ol>&lt;li>&lt;strong>Your baseline rate&lt;/strong> during normal operation (requests per second)&lt;/li>&lt;li>&lt;strong>Your peak rate&lt;/strong>, which is the highest sustained load you've actually seen&lt;/li>&lt;/ol>&lt;p>That peak number is your load test target. Not a guess or just 2x your average, but the actual peak your system has handled (or needs to handle).&lt;/p>&lt;p>If your baseline is 120 req/s and your peak is 340 req/s, you now have real numbers to work with.&lt;/p>&lt;h2>&lt;strong>Step 2: Get your latency baseline&lt;/strong>&lt;/h2>&lt;p>This becomes your threshold. Run this query to get p95 latency under normal production load:&lt;/p>&lt;pre>&lt;code>histogram_quantile(
0.95,
sum(rate(http_server_request_duration_seconds_bucket{app="leaderboard-api"}[5m])) by (le)
)
&lt;/code>&lt;/pre>&lt;p>This value becomes the check in your Grafana Cloud k6 script, for example: &lt;code>http_req_duration: ["p(95)&lt;280"]&lt;/code>.&lt;/p>&lt;p>Whatever this returns, such as 280ms, that's your threshold. If your load test pushes p95 above 280ms, the test should fail. &lt;/p>&lt;p>Also pull p99 while you're here. A gap between p95 and p99 tells you something about tail latency behavior that will matter under load.&lt;/p>&lt;h2>&lt;strong>Step 3: Read your traffic shape&lt;/strong>&lt;/h2>&lt;p>Effective traffic modeling isn't just about mimicking a 24-hour cycle; it's about modeling your test configuration after the specific performance question you want to answer. Pull your request rate panel and look at the line's shape, then choose one of these four customer-modeled patterns to guide your &lt;code>stages&lt;/code> array:&lt;/p>&lt;ul>&lt;li>&lt;strong>Constant load&lt;/strong>: Stay at a small percentage above your baseline for an extended time to verify steady-state stability.&lt;/li>&lt;li>&lt;strong>Stress&lt;/strong>: Step through increasing load levels to identify the exact breakdown point where latency or error rates spike.&lt;/li>&lt;li>&lt;strong>Spike&lt;/strong>: Ramp to your observed peak quickly for a short duration to test how the system handles sudden bursts of traffic.&lt;/li>&lt;li>&lt;strong>Endurance&lt;/strong>: Maintain a baseline load for a very long period to uncover resource leaks or performance degradation over time.&lt;/li>&lt;/ul>&lt;p>For example, a &lt;strong>stress&lt;/strong> test designed to find the breaking point might use stages that step through increasing load levels:&lt;/p>&lt;pre>&lt;code>stages: [
{ target: 60, duration: "5m" }, // 50% of baseline
{ target: 120, duration: "5m" }, // baseline
{ target: 240, duration: "5m" }, // 2x baseline
{ target: 340, duration: "5m" }, // observed peak
{ target: 0, duration: "5m" }, // ramp down and cool off
]
&lt;/code>&lt;/pre>&lt;p>In this case, you're not making this up; you're reading it off the dashboard.&lt;/p>&lt;h2>&lt;strong>Step 4: Wire it all together in Grafana Cloud k6&lt;/strong>&lt;/h2>&lt;p>Here's what a full scenario config looks like when built from real data instead of assumptions:&lt;/p>&lt;pre>&lt;code>import http from "k6/http";
import { check } from "k6";
export const options = {
scenarios: {
realistic_load: {
executor: "ramping-arrival-rate",
// Derived from observed peak concurrency in Grafana
preAllocatedVUs: 50,
maxVUs: 200,
// Derived from your traffic shape panel
stages: [
{ target: 120, duration: "5m" }, // ramp to baseline (req/s)
{ target: 340, duration: "10m" }, // push to observed peak
{ target: 120, duration: "5m" }, // return to baseline
],
},
},
thresholds: {
// Derived from production p95 — not invented
http_req_duration: ["p(95)&lt;280"],
// Derived from observed error rate in Grafana
http_req_failed: ["rate&lt;0.01"],
},
};
export default function () {
const res = http.get("https://your-service.example.com/api/health");
check(res, { "status 200": (r) => r.status === 200 });
}
&lt;/code>&lt;/pre>&lt;p>Every number in that config came from a dashboard query. When a teammate asks why you chose these values, you have an answer.&lt;/p>&lt;h2>&lt;strong>Step 5: Close the loop back in Grafana Cloud&lt;/strong>&lt;/h2>&lt;p>Run the test and send results to Grafana Cloud k6. Now open the same Grafana Cloud dashboards you pulled the baseline from and add an &lt;u>&lt;a href="https://grafana.com/docs/grafana/latest/visualizations/dashboards/build-dashboards/annotate-visualizations/">annotation&lt;/a>&lt;/u> for the time the test ran.&lt;/p>&lt;p>You're now looking at the same panels for request rate, p95 latency, and error rate during a load test that was designed to match production conditions. If the test surfaces a problem, you already know which dashboard to go to. The debugging context is built in.&lt;/p>&lt;p>This is what the observability and testing loop is supposed to look like: production data informs the test design, and test results feed back into your observability stack. Neither side is isolated.&lt;/p>&lt;h2>&lt;strong>A simpler way: Generate your script with Grafana Assistant&lt;/strong>&lt;/h2>&lt;p>While calculating all the values manually (request rate, p95 latency, stages) provides maximum control and fidelity, there is an easier path to achieving a production-ready script.&lt;/p>&lt;p>&lt;u>&lt;a href="https://grafana.com/products/cloud/ai-assistant/">Grafana Assistant&lt;/a>&lt;/u>, the AI-powered agent in Grafana Cloud, includes a &lt;strong>k6 script authoring&lt;/strong> capability that allows you to generate performance tests using natural language. This feature automates the heavy lifting of performance engineering by programmatically translating production-observed metrics—such as the exact request rates and latency thresholds identified in your panels—directly into a valid k6 configuration. Instead of manually parsing Prometheus queries and hard-coding values into a script, you can simply ask Assistant to derive a test from your service's existing traffic patterns.&lt;/p>&lt;p>Grafana Assistant analyzes your service context, including relevant metrics, logs, and traces, to mirror observed behaviors like morning ramps or evening peaks. &lt;/p>&lt;p>By automatically mapping these real-world traffic patterns to k6 arrival-rate executors and defined stages, Assistant significantly reduces manual effort while ensuring higher fidelity to production conditions. This ensures your tests aren't just syntactically correct, but are grounded in the actual pressure your systems face every day.&lt;/p>&lt;p>To learn more about k6 script authoring in Assistant, please check out our &lt;u>&lt;a href="https://grafana.com/blog/generate-test-scripts-from-natural-language-with-grafana-assistant-introducing-k6-script-authoring/">blog post&lt;/a>&lt;/u> and &lt;u>&lt;a href="https://grafana.com/docs/grafana-cloud/testing/k6/author-run/k6-script-authoring-mode/?pg=generate-test-scripts-from-natural-language-with-grafana-assistant-introducing-k6-script-authoring&amp;plcmt=in-text">technical docs&lt;/a>&lt;/u>. &lt;/p>&lt;h2>&lt;strong>Wrapping up: what this all means for shift-left testing &lt;/strong>&lt;/h2>&lt;p>&lt;u>&lt;a href="https://grafana.com/products/cloud/performance-load-testing-k6/#why-scale-metrics-with-grafana-cloud#improve-reliability-with-shift-left-testing">Shift-left testing&lt;/a>&lt;/u> is an approach where testing happens earlier into the development pipeline to catch issues before they impact users. While you can apply this approach to load testing, it doesn’t really help you if you’re testing on assumptions.&lt;/p>&lt;p>Shifting left with fidelity, however, means your pre-production tests reflect the pressure that’s actually applied by production workloads. That requires real data, which Grafana Cloud gives you, and then k6 gives you the test framework to act on it. &lt;/p>&lt;p>&lt;u>&lt;em>&lt;a href="https://grafana.com/products/cloud/?pg=blog&amp;plcmt=body-txt">Grafana Cloud&lt;/a>&lt;/em>&lt;/u>&lt;em> is the easiest way to get started with k6 and performance testing. We have a generous forever-free tier and plans for every use case. &lt;/em>&lt;u>&lt;em>&lt;a href="https://grafana.com/auth/sign-up/create-user/?src=k6io&amp;redirectPath=k6&amp;pg=blog&amp;plcmt=body-txt">Sign up for free now!&lt;/a>&lt;/em>&lt;/u>&lt;em> &lt;/em>&lt;/p></description></item><item><title>Tempo 3.0 release: a new architecture for scale and lower TCO, TraceQL metrics GA, and more</title><link>https://grafana.com/blog/tempo-3-0-release-all-the-latest-features/</link><pubDate>Mon, 01 Jun 2026 17:59:54</pubDate><author>Tiffany Jernigan</author><guid>https://grafana.com/blog/tempo-3-0-release-all-the-latest-features/</guid><description>&lt;p>Tempo started with a simple goal: make distributed tracing easier to run at scale. As tracing adoption has grown, however, so have the challenges, including higher data volumes, more complex architectures, and increasing demand for real-time insights directly from traces.&lt;/p>&lt;p>Over the last year, we’ve been evolving Tempo’s architecture to meet that moment. And today, we’re sharing the results of those efforts with the release of &lt;u>&lt;a href="https://github.com/grafana/tempo/releases/tag/v3.0.0">Tempo 3.0&lt;/a>&lt;/u>.&lt;/p>&lt;p>The latest major release of the &lt;u>&lt;a href="https://grafana.com/docs/tempo/latest/">open source distributed tracing backend&lt;/a>&lt;/u> introduces a &lt;u>&lt;a href="https://kafka.apache.org/intro">Kafka&lt;/a>&lt;/u>-compatible architecture for microservices deployments that changes how Tempo ingests, stores, and queries trace data at scale.&lt;/p>&lt;p>With the general availability of &lt;u>&lt;a href="https://grafana.com/docs/tempo/latest/metrics-from-traces/metrics-queries/">TraceQL metrics&lt;/a>&lt;/u>, Tempo 3.0 also makes it easier to create metrics and gain insights directly from tracing data. Alongside improvements to query performance, trace redaction, and more, these updates help teams deploy tracing more efficiently and cost effectively.&lt;/p>&lt;p>You can continue reading and check out the video below to learn more about these and other updates. The Tempo 3.0 &lt;u>&lt;a href="https://grafana.com/docs/tempo/latest/release-notes/v3-0/">release notes&lt;/a>&lt;/u> and &lt;u>&lt;a href="https://github.com/grafana/tempo/releases/tag/v3.0.0">changelog&lt;/a>&lt;/u> also provide more in-depth details and include all the changes included in this release.&lt;/p>&lt;h2>New Tempo architecture for lower TCO &lt;/h2>&lt;p>In the &lt;u>&lt;a href="https://grafana.com/blog/grafana-tempo-2-9-release-mcp-server-support-traceql-metrics-sampling-and-more/">Tempo 2.9 release blog&lt;/a>&lt;/u> and &lt;u>&lt;a href="https://grafana.com/docs/tempo/latest/release-notes/v2-9/?pg=blog&amp;plcmt=body-txt/#project-rhythm-new-tempo-architecture">release notes&lt;/a>&lt;/u>, we introduced a new Tempo architecture, codenamed Project Rhythm, with three goals: remove the replication factor 3 (RF3) requirement, decouple the read and write paths, and lay the groundwork for lower total cost of ownership (TCO). In practice, this means storing one copy of your trace data instead of three, which directly reduces object storage costs.&lt;/p>&lt;p>With Tempo 3.0, that architecture is now the default.&lt;/p>&lt;h3>&lt;strong>Why we evolved the Tempo architecture &lt;/strong>&lt;/h3>&lt;p>The previous Tempo architecture centered on ingesters. Distributors sent trace data to ingesters, ingesters buffered recent data and flushed blocks to object storage, and compactors handled backend block maintenance.&lt;/p>&lt;p>This created some challenges, including:&lt;/p>&lt;ol>&lt;li>Tempo needed RF3 replication across ingesters for durability, and TraceQL metrics added a second set of blocks through the &lt;u>&lt;a href="https://grafana.com/docs/tempo/latest/metrics-from-traces/metrics-generator/">metrics-generator&lt;/a>&lt;/u> and local-blocks processor. This meant storing 3 to 3.5 times the actual volume of tracing data. &lt;/li>&lt;li>TraceQL metrics had availability limitations. Metric data was tied to individual generators, so if a generator went down, the data it held became temporarily unavailable.&lt;/li>&lt;li>The read and write paths shared the same components, so one bad workload could destabilize both. Real incidents made this clear: Mimir queries took down ingestion, slow queries starved fast ones, and large Tempo trace attributes OOMed queriers.&lt;/li>&lt;/ol>&lt;p>To learn more about these and other factors leading up to Tempo’s re-architecture, check out the GrafanaCON 2026 talk, &lt;u>&lt;a href="https://www.youtube.com/watch?v=iFyGv5nOgr4&amp;list=PLDGkOdUX1UjoSfz1IRj5c0xetw8tl8iin&amp;index=16">“When theory meets scale: Operating Mimir and Tempo in production.”&lt;/a>&lt;/u>&lt;br>&lt;br>&lt;/p>&lt;p>These updates to Tempo also align with other recent architectural changes we’ve made to &lt;u>&lt;a href="https://grafana.com/blog/grafana-mimir-3-0-release-all-the-latest-updates/#a-new-decoupled-architecture">Mimir&lt;/a>&lt;/u>, &lt;u>&lt;a href="https://grafana.com/blog/grafanacon-2026-announcements/#the-evolution-of-loki">Loki&lt;/a>&lt;/u>, and &lt;u>&lt;a href="https://grafana.com/blog/pyroscope-2-0-release/">Pyroscope&lt;/a>&lt;/u> to optimize performance across the OSS projects.&lt;/p>&lt;h3>&lt;strong>Splitting the read and write path &lt;/strong>&lt;/h3>&lt;p>The solution to these issues was to fully separate the read and write paths so neither can affect the other. As a result, one of the main architectural changes in Tempo 3.0 is that the read and write paths are now split while in &lt;u>&lt;a href="https://grafana.com/docs/tempo/latest/set-up-for-tracing/setup-tempo/plan/deployment-modes/#microservices-mode">microservices mode&lt;/a>&lt;/u>, reducing storage and overhead while improving performance and reliability. &lt;/p>&lt;p>The distributor still receives trace data first, but what happens next is different. In microservices mode, the distributor writes that data to a Kafka-compatible system, which becomes the durable buffer between ingestion and the rest of the system. From there, block-builders consume trace data and write &lt;u>&lt;a href="https://parquet.apache.org/">Parquet&lt;/a>&lt;/u> blocks to object storage, while live-stores, a read-path component, consume the same stream to serve recent traces before those blocks are available in storage. This is what lets newly ingested traces show up in queries quickly, without tying recent reads to the old ingester path.&lt;/p>&lt;p>The metrics-generator now consumes trace data from the same &lt;u>&lt;a href="https://kafka.apache.org/intro/#main-concepts-and-terminology">Kafka topic&lt;/a>&lt;/u> (a named stream that multiple components can read from independently), instead of receiving it through the previous distributor push path. Compaction also moves into the new architecture: compactors are replaced by the backend scheduler and backend worker, which coordinate compaction, retention, blocklist maintenance, and redaction jobs.&lt;/p>&lt;p>For smaller deployments, &lt;u>&lt;a href="https://grafana.com/docs/tempo/latest/set-up-for-tracing/setup-tempo/plan/deployment-modes/#monolithic-mode">monolithic mode&lt;/a>&lt;/u> is still the simpler option with no Kafka required. Everything runs in a single binary.&lt;/p>&lt;p>To learn more, check out our &lt;u>&lt;a href="https://grafana.com/docs/tempo/latest/introduction/architecture/">Tempo architecture documentation&lt;/a>&lt;/u>.&lt;/p>&lt;h2>Deeper insights: TraceQL metrics is now GA&lt;/h2>&lt;p>TraceQL metrics, now generally available, lets you query ad-hoc metrics directly from trace data, making it easy to answer questions about performance, error rates, and service behavior across distributed systems. &lt;/p>&lt;p>Using TraceQL metrics, you can compute time series directly from trace data using queries such as:&lt;/p>&lt;p>&lt;/p>&lt;pre>&lt;code>{} | rate() by (resource.service.name)&lt;/code>&lt;/pre>&lt;p>With the new architecture, TraceQL metrics no longer depends on the &lt;u>&lt;a href="https://grafana.com/docs/tempo/latest/metrics-from-traces/metrics-generator/">metrics-generator component&lt;/a>&lt;/u> and is enabled by default. Recent TraceQL metrics queries are served by the live-store, and historical TraceQL metrics queries read RF1 blocks from object storage.&lt;/p>&lt;p>This matters during migration: Tempo 3.0 can query old RF3 blocks for normal trace search and trace-by-ID lookups, but TraceQL metrics queries only read RF1 blocks. If your 2.x deployment didn’t use the local-blocks processor, TraceQL metrics results cover data ingested after the upgrade.&lt;/p>&lt;p>&lt;em>Note: Alerting on TraceQL metrics remains experimental.&lt;/em>&lt;/p>&lt;h3>&lt;strong>Comparison operators in TraceQL metrics&lt;/strong>&lt;/h3>&lt;p>TraceQL metrics queries now support comparison operators: &lt;code>>&lt;/code>, &lt;code>&lt;&lt;/code>, &lt;code>>=&lt;/code>, &lt;code>&lt;=&lt;/code>, &lt;code>=&lt;/code>, and &lt;code>!=&lt;/code>.  This lets you filter metric query results directly, so you can focus on the series that cross a threshold instead of scanning every returned series manually.&lt;/p>&lt;p>For example, this query returns the per-second span rate for every service except &lt;code>tempo-all&lt;/code>, grouped by service name:&lt;/p>&lt;p>&lt;/p>&lt;pre>&lt;code>{ resource.service.name != "tempo-all" } | rate() by (resource.service.name)&lt;/code>&lt;/pre>&lt;p>Previously, you'd get results back for every service and filter from there. Now you can add a threshold directly to the query. For example, to return only services where the rate exceeds 1 span per second:&lt;/p>&lt;p>&lt;/p>&lt;pre>&lt;code>{ resource.service.name != "tempo-all" } | rate() by (resource.service.name) > 1&lt;/code>&lt;/pre>&lt;p>See the &lt;u>&lt;a href="https://github.com/grafana/tempo/pull/6474">PR&lt;/a>&lt;/u> and &lt;u>&lt;a href="https://grafana.com/docs/tempo/latest/release-notes/v3-0/#comparison-operators-in-metrics-queries">release notes&lt;/a>&lt;/u> for more information. To mirror the specific example above, you can run the single-binary example.&lt;/p>&lt;h2>Enhancements to the metrics-generator&lt;/h2>&lt;p>Metrics-generator is an optional Tempo component that derives metrics from ingested traces, including span metrics, service graphs, and host info metrics.&lt;/p>&lt;p>In the 3.0 release, we focused on making it easier to manage cardinality, filter generated metrics, and compensate for sampling.&lt;/p>&lt;h3>&lt;strong>Per-label cardinality limiting&lt;/strong>&lt;/h3>&lt;p>High-cardinality labels can quickly cause a cardinality explosion, consume the active series budget, and overwhelm the downstream systems (e.g, Prometheus, Mimir, or Cortex) where you remote write metrics. A common example is a label like &lt;code>http.url&lt;/code>, where every unique URL can create another set of series.&lt;/p>&lt;p>The new per-label limiter gives operators more control on the upper bound of the cardinality per label. You can set the per-tenant &lt;code>max_cardinality_per_label&lt;/code> override, and when a label exceeds that threshold, only that label’s value is replaced with &lt;code>__cardinality_overflow__&lt;/code>. The rest of the label set is preserved. The new per-label limiter runs before the existing &lt;code>max_active_series&lt;/code> / &lt;code>max_active_entities&lt;/code> limiters, which remain as the last-resort safety net. This gives you a way to limit one problematic label without dropping all the other useful dimensions. &lt;/p>&lt;p>New metrics also help estimate per-label cardinality demand. Cardinality demand is tracked using &lt;u>&lt;a href="https://grafana.com/blog/generating-metrics-from-traces-with-cardinality-control-a-closer-look-at-hyperloglog-in-tempo/">HyperLogLog&lt;/a>&lt;/u> sketches, so the new metrics that help estimate per-label cardinality demand are estimates rather than exact counts. If cardinality drops back below the threshold, the limiter recovers automatically once the label ages out.&lt;/p>&lt;p>Check out the &lt;u>&lt;a href="https://github.com/grafana/tempo/pull/6414">PR&lt;/a>&lt;/u>, &lt;u>&lt;a href="https://grafana.com/docs/tempo/latest/configuration/#standard-overrides">configuration docs&lt;/a>&lt;/u>, &lt;u>&lt;a href="https://grafana.com/docs/tempo/latest/troubleshooting/metrics-generator/#per-label-cardinality-limiting">troubleshooting docs&lt;/a>&lt;/u>, and &lt;u>&lt;a href="https://grafana.com/blog/generating-metrics-from-traces-with-cardinality-control-a-closer-look-at-hyperloglog-in-tempo/">HyperLogLog blog&lt;/a>&lt;/u> for more information. Benchmarks can also be found in the PR.&lt;/p>&lt;h3>&lt;strong>DRAIN-based span_name sanitizer&lt;/strong>&lt;/h3>&lt;p>Span names with dynamic values (for example: &lt;code>GET /users/123&lt;/code>) can create an active series per value and quickly drive up cardinality. The &lt;u>&lt;a href="https://jiemingzhu.github.io/pub/pjhe_icws2017.pdf">DRAIN&lt;/a>&lt;/u>-based span name sanitizer clusters similar span name values and uses heuristics to replace data-like tokens (e.g. numbers, uuids, etc) with a placeholder, so those span name values are grouped under a stable placeholder pattern instead of exploding active series. &lt;/p>&lt;p>In practice, this reduces total metrics cardinality, active series, and keeps metrics-generator and downstream systems stable from the cardinality explosion due to dynamic values. It’s disabled by default and can be enabled with &lt;code>span_name_sanitization&lt;/code>, including a &lt;code>dry_run&lt;/code> mode to estimate impact before applying changes.&lt;/p>&lt;p>See the &lt;u>&lt;a href="https://github.com/grafana/tempo/pull/6098">PR&lt;/a>&lt;/u> for more details.&lt;/p>&lt;h3>&lt;strong>More flexible filtering with a new &lt;/strong>&lt;code>include_any&lt;/code>&lt;strong> &lt;/strong>&lt;strong>policy&lt;/strong>&lt;/h3>&lt;p>Span metrics filters now support a new &lt;code>include_any policy&lt;/code>, giving you more flexibility to express OR-style inclusion rules that were previously impossible to configure.&lt;/p>&lt;p>Previously, multiple &lt;code>include&lt;/code> policies behaved like a logical AND. That worked for strict filtering, but it made some configurations hard or impossible to express. &lt;/p>&lt;p>With &lt;code>include_any&lt;/code>, you can express OR-style logic while keeping the existing &lt;code>include&lt;/code> behavior unchanged.&lt;/p>&lt;p>&lt;/p>&lt;pre>&lt;code>metrics_generator:
processor:
span_metrics:
filter_policies:
# Main rule: only public/server traffic.
- include:
match_type: strict
attributes:
- key: kind
value: SPAN_KIND_SERVER
# Exception: also keep one critical internal span.
- include_any:
match_type: strict
attributes:
- key: kind
value: SPAN_KIND_INTERNAL
- key: resource.service.name
value: auth-service
- key: name
value: check-token&lt;/code>&lt;/pre>&lt;p>See the &lt;u>&lt;a href="https://github.com/grafana/tempo/pull/6392">PR&lt;/a>&lt;/u> and &lt;u>&lt;a href="https://grafana.com/docs/tempo/latest/metrics-from-traces/span-metrics/span-metrics-metrics-generator/#filtering">documentation&lt;/a>&lt;/u> for configuration and more information.&lt;/p>&lt;h3>&lt;strong>Sampling compensation with probability sampling state &lt;/strong>&lt;/h3>&lt;p>For environments that use &lt;u>&lt;a href="https://grafana.com/docs/tempo/latest/set-up-for-tracing/instrument-send/set-up-collector/tail-sampling/#head-and-tail-sampling">head-based sampling&lt;/a>&lt;/u>, generated metrics can undercount real traffic unless the processor knows how to scale back the metrics from the sampled spans.&lt;/p>&lt;p>Tempo has support for &lt;code>span_multiplier_key&lt;/code> config to configure a span or resource attribute that has the sampling ratio and scales the metrics accordingly. For example, if spans have attribute &lt;code>X-SampleRatio=0.1&lt;/code> (10% sampling), setting &lt;code>span_multiplier_key: "X-SampleRatio"&lt;/code> results in each span counting as 10 requests.&lt;/p>&lt;p>New in 3.0, we also added the &lt;code>enable_tracestate_span_multiplier&lt;/code> config on span metrics and service graphs to provide an alternative approach that extracts the multiplier from the W3C tracestate header using the &lt;u>&lt;a href="https://opentelemetry.io/docs/specs/otel/trace/tracestate-probability-sampling/">OpenTelemetry probability sampling threshold&lt;/a>&lt;/u>.&lt;/p>&lt;p>When enabled,  &lt;code>enable_tracestate_span_multiplier &lt;/code>takes priority over &lt;code>span_multiplier_key&lt;/code>.&lt;/p>&lt;p>For more information, check out the &lt;u>&lt;a href="https://github.com/grafana/tempo/pull/6684">PR&lt;/a>&lt;/u> and &lt;u>&lt;a href="https://grafana.com/docs/tempo/latest/metrics-from-traces/span-metrics/span-metrics-metrics-generator/#handling-sampled-traces">documentation&lt;/a>&lt;/u> for more information.&lt;/p>&lt;h2>Trace redaction to protect sensitive data &lt;/h2>&lt;p>Trace data can include sensitive data, including personally identifiable information (PII). This release includes new redaction capabilities to help operators remove trace data when needed.&lt;/p>&lt;p>The new &lt;code>tempo-cli redact&lt;/code> command submits redaction jobs to the backend scheduler over gRPC. Operators provide a tenant and one or more trace IDs, and Tempo creates backend jobs to rewrite the affected blocks in object storage, removing the matching trace data.&lt;/p>&lt;p>&lt;/p>&lt;pre>&lt;code>tempo-cli redact \
--tenant=STRING \
--trace-id=TRACE-ID,... \
&lt;scheduler-addr>&lt;/code>&lt;/pre>&lt;p>See the &lt;u>&lt;a href="https://github.com/grafana/tempo/pull/6832">PR&lt;/a>&lt;/u> and the &lt;u>&lt;a href="https://grafana.com/docs/tempo/latest/operations/tempo_cli/#redact-traces">documentation&lt;/a>&lt;/u> for more information.&lt;/p>&lt;h2>Improvements to query performance&lt;/h2>&lt;p>Tempo 3.0 delivers major query performance improvements across both vParquet5 and TraceQL execution. &lt;/p>&lt;h3>&lt;strong>vParquet5&lt;/strong>&lt;/h3>&lt;p>Tempo uses the &lt;u>&lt;a href="https://grafana.com/docs/tempo/latest/configuration/parquet/?pg=blog&amp;plcmt=body-txt/#apache-parquet-block-format">Apache Parquet columnar block format&lt;/a>&lt;/u> to store traces, and TraceQL queries operate on traces stored in those Parquet blocks. &lt;/p>&lt;p>vParquet5, the latest iteration of the Parquet-based columnar block format in Tempo, now supports up to 20 dedicated string columns per scope (span, resource, and event), doubled from 10, and up to 5 integer columns. It also does this dynamically, to reduce overhead to only what you actually use.&lt;/p>&lt;p>Dedicated columns can improve query performance by storing commonly used or queried attributes in their own columns, to put like-data together for best efficiency, and targeted reads when you query them. The Tempo CLI analysis and suggestion commands were also updated to help determine the best attributes to dedicate for your data.  It's not always the case that more is better, and your optimal configuration might not need to use all available columns.&lt;/p>&lt;p>Please see the &lt;u>&lt;a href="https://github.com/grafana/tempo/pull/6282">PR&lt;/a>&lt;/u> for more details.&lt;/p>&lt;h3>&lt;strong>TraceQL AST optimization&lt;/strong>&lt;/h3>&lt;p>Tempo now automatically rewrites supported repeated conditions on the same attribute into a faster internal array operation, speeding up query execution. For example, a query like  &lt;code>{ resource.service.name="frontend" || resource.service.name="backend" || resource.service.name="api" }&lt;/code> can be rewritten internally before execution, with no changes required from users. &lt;/p>&lt;p>This is a breaking change. The implicit array matching semantics of the operators &lt;code>!=&lt;/code> and &lt;code>!~&lt;/code> have changed: &lt;code>!=&lt;/code> now means &lt;code>NOT IN&lt;/code>, and &lt;code>!~&lt;/code> now means &lt;code>MATCH NONE&lt;/code>. Regex operands must now be strings or string arrays; values that were previously accepted because they could be converted to regex strings may now be rejected. Review the &lt;u>&lt;a href="https://grafana.com/docs/tempo/v3.0.x/set-up-for-tracing/setup-tempo/upgrade/#traceql-array-matching-changes">upgrade notes&lt;/a>&lt;/u> before moving to Tempo 3.0 if you use complex TraceQL queries.&lt;/p>&lt;p>TraceQL AST (Abstract Syntax Tree) optimizations can be disabled by setting &lt;code>skip_ast_transformations=all&lt;/code> in the query frontend config. You can opt out per query with the hint &lt;code>skip_ast_transformations=all&lt;/code>.&lt;/p>&lt;p>See the &lt;u>&lt;a href="https://github.com/grafana/tempo/pull/6353">PR #6353&lt;/a>&lt;/u> and &lt;u>&lt;a href="https://github.com/grafana/tempo/pull/7012">PR #7012&lt;/a>&lt;/u> for more details, including limitations and breaking changes.&lt;/p>&lt;h2>A new way to enable Tempo MCP server&lt;/h2>&lt;p>In Tempo 2.9, we &lt;u>&lt;a href="https://grafana.com/blog/grafana-tempo-2-9-release-mcp-server-support-traceql-metrics-sampling-and-more/#analyze-tracing-data-with-llms-mcp-server-support">introduced MCP server support &lt;/a>&lt;/u>as an experimental way for LLMs and AI agents to query and analyze Tempo data through TraceQL and other endpoints. In addition to enabling the feature via configuration, you can &lt;u>&lt;a href="https://github.com/grafana/tempo/releases/tag/v2.10.4">now use a dedicated flag&lt;/a>&lt;/u> to enable it as well:&lt;/p>&lt;p>&lt;/p>&lt;pre>&lt;code>--query-frontend.mcp-server.enabled=true&lt;/code>&lt;/pre>&lt;p>See the &lt;u>&lt;a href="https://github.com/grafana/tempo/pull/6903">PR&lt;/a>&lt;/u> and &lt;u>&lt;a href="https://grafana.com/docs/tempo/latest/api_docs/mcp-server/">documentation&lt;/a>&lt;/u> for more details.&lt;/p>&lt;h2>Migrating to Tempo 3.0&lt;/h2>&lt;p>Because Tempo 3.0 includes major architectural changes, we created a &lt;u>&lt;a href="https://grafana.com/docs/tempo/latest/set-up-for-tracing/setup-tempo/migrate-to-3/">migration guide&lt;/a>&lt;/u> to help you move your Tempo 2.x configurations to 3.0. Also, see &lt;u>&lt;a href="https://grafana.com/docs/tempo/latest/set-up-for-tracing/setup-tempo/upgrade/">Upgrade your Tempo installation&lt;/a>&lt;/u> for breaking changes and migration guidance before upgrading to Tempo 3.0.&lt;/p>&lt;p>We also added a new migration command to the &lt;u>&lt;a href="https://grafana.com/docs/tempo/latest/operations/tempo_cli/">Tempo CLI&lt;/a>&lt;/u>.&lt;/p>&lt;p>For monolithic deployments, we recommend using:&lt;/p>&lt;p>&lt;/p>&lt;pre>&lt;code>tempo-cli migrate config old-config.yaml > new-config.yaml
&lt;/code>&lt;/pre>&lt;p>For microservices deployments, we recommend using:&lt;/p>&lt;p>&lt;/p>&lt;pre>&lt;code>tempo-cli migrate config \
--kafka-address=&lt;KAFKA_BROKER_ADDRESS> \
--kafka-topic=&lt;YOUR_TOPIC> \
old-config.yaml > new-config.yaml&lt;/code>&lt;/pre>&lt;p>&lt;br>The &lt;code>tempo-cli migrate&lt;/code> command:&lt;/p>&lt;ul>&lt;li>Removes obsolete config sections: &lt;code>ingester&lt;/code>, &lt;code>ingester_client&lt;/code>, and &lt;code>compactor&lt;/code>&lt;/li>&lt;li>Adds Kafka (microservices mode only)&lt;/li>&lt;li>Removes the local-blocks processor&lt;/li>&lt;li>Disables compaction during parallel migration&lt;/li>&lt;/ul>&lt;p>For step-by-step instructions and more details, refer to the &lt;u>&lt;a href="https://grafana.com/docs/tempo/latest/set-up-for-tracing/setup-tempo/migrate-to-3/">migration guide&lt;/a>&lt;/u> and the &lt;u>&lt;a href="https://github.com/grafana/tempo/pull/6982">PR&lt;/a>&lt;/u>.&lt;/p>&lt;h2>How to learn more&lt;/h2>&lt;p>To see the full list of improvements, bug fixes, and breaking changes in Tempo 3.0, refer to the &lt;u>&lt;a href="https://grafana.com/docs/tempo/latest/release-notes/v3-0/">release notes&lt;/a>&lt;/u> and &lt;u>&lt;a href="https://github.com/grafana/tempo/releases/tag/v3.0.0">changelog&lt;/a>&lt;/u>.&lt;/p>&lt;p>If you’re upgrading from Tempo 2.x, start with the &lt;u>&lt;a href="https://grafana.com/docs/tempo/latest/set-up-for-tracing/setup-tempo/migrate-to-3/">migration guide&lt;/a>&lt;/u> and &lt;u>&lt;a href="https://grafana.com/docs/tempo/latest/set-up-for-tracing/setup-tempo/upgrade/">upgrade notes&lt;/a>&lt;/u> and review the breaking changes carefully, especially the removal of ingesters, compactors, the local-blocks processor, v2 block format, the OpenCensus receiver, and the scalable single binary target.&lt;/p>&lt;p>If you are interested in learning more about Tempo, please join us on the &lt;u>&lt;a href="https://slack.grafana.com/">Grafana Labs Community Slack&lt;/a>&lt;/u> channel #tempo, post a question in &lt;u>&lt;a href="https://community.grafana.com/c/grafana-tempo/40">our community forums&lt;/a>&lt;/u>, reach out on &lt;u>&lt;a href="https://twitter.com/actually_chores">X (formerly Twitter&lt;/a>&lt;/u>), or join our &lt;u>&lt;a href="https://docs.google.com/document/d/1yGsI6ywU-PxZBjmq3p3vAXr9g5yBXSDk4NU8LGo8qeY/edit#">monthly Tempo community call&lt;/a>&lt;/u>. See you there!&lt;/p></description></item></channel></rss>