OurCrowd

How OurCrowd moved from reactive troubleshooting to proactive observability with Grafana

When Or Angrest, CTO at OurCrowd, joined the company more than eight years ago as a software developer, observability meant logging into individual AWS consoles, tailing log files on production servers, and relying on a handful of engineers who knew how to do it. Today, Grafana sits at the center of how OurCrowd monitors its platform, investigates incidents, and automates root cause analysis—with developers, support, and QA all working from the same observability stack.

“Using Grafana, I can find in a few minutes very complex and hard issues,” Or said. “In the past, it took hours to days.”

OurCrowd operates a platform where investors invest in startups and private companies. Behind that experience, its teams run workflows backed by Salesforce, Node.js services on AWS, Kubernetes, and MongoDB. For Or, Grafana OSS became the foundation for moving from manual troubleshooting to real-time visibility and AI-driven automation.

Challenge

Before Grafana, OurCrowd had no unified way to see what was happening across its production environment.

“We did everything manually. We used the AWS basic dashboards for Elastic Beanstalk, EC2 instances, and AWS Lambda. We went to each service and used its dashboard. It was very hard to maintain.”

Log investigation was equally painful. Engineers SSH’d into servers and tailed log files by hand. It was work that demanded specialized skills that couldn’t scale across the team.

“In terms of logs, we did everything manually,” Or said. “We logged into the servers, checked everything, and tailed the log files. It was very, very hard to maintain and manage. And it required very specific skills.”

Without real-time alerting, the team stayed in reactive mode. As OurCrowd grew, that gap became harder to ignore. On-call developers spent hours on investigations that varied widely depending on who was on shift. Or saw an opportunity to standardize how the team detected problems, searched logs, and responded—without adding expensive, rigid tooling.

Solution

Or first adopted Grafana for a simple reason: one place to see everything. The team suddenly had dashboards exposing all of its services in one place. It could track metrics and alerts for production services and know when something was happening in real time.

“It’s something that was very new for us,” Or said. “And we were kind of getting a one-stop shop for our logs analysis and our service health page. It was very, very easy to become a fan of Grafana.”

The first breakthrough came when Or centralized logs. At the time, Amazon CloudWatch did not yet offer a logs data source in Grafana, so the team built a Python pipeline on AWS Lambda to stream production log files into Grafana Loki via InfluxDB—making log search available in one place for the first time.

“Based on that, we could query the logs in real time,” Or said. “It was a huge milestone. It was a game changer for us—one place where you can search and query the logs very easily. It’s something that was very hard for developers to do.”

Or then used Grafana as the foundation for two automation projects. The first identifies IP addresses associated with scans against OurCrowd’s servers and automates a process that developers previously performed manually. The second uses AI agents to query logs through Grafana in real time, analyze signals such as slow queries and web requests, and deliver root cause reports to on-call developers in Slack.

“The agents know how to query Grafana in real time,” Or said. “They cross the information and run a live analysis for the on-call developer to understand the root cause. The root cause analysis is fully based on Grafana.”

Outcome

By combining Grafana alerts, APIs, and AI agents, OurCrowd reduced investigations that once took hours to a process that takes seconds or minutes. The automation also made incident response less dependent on an individual developer’s experience.

“By far, it’s the efficiency,” Or said. “It’s the efficiency of automated, non-human, proactive actions triggered by Grafana using Grafana’s API—and, of course, less time spent by developers on analysis. That’s what makes the impact here.”

Grafana has also changed how the team develops new features. Developers now consider metrics, alerts, and observability during design instead of waiting for problems in production. And Grafana’s explore mode makes logs across its environment accessible to developers, support, and QA.

Next, OurCrowd plans to use Tempo with OpenTelemetry for request-level tracing and Faro for frontend observability. Together, Or said, those additions could create “a fully one-stop-shop solution for our company” across the backend and frontend.

OurCrowd logo
Industry
Financial Services
Company Size
200+
Headquarters
Jerusalem, Israel