Grafana Cloud

Troubleshoot CloudWatch metrics scrape jobs

This page describes common issues you might encounter with Amazon CloudWatch metrics scrape jobs and how to resolve them.

Scrape job resource limit reached

If you receive the error You reached the resource limit for a single scrape job, it means Grafana Cloud has imposed a resource limit to prevent surprise bills from CloudWatch when monitoring a large AWS environment. This limit is applied per service across all configured regions for a job.

If you have a large AWS environment and the potential CloudWatch bill is acceptable, you can do one of the following:

  • If you have one job that covers multiple regions, split the job into a single job for each region.
  • Open a support ticket to have the limit increased. When you open the ticket, choose the topic “Cloud Integrations” and the subject “Increase AWS Metrics Resource Limits”.

Metrics continue to appear for terminated or deleted AWS resources

Symptom

After you delete an AWS resource (for example, an EC2 instance), its metrics can continue to appear in queries for up to approximately three hours.

Cause

Grafana Cloud’s CloudWatch metrics integration uses the AWS RecentlyActive API parameter when it queries CloudWatch. This parameter limits how long stale metrics can appear, but it doesn’t eliminate the delay entirely. Without it, metrics for removed resources could otherwise persist for weeks.

AWS defines this behavior as an approximation, and stale metrics can occasionally persist up to 50 minutes longer than the three-hour window.

Note

A stale-metric window of up to three hours after a resource is removed is expected behavior, not a bug. There’s currently no way to remove this delay entirely.

Resolution

To stop metrics from being reported for a resource you’re about to remove, remove all tags first. Grafana Cloud’s CloudWatch integration only scrapes metrics for tagged resources, so removing the tag stops the resource from being scraped going forward.

Removing the tag after the resource is gone doesn’t help, because the stale metrics have already been captured during the scrape window.