Smarter Week

How to automate it

How to automate “tune noisy alerts and dashboards”

Here are 2 ways to spend less time on this, best first. Each comes with steps you can follow today.

45 min
typically, once a week
40%
of the time can be automated
Some setup
to set up

Fix 1 of 2

Process fixBest fix

Delete or downgrade alerts nobody acts on

Noisy alerts waste time and hide real problems. A quarterly cleanup that keeps only actionable alerts makes on-call calmer and faster.

Typically saves about 35% of the time3 h to set up
  1. 1Export alerts from the last 90 days with how often each fired and whether anyone acted.
  2. 2Delete alerts with no action taken; turn informational ones into dashboard panels.
  3. 3Make every paging alert link to a runbook.
  4. 4Alert on customer-facing symptoms (SLOs) rather than every cause.

Tools: Datadog, Grafana, New Relic or Prometheus Alertmanager · PagerDuty or Opsgenie analytics

Fix 2 of 2

Software feature

Use the AI in your incident and observability tools

incident.io, PagerDuty, Rootly, Datadog and Grafana now summarize incidents, group related alerts and suggest likely causes, which shortens the first confused half hour.

Typically saves about 25% of the time1 h to set up
  1. 1Check which AI features your incident and monitoring tools include (incident.io AI SRE, PagerDuty Advance, Rootly AI, Datadog Bits AI, Grafana Assistant).
  2. 2Turn on alert grouping and automatic incident summaries.
  3. 3During an incident, ask the assistant for recent deploys, related alerts and similar past incidents.
  4. 4Treat its root-cause suggestions as hypotheses to check, not answers.

Tools: incident.io · PagerDuty · Rootly · Datadog Bits AI · Grafana Assistant

Who does this task

Roles in our library that list this as one of their common tasks. Each guide covers the rest of that role’s week.

HourLeak · the 8-minute work audit

How many hours does this cost you?

The free 8-minute check works out where your week goes and gives you your top fixes. The team scan does the same for everyone and adds it up, so you know which leaks to fix first.

Answers are anonymous. Leaders only see team totals.

Other common tasks for DevOps / Site reliability engineers