Monitoring & Alerting Setup Checklist Template

Most alerting problems are not missing alerts. They are the disk warning that has fired every night for a year, the page routed to someone who left in spring, and the outage customers reported before any alert did.

This free monitoring and alerting checklist brings a service under monitoring properly, then keeps it that way. Operations teams, SREs and MSPs use it when a service goes live and every quarter after. It covers the inventory of what to watch, golden signals, SLOs and burn-rate alerts, severity and routing to the on-call rotation, a runbook for every page, a test that proves each page arrives, and a review that removes the alerts nobody acts on.

Use This Template Free See Live Example
No Credit Card Required

What to Measure: Golden Signals, RED and USE

Three well-known methods tell you what to measure. They overlap but answer different questions: is the service working for users, is each component healthy, and is each resource running out. A well-monitored service uses all three.

Four golden signals

The user-facing service

Measures: latency, traffic, errors and saturation.

Source: Google’s Site Reliability Engineering book, chapter 6.

Where: at the edge, where users meet the service.

RED method

Each request-driven component

Measures: rate, errors and duration of requests.

Source: Tom Wilkie, then at Weaveworks, for microservices.

Where: every API, endpoint and queue consumer.

USE method

Each resource underneath

Measures: utilisation, saturation and errors.

Source: Brendan Gregg’s performance methodology.

Where: CPU, memory, disk, network and connection pools.

Measuring is not the same as alerting. The SRE book separates symptoms (what is broken: users get errors or slow pages) from causes (why: a database refusing connections, a CPU at its limit). It says every page should be actionable, and that cause-based pages should be kept for causes that are definite and imminent. This checklist applies that rule: golden signals and SLOs drive pages; most RED and USE metrics feed dashboards and tickets. Writing the runbook behind each page is covered in more depth in our guide to writing an IT runbook.

What the Monitoring & Alerting Setup Checklist Covers

Seven phases take a service from inventory to a tested, reviewed alert set. Phases 2 to 4 appear for a new service, Phase 6 for the quarterly alert review, and a failed end-to-end test adds a fix task before go-live is approved.

Phase 1

Phase 1: Scope, Owners & Inventory

The scope answer on the first task decides whether the build phases or the review phase appear.

  • Open the record, name the service and choose the scope — new service or quarterly alert review, with the service owner and on-call lead picked from your members
  • List every component the service depends on — hosts, containers, databases, queues, load balancers, DNS, certificates and third-party APIs
  • Write down the user journeys that define “working” — usually two or three, such as sign in, place an order or call the API
  • Check which components already report to monitoring — agents installed, exporters running, logs shipped; anything missing goes on a gap list
  • Confirm an on-call rotation and escalation policy cover this service — an alert with no rotation behind it pages nobody
  • Note the monitoring and paging tools in use and any end-of-life dates — a tool migration changes every route and needs the same testing
Phase 2 — New service

Phase 2: Signals, SLIs & SLOs

Tasks appear only when the scope is a new service.

  • Instrument the four golden signals at the service edge — latency, traffic, errors and saturation, measured at the load balancer or API gateway
  • Add RED metrics for each internal component — request rate, errors and duration per endpoint or queue consumer
  • Apply USE to each resource — utilisation, saturation and errors for CPU, memory, disk, network and connection pools
  • Define an SLI for each user journey — for example, the proportion of sign-in requests served successfully in under 500 ms
  • Agree the SLO and error budget with the service owner — 99.9% over 30 days allows about 43 minutes of total failure
  • Add synthetic checks for each journey from outside your network — they catch DNS, certificate and CDN failures internal metrics never see
Phase 3 — New service

Phase 3: Thresholds, Severity & Routing

Tasks appear only when the scope is a new service. The alert table lives in this phase.

  • Agree what each severity level does — for example, page now, ticket for the next working day, or dashboard only
  • Base pages on symptoms wherever you can — error rate against the SLO rather than CPU above 80%; keep cause alerts for definite causes such as a disk filling within hours
  • Add burn-rate alerts to each SLO — the SRE Workbook pages on a 14.4× burn over one hour or 6× over six hours, and raises a ticket at 1× over three days
  • Give every threshold a duration — a for: 5m clause in Prometheus, or the equivalent elsewhere, so one spike does not wake anyone
  • Record every alert in the alert table — name, condition and threshold, severity, route and runbook link, one row each
  • Configure grouping and inhibition — related alerts arrive as one notification, and warnings stay quiet while the matching critical is firing
  • Route pages to the rotation, not to named people — with escalation to the secondary if a page is not acknowledged in time
Phase 4 — New service

Phase 4: Runbooks & Dashboards

Tasks appear only when the scope is a new service.

  • Write a runbook for every paging alert — what it means, who is affected, the first three checks, how to mitigate and when to escalate
  • Put the runbook and dashboard links in the notification itself — at 03:00 nobody searches the wiki
  • Build the service dashboard from the top down — golden signals and SLO first, error budget remaining next, resources last
  • Add deployment and change markers to the dashboard — most incidents follow a change, and the marker shows which one
  • Demote any alert whose runbook says “wait and see” — if there is nothing to do, it is a ticket or a panel, not a page
Phase 5

Phase 5: Test That Every Page Arrives

A Fail result shows the fix task. The on-call lead’s approval on the last task halts the checklist until they decide.

  • Validate rules and routing before deploying them — promtool check rules, amtool check-config, and amtool config routes test with each alert’s labels
  • Fire each paging alert once under control — stop a test instance, break a synthetic check or inject errors in staging, then confirm the resolve notification too
  • Send a test notification to everyone on the rotation — to the phone, not just a chat channel, including any SMS or voice fallback
  • Add a watchdog alert that always fires to an external heartbeat service — if the monitoring pipeline dies, the heartbeat stops and something else pages
  • Record the end-to-end test result — pass, or fail with the alerts that did not arrive and where they went instead
  • Fix the failed routes and test them again — go-live waits until every paging alert reaches the rotation
  • Approve the alert set — the on-call lead records approved or not approved, with the reason
Phase 6 — Quarterly review

Phase 6: Quarterly Alert Review

Tasks appear only when the scope is the quarterly alert review.

  • Export every alert that fired this quarter — count, time to acknowledge and what the responder did
  • Tune, demote or delete alerts that fired with no action taken — a page nobody acted on is noise, and noise trains people to ignore pages
  • Count pages per on-call shift against a target — the SRE book sets a maximum of two incidents per 12-hour shift
  • Find incidents no alert caught — from postmortems and customer-reported tickets; add the missing symptom alert
  • Check every alert still has an owner, a working route and a runbook that opens — people leave and teams reorganise
  • Remove silences and maintenance windows that outlived their purpose — an open-ended silence is an alert switched off
Phase 7

Phase 7: Handover & Close

  • Update the alert table and the service catalogue entry — the table on this run becomes the baseline for the next review
  • Brief the rotation on new, changed and deleted alerts — at the next on-call handover, so nobody meets a new page cold
  • Raise tickets for every gap found — missing instrumentation, runbooks to write, dashboards to fix, each with an owner and a date
  • Confirm the date of the next quarterly review — the recurring schedule opens it automatically
  • Sign off and close — the service owner confirms the alert set, with the routing changes and test evidence attached

The Alert Table, Row by Row

Phase 3 records every alert as one row, and Phase 6 reviews the same table each quarter. These example rows are for a web shop’s checkout service. The thresholds are common starting points, not rules; tune them to your own SLOs and traffic.

Alert Condition Severity Route Runbook
Checkout errors, fast burnError budget burning at 14.4× over 1 hour and over the last 5 minutesPagePrimary on-call, then secondary if not acknowledgedCheckout errors
Checkout errors, slow burnBurning at 6× over 6 hours and over the last 30 minutesPagePrimary on-callCheckout errors
Checkout budget driftBurning at 1× over 3 days and over the last 6 hoursTicketTeam queueSLO review
Sign-in journey failingSynthetic check failing from 2 of 3 locations for 5 minutesPagePrimary on-callSign-in failures
Database disk fillingPredicted to fill within 4 hours at the current ratePagePrimary on-callDisk space
Certificate expiringFewer than 14 days to expiry on any endpointTicketPlatform team queueCertificate renewal
WatchdogAlways firingHeartbeatExternal heartbeat serviceMonitoring pipeline down

The burn-rate rows come from the SRE Workbook’s chapter on alerting on SLOs. A 14.4× burn sustained for an hour uses 2% of a 30-day error budget; requiring the short window as well means the alert clears soon after the problem stops. Azure Monitor offers five severities (Sev 0 Critical to Sev 4 Verbose); if two levels lead to the same response, merge them.

Prometheus Alertmanager waits 30 seconds by default before the first notification for a group, 5 minutes before telling you about new alerts in that group, and 4 hours before repeating one still firing. Escalation when nobody acknowledges belongs in the paging tool’s escalation policy. If that tool is Opsgenie, note that Atlassian ended new sales on 4 June 2025 and switches the product off on 5 April 2027; the move to Jira Service Management or Compass is a routing change, and Phase 5’s tests apply to it in full.

Why Run Monitoring Setup and Alert Reviews in CheckFlow?

1

Set up once, reviewed every quarter

One scope answer shows the build phases for a new service or the review phase for an existing one. A quarterly recurring schedule opens the alert review for each service, assigned to its owner with a due date, so pruning noise does not wait for the next bad week on call.

2

Every alert in one table

The alert table inside the routing task holds every alert’s threshold, severity, route and runbook, and the next review starts from it. A data set can hold your services with owner and rotation, and tags keep each client’s estate separate for MSPs.

3

No go-live without a test page

The end-to-end test result is a required dropdown: record Fail and the fix task appears. The on-call lead’s approval halts the checklist until they decide, and the audit trail shows who tested which route and when, with screenshots attached to the task.

Alerts feed the people who answer them. Pass changed and silenced alerts across at each rotation change with the On-Call Handover Checklist, run the response through the Incident Management Process Checklist, and feed incidents no alert caught back into Phase 6 from the Incident Postmortem Template.

Certificate expiry is one of the alerts most teams add after an outage; the SSL/TLS Certificate Renewal Checklist covers the renewal behind it. To see how quarterly reviews run on a schedule without anyone remembering to start them, read about CheckFlow’s recurring checklist software.

Frequently Asked Questions

What are the four golden signals of monitoring?

+

Latency, traffic, errors and saturation. They come from the “Monitoring Distributed Systems” chapter of Google’s Site Reliability Engineering book, which advises that if you can measure only four metrics of a user-facing system, these are the four. Track latency as percentiles, and separately for failed requests: a fast error can make latency look healthy while users are failing.

Should I use golden signals, RED or USE?

+

All three, at different layers. Golden signals describe the service as users experience it. RED (rate, errors, duration) applies the same idea to every request-driven component inside it. USE (utilisation, saturation, errors) checks each resource such as CPU, disk and connection pools. Page on the first; use RED and USE mainly to diagnose, on dashboards and in tickets.

What is a burn-rate alert?

+

An alert on how fast a service is spending its error budget. A burn rate of 1 uses the whole budget in exactly the SLO period; 14.4 sustained for an hour uses 2% of a 30-day budget. The SRE Workbook recommends checking a long and a short window together, which catches real problems quickly and stops alerting soon after they end, with fewer false pages than a fixed error threshold.

How many pages per on-call shift is too many?

+

Google’s SRE book sets a maximum of two incidents per 12-hour shift, because handling one properly, from diagnosis to postmortem and follow-up fixes, takes about six hours. Your target may differ, but count pages per shift at every quarterly review. When the number climbs, the fix is usually to delete or demote alerts, not to add people to the rotation.

How do you test that alerts actually fire?

+

In three layers. Check configuration before it ships, with promtool for rules and amtool config routes test for routing. Then fire each paging alert once in a controlled way and confirm it reaches the right phone and resolves. Finally, run a watchdog alert that always fires to an external heartbeat service, so a broken monitoring pipeline raises its own alarm instead of going quiet.

Is CheckFlow free for this template?

+

14-day free trial, no card required. The Business plan is $10 per user per month after the trial. Full details at checkflow.io/pricing.

Alerts Your On-Call Team Can Trust

Free trial — no credit card required.