Most alerting problems are not missing alerts. They are the disk warning that has fired every night for a year, the page routed to someone who left in spring, and the outage customers reported before any alert did.
This free monitoring and alerting checklist brings a service under monitoring properly, then keeps it that way. Operations teams, SREs and MSPs use it when a service goes live and every quarter after. It covers the inventory of what to watch, golden signals, SLOs and burn-rate alerts, severity and routing to the on-call rotation, a runbook for every page, a test that proves each page arrives, and a review that removes the alerts nobody acts on.
Three well-known methods tell you what to measure. They overlap but answer different questions: is the service working for users, is each component healthy, and is each resource running out. A well-monitored service uses all three.
Four golden signals
The user-facing service
Measures: latency, traffic, errors and saturation.
Source: Google’s Site Reliability Engineering book, chapter 6.
Where: at the edge, where users meet the service.
RED method
Each request-driven component
Measures: rate, errors and duration of requests.
Source: Tom Wilkie, then at Weaveworks, for microservices.
Where: every API, endpoint and queue consumer.
USE method
Each resource underneath
Measures: utilisation, saturation and errors.
Source: Brendan Gregg’s performance methodology.
Where: CPU, memory, disk, network and connection pools.
Measuring is not the same as alerting. The SRE book separates symptoms (what is broken: users get errors or slow pages) from causes (why: a database refusing connections, a CPU at its limit). It says every page should be actionable, and that cause-based pages should be kept for causes that are definite and imminent. This checklist applies that rule: golden signals and SLOs drive pages; most RED and USE metrics feed dashboards and tickets. Writing the runbook behind each page is covered in more depth in our guide to writing an IT runbook.
What the Monitoring & Alerting Setup Checklist Covers
Seven phases take a service from inventory to a tested, reviewed alert set. Phases 2 to 4 appear for a new service, Phase 6 for the quarterly alert review, and a failed end-to-end test adds a fix task before go-live is approved.
Phase 1
Phase 1: Scope, Owners & Inventory
The scope answer on the first task decides whether the build phases or the review phase appear.
Open the record, name the service and choose the scope — new service or quarterly alert review, with the service owner and on-call lead picked from your members
List every component the service depends on — hosts, containers, databases, queues, load balancers, DNS, certificates and third-party APIs
Write down the user journeys that define “working” — usually two or three, such as sign in, place an order or call the API
Check which components already report to monitoring — agents installed, exporters running, logs shipped; anything missing goes on a gap list
Confirm an on-call rotation and escalation policy cover this service — an alert with no rotation behind it pages nobody
Note the monitoring and paging tools in use and any end-of-life dates — a tool migration changes every route and needs the same testing
Phase 2 — New service
Phase 2: Signals, SLIs & SLOs
Tasks appear only when the scope is a new service.
Instrument the four golden signals at the service edge — latency, traffic, errors and saturation, measured at the load balancer or API gateway
Add RED metrics for each internal component — request rate, errors and duration per endpoint or queue consumer
Apply USE to each resource — utilisation, saturation and errors for CPU, memory, disk, network and connection pools
Define an SLI for each user journey — for example, the proportion of sign-in requests served successfully in under 500 ms
Agree the SLO and error budget with the service owner — 99.9% over 30 days allows about 43 minutes of total failure
Add synthetic checks for each journey from outside your network — they catch DNS, certificate and CDN failures internal metrics never see
Phase 3 — New service
Phase 3: Thresholds, Severity & Routing
Tasks appear only when the scope is a new service. The alert table lives in this phase.
Agree what each severity level does — for example, page now, ticket for the next working day, or dashboard only
Base pages on symptoms wherever you can — error rate against the SLO rather than CPU above 80%; keep cause alerts for definite causes such as a disk filling within hours
Add burn-rate alerts to each SLO — the SRE Workbook pages on a 14.4× burn over one hour or 6× over six hours, and raises a ticket at 1× over three days
Give every threshold a duration — a for: 5m clause in Prometheus, or the equivalent elsewhere, so one spike does not wake anyone
Record every alert in the alert table — name, condition and threshold, severity, route and runbook link, one row each
Configure grouping and inhibition — related alerts arrive as one notification, and warnings stay quiet while the matching critical is firing
Route pages to the rotation, not to named people — with escalation to the secondary if a page is not acknowledged in time
Phase 4 — New service
Phase 4: Runbooks & Dashboards
Tasks appear only when the scope is a new service.
Write a runbook for every paging alert — what it means, who is affected, the first three checks, how to mitigate and when to escalate
Put the runbook and dashboard links in the notification itself — at 03:00 nobody searches the wiki
Build the service dashboard from the top down — golden signals and SLO first, error budget remaining next, resources last
Add deployment and change markers to the dashboard — most incidents follow a change, and the marker shows which one
Demote any alert whose runbook says “wait and see” — if there is nothing to do, it is a ticket or a panel, not a page
Phase 5
Phase 5: Test That Every Page Arrives
A Fail result shows the fix task. The on-call lead’s approval on the last task halts the checklist until they decide.
Validate rules and routing before deploying them — promtool check rules, amtool check-config, and amtool config routes test with each alert’s labels
Fire each paging alert once under control — stop a test instance, break a synthetic check or inject errors in staging, then confirm the resolve notification too
Send a test notification to everyone on the rotation — to the phone, not just a chat channel, including any SMS or voice fallback
Add a watchdog alert that always fires to an external heartbeat service — if the monitoring pipeline dies, the heartbeat stops and something else pages
Record the end-to-end test result — pass, or fail with the alerts that did not arrive and where they went instead
Fix the failed routes and test them again — go-live waits until every paging alert reaches the rotation
Approve the alert set — the on-call lead records approved or not approved, with the reason
Phase 6 — Quarterly review
Phase 6: Quarterly Alert Review
Tasks appear only when the scope is the quarterly alert review.
Export every alert that fired this quarter — count, time to acknowledge and what the responder did
Tune, demote or delete alerts that fired with no action taken — a page nobody acted on is noise, and noise trains people to ignore pages
Count pages per on-call shift against a target — the SRE book sets a maximum of two incidents per 12-hour shift
Find incidents no alert caught — from postmortems and customer-reported tickets; add the missing symptom alert
Check every alert still has an owner, a working route and a runbook that opens — people leave and teams reorganise
Remove silences and maintenance windows that outlived their purpose — an open-ended silence is an alert switched off
Phase 7
Phase 7: Handover & Close
Update the alert table and the service catalogue entry — the table on this run becomes the baseline for the next review
Brief the rotation on new, changed and deleted alerts — at the next on-call handover, so nobody meets a new page cold
Raise tickets for every gap found — missing instrumentation, runbooks to write, dashboards to fix, each with an owner and a date
Confirm the date of the next quarterly review — the recurring schedule opens it automatically
Sign off and close — the service owner confirms the alert set, with the routing changes and test evidence attached
Phase 3 records every alert as one row, and Phase 6 reviews the same table each quarter. These example rows are for a web shop’s checkout service. The thresholds are common starting points, not rules; tune them to your own SLOs and traffic.
Alert
Condition
Severity
Route
Runbook
Checkout errors, fast burn
Error budget burning at 14.4× over 1 hour and over the last 5 minutes
Page
Primary on-call, then secondary if not acknowledged
Checkout errors
Checkout errors, slow burn
Burning at 6× over 6 hours and over the last 30 minutes
Page
Primary on-call
Checkout errors
Checkout budget drift
Burning at 1× over 3 days and over the last 6 hours
Ticket
Team queue
SLO review
Sign-in journey failing
Synthetic check failing from 2 of 3 locations for 5 minutes
Page
Primary on-call
Sign-in failures
Database disk filling
Predicted to fill within 4 hours at the current rate
Page
Primary on-call
Disk space
Certificate expiring
Fewer than 14 days to expiry on any endpoint
Ticket
Platform team queue
Certificate renewal
Watchdog
Always firing
Heartbeat
External heartbeat service
Monitoring pipeline down
The burn-rate rows come from the SRE Workbook’s chapter on alerting on SLOs. A 14.4× burn sustained for an hour uses 2% of a 30-day error budget; requiring the short window as well means the alert clears soon after the problem stops. Azure Monitor offers five severities (Sev 0 Critical to Sev 4 Verbose); if two levels lead to the same response, merge them.
Prometheus Alertmanager waits 30 seconds by default before the first notification for a group, 5 minutes before telling you about new alerts in that group, and 4 hours before repeating one still firing. Escalation when nobody acknowledges belongs in the paging tool’s escalation policy. If that tool is Opsgenie, note that Atlassian ended new sales on 4 June 2025 and switches the product off on 5 April 2027; the move to Jira Service Management or Compass is a routing change, and Phase 5’s tests apply to it in full.
Why Run Monitoring Setup and Alert Reviews in CheckFlow?
1
Set up once, reviewed every quarter
One scope answer shows the build phases for a new service or the review phase for an existing one. A quarterly recurring schedule opens the alert review for each service, assigned to its owner with a due date, so pruning noise does not wait for the next bad week on call.
2
Every alert in one table
The alert table inside the routing task holds every alert’s threshold, severity, route and runbook, and the next review starts from it. A data set can hold your services with owner and rotation, and tags keep each client’s estate separate for MSPs.
3
No go-live without a test page
The end-to-end test result is a required dropdown: record Fail and the fix task appears. The on-call lead’s approval halts the checklist until they decide, and the audit trail shows who tested which route and when, with screenshots attached to the task.
Latency, traffic, errors and saturation. They come from the “Monitoring Distributed Systems” chapter of Google’s Site Reliability Engineering book, which advises that if you can measure only four metrics of a user-facing system, these are the four. Track latency as percentiles, and separately for failed requests: a fast error can make latency look healthy while users are failing.
Should I use golden signals, RED or USE?
+
All three, at different layers. Golden signals describe the service as users experience it. RED (rate, errors, duration) applies the same idea to every request-driven component inside it. USE (utilisation, saturation, errors) checks each resource such as CPU, disk and connection pools. Page on the first; use RED and USE mainly to diagnose, on dashboards and in tickets.
What is a burn-rate alert?
+
An alert on how fast a service is spending its error budget. A burn rate of 1 uses the whole budget in exactly the SLO period; 14.4 sustained for an hour uses 2% of a 30-day budget. The SRE Workbook recommends checking a long and a short window together, which catches real problems quickly and stops alerting soon after they end, with fewer false pages than a fixed error threshold.
How many pages per on-call shift is too many?
+
Google’s SRE book sets a maximum of two incidents per 12-hour shift, because handling one properly, from diagnosis to postmortem and follow-up fixes, takes about six hours. Your target may differ, but count pages per shift at every quarterly review. When the number climbs, the fix is usually to delete or demote alerts, not to add people to the rotation.
How do you test that alerts actually fire?
+
In three layers. Check configuration before it ships, with promtool for rules and amtool config routes test for routing. Then fire each paging alert once in a controlled way and confirm it reaches the right phone and resolves. Finally, run a watchdog alert that always fires to an external heartbeat service, so a broken monitoring pipeline raises its own alarm instead of going quiet.
Is CheckFlow free for this template?
+
14-day free trial, no card required. The Business plan is $10 per user per month after the trial. Full details at checkflow.io/pricing.
Alerts Your On-Call Team Can Trust
Free trial — no credit card required.
Do you like cookies? 🍪 We use cookies to ensure you get the best experience on our website. Learn more