A NOC shift is rarely let down by the alert it worked on. It is let down by the one nobody acknowledged, the device left in maintenance mode, and the handover note that said “all quiet” while a client site was still down.
A staffed network operations centre exists so that somebody is watching every client’s estate while the rest of the MSP is asleep or busy. That only works if every shift does the same things in the same order. This free NOC shift checklist gives NOC leads and shift engineers a routine for every shift, whether you run round the clock or extended hours. It starts with the previous shift’s log, open incidents and tonight’s maintenance windows, then sweeps the RMM alert queue, network and uptime monitoring, backup failures, provider status pages and certificate expiry. Alerts are acknowledged within your target, de-duplicated, ticketed in the PSA by priority and escalated by each client’s runbook. If a major incident is declared, a conditional phase runs the bridge and client updates, and every shift ends with a handover note and a closed log.
Many MSPs start with an on-call engineer who carries a phone overnight and gets woken by the alerts that matter. That model has its own routine, and the On-Call Handover Checklist covers it: one engineer passing the pager and the context to the next at the end of a rotation. A NOC is different. Someone is at a desk for the whole shift with the consoles open, so the job is not to be reachable. It is to look, all the time, and to act before a client notices.
So the checklist does a different job. A NOC shift has a start, a full sweep, continuous triage, hourly checks and a handover, often three times a day. It also sits between two other routines. The IT Help Desk Daily Checklist runs the user-facing ticket queue during business hours, and the MSP Daily Backup Monitoring Checklist owns the detailed morning review of every backup job. The NOC catches failures as they happen and tickets them, so neither routine is duplicated.
On-call engineer
Reachable when something breaks
Where: at home, with a phone and a laptop.
Trigger: a page from monitoring or the service desk.
Handover: at the end of a rotation, usually weekly.
Risk: anything that does not page is not seen until morning.
NOC shift
Watching before anything breaks
Where: at a desk, with every console open for the whole shift.
Trigger: the shift itself, plus a sweep every hour.
Handover: at the end of every shift, from lead to lead.
Risk: alert fatigue, and a log nobody reads at the next shift.
What the NOC Shift Checklist Covers
Five phases run on every shift, in order. Phase 5 appears only when the shift lead records a major incident at the end of triage.
Phase 1
Phase 1: Start the Shift
Done in the first fifteen minutes, before the outgoing shift leaves.
Open the shift log and record who is on — the shift lead, each engineer, start time and any gap in staffing against the rota
Read the previous shift’s handover note — every open item, its owner and what the outgoing lead expected to happen next
Review open incidents and P1 and P2 tickets — status, next client update due and who is working each one
Check tonight’s maintenance windows and change freezes — which clients, which devices, and who is doing the work
Confirm tonight’s escalation contacts — the on-call engineer, the duty service manager and each client’s after-hours contact
Phase 2
Phase 2: Console Sweep
A full pass of every console, so triage starts from a known state.
Work the RMM alert queue down to a known state — new, acknowledged and ticketed, with nothing older than the last shift left unread
Check network and uptime monitoring — sites down, firewalls and switches offline, flapping links and ISP circuits with packet loss
Check backup failures since the last shift — ticket new failures so the morning backup check finds an owner already assigned
Check provider status pages — ISPs, Microsoft 365 and Azure service health, and any cloud platform your clients depend on
Check certificate and domain expiry alerts — anything expiring in the next 30 days without a renewal ticket
Confirm the monitoring itself is healthy — probes, collectors and agents reporting, with no client site that has gone silent
Phase 3
Phase 3: Alert Triage
Runs continuously through the shift. The last task is answered once the start-of-shift backlog is cleared, and it decides whether Phase 5 appears.
Acknowledge new alerts within the target — the acknowledgement time your agreements and runbooks set for each priority
De-duplicate and correlate before ticketing — a site that drops takes every device with it, and that is one ticket, not forty
Ticket each genuine alert in the PSA — against the right client and agreement, with the priority the matrix gives it
Escalate by the client’s runbook — to the on-call engineer, the client’s named contact or both, and log the time you did
Suppress an alert only with a reason and an expiry — and add it to the handover note so the next shift knows it is muted
Record whether a major incident has been declared — against the criteria in your runbook; a Yes opens the major incident phase
Phase 4
Phase 4: In-Shift Scheduled Checks
Sweep the alert and ticket queues every hour — log the time and the count, so a quiet hour is recorded rather than assumed
Watch scheduled maintenance and patch jobs — confirm each one started, finished and the devices came back online
Take devices out of maintenance mode when windows close — a device left muted is a device nobody is monitoring
Make after-hours callbacks — clients who rang or emailed out of hours get a call within the time their agreement sets
Recheck suppressed alerts and their expiry — renew with a reason or let them fire again
Phase 5 — Major Incident Only
Phase 5: Major Incident
Shown only when a major incident is declared. The NOC runs communications and coordination; the fix belongs to the engineers on the bridge.
Declare the major incident and name an incident lead — the time, the clients and services affected and the impact on their users
Open a bridge — a call or channel with the on-call engineer, the duty service manager and anyone else the runbook names
Send the first client update — what is affected, what you are doing and when the next update will come
Update clients at the agreed interval — every update on time, even when the only news is that work continues
Post status updates where clients look — the status page or client portal, and the service desk’s phone message
Hand over to incident management — open the incident record, attach the timeline and book the post-incident review
Phase 6
Phase 6: Shift End & Handover
Bring every open alert to a state — resolved, ticketed with an owner, or escalated with a time recorded
Write the handover note — open incidents, muted alerts with expiry, maintenance still running and clients awaiting a callback
Walk the incoming shift lead through the note — five or ten minutes, live, before the outgoing shift signs off
Close the shift log — alerts received, tickets raised, escalations made and the slowest acknowledgement of the shift
Flag noisy alerts for tuning — any alert that fired repeatedly without needing action goes to the monitoring owner
Agreeing the first move for each kind of alert means two engineers on different shifts handle the same alert the same way. The targets below are common practice, not a standard. Each client’s agreement and runbook take precedence.
Alert
First move
Escalate when
Whole site offline
Check the ISP status page and the firewall, then raise one ticket for the site
The site is still down after the runbook’s wait time, or it is a 24/7 client
Server down or unresponsive
Check whether it is in a maintenance window, then try a remote check of the host
It is a production server outside a window, or it does not return after a restart
Disk space critical
Find what is growing and clear what is safe to clear
The volume holds a database or mail store and the growth will not stop
Backup job failed
Ticket it for the morning backup check, and retry if the window allows
The same job failed last night too, or the system has no other recent copy
Provider or cloud service incident
Confirm on the provider’s status page and note which clients use the service
Clients are affected and need a message before they start calling
Certificate expiring
Check for an existing renewal ticket and raise one if there is none
The certificate expires within days on a client-facing service
Monitoring silent for a site
Check the collector or probe and its network path
The site cannot be seen at all, so nothing else on it can alert
Certificate expiry deserves its place in the sweep. Under CA/Browser Forum Ballot SC-081, the maximum life of a publicly trusted TLS certificate fell to 200 days on 15 March 2026, and it drops to 100 days in March 2027 and 47 days in March 2029. Annual renewals now come round several times a year, and each manual one is a chance to miss one.
What a NOC shift log should record
The next shift, the service manager and sometimes a client will read it, so keep it factual and timestamped:
Who was on, when the shift started and any staffing gap.
Each hourly sweep, with the time and the number of open alerts.
Every escalation: who was contacted, how, when and what they said.
Every alert muted or device put into maintenance mode, with its expiry.
Maintenance and patch jobs watched, and whether they finished cleanly.
Microsoft partners can check the service health of a customer’s Microsoft 365 and Azure services from Partner Center, which links through to each tenant’s admin centre as delegated admin. For a NOC covering many tenants, that is a quicker first stop than signing in to each one.
Why Run NOC Shifts in CheckFlow?
1
A fresh checklist for every shift
A recurring schedule creates the checklist at the start of each shift and assigns it to the NOC group. The shift lead, on-call engineer and duty service manager are chosen on the first task, and every escalation names a real person.
2
Major incidents that follow the same path
One answer at the end of triage decides whether conditional logic adds the major incident phase. When it appears, the bridge, client updates and handover to incident management are tasks with owners.
3
A log the next shift can trust
Hourly sweeps, escalations and muted alerts are recorded against tasks with comments and attachments, with a time and a name on each. The service manager can see across every shift which alerts keep firing.
The MSP process management guide explains why monitoring tools automate the technical layer but not the human process around it. A NOC shift is that human layer, and CheckFlow for MSPs runs it alongside your service desk, backup and reporting routines.
A major incident that starts on a NOC shift continues in the Incident Management Checklist, which carries it through diagnosis, resolution and the post-incident review. The runbooks each escalation depends on are easier to keep current with the IT runbook template guide.
A network operations centre watches the infrastructure the MSP manages for its clients: servers, network devices, internet links, backups and cloud services. Engineers on shift work through monitoring alerts, open tickets for genuine problems, fix what the runbook allows them to fix and escalate the rest to the right engineer or client contact. The aim is to find and act on failures before a client’s users do.
What is the difference between a NOC and a SOC?
+
A NOC looks after availability and performance: is the server up, is the link working, did the backup run. A security operations centre looks for threats: suspicious sign-ins, malware detections and signs of an attacker. Smaller MSPs often have the same people covering both. Where they are separate, agree which alerts go where and how a NOC engineer hands a suspected security incident to the SOC.
How quickly should a NOC acknowledge an alert?
+
There is no single standard. Common practice is to acknowledge critical alerts within minutes and lower priorities within the hour, but the right target is whatever your agreements and runbooks commit to. Acknowledgement only means someone owns the alert, so set the target tight. The shift log should record the slowest acknowledgement each shift, because averages hide the alert that waited forty minutes.
Does an MSP need a 24/7 NOC?
+
Not always. A staffed NOC makes sense when enough clients have round-the-clock services or tight agreements that an on-call engineer would be woken every night. Many MSPs staff extended hours and use on-call cover overnight, or buy in an outsourced NOC. Either way, the checklist defines what any shift must do, whoever staffs it.
When should a NOC shift declare a major incident?
+
When the criteria written in your runbook are met, not when someone feels it is serious. Typical criteria include a whole client site or a critical service down, several clients affected by the same fault, or a security event with business impact. The NOC should be allowed to declare one without waiting for permission. Declaring early and standing down costs little; declaring late costs you the first client updates.
How do you reduce alert fatigue in a NOC?
+
Start with the shift logs. Any alert that fires repeatedly without anyone needing to act should be tuned, re-thresholded or removed by the monitoring owner. Correlate alerts so one site outage produces one ticket. Use maintenance windows so planned work stays quiet, and review the noisiest alerts monthly.
Is CheckFlow free for this template?
+
14-day free trial, no card required. The Business plan is $10 per user per month after the trial. Full details at checkflow.io/pricing.
Run Every NOC Shift the Same Way, Whoever Is On
Free trial — no credit card required.
Do you like cookies? 🍪 We use cookies to ensure you get the best experience on our website. Learn more