MSP Daily Backup Monitoring Checklist Template

Most backup failures are found on the day someone needs a restore, but they start weeks earlier: a job that stopped running, a new server nobody added, a SaaS connector whose permission expired. The morning check is where they are cheapest to catch.

Every MSP backs up its clients. Fewer can say, on any given morning, which jobs failed overnight across every client, which ones never ran at all, and whether the offsite copy landed. The backup consoles send alerts, but alerts get filtered, missed jobs often raise none, and an alert nobody acts on protects nothing. This free MSP daily backup monitoring checklist turns the morning backup check into a short, repeatable routine that any technician can run. It covers overnight server and endpoint jobs, Microsoft 365 and Google Workspace backups, offsite and immutable copies, coverage gaps from new devices and users, one rotating spot-restore, a ticket for every failure and a daily log. When the same job has failed on consecutive runs, or a system’s last good backup is older than its recovery point objective, a conditional escalation phase brings in the service manager the same day.

Use This Template Free See Live Example
No Credit Card Required

The Morning Check and the Restore Test Answer Different Questions

The daily check asks a narrow question: did last night’s backups run, for every client and every system they should cover, and did the copies land where they should? It takes minutes per client, it runs every working day, and its job is to make sure no failure goes more than a day without an owner. It doesn’t prove that a server can be recovered within its target time. That needs a real restore, and the Backup Verification & Restore Test Checklist runs one every month, with deeper quarterly tests and measured times against RTO and RPO.

The UK National Cyber Security Centre’s principles for ransomware-resistant backups make the same point about alerts. Backups stopping, mass deletions and changes to retention or administrator accounts should all raise an alert, but alerts only work if someone starts a follow-up process when they fire. The daily checklist is that process, written down and assigned.

Daily backup monitoring

Did it run, and is it covered?

Cadence: every working day, across all clients.

Checks: job status, missed jobs, offsite copies, coverage gaps and one small spot-restore.

Output: a ticket for every failure and a dated daily log.

Restore testing

Can it come back in time?

Cadence: monthly, with deeper quarterly tests.

Checks: full restores of files, servers, databases and SaaS data into an isolated network.

Output: measured restore times against RTO and RPO, signed off per system.

What the Daily Backup Monitoring Checklist Covers

Six phases run every working day across your client base. Phase 7 appears only when a failure has repeated or a system’s recovery point is at risk.

Phase 1

Phase 1: Start the Day’s Check

On Mondays, the check covers every job since Friday’s check, not only last night’s.

  • Open today’s log — carry forward yesterday’s open failures and their tickets
  • Confirm every backup console is reachable — and that alerting still works, because a quiet console can mean alerts have stopped
  • Check platform notices — service status, storage capacity warnings and any subscription or licence close to expiry
  • Check for new clients and service changes — a client onboarded or a backup service added since yesterday belongs on today’s list
Phase 2

Phase 2: Server & Endpoint Jobs

  • Review overnight server jobs for every client — failed, completed with warnings and still running at check time
  • Look for jobs that never started — a missed schedule often raises no failure alert at all
  • Review endpoint backups — devices with no successful backup for longer than your agreed threshold
  • Check the last good backup age for each critical system — against the recovery point objective agreed with the client
  • Check backup storage capacity — repositories and appliances approaching the level where jobs will start to fail
Phase 3

Phase 3: Microsoft 365 & Google Workspace

  • Check Microsoft 365 backup jobs — Exchange Online mailboxes, OneDrive, SharePoint sites and Teams
  • Check Google Workspace backup jobs — Gmail, Drive and shared drives
  • Check the backup service still has access — an expired app consent or a changed service account stops SaaS backups without an obvious error
  • Confirm new users, sites and shared drives are protected — automatic inclusion rules picked them up, or add them now
Phase 4

Phase 4: Offsite Copies & Backup Security

  • Confirm the offsite or cloud copy completed — and how far it lags behind the local backup
  • Confirm the immutable or offline copy is in place — retention locks still on, and the latest copy within the expected window
  • Review security alerts from the backup platforms — deletions, retention changes, new administrator accounts and encryption or anomaly warnings
  • Investigate anything unexplained the same day — a deleted backup set or a shortened retention period is a possible incident, not a housekeeping note
Phase 5

Phase 5: Coverage & Spot-Restore

  • Check new servers and devices are protected — compare yesterday’s new RMM devices with the backup consoles
  • Check new starters in the PSA are in SaaS backup — onboarding tickets closed yesterday should have a protected mailbox and OneDrive or Drive
  • Review exclusions added since yesterday — each has a reason and an approver
  • Run today’s spot-restore from the rotation — one file, folder or mailbox item for one client, restored to an alternate location and opened
Phase 6

Phase 6: Tickets, Log & Handover

  • Raise or update a ticket for every failure — one ticket per issue in the PSA, against the client, linked to the failed job
  • Record today’s results in the log — clients checked, failures, tickets raised and the spot-restore result
  • Answer the escalation question — has any job failed on consecutive runs, or has any system’s last good backup passed its RPO?
  • Hand over open issues — to the next shift or tomorrow’s checker, with the ticket numbers
Phase 7 — Repeated Failure or RPO at Risk

Phase 7: Escalate RPO Risk

Shown only when a job has failed on consecutive runs or a recovery point objective is at risk. It is assigned to the service manager.

  • Escalate to the service manager — system, client, the last good backup and how long data has been unprotected
  • Take an interim backup — a manual job, a snapshot or a copy to another target until the cause is fixed
  • Tell the client where the agreement requires it — what is at risk, what you are doing and when you expect it to be resolved
  • Record any breach of the agreement’s backup terms — for the monthly report and any service credit due
  • Close only on a successful run — record the root cause and the date protection was restored

What Each Backup Status Means, and What to Do Today

Different consoles use different words, but every result falls into one of the rows below. Agree the action for each with your team, so two technicians looking at the same status do the same thing.

Status What it often means Action today
SuccessThe job completed as configuredNone, but check the configuration still covers what it should
Completed with warningsSome files were skipped, often locked or open filesRead the warning; ticket it if the skipped data matters
FailedThe job ran and did not completeTicket it; retry if the cause is clear and the window allows
Missed or not startedThe device was off, the agent stopped or the schedule changedTicket it; check the agent and the device’s status in the RMM
Still runningThe job is overrunning, often after a change in data volumeNote it and check again before midday
Last good backup older than the RPOSeveral failures in a row, whatever today’s status saysEscalate through Phase 7

Write the thresholds down once, so nobody decides them at eight in the morning. Common choices are an endpoint threshold of a few days without a good backup, escalation after two consecutive failures of the same job, and an offsite lag that still leaves the recovery point objective achievable if the primary site is lost. Whatever you choose, align it with what each client’s agreement says, because that is the standard the client will hold you to.

Rotating the daily spot-restore

The spot-restore is a smoke test, not a recovery test. It proves that a recent restore point is readable and that the technician knows the console, in a few minutes. A rotation keeps it from becoming the same easy file every day:

  • Rotate through clients so each one gets at least one spot-restore a month.
  • Rotate the data type: a file from a server, a mailbox item, a OneDrive or Drive file, an endpoint file.
  • Restore to an alternate location, open the item, then delete the copy.
  • Log the client, the item, the restore point used and how long it took.

CISA recommends the 3-2-1 approach of three copies, on two types of storage, with one offsite, and its ransomware guidance calls for offline, encrypted backups that are tested regularly. The daily check confirms the copies exist and are current. The monthly restore test proves they work.

Why Run the Daily Backup Check in CheckFlow?

1

There every morning, for whoever is on

A recurring schedule creates the check each weekday and assigns it to the technician on backup duty, or to a group so whoever is in picks it up. Missed checks show as overdue, not as silence.

2

Escalation that can’t be skipped

One yes/no answer at the end of the check decides whether conditional logic adds the escalation phase. When it appears, it goes to the service manager with a due date the same day.

3

A log that answers the client’s question

Each day’s check is a dated record of what was looked at, what failed and which ticket owns it, with console exports attached. When a client or auditor asks how backups are monitored, you can show them every morning of the last year.

The daily check is one of the recurring processes that the MSP process management guide recommends standardising, because it depends on discipline rather than skill. See how CheckFlow for MSPs runs daily, monthly and quarterly routines across every client from one place.

At month end, the daily logs feed the backup section of the Monthly Managed Services Report Checklist, where every failed job and its fix is explained to the client before they ask. For in-house teams, the Server Maintenance Checklist reviews backup job status weekly alongside server health. Recovery planning beyond the daily check, from recovery tiers to DR testing, is covered in our disaster recovery checklist guide.

Frequently Asked Questions

What should an MSP check in its daily backup review?

+

For every client: last night’s server and endpoint jobs, including jobs that never started; Microsoft 365 and Google Workspace backups; whether the offsite and immutable copies completed; any security alerts from the backup platforms; and whether new devices and users are covered. Then raise a ticket for each failure, run one small spot-restore from a rotation and log the results.

Why check backups every day if the software sends alerts?

+

Because alerts only report what the software notices, and only help if someone acts on them. A job that never started may raise no alert, alert emails get filtered or ignored, and a new server that was never added to a job produces no alert at all. The daily check looks for the absence of a good backup, not just the presence of a failure.

How long does the daily backup check take?

+

It depends on the number of clients and consoles, but a well-run check is usually measured in minutes per client, not hours. Most of the time goes on failures, so it shrinks as recurring causes are fixed. If it regularly takes much longer, look for repeat offenders: the same job failing every week usually needs a problem ticket, not another retry.

Do Microsoft 365 and Google Workspace data need their own backup?

+

Most MSPs treat them as needing one. The platforms keep deleted items for a limited period and offer retention features, but those copies sit in the same tenant, behind the same administrator accounts, and are not an independent copy under your control. Whatever the client chooses, record the decision, and if a SaaS backup is in place, check it daily like any other job.

When should a backup failure be escalated?

+

Set the rule in advance so the technician doesn’t have to judge it. A common rule is to escalate when the same job has failed on two or more consecutive runs, or when a system’s last good backup is older than its agreed recovery point objective. Escalation brings in someone who can authorise an interim backup, extra effort or a conversation with the client the same day.

Who should run the daily backup check?

+

A named technician on a rota, with a backup named for holidays and sickness, or a group so whoever is on shift picks it up. Rotating the duty spreads console knowledge across the team, which matters on the day a restore is urgent. The service manager should review the log weekly and own escalations, not the routine check.

Is CheckFlow free for this template?

+

14-day free trial, no card required. The Business plan is $10 per user per month after the trial. Full details at checkflow.io/pricing.

Catch Every Failed Backup the Morning After, Not the Day You Need It

Free trial — no credit card required.