Production Deployment & Rollback Checklist Template

The deploy command is rarely what breaks production. It is the migration that cannot be reversed, the rollback nobody has tried, and the forty minutes spent watching graphs before anyone decides to go back.

This free production deployment checklist takes one release from a tested build to a closed change. Engineering, platform and DevOps teams use it for service and application deploys: readiness checks on the artefact, migrations and feature flags, a go/no-go approval, steps for rolling, blue-green, canary or all-at-once strategies, verification against your SLOs, and a rollback decision with a time limit set before the deploy starts.

Use This Template Free See Live Example
No Credit Card Required

Deploying Is Not the Same as Releasing

Google’s SRE book reports that roughly 70% of outages are due to changes in a live system. Three processes touch every one of those changes, and teams that run them as one list either drown the deployer in marketing tasks or skip the rollback plan. Keep them apart and link them.

Feature release

Getting it to customers

Covers: specification sign-off, QA, release notes, support briefing, launch timing.

Owner: product, with engineering.

Rollback: often a feature flag, turned off.

Production deployment

This checklist

Covers: artefact, migrations, strategy, verification and the rollback decision.

Owner: the deployer and the on-call engineer.

Rollback: planned, triggered and time-boxed.

Change record

Permission and history

Covers: risk, approval, scheduling and the outcome of the change.

Owner: the change owner and approvers.

Rollback: referenced, not run.

The go-to-market side runs through the Feature Release Process Checklist, and the approval record through the IT Change Management Checklist. This page is the technical runbook in between. Separating deployment from release also means you can ship code with the feature off, then release it with a flag when product is ready.

What the Production Deployment Checklist Covers

Seven phases take a deployment from readiness to close. The strategy chosen on the first task decides which Phase 4 tasks appear, the go/no-go approval halts the checklist before anything ships, and a failed verification opens the rollback phase.

Phase 1

Phase 1: Change, Artefact & Window

The deploy strategy on the first task shows the matching tasks in Phase 4. The approver is named here for Phase 3.

  • Record the service, version, deployer, approver and deploy strategy — Rolling, Blue-green, Canary or All-at-once
  • Deploy the exact artefact that passed CI and staging — identified by version or image digest; never rebuild for production
  • Link the change record and confirm its status — a standard change may be pre-approved; a normal change needs its approval first
  • Check the window against freezes and other changes — peak trading, month-end and deploys to shared dependencies
  • Confirm the deployer and on-call engineer are available for the deploy and the watch period — nobody deploys and then leaves
  • Tell support and dependent teams when it starts — and update the status page if users will see downtime
Phase 2

Phase 2: Migrations, Flags & Rollback Plan

  • Check every migration is backward compatible — the old version must run against the new schema, or rollback stops being an option
  • Split breaking schema changes into expand, migrate and contract — add the new column now; drop the old one in a later release
  • Check locking on large tables — in PostgreSQL, CREATE INDEX CONCURRENTLY avoids blocking writes but cannot run inside a transaction
  • Put risky behaviour behind a feature flag, off by default — turning a flag off is faster than any redeploy
  • Write the rollback steps and what they cannot undo — previous artefact, flag off or traffic switch; sent emails and data written in a new format stay
  • Set rollback triggers and the decision time-box — for example error rate, p95 latency or SLO burn rate above an agreed level, decided within 30 minutes
Phase 3

Phase 3: Go/No-Go

The approver named on the first task records Approved or Not approved on the last task. The checklist halts there, and nothing is deployed until they decide.

  • Confirm CI, staging tests and security scans passed for this build — not for an earlier one with the same version number
  • Confirm dashboards and alerts for the service work — and note the current error rate and latency as the baseline
  • Confirm the rollback plan is attached and its operator is present — a plan only one absent person can run is not a plan
  • Confirm dependent teams are ready — database, platform, support and any team whose service you call
  • Record the go/no-go decision — Approved or Not approved, with the reason
Phase 4

Phase 4: Strategy Steps

Only the tasks for the strategy chosen on the first task appear.

  • Rolling: set surge and unavailability limits and readiness checks — Kubernetes defaults to 25% for both; a stalled rollout is reported after 600 seconds but not rolled back
  • Blue-green: deploy to the idle environment, warm it up and test it directly — before any production traffic reaches it
  • Blue-green: switch traffic and keep the old environment running — until the watch period ends, so going back is one switch
  • Canary: send a small share of traffic to the new version — and compare it with a control group running the old version at the same time
  • Canary: promote in steps only while the canary matches the control — each step recorded with its result
  • All-at-once: take a backup or snapshot and post a maintenance notice — reserve it for services that can tolerate a short outage
Phase 5

Phase 5: Execute & Verify

The verification result is a required Pass or Fail. Fail shows Phase 6.

  • Run the expand migrations and confirm they finished — before the new code that depends on them
  • Start the deploy and record the time — the rollback time-box counts from here
  • Run production smoke tests on the critical user journeys — sign-in, the main transaction and anything this change touched
  • Check health and readiness across every instance — a partial rollout can look healthy in the average
  • Compare error rate, latency and saturation with the baseline and the SLO — on the dashboard noted at go/no-go
  • Record the verification result — Pass, or Fail with what breached
Phase 6 — Rollback

Phase 6: Rollback Decision

Tasks appear only when verification is recorded as Fail.

  • Decide within the time-box: roll back or fix forward — default to rollback when the cause is not yet understood
  • Roll back by the fastest safe route — flag off, traffic switch, kubectl rollout undo or redeploy of the previous artefact
  • Leave expand migrations in place — the old code was built to run against them; reverse schema only if the plan says so
  • Confirm recovery with the same checks used to verify — and record the time; it is your failed deployment recovery time
  • Open an incident if users were affected — and tell support what they will hear about
  • Schedule the postmortem — blameless, within a few working days, while logs and memories are fresh
Phase 7

Phase 7: Watch & Close

  • Hold the agreed watch period before closing — slow leaks and daily jobs show up after the first hour
  • Publish release notes or the internal changelog — what changed, for whom, and any action needed
  • Close the change record with the outcome — successful, rolled back or failed, with the timeline attached
  • Retire the old environment and plan the contract migration — in a later release, once nothing reads the old schema
  • Record deployment data for your delivery metrics — commit and deploy times, planned or unplanned fix, whether it failed and how long recovery took
  • Remove flags that are fully rolled out — each stale flag is a code path nobody tests

Choosing a Deployment Strategy

The strategy decides how much of production sees a bad build and how fast you can take it back. Each option trades cost and complexity for blast radius.

Strategy How traffic moves Rollback Watch out for
RollingInstances replaced in batches; Kubernetes Deployments default to 25% surge and 25% unavailablekubectl rollout undo, which also aborts a rollout in progress; 10 old revisions kept by defaultOld and new versions serve traffic together, so both must work with the same schema and APIs
Blue-greenA full second environment, switched over at onceSwitch back; an Azure App Service slot swap reverses with another swapDouble capacity; AWS CodeDeploy terminates the old EC2 instances after your wait time, up to 2,880 minutes, unless set to keep them
CanaryA small share of traffic first, then larger stepsRoute traffic back to the old versionNeeds per-version metrics and a control group, or the comparison means nothing
All-at-onceEvery instance updated togetherRedeploy the previous artefactDowntime and full exposure; keep it for internal tools and low-traffic windows

Whatever the strategy, define rollback triggers in numbers before you start. Google’s SRE workbook suggests paging when a service burns 2% of a 30-day error budget in an hour, a burn rate of 14.4. The same arithmetic gives a deploy a clear line: if the new version burns budget that fast, roll back first and investigate afterwards.

Measure the process too. DORA’s software delivery metrics are deployment frequency, change lead time, change fail rate and failed deployment recovery time, the name that replaced “time to restore service” in 2023. In 2024 DORA added a fifth, deployment rework rate: the share of deployments that are unplanned fixes for production problems. Phase 7 records the data all five need.

Why Run Production Deployments in CheckFlow?

1

A go/no-go that actually stops

The approver picked on the first task records Approved or Not approved, and the checklist halts until they do. Nobody starts the deploy because the window opened and the approver was in another meeting.

2

One template, four strategies

A required dropdown picks rolling, blue-green, canary or all-at-once, and only that strategy’s steps appear. A data set holds your services with owner, strategy and dashboard link, and tags separate services or clients.

3

Rollbacks leave a record

A Fail on verification opens the rollback phase. Canary steps go in a table inside the task, graphs and smoke test output attach as evidence, and the audit trail timestamps each step, which is the raw data for change fail rate and recovery time.

The approval record for each deployment belongs in the IT Change Management Checklist, and CheckFlow’s change management checklist software shows how change approvals work across engineering and IT.

Every rollback should end in the Incident Postmortem Template. Upstream, the Code Review Checklist flags high-risk changes such as data migrations before they reach this checklist.

Frequently Asked Questions

What is the difference between a deployment and a release?

+

A deployment puts a new version of the code into production. A release makes a change available to users. With feature flags they can happen on different days: deploy with the feature off, check the system is healthy, then release by turning the flag on for a group of users. That keeps the risky technical step and the customer-facing step separate, each with its own rollback.

What should a rollback plan include?

+

The exact steps and who runs them, the triggers that start it in measurable terms, the time limit for deciding, what the rollback cannot undo, and how you will confirm recovery. Write it before the go/no-go meeting and check that the previous artefact still exists. Kubernetes keeps 10 old revisions by default; a pipeline that cleans up old images may not.

How do you make a database migration safe to roll back?

+

Use expand and contract, which Martin Fowler’s site describes as parallel change. First add the new structure without removing the old, so both code versions work. Then move the code and data across. Only in a later release, once nothing reads the old structure, remove it. Rolling back the code then never needs a schema rollback, which is the step most likely to lose data.

Should you use blue-green or canary deployments?

+

Blue-green gives the fastest rollback, one traffic switch, but every user meets the new version at once and you pay for two environments. Canary exposes only a slice of traffic first, so a bad build hurts fewer users, but it needs metrics split by version and a control group to compare against. Blue-green suits services where a full duplicate is affordable; canary suits high-traffic services where a few per cent of requests is still a meaningful sample.

What are the DORA metrics today?

+

DORA measures software delivery with five metrics: deployment frequency, change lead time, change fail rate, failed deployment recovery time and deployment rework rate. The recovery metric was renamed from time to restore service in 2023, and rework rate was added in 2024. A deployment checklist that records start time, outcome and recovery time gives you the data for most of them.

Is CheckFlow free for this template?

+

14-day free trial, no card required. The Business plan is $10 per user per month after the trial. Full details at checkflow.io/pricing.

Decide on Rollback Before You Need It

Free trial — no credit card required.