Problem Management Checklist Template

The same ticket gets closed four times a quarter, each time by a different engineer, each time with a restart. Nobody owns the question of why it keeps happening, so it keeps happening.

Incident management restores service. Problem management removes the reason it failed, and in many IT teams it exists on the org chart but not in the working week. This free problem management checklist follows the three phases of the ITIL 4 practice: problem identification, problem control and error control. It is for service desk managers, problem managers, infrastructure leads and managed service providers who need problem records to reach “fixed and verified”. Each run produces a problem statement backed by linked incidents, a root cause with the evidence and method attached, a known error record and workaround where one exists, a change request for the permanent fix, and a closure check that proves the incidents stopped. A separate monthly run handles trend analysis and the review of every open known error.

Use This Template Free See Live Example
No Credit Card Required

Problems, Known Errors and Workarounds: What Each Record Is For

ITIL 4 defines a problem as a cause, or potential cause, of one or more incidents. The purpose of the practice is to reduce both the likelihood and the impact of incidents by finding those causes and by managing workarounds and known errors. A known error is a problem that has been analysed but not yet resolved. A workaround is anything that reduces or removes the impact of an incident or problem while a full resolution is not yet available. Keeping them separate matters: the service desk needs the workaround today, the change authority needs the root cause evidence next week, and the problem owner needs to see which known errors have waited longest.

The practice also works at a different scale from a single incident review. An incident postmortem looks back at one outage in depth and produces actions for that event. Problem management looks across many incidents, including small ones nobody would write a postmortem for, and asks what they share. A postmortem often opens a problem record, and the investigation carries on there after the meeting ends.

Post-incident review

One incident, looked at closely

Starts from: a single SEV-1 or SEV-2, or any event that meets the review criteria.

Question: what happened, and what should we change after this outage?

Timescale: days, while memories and logs are fresh.

Output: a written review and a list of owned actions.

Problem record

One cause, traced across many incidents

Starts from: trend analysis, a recurring ticket pattern, a major incident with no confirmed cause, or a supplier notice.

Question: why does this keep happening, and what would stop it for good?

Timescale: weeks, sometimes longer if the fix needs a release or a hardware change.

Output: a root cause, a known error and workaround, and a change that removes the cause.

What the Problem Management Checklist Covers

Six phases take each problem from identification to verified closure. Answers on the record decide which workaround and fix tasks appear. A seventh phase runs on its own as the monthly problem review.

Phase 1

Phase 1: Identify & Log the Problem

ITIL 4 problem identification is both reactive and proactive. Record the source: a trend needs different evidence from a major incident.

  • Record how the problem was found — trend analysis, a major incident, a pattern the service desk spotted, or a notice from a supplier or developer
  • Link every related incident record — search the last 90 days for the same configuration item, symptom or error text; the count is your evidence for priority
  • Check the known error database first — if the cause is already recorded, link the new incidents to that record instead of opening a duplicate
  • Write a problem statement that describes the fault, not a guess at the cause — “payroll export fails between 01:00 and 02:00 on the first working day”, not “database issue”
  • Name the affected services and configuration items — from the CMDB or service catalogue, so the investigation starts from the right components
Phase 2

Phase 2: Prioritise & Assign

  • Set priority from impact and likelihood of recurrence — ten minor incidents a week on a payment service can outrank one noisy outage on an internal wiki
  • Assign a problem owner and investigators — one owner accountable for the record, with engineers from each team that runs an affected component
  • Set a target date for a confirmed root cause — driven by priority, with a review date if the target is missed rather than a silent slip
  • Decide whether the service desk needs an interim workaround now — if incidents are still arriving, first line needs something to apply before the cause is known
  • Tell the service desk the problem is open — new incidents with the same signature get linked to the record instead of being worked from scratch
Phase 3

Phase 3: Investigate & Find the Root Cause

  • Build a timeline of occurrences and changes — every linked incident alongside deployments, patches, configuration changes and capacity events in the same window
  • Specify the problem as IS and IS NOT — which servers, times, users and transactions are affected, and which comparable ones are not; the differences point at the cause
  • Map candidate causes on an Ishikawa diagram — group them under people, process, technology, suppliers and environment so no category is skipped
  • Drill each plausible cause with 5 Whys — stop at a condition the organisation can change, never at a person
  • Test the most likely cause against the evidence — a true cause explains every IS and every IS NOT; reproduce it in a test environment where you can
  • Record the root cause, the evidence and the method used — attach the diagram, log extracts and test results to the task
Phase 4

Phase 4: Known Error & Workaround

Tasks 2–4 appear only when the record answers Yes to “Is a workaround available?”. Task 5 appears only when it answers No.

  • Record the problem as a known error — the cause is analysed and the fix is not yet in place, so the record changes status and stays open
  • Document the workaround step by step — written so a first-line analyst can apply it without escalating to the team that found it
  • Publish the workaround to the service desk knowledge base — linked to the known error, so incidents can cite it and be resolved faster
  • Measure how well the workaround works — time to apply, and whether incidents that use it still escalate to second line
  • Record the business impact of having no workaround — it raises the priority of the permanent fix in Phase 5
Phase 5

Phase 5: Error Control & the Permanent Fix

The “Permanent fix route” answer controls this phase. Tasks 2–4 appear for “Change required”, task 5 for “No change needed” and task 6 for “Accept the known error”.

  • Compare the options for a permanent fix — cost, risk and time to deliver, set against the cost of living with the workaround
  • Raise a request for change for the fix — the problem record is the justification, and the change follows your normal approval route
  • Attach the root cause evidence to the change — so the change authority can weigh the risk of fixing against the risk of waiting
  • Track the change through to implementation — the problem stays open until the change is deployed, not when it is approved
  • Apply and document a fix that needs no change — such as a corrected knowledge article or a revised operating procedure
  • Record the decision to accept the known error — who accepted the risk, why, and the date it will be reassessed
Phase 6

Phase 6: Verify & Close

  • Watch for recurrence over an agreed period — for example 30 days, or through the next month-end run for a cyclical fault
  • Confirm the fix with the service owner — the people who raised the incidents agree the symptom has gone
  • Close the known error and retire its workaround article — a stale workaround sends future analysts down the wrong path
  • Record what made the cause hard to find — missing logs, weak monitoring or unclear ownership become improvement actions
  • Close the problem record with owner sign-off — the linked incidents, root cause and change stay attached as the audit trail
Phase 7 — Monthly Review Only

Phase 7: Monthly Problem Review

Shown only when the checklist is started as the monthly review. Phases 1–6 are hidden for that run, and a recurring schedule starts it on the first working day of the month.

  • Run trend analysis on last month’s incidents — group by configuration item, category and error text; any cluster above your threshold becomes a candidate problem
  • Check every major incident from the month — each one whose review left the cause unconfirmed should have a problem record
  • Reassess every open known error — current impact, the cost of a permanent fix and how well the workaround is holding up
  • Chase problems past their root cause target date — agree a new date with the owner or escalate to the service owner
  • Review supplier notices and published defects — known bugs in products you run are problems you can log before they cause an incident
  • Report the problem KPIs — open problems by age, incidents linked to known errors, and problems closed with a verified fix

Choosing a Root Cause Analysis Technique

A repeatable fault on one component rarely needs more than a timeline and 5 Whys. An intermittent fault that affects some servers and not others is where a structured comparison earns its time. Phase 3 asks the investigator to record which method was used, so the monthly review can see which approaches actually produce confirmed causes.

Technique Best for Watch out for Attach to the record
Timeline and change correlationAny problem; always the first stepChanges made outside the change process never appear in the logTimeline with incident and change references
5 WhysSingle-path faults with a clear chain of eventsStopping at “human error”, or following one chain when several causes combinedThe question chain, with evidence for each answer
Ishikawa (fishbone) diagramGroup sessions where many causes are plausibleA long list of possibilities with nothing testedThe diagram, with each cause marked tested or ruled out
Kepner-Tregoe problem analysisIntermittent faults that affect some systems, times or users and not othersThin IS NOT data, which leaves no distinctions to work fromThe IS / IS NOT specification, distinctions, changes and the cause tested against both columns

Why Run Problem Management in CheckFlow?

1

Proactive work that starts itself

Trend analysis is the first thing dropped in a busy month. A recurring schedule starts the monthly review on the first working day and assigns it to the problem manager, so a skipped month shows as an overdue checklist.

2

The record shapes itself to the problem

Answer whether a workaround exists and how the fix will be delivered, and conditional logic shows only the tasks that apply. Linked incident counts, root cause method and target dates are captured as form fields, and the fishbone diagram and log extracts are attached as evidence.

3

Known errors that do not go stale

A dynamic due date sets the reassessment date when a known error is accepted, and the monthly review brings every open one back in front of an owner. The activity trail shows who accepted each risk, who approved each fix and when the problem was verified closed.

Most permanent fixes reach production as a change. CheckFlow’s change management checklist software shows how the request, approval and implementation run with the same evidence trail, and the IT Change Management Checklist is the template to raise from Phase 5.

Problem management starts where live response stops. The Incident Management Checklist restores service and checks the known error database before anyone starts diagnosing, which is where this template’s workarounds pay off.

Frequently Asked Questions

What is problem management in ITIL 4?

+

It is the ITIL 4 practice that reduces the likelihood and impact of incidents by finding their actual and potential causes and managing workarounds and known errors. It has three phases. Problem identification finds and logs problems, from incident trends, major incidents, recurring tickets and supplier information. Problem control analyses the problem and documents workarounds and known errors. Error control manages known errors through to a permanent fix, usually by raising a change.

What is the difference between a problem and a known error?

+

A problem is the cause, or potential cause, of one or more incidents, and it may not be understood when it is first logged. A known error is a problem that has been analysed but not yet resolved. The record usually gains a documented workaround at that point, and stays a known error until the permanent fix is verified or the organisation formally accepts it.

What is the difference between incident management and problem management?

+

Incident management restores normal service as quickly as possible, and a restart or a workaround is a perfectly good incident resolution. Problem management asks why the incident happened and removes the cause so it does not happen again. The two are linked: incidents feed problem records, and the known errors and workarounds that problem management publishes help the service desk resolve future incidents faster.

When should a problem record be opened?

+

Open one when the same symptom or configuration item keeps generating incidents, when a major incident was resolved without a confirmed cause, when trend analysis shows a cluster above your agreed threshold, or when a supplier publishes a defect in something you run. Set the threshold in advance, for example three incidents on one component in 30 days, so nobody has to notice the pattern by chance.

What is proactive problem management?

+

Reactive problem management starts from incidents that have already happened. Proactive problem management looks for causes before they produce one: trend analysis of incident records, monitoring data, and known defects reported by suppliers, developers and test teams. In this template that work lives in the scheduled monthly review.

Which root cause analysis techniques are used in problem management?

+

The most common are 5 Whys, Ishikawa (fishbone) diagrams and Kepner-Tregoe problem analysis, which compares what the problem IS with what it IS NOT to isolate the cause. Most investigations also start from a timeline of incidents and changes. The test at the end is the same for all of them: a confirmed root cause explains every occurrence, and removing it stops the incidents.

Is CheckFlow free for this template?

+

14-day free trial, no card required. The Business plan is $10 per user per month after the trial. Full details at checkflow.io/pricing.

Fix the Cause Once Instead of the Symptom Every Week

Free trial — no credit card required.