All guides

How to write a blameless incident postmortem

A postmortem explains what broke, why it broke and what the team will change so it does not break the same way. It looks at the system and the process, not at who was on call.

Topic
Team practice
Reading time
3 min
Updated

What a postmortem is

An incident postmortem, or incident review, is a written account of something that went wrong in production. It exists to learn, not to assign fault. A team that punishes the person nearest the failure learns to hide failures. A team that studies them gets safer.

Write it soon, and write it together

Start within a few days, while the details are fresh. The people who were involved should write or review it together. One person drafts, the others correct. The first draft is almost always missing something that someone else remembers.

The sections

  1. Summary: what broke, who was affected and for how long, in two or three sentences.
  2. Impact: when it started, when it was detected, when it was resolved, who was affected and how severe it was.
  3. Timeline: what happened, what was seen and what was done, in time order, in plain words.
  4. Root cause: the chain of causes, down to the one that would have prevented the incident if it had been removed.
  5. What went well, what went badly and where we got lucky: the last one is the most honest.
  6. Actions: concrete changes, each with an owner and a date.

Finding the root cause

Ask “why?” until the answer is something you can change. “The deploy broke the database” is a symptom. “The migration was run without a dry run, because the dry run step is not in the release checklist” is a cause you can fix. There is rarely one cause. Write the chain, and put the actions at the links you can break.

Write actions that will happen

  • Each action has an owner and a date, and goes on the team’s board, not only in the document.
  • Prefer changes that make the failure harder to repeat over reminders to be careful.
  • Keep the list short. Three actions that get done beat twelve that do not.

Mistakes to avoid

  • Naming a person as the cause.
  • Stopping at the first answer to “why?”.
  • Writing it, filing it and never reading it again. Review old postmortems in the retrospective.

How serious was it?

Agree a small scale for severity before you need it, so the first minutes of an incident are not spent arguing about words.

  • 1

    What it looks like

    The product is down, or data is lost or exposed

    What follows

    A full postmortem, reviewed by the whole team

  • 2

    What it looks like

    A main feature is broken for many people

    What follows

    A postmortem, shared with the team

  • 3

    What it looks like

    A feature is degraded or broken for a few people

    What follows

    A short written note and a task

  • 4

    What it looks like

    A minor problem with a workaround

    What follows

    A task on the board

A short example

Share what you learn

A postmortem helps only the people who read it. Send it to the whole team, keep every one in one place and look back through them in the retrospective. Patterns show up across incidents that no single review can see: the same missing test, the same confusing alert, the same unowned system.

A postmortem in Kanso

The incident postmortem template has the summary, impact table, timeline, root cause and actions ready. Turn each action into a card on the board, and link the incident thread from team chat so the conversation and the review stay together.

Common questions

A postmortem that looks at what the system and the process allowed to happen, instead of who made the mistake. People are more honest about failures when they do not fear being blamed.

Keep reading

All guides

Put this guide to work.

Free does not expire. No card required.