Skip to content

Engineering · For your bots

Incident postmortem

Write a blameless postmortem from an incident's alerts, logs and chat: timeline, impact, root cause, and actions with owners.

In Lobstack: Skills › Library › Add

Download Lobstack

When to use it

An outage or serious bug is resolved and the team needs it written down before memories drift — ideally within two working days.

What your bot will do

  1. 01

    Collect evidence first: alerts, error logs, deploy history, the incident channel, the fix that merged. Every timestamp you write comes from one of these.

  2. 02

    Timeline in UTC: first bad event, detection, response, mitigation, resolution. Call out the gap between start and detection; it is usually the main finding.

  3. 03

    Impact in numbers: duration, users or requests affected, money or data lost. "Some users" is not an impact.

  4. 04

    Root cause: ask why until the answer is a system or a process, not a person. Name the triggering change and the condition that let it through.

  5. 05

    What went well and what did not, briefly and honestly.

  6. 06

    Actions as a table: action, owner, due date, and whether it prevents a repeat or only speeds up detection. Three solid actions beat ten wishes.

  7. 07

    Blameless throughout: what people did and what they knew at the time, never who was careless.

  8. 08

    Post it as a document marked draft, and ask the incident lead to review it before it is shared.

Where the evidence cannot settle the cause, say so rather than picking the likeliest story.

The file

incident-postmortem/SKILL.md36 lines
---name: incident-postmortemdescription: "Write a blameless postmortem from an incident's alerts, logs and chat: timeline, impact, root cause, and actions with owners."metadata:  title: "Incident postmortem"  category: engineering  tags: [incidents, reliability]--- ## When to use this An outage or serious bug is resolved and the team needs it written down beforememories drift -- ideally within two working days. ## How 1. Collect evidence first: alerts, error logs, deploy history, the incident   channel, the fix that merged. Every timestamp you write comes from one of   these.2. Timeline in UTC: first bad event, detection, response, mitigation,   resolution. Call out the gap between start and detection; it is usually the   main finding.3. Impact in numbers: duration, users or requests affected, money or data   lost. "Some users" is not an impact.4. Root cause: ask why until the answer is a system or a process, not a   person. Name the triggering change and the condition that let it through.5. What went well and what did not, briefly and honestly.6. Actions as a table: action, owner, due date, and whether it prevents a   repeat or only speeds up detection. Three solid actions beat ten wishes.7. Blameless throughout: what people did and what they knew at the time, never   who was careless.8. Post it as a document marked draft, and ask the incident lead to review it   before it is shared. Where the evidence cannot settle the cause, say so rather than picking thelikeliest story.

Get Lobstack.