Incident postmortem methodology and templates. Use when conducting incident postmortems, writing postmortem reports, establishing postmortem processes, or performing post-incident analysis.
Scope: Technical/engineering incident postmortem templates, methodologies, and best practices Load if: Conducting incident postmortems, writing postmortem reports, establishing postmortem processes, incident response workflows, post-incident analysis Prerequisites: None (standalone guideline)
Postmortems are structured reviews after incidents to understand what happened, why it happened, and how to prevent recurrence. Principles: blameless culture (systems, not people), learning focus, timely execution (48-72 hours), actionable outcomes (action items with timelines).
Include these sections in order:
Include: title, ID, date, duration (ISO 8601 time range), severity (P0/P1/P2), brief description (2-3 sentences), key metrics (downtime, affected users, error rates)
Include: customer impact (users, regions, services), business impact (revenue, SLA violations, reputation), technical impact (degradation, data loss, performance), duration
Include: discovery time/method, key events chronologically (local timezone, ISO 8601), response actions, resolution time, post-resolution verification
YYYY-MM-DDTHH:MM:SS±HH:MM - Alert triggered: «Alert description»
YYYY-MM-DDTHH:MM:SS±HH:MM - On-call engineer paged, investigation started
YYYY-MM-DDTHH:MM:SS±HH:MM - Root cause identified: «Root cause description»
YYYY-MM-DDTHH:MM:SS±HH:MM - Mitigation applied: «Mitigation action»
YYYY-MM-DDTHH:MM:SS±HH:MM - Service restored, monitoring confirmed normal operation
Include: primary root cause, contributing factors (system design, process gaps, monitoring gaps, documentation gaps, training gaps, environmental factors), analysis methodology (Five Whys, fishbone diagram, timeline analysis), evidence/data
Include: immediate mitigation actions, long-term fixes, verification steps, rollback procedures (if applicable)
Include: ID, description, owner (individual or team), priority (P0/P1/P2 or High/Medium/Low), target completion date, success criteria
Tracking: Use structured lists, issue trackers, or project management tools.
Include: what went well, what could be improved, process improvements, tooling improvements, knowledge gaps
Include: internal notifications, customer communications (if applicable), status page updates, post-incident review meetings, documentation updates
Ask "why" five times to drill down to root cause:
Categorize contributing factors:
Identify: trigger events, cascade failures, response delays, resolution bottlenecks
Required: Incident commander, primary responders, on-call engineers involved, team leads from affected systems, product/engineering managers (if customer impact)
Optional: SRE/DevOps team members, security team (if security-related), customer support (if customer impact), executive stakeholders (for high-severity incidents)
Core principle: Focus on systems, not people. Incidents are system failures; blame prevents learning.
Guidelines: Use "we" not "they". Focus on "what" and "why" not "who".
When conducting postmortems:
@smith-clarity/SKILL.md - Root cause analysis techniques (Five Whys, fishbone)@smith-validation/SKILL.md - Hypothesis testing