An incident postmortem is the document a team writes after production breaks: what happened, why, and what changes so the same failure mode doesn't repeat. The standard references for running one (Atlassian's incident management guide, Google's SRE book, Rootly's and incident.io's template pages) write for a single engineering org: one on-call rotation, one Slack workspace, one payroll. They don't cover who owns a postmortem or its action items once the responding team is blended, part in-house and part nearshore or external. That gap shows up the moment a follow-up ticket needs an owner who works for a different company than the person running the review.
What is an incident postmortem?
Atlassian's incident management guide is direct about what the document is for: it's "blameless by design: the goal is to understand the system and the sequence of events, not to find someone at fault." Atlassian traces the idea to former Etsy CTO John Allspaw, whose writing on blameless postmortems it credits with the approach. The document itself is procedural: a written record produced after an incident closes, covering what happened, the timeline, who was affected, the root cause, and what changes as a result.
The references above agree on that definition. Each one writes for one engineer reviewing another engineer's incident, both on the same payroll, in the same time zone. They don't address what happens when the engineer who owns a follow-up ticket reports to a staffing vendor instead of the company running the postmortem.
What makes a postmortem blameless?
Google's SRE book states the requirement precisely: "For a postmortem to be truly blameless, it must focus on identifying the contributing causes of the incident without indicting any individual or team for bad or inappropriate behavior." That's a specific, testable requirement, and it's easy to break under pressure. Once an incident gets expensive, it's tempting to write the shortest true sentence, and the shortest true sentence often names a person.
Rootly's incident postmortem guide gives a clear before-and-after example of what changes on the page. Poor framing reads: "Alice deployed broken code." The same event written blamelessly reads: "A lack of automated tests allowed faulty code to be deployed without detection." Same incident, different question. The first names a person. The second names a gap in the system, which is the only kind of finding a follow-up ticket can actually fix.
How to run an incident postmortem
1. Set the trigger threshold before the next incident happens
Decide what counts as postmortem-worthy in advance, before the pressure of a live outage forces a rushed call. The SRE book's postmortem culture chapter lists concrete triggers: user-visible downtime or degradation past a threshold, any data loss, on-call intervention such as a rollback or a traffic reroute, a resolution time above a set threshold, or a monitoring failure that meant the incident was found manually. Any stakeholder can also request a postmortem for an event that misses every one of those criteria. Write the thresholds into the runbook once, so nobody argues about it in the incident channel while the outage is still open.
2. Assign a single owner, and name who that is when the owner sits on a vendor team
One person runs the postmortem: schedules it, drives the review, and signs off on the finished document. On a single-employer team, that's usually whoever is on-call that week, and it's rarely written down because it's obvious. On a blended team it isn't obvious, and skipping the decision is how a postmortem ends up with two half-finished drafts, one from the in-house lead and one from the nearshore engineer who actually fixed the bug. Name the owner in the runbook ahead of the next incident, before the retro that follows it forces the question.
3. Schedule the review inside the 24-48 hour window, timezone gaps included
Atlassian's guidance on post-incident review timing is specific: a postmortem is "ideally... drafted immediately after a post-incident review meeting to be held within 24-48 hours of the incident resolving, and not more than five business days." That window assumes everyone who needs to be in the room can actually get there. On a team split between CET and US Pacific, a 24-hour deadline can mean asking someone to join at 6am, or running the review without the person who actually diagnosed the issue. Set the window against your team's real overlap hours, the same adjustment setting review-turnaround SLAs across a blended team already requires for ordinary code review.
4. Build the timeline on one shared clock across the team
incident.io's postmortem template recommends logging timestamps "in the dominant timezone for the incident or UTC if spanning many," rather than whatever local time each contributor happened to be working in. That's easy to skip when half the timeline comes from a support engineer in one city and half from a backend team in another, and skipping it is exactly what makes a cross-team timeline unreadable a week later. Pick one clock before anyone starts filling in timestamps.
5. Run root cause analysis blamelessly
Atlassian's guide lists the 5 Whys as one route to a root cause. The mechanics matter less than the discipline behind them: keep asking why until the answer names a missing safeguard rather than a person. If the analysis stalls on "the engineer made a mistake," it hasn't reached a root cause yet. That's usually the exact point where blameless framing gets tested for real, on an incident that already cost the company something.
6. Write action items with an owner, a tracking ticket, and a verifiable end state
Google's postmortem workbook sets a clear bar: "All action items have both an owner and a tracking number," and "The action items have a verifiable end state." The same source notes that on well-run teams, the postmortem itself gets "written and circulated less than a week after the incident was closed." An action item with no owner, no ticket, and no defined "done" rarely survives past the meeting where it got written down. It ages into a line in a document nobody opens again.
7. Review the backlog of open action items on a fixed cadence
A single postmortem review isn't the finish line. Put open action items from every past postmortem, across every team that holds one, on a recurring agenda: a monthly engineering review, a standing item in a leadership sync, whatever cadence already exists. An action item still open after two review cycles is worth asking about directly, before it ages quietly into next quarter.
What changes when part of the team is nearshore or external
The single-employer assumption behind those guides breaks in three specific places once part of the team is nearshore or externally staffed.
Ownership across a company boundary is the first. A postmortem action item assigned to "the backend team" works fine when everyone on that team reports to the same VP Eng. It stops working when half the backend team works for a staffing vendor and the ticket needs someone with write access to a repo they don't administer, or a decision only an in-house engineer is positioned to make. Assign each action item to a named person confirmed to have the access and authority to close it, rather than to a team name that spans two companies.
Time zone is the second. incident.io's template treats time zone as a question of which clock to log timestamps in, and leaves open whether the 24-48 hour review window from step 3 is actually reachable for a contributor nine hours offset. A team blended from Central and Eastern Europe and Spain, the geography HighCircl staffs engineers from, usually sits close enough to a CET workday that the window holds without much adjustment. A team split between CET and US Pacific is a different case: the engineer who owns the timeline is often asleep by the time the review call starts.
A shadow process is the third, and the quietest to develop. If augmented or nearshore engineers run their own review on their own template, separate from what the in-house team uses, the company ends up with two postmortem processes and no shared record of what either one decided. HighCircl's broader position on shared ceremonies, one set of shared ceremonies for the whole team, applies here directly: run a single postmortem process across every engineer who touched the incident, regardless of who they're contracted through. It also makes postmortem fluency worth checking before someone joins the team, which is why HighCircl's DevOps hiring guide includes a postmortem walkthrough in its vetting.
What to put in the postmortem document
Rootly's list of postmortem components matches the other templates closely: a one-to-two-sentence summary, a chronological timeline with timestamps, who and what was affected, a root cause analysis covering both technical and process-level causes, the steps that resolved the incident, action items with clear ownership, and metrics like MTTR, duration, severity, and cost where it applies. Google's SRE book worked example follows the same shape, with one addition: a dedicated detection section, how the team found out about the incident, kept separate from how they fixed it.
The summary deserves more care than its length suggests. incident.io's guidance on writing the summary is blunt: "We find it useful to pitch it so your boss's boss would understand!" A summary written for the responding engineers, full of internal system names and acronyms, fails the document's actual audience: whoever is deciding six months from now whether a similar risk still exists.
The metrics line is worth taking seriously rather than filling in as a formality. How failed deployment recovery time gets measured draws on the same clock a postmortem's timeline is already reconstructing. If the team tracks DORA metrics anywhere, the postmortem timeline is a natural source for the recovery-time number, already reconstructed, rather than a separate calculation done later from memory.
How postmortem action items turn into technical debt if nobody owns them
An action item with no owner and no deadline doesn't disappear. It sits in whatever tracker it landed in, unassigned, until someone rediscovers the same gap during the next incident. By then it has become technical debt that never got logged as debt.
The fix is the same discipline a debt register with an owner and a cost-to-fix estimate already applies to everything else a team knows is wrong but hasn't fixed yet. Treat unshipped postmortem action items as entries in that same register, rather than a separate list that only gets reviewed after the next outage reminds someone it exists. On a blended team this matters more: an action item owned by an engineer rotating off an engagement in three months needs a named successor, or it becomes debt with no owner at all.
FAQ
What is an incident postmortem?
A written record produced after an incident closes: what happened, the timeline, who or what was affected, the root cause, the steps that resolved it, and the action items meant to prevent a repeat. The document exists to answer those questions once, in one place, rather than relying on whoever remembers the incident best months later.
What makes a postmortem "blameless"?
The analysis has to identify contributing causes in the system rather than naming an individual or team as being at fault. In practice that shows up in how findings get written: a sentence like "a lack of automated tests allowed faulty code to reach production" rather than naming who deployed it. Blameless doesn't mean vague. A postmortem can be specific about what failed without being specific about who to blame for it.
Who owns postmortem action items when part of the team is nearshore or outsourced?
Assign the action item to one named engineer with the access and authority to close it, decided before the review ends. That holds whether the owner is in-house or works for a staffing vendor. The failure mode to avoid is handing an action item to "the backend team" when half that team reports to a different company and nobody confirmed they can actually make the fix.
How soon after an incident should you hold the postmortem?
Atlassian recommends holding the review within 24-48 hours of resolution, with the document drafted soon after and circulated within a week at the latest. On a team split across time zones, schedule that window against actual overlap hours rather than a blanket deadline that assumes everyone who needs to attend is awake and reachable.
Is a postmortem the same thing as a root cause analysis?
No. Root cause analysis is one section inside a postmortem. The document also covers the timeline, the impact, the resolution steps, and the action items that follow from the root cause, the part most likely to get skipped if the write-up stops at "here's why it happened."
