An alert fires at 02:10. An engineer pushes a configuration change, the checkout service starts returning errors, and the on-call lead rolls it back 40 minutes later. In the review meeting, the first question is "Who pushed the change?" Within minutes the conversation is about one person's judgement at 2am, and nobody asks why the change was able to reach production without a canary, why the alert took so long to page, or why the runbook was out of date.
This is an illustrative scenario, not a customer case. It is recognisable because it happens: an incident review that fixates on the last person to act produces a tidy story and very little learning.
This guide is for engineering managers and enablement leads who want people to get better at analysing incident evidence and agreeing improvements that can be tested. It covers how to separate skill gaps from the conditions around them, how to design practice, how to debrief, and how to measure whether the reviews are working.
What "blameless" does and does not mean
Google's Site Reliability Engineering book puts it plainly: for a postmortem to be truly blameless, it must focus on identifying the contributing causes of the incident without indicting any individual or team for bad or inappropriate behaviour. It also describes the assumption underneath: everyone involved acted with good intentions and did the right thing with the information they had. (Google SRE book, "Postmortem Culture: Learning from Failure".)
Blameless does not mean consequence-free. The same chapter expects postmortems to end in action items with a priority and an owner. Blameless inquiry still requires accountability, but the accountability moves from "who caused this" to "who will make the system safer, and by when". If your reviews skip the second half, you have a pleasant conversation, not a learning process.
1. Identify the behaviour and its context
Start with the specific behaviour you want to change. "Our incident reviews focus on the last person to act" is observable. "Our culture is blame-heavy" is not.
Then ask whether this is a skill problem at all. Training can only fix some causes of poor performance. Before you build anything, separate three possibilities:
- Isang nawawalang kasanayan. Reviewers do not know how to build a timeline, distinguish a trigger from a contributing factor, or ask "what made this action look reasonable at the time?"
- Mga Insentibo. Reviews feed into performance conversations, so people shape their accounts to protect themselves.
- Authority gradients and past responses. A junior engineer who raised a concern last quarter and was told to stop being negative will not raise one in this meeting, however well trained they are.
Only the first is a training problem. The other two are conditions. If you train the skill and leave the conditions alone, people will rehearse something the workplace punishes. Treat your current view of the cause as a hypothesis, and test it with a few conversations with engineers and a look at recent review documents before you commit to a programme.
2. Set the operational conditions for practice
Practice only transfers if the workplace answers it. Before anyone rehearses, agree these with the people who run reviews:
- Mga inaprubahang responsibilidad. Who facilitates, who writes the timeline, who owns follow-up actions? Make the roles explicit so a facilitator is not improvising authority.
- How leaders receive concerns. If an engineer says "the deploy tool let me skip the canary", what does the director do next? Decide the answer in advance, for example: thank them, log it as a contributing factor, and assign an action.
- The supervisor response. Staff rehearsal needs a matching supervisor response. If you train engineers to speak up in reviews but their managers still ask "whose fault was it?", the training will lose.
A short briefing for managers on the review process is as important as the training for participants. Include a plain statement of how review documents are and are not used in performance management.
3. Design a behaviour-based learning journey
The Carnegie Mellon Eberly Center's guidance on learning objectives makes the useful point that objectives, assessments and instructional strategies should line up with each other (Eberly Center, Learning Objectives). That is design guidance, not evidence about any particular tool or outcome, but it is a good test for this programme: if the objective is to analyse evidence and agree testable improvements, then recall questions alone cannot be the assessment.
Bloom's taxonomy helps here as a way to describe the cognitive demand of the task: analyse, evaluate, create. It describes what the task asks of a person. It does not explain why they do or do not do it, so do not use it to diagnose motivation.
A workable journey has three stages.
| Stage | Aksyon ng mag-aaral | paghahatid |
|---|---|---|
| Paghahanda | Read your incident review policy and a short glossary (trigger, contributing factor, mitigation, action item). A brief diagnostic checks prerequisites. | Self-paced study. A recall check confirms prerequisites only, not review skill. |
| Pagsasanay | Analyse a prepared incident case and agree testable improvements. A reasoned response is required, not a multiple-choice pick. | Interactive or facilitated case practice, with feedback on the quality of contributing factor analysis and follow-through. |
| Pagsusuri at paglilipat | Discuss one ambiguous choice, then complete a new case or a supervised real review. | Live coaching or asynchronous review. Check transfer after people have had time to use the skill at work. |
Keep the case realistic and have an experienced engineer validate it. A case that is technically wrong teaches reviewers to distrust the exercise.
4. Work through the scenario and debrief
Here is how a practice session could run with the 02:10 scenario. The evidence pack contains the deploy log, the alert timeline, the paging policy, the pull request and the runbook, all fictional.
Prompt to the participant. "Explain why the engineer pushed the change when they did. Cite the evidence you used. Tell us what further information or support you would need before you could be confident."
A weak response sounds like this: "The engineer should have checked staging." It names a person, cites nothing, and proposes no change.
A stronger response sounds like this: "The deploy log shows the change went out directly because the canary step is optional in the tool. The pull request had one reviewer who approved after a 20-minute window. The page fired 18 minutes after errors started because the alert threshold was set for a five-minute average. I would want to know how many other changes this month skipped the canary, and whether on-call engineers have ever been told the step is skippable. Proposed actions: make the canary mandatory for this service tier, owner: platform team, check: attempt a skipped-canary deploy in staging and confirm it is blocked; lower the alert window, owner: on-call lead, check: replay last week's errors against the new threshold."
Notice what makes it better. It uses evidence, names constraints and uncertainty, and ends in actions that can be tested.
Now the debrief. Do not just score it. Ask:
- What did the evidence support, and what did you assume?
- Which of your actions would the team actually have authority to make?
- What would make it easier to repeat this behaviour in your next real review? What would make it harder?
The third question matters most. It surfaces the conditions from section 2. If three participants say "my manager will ask who did it", you have found a constraint training cannot remove.
5. Measure behaviour and follow-through
The key question is what evidence would show that learners can analyse incident evidence and agree testable improvements. Here is a rubric to start from. It is a proposal to adapt with an experienced engineer, not a validated readiness standard.
| Saligan | Napapansing ebidensya | 0 | 1 | 2 |
|---|---|---|---|---|
| Pagpapatupad ng gawain | Analyses the evidence and agrees improvements; records the action and the evidence behind it | Wala o walang suporta | Bahagyang, may mga kaugnay na pagkukulang | Complete and justified against agreed criteria |
| pangangatwiran | Explains constraints, alternatives and uncertainty; contributing factors go beyond the last action | Wala o walang suporta | Bahagyang, may mga kaugnay na pagkukulang | Kumpleto at makatwiran |
| Mga Hangganan | Uses approved procedures; asks for help when information or authority is missing | Wala o walang suporta | Bahagyang, may mga kaugnay na pagkukulang | Kumpleto at makatwiran |
Define critical errors separately, for example naming an individual as the cause or proposing an action with no owner. A high total score must never hide a critical failure.
For the measurement itself:
- Quality of contributing factor analysis. Compare reviews written before the programme with reviews written afterwards, scored against the same rubric by someone who does not know which is which. Use unseen follow-up cases as well as real ones.
- Follow-through. Count the share of action items that have an owner, a test, and were completed, out of all action items raised. State the denominator and the time window.
- Pair observation with process changes. Sit in a real review. Check what happened to the action items afterwards.
Be careful with reporting volume. If more concerns and near-misses are logged after the programme, that could mean safety has improved, trust has improved, or a new tool made logging easier. A fall could mean fewer problems or less willingness to speak. Look at what changed in the same period before attributing any movement to training, and avoid claiming an effect size you have not measured.
The Eberly guidance above frames alignment; it is not proof that this approach reduces incidents. Keep your own baseline and follow-up data and treat them as local evidence.
Which constraints may remain even after skills improve
Be honest about these with sponsors:
- Review documents may still feed into performance decisions.
- Teams under delivery pressure may skip actions that are costly but sensible.
- Cross-team actions may stall because nobody owns the boundary.
- Senior engineers may still dominate the discussion.
Better analysis does not remove these. Plan separate process changes for them.
Putting it into practice with AhaSlides
AhaSlides supports adaptive learning, interactive learning, and moving between live and self-paced delivery. In this journey that could mean a self-paced preparation step, a live or facilitated case session with open-response prompts so everyone commits to an analysis before the group discusses, and a follow-up task after the session. Specific authoring, scoring and reporting functions should be checked in a demonstration against your own case before you rely on them, and anything involving AI, simulators or integrations with your incident tooling should be treated as external until shown.
Susunod na hakbang: take the rubric above and adapt it with one of your senior engineers, then build the case you will use. If you would like to see how a live and self-paced version of this journey could work, book an AhaSlides workflow demo once you have the case and rubric drafted.








