True Causes
← lobby
Why does the same outage keep recurring?
tap any completed stage to explore it
Generate
Clarify
Cluster
Select
Structure
Map
7
Actions
walkthrough · not counted
anyone with the link · listed publicly · federation on
30 people · 30 ideas · 2 groups
Logic → · record
THE QUESTION

Why does the same outage keep recurring?

Spend your stickers on what matters most. An idea with evidence can take a held seat.

closed · back to Actions
Viewing as guest. Enter a handle to join and take part.
Your sticker budget
One per idea, 5 total. Floor is one vote.
5 places are held for overlooked ideas — nominate one with a link, and it counts once 2 people agree the source is relevant.
why
At stage close one nominee enters by random draw and the rest by nomination count, at most one per stakeholder class. A model checks only whether the page is about the same subject — never whether it proves the claim. That judgement is the room's, and a nomination nobody else stands behind takes no seat.
close ✕
The service map in the wiki is two years old
cluster 6
Nothing said yet.
CLUSTER 1
The on-call rota has the same two people on it most weeks
7 votesdiscuss
The same three services cause most of the pages
14 votesdiscuss
CLUSTER 2
Post-mortem actions are written up and then never scheduled
4 votesdiscuss
Feature flags are never cleaned up, so nobody knows which paths are live
Meaning the repeat outages, not the one-off hardware failure in March.
7 votesdiscuss
CLUSTER 3
Migrations are run by hand from a laptop
2 votesdiscuss
Config lives in five places and drifts between them
I mean the pattern over the last two quarters, not a single incident.
3 votesdiscuss
CLUSTER 4
Every team has its own logging format
3 votesdiscuss
Reliability has no budget line of its own
To be precise: this is about the process, not the person on call that night.
11 votesdiscuss
CLUSTER 5
Alerts fire so often that on-call mutes them
I mean the pattern over the last two quarters, not a single incident.
6 votesdiscuss
Customers report outages before monitoring does
I mean the pattern over the last two quarters, not a single incident.
10 votesdiscuss
CLUSTER 6
Runbooks are out of date within a month of being written
To be precise: this is about the process, not the person on call that night.
The service map in the wiki is two years old
NOT CLUSTERED
Success is measured in features shipped per quarter
1 votesdiscuss
Post-mortems assign actions to people who were not in the room
Scope: the services this team runs, not the whole platform.
Feature work always outranks reliability work in planning
Scope: the services this team runs, not the whole platform.
4 votesdiscuss
Load tests were last run before the customer base doubled
1 votesdiscuss
There is no time set aside after an incident, so the write-up is done at midnight
2 votesdiscuss
One engineer knows how the payment path actually works
Scope: the services this team runs, not the whole platform.
5 votes
↳ nominated for a seat: https://example.org/ops-incident-review.pdf could not read it (HTTPError) 1 of 2 behind it
discuss
Dependencies are upgraded only when something breaks
Meaning the repeat outages, not the one-off hardware failure in March.
0 votesdiscuss
Tests that fail intermittently are retried until green
10 votesdiscuss
Error budgets exist on a slide and nowhere else
8 votesdiscuss
The staging environment does not resemble production
7 votesdiscuss
Rollbacks take longer than the outage they are meant to end
To be precise: this is about the process, not the person on call that night.
3 votesdiscuss
Deploys go out on Friday afternoons because that is when the sprint ends
2 votesdiscuss
Capacity is added after the outage, never before
0 votesdiscuss
Leadership asks for a root cause within the hour, so the first plausible one wins
I mean the pattern over the last two quarters, not a single incident.
1 votesdiscuss
The incident channel fills with people asking for status instead of giving it
3 votesdiscuss
Retries are unbounded, so a slow dependency becomes a flood
6 votesdiscuss
Nobody owns the shared database schema
To be precise: this is about the process, not the person on call that night.
2 votesdiscuss
Incidents are declared late because declaring one feels like blame
1 votesdiscuss