THE QUESTION
Why does the same outage keep recurring?
Spend your stickers on what matters most. An idea with evidence can take a held seat.
closed · back to Actions
Viewing as guest. Enter a handle to join and take part.
Your sticker budget
One per idea, 5 total. Floor is one vote.
5 places are held for overlooked ideas — nominate one with a link, and it counts once 2 people agree the source is relevant.
why
At stage close one nominee enters by random draw and the rest by nomination count, at most one per stakeholder class. A model checks only whether the page is about the same subject — never whether it proves the claim. That judgement is the room's, and a nomination nobody else stands behind takes no seat.
CLUSTER 1
The on-call rota has the same two people on it most weeks
7 votesdiscuss
The same three services cause most of the pages
14 votesdiscuss
CLUSTER 2
Post-mortem actions are written up and then never scheduled
4 votesdiscuss
Feature flags are never cleaned up, so nobody knows which paths are live
Meaning the repeat outages, not the one-off hardware failure in March.
7 votesdiscuss
CLUSTER 3
Migrations are run by hand from a laptop
2 votesdiscuss
Config lives in five places and drifts between them
I mean the pattern over the last two quarters, not a single incident.
3 votesdiscuss
CLUSTER 4
Every team has its own logging format
3 votesdiscuss
Reliability has no budget line of its own
To be precise: this is about the process, not the person on call that night.
11 votesdiscuss
CLUSTER 5
Alerts fire so often that on-call mutes them
I mean the pattern over the last two quarters, not a single incident.
6 votesdiscuss
Customers report outages before monitoring does
I mean the pattern over the last two quarters, not a single incident.
10 votesdiscuss
CLUSTER 6
Runbooks are out of date within a month of being written
To be precise: this is about the process, not the person on call that night.
2 votesdiscussing (1)
The service map in the wiki is two years old
5 votesin focus ↑ · close ✕
NOT CLUSTERED
Success is measured in features shipped per quarter
1 votesdiscuss
Post-mortems assign actions to people who were not in the room
Scope: the services this team runs, not the whole platform.
8 votesdiscussing (1)
Feature work always outranks reliability work in planning
Scope: the services this team runs, not the whole platform.
4 votesdiscuss
Load tests were last run before the customer base doubled
1 votesdiscuss
There is no time set aside after an incident, so the write-up is done at midnight
2 votesdiscuss
One engineer knows how the payment path actually works
Scope: the services this team runs, not the whole platform.
5 votes
↳ nominated for a seat: https://example.org/ops-incident-review.pdf could not read it (HTTPError) 1 of 2 behind it
discussDependencies are upgraded only when something breaks
Meaning the repeat outages, not the one-off hardware failure in March.
0 votesdiscuss
Tests that fail intermittently are retried until green
10 votesdiscuss
Error budgets exist on a slide and nowhere else
8 votesdiscuss
The staging environment does not resemble production
7 votesdiscuss
Rollbacks take longer than the outage they are meant to end
To be precise: this is about the process, not the person on call that night.
3 votesdiscuss
Deploys go out on Friday afternoons because that is when the sprint ends
2 votesdiscuss
Capacity is added after the outage, never before
0 votesdiscuss
Leadership asks for a root cause within the hour, so the first plausible one wins
I mean the pattern over the last two quarters, not a single incident.
1 votesdiscuss
The incident channel fills with people asking for status instead of giving it
3 votesdiscuss
Retries are unbounded, so a slow dependency becomes a flood
6 votesdiscuss
Nobody owns the shared database schema
To be precise: this is about the process, not the person on call that night.
2 votesdiscuss
Incidents are declared late because declaring one feels like blame
1 votesdiscuss