THE QUESTION
Why does the same outage keep recurring?
Group cards you judge alike. Groups organise reading; they decide nothing.
closed · back to Actions
Viewing as guest. Enter a handle to join and take part.
Clusters form from similarity, names come last
Canonical order: link pairs of ideas you judge similar — links are symmetric (A~B is B~A; direction never matters) and a tie holds once 2 members make it.
why
Links are symmetric — A~B is B~A. Clusters are the groups that emerge from held ties, and only a formed cluster can be named, by a person. Your own links carry a ✕ to withdraw; you cannot remove someone else's. Clusters organise the wall and never gate the vote — two cards making the same CLAIM are a merge back in Clarify.
Links and names are made while Cluster is live; the steward’s ← reopens it.
Feature work always outranks reliability work in planning
Scope: the services this team runs, not the whole platform.
Which service are we talking about here?
unnamed cluster — name it if the group has a common thread, or leave it unnamed (2)
The same three services cause most of the pages
~ The on-call rota has … · helddiscuss
The on-call rota has the same two people on it most weeks
~ The same three servic… · helddiscuss
unnamed cluster — name it if the group has a common thread, or leave it unnamed (2)
Runbooks are out of date within a month of being written
~ The service map in th… · helddiscussing (1)
The service map in the wiki is two years old
~ Runbooks are out of d… · helddiscuss
unnamed cluster — name it if the group has a common thread, or leave it unnamed (2)
Post-mortem actions are written up and then never scheduled
Feature flags are never cleaned up, so nobody knows which paths are live
~ Post-mortem actions a… · helddiscuss
unnamed cluster — name it if the group has a common thread, or leave it unnamed (2)
Every team has its own logging format
~ Reliability has no bu… · helddiscuss
Reliability has no budget line of its own
~ Every team has its ow… · helddiscuss
unnamed cluster — name it if the group has a common thread, or leave it unnamed (2)
Config lives in five places and drifts between them
~ Migrations are run by… · helddiscuss
Migrations are run by hand from a laptop
~ Config lives in five … · helddiscuss
unnamed cluster — name it if the group has a common thread, or leave it unnamed (2)
Alerts fire so often that on-call mutes them
~ Customers report outa… · helddiscuss
Customers report outages before monitoring does
~ Alerts fire so often … · helddiscuss
· not yet clustered (18)
Success is measured in features shipped per quarter
Post-mortems assign actions to people who were not in the room
~ Post-mortem actions a… · helddiscuss
Feature work always outranks reliability work in planning
Load tests were last run before the customer base doubled
There is no time set aside after an incident, so the write-up is done at midnight
One engineer knows how the payment path actually works
Dependencies are upgraded only when something breaks
Tests that fail intermittently are retried until green
Error budgets exist on a slide and nowhere else
The staging environment does not resemble production
Rollbacks take longer than the outage they are meant to end
Deploys go out on Friday afternoons because that is when the sprint ends
Capacity is added after the outage, never before
Leadership asks for a root cause within the hour, so the first plausible one wins
The incident channel fills with people asking for status instead of giving it
Retries are unbounded, so a slow dependency becomes a flood
Nobody owns the shared database schema
Incidents are declared late because declaring one feels like blame