Every system fails eventually — that's been the working assumption for this entire phase of the series, from MTTR (Day 14) to blameless postmortems (Day 16). What's left is the part that happens in real time, while the incident is still live: how it gets detected, who gets pulled in, how fast a fix gets found, and how clearly everyone involved understands what's happening. Two teams with identical systems and identical failure rates can have wildly different outcomes from the same category of incident, and the difference is almost entirely process, not luck.
The incident lifecycle
1. Detect → an alert fires, ideally before a user reports it
2. Acknowledge → on-call confirms they're on it, quickly
3. Triage → assess severity and actual impact
4. Respond → mitigate or roll back — stop the bleeding first
5. Resolve → confirm the fix, close the incident
6. Postmortem → what happened, why, what's next (Day 16)
7. Learn → turn the postmortem's findings into shipped fixesThe loop closes back to step 1 not because incidents are supposed to repeat, but because each pass through steps 6 and 7 should make some future incident detectable, or preventable, or at least faster to resolve. A team that skips 6 and 7 is running the same five-step loop forever, with the ceiling on improvement being whatever individual on-call engineers happen to remember from last time.
Severity matters because response should scale with impact
Treating every incident like a five-alarm fire burns out a team fast; treating a checkout outage like a minor cosmetic bug is its own kind of failure. A basic severity scale gives everyone a shared, fast way to calibrate response:
| Severity | Impact | Example |
|---|---|---|
| P1 – Critical | Major outage, service down | Checkout completely unavailable |
| P2 – High | Partial outage, major feature broken | Search returns no results |
| P3 – Medium | Degraded performance, workaround exists | Reports load slowly |
| P4 – Low | Minor issue, no real user impact | Typo in an error message |
The value of the scale isn't precision — it's speed of agreement. "This is a P1" tells everyone in the incident channel exactly how much urgency and how many people to pull in, without a debate.
A real timeline, and why 18 minutes is a good outcome
10:02 Alert fires — 5xx error rate spiking
10:03 Acknowledged — on-call confirms, declares P1
10:05 Investigate — dashboards + traces point at a recent deploy
10:12 Mitigate — kubectl rollout undo, rollback triggered
10:20 Resolved — error rate back to baseline, incident closedEighteen minutes of partial outage sounds bad in isolation. It's actually a strong outcome, because every stage was fast: three minutes to acknowledge, two more to diagnose using the tracing this series covered on Day 13, and the fix was a rollback — a rehearsed, low-risk, one-command action, not an improvised patch written under pressure. Compare that to a team without tracing, without a clear on-call owner, and without a tested rollback path: the "investigate" step alone can eat the better part of an hour.
Tools support the process, they don't replace it
PagerDuty, Opsgenie, Slack and Grafana all show up in a mature incident workflow, and none of them fix an incident by existing. A team with expensive tooling and no clear on-call ownership will still flounder; a team with clear roles, tested runbooks, and a shared severity scale can run an effective incident response over a phone and a shared doc. People and process first — tools accelerate a process that already works.
The mistakes that turn a small incident into a long one
- Slow acknowledgment — an alert sitting unacknowledged for ten minutes is ten minutes of pure loss, before anyone's even started diagnosing.
- No clear owner — "who's driving this?" asked mid-incident is a sign the on-call rotation or escalation path isn't actually clear.
- Poor communication — stakeholders finding out about an outage from customers instead of from the team is a trust problem that outlasts the incident itself.
- Skipping the postmortem — treats every incident as a one-off, which guarantees the same failure mode returns, undiagnosed, unfixed, on someone else's shift.
Closing out SRE Foundations
This is the last article in the SRE Foundations phase that opened on Day 10. The throughline across all eight days: reliability isn't hoped for, it's measured (SLOs), protected (toil reduction), made visible (observability), recovered from quickly (MTTR), investigated honestly (blameless postmortems), and responded to with a real process (today). Tomorrow, Day 18 asks who's actually responsible for making all of that the default experience for every team in an organization — not something each team has to rediscover independently. That's platform engineering, and it's the next phase of this series.





