Skip to content
SRE

Day 17: Incidents Will Happen. How You Respond Determines the Damage.

The outage isn't the whole story — the eighteen minutes between the alert firing and the fix landing are. Closing out the SRE Foundations phase with the operational discipline that decides how bad an incident actually gets.

Thamunkpillai 5 min read

Every system fails eventually — that's been the working assumption for this entire phase of the series, from MTTR (Day 14) to blameless postmortems (Day 16). What's left is the part that happens in real time, while the incident is still live: how it gets detected, who gets pulled in, how fast a fix gets found, and how clearly everyone involved understands what's happening. Two teams with identical systems and identical failure rates can have wildly different outcomes from the same category of incident, and the difference is almost entirely process, not luck.

The incident lifecycle

Text
1. Detect      → an alert fires, ideally before a user reports it
2. Acknowledge → on-call confirms they're on it, quickly
3. Triage      → assess severity and actual impact
4. Respond     → mitigate or roll back — stop the bleeding first
5. Resolve     → confirm the fix, close the incident
6. Postmortem  → what happened, why, what's next (Day 16)
7. Learn       → turn the postmortem's findings into shipped fixes

The loop closes back to step 1 not because incidents are supposed to repeat, but because each pass through steps 6 and 7 should make some future incident detectable, or preventable, or at least faster to resolve. A team that skips 6 and 7 is running the same five-step loop forever, with the ceiling on improvement being whatever individual on-call engineers happen to remember from last time.

Severity matters because response should scale with impact

Treating every incident like a five-alarm fire burns out a team fast; treating a checkout outage like a minor cosmetic bug is its own kind of failure. A basic severity scale gives everyone a shared, fast way to calibrate response:

SeverityImpactExample
P1 – CriticalMajor outage, service downCheckout completely unavailable
P2 – HighPartial outage, major feature brokenSearch returns no results
P3 – MediumDegraded performance, workaround existsReports load slowly
P4 – LowMinor issue, no real user impactTypo in an error message

The value of the scale isn't precision — it's speed of agreement. "This is a P1" tells everyone in the incident channel exactly how much urgency and how many people to pull in, without a debate.

A real timeline, and why 18 minutes is a good outcome

Text
10:02  Alert fires        — 5xx error rate spiking
10:03  Acknowledged       — on-call confirms, declares P1
10:05  Investigate        — dashboards + traces point at a recent deploy
10:12  Mitigate           — kubectl rollout undo, rollback triggered
10:20  Resolved           — error rate back to baseline, incident closed

Eighteen minutes of partial outage sounds bad in isolation. It's actually a strong outcome, because every stage was fast: three minutes to acknowledge, two more to diagnose using the tracing this series covered on Day 13, and the fix was a rollback — a rehearsed, low-risk, one-command action, not an improvised patch written under pressure. Compare that to a team without tracing, without a clear on-call owner, and without a tested rollback path: the "investigate" step alone can eat the better part of an hour.

Tools support the process, they don't replace it

PagerDuty, Opsgenie, Slack and Grafana all show up in a mature incident workflow, and none of them fix an incident by existing. A team with expensive tooling and no clear on-call ownership will still flounder; a team with clear roles, tested runbooks, and a shared severity scale can run an effective incident response over a phone and a shared doc. People and process first — tools accelerate a process that already works.

The mistakes that turn a small incident into a long one

  • Slow acknowledgment — an alert sitting unacknowledged for ten minutes is ten minutes of pure loss, before anyone's even started diagnosing.
  • No clear owner — "who's driving this?" asked mid-incident is a sign the on-call rotation or escalation path isn't actually clear.
  • Poor communication — stakeholders finding out about an outage from customers instead of from the team is a trust problem that outlasts the incident itself.
  • Skipping the postmortem — treats every incident as a one-off, which guarantees the same failure mode returns, undiagnosed, unfixed, on someone else's shift.

Closing out SRE Foundations

This is the last article in the SRE Foundations phase that opened on Day 10. The throughline across all eight days: reliability isn't hoped for, it's measured (SLOs), protected (toil reduction), made visible (observability), recovered from quickly (MTTR), investigated honestly (blameless postmortems), and responded to with a real process (today). Tomorrow, Day 18 asks who's actually responsible for making all of that the default experience for every team in an organization — not something each team has to rediscover independently. That's platform engineering, and it's the next phase of this series.

Written by Thamunkpillai · Have a question or a correction? Reach out via email.

Get the useful stuff, not the noise.

Occasional notes on engineering, Platform Engineering, AI, cloud and things I’m learning along the way.

No spam. Unsubscribe anytime.