The first nine days of this series were about DevOps culture and getting infrastructure under control — silos, automation, code as the source of truth. Today starts a different phase: Site Reliability Engineering, the discipline of making sure the systems that culture and automation produce actually stay up. It's worth being direct about why this gets nine days of its own rather than a paragraph inside the platform engineering material: reliability isn't a property that shows up automatically once you've automated enough. It has to be engineered in, on purpose, the same way a feature does.
Most teams don't disagree with that in principle. They disagree with it in practice, every time a roadmap gets prioritized and "make the checkout flow faster" beats "add a retry budget to the payment integration" because one of those has a visible demo and the other doesn't — until the day it does, in production, in front of customers.
Reliability is the feature users notice by its absence
Nobody opens a support ticket that says "thanks for the 99.95% uptime this quarter." They open one the moment the number dips below whatever they'd silently assumed was a given. That asymmetry is the whole argument: reliability doesn't earn credit when it's present, but it spends credit fast when it's missing.
Note
This is the same shape as security and performance — features that don't show up on a demo but absolutely show up in churn, trust, and word of mouth. The difference with reliability is how directly it compounds: an unreliable system doesn't just lose the user who hit the outage, it slows down every future feature because the team's now firefighting instead of building.
What it costs to treat reliability as an afterthought
| Without reliability engineered in | With it |
|---|---|
| Incidents are frequent and mostly reactive | Incidents are rarer, and mostly caught before users notice |
| Every outage is an all-hands fire drill | Response is defined, rehearsed, and boring on purpose |
| Trust erodes a little more each incident | Trust compounds — the product is dependable by reputation |
| Engineers spend increasing time firefighting | Engineers spend most of their time building |
| Downtime cost climbs as the business scales | Downtime cost is bounded by design, not luck |
The multiplier in that last row is easy to underestimate early on. A two-hour outage costs a five-person startup an afternoon of apologies. The same two-hour outage at a company processing payments for ten thousand merchants is a different order of magnitude entirely — the cost of unreliability doesn't grow linearly with scale, it grows with however much now depends on the system staying up.
Reliability is designed, not hoped for
The instinct to treat reliability as something you add later — "we'll get to hardening this once we've shipped the feature" — comes from a real, well-intentioned trade-off: ship the thing people are waiting for. What that instinct misses is that most of what makes a system reliable is cheap to design in from the start and expensive to retrofit:
- Design for failure — assume dependencies will time out, disks will fill, and nodes will die, because on a long enough timeline they will.
- Observability — you cannot fix what you cannot see; instrumenting after an incident means the next one looks just as opaque as the last.
- SLOs and error budgets — an explicit, agreed number for "how much unreliability is acceptable" turns an endless debate into a shared budget. Tomorrow's article is entirely about this.
- Automation — most reliability work is removing manual steps from recovery, because manual steps are exactly where 2am mistakes happen.
- Fast, practiced recovery — a system that fails and heals in ninety seconds behaves, from a user's perspective, almost like a system that didn't fail.
- Blameless learning culture — an org where postmortems hunt for "why did the system allow this" instead of "who broke it" actually fixes root causes instead of just consequences. Day 16 of this series goes deep on that specifically.
None of these are exotic. They're ordinary engineering decisions, made deliberately, early, instead of accidentally, late, under pressure.
Reliability compounds like technical debt, just in the other direction
A small investment now — a retry with backoff, a health check that actually checks something meaningful, a runbook written before the incident instead of during it — is cheap today and valuable for years. Skipping it is also cheap today, and that's exactly the trap: the cost doesn't disappear, it just gets deferred to whichever engineer is on call when it finally comes due.
What the next few days build on this
Today's argument is deliberately abstract — reliability matters, here's why, here's the shape of what building it in looks like. The next several days in this series make it concrete: SLIs, SLOs and error budgets as the actual measurement system (Day 11), toil as the tax you pay for not automating (Day 12), monitoring versus observability as genuinely different capabilities, not synonyms (Day 13), and blameless postmortems as the mechanism that turns incidents into permanent fixes instead of a monthly rerun of the same page.
The through-line across all of it: reliability isn't a phase you complete once. It's a discipline you keep practicing, and the teams that treat it as a first-class feature from the start spend far less of their lives firefighting than the ones who bolt it on after the outage that finally made it unavoidable.





