Yesterday's article argued that reliability has to be engineered in deliberately, not hoped for. The tool that turns "be reliable" from a vague aspiration into something a team can actually manage is a small, specific vocabulary: SLI, SLO, and SLA. They get used interchangeably in casual conversation, which is exactly the problem — they measure three different things, and conflating them is how "is the service reliable enough?" turns into an argument instead of a number everyone already agreed on.
The three terms, in the order you'd actually use them
SLI — Service Level Indicator. The raw measurement. A number you actually collect: request latency, error rate, availability. The SLI is just data; it doesn't say whether the data is good or bad.
SLO — Service Level Objective. The internal target for that indicator. "99.9% of requests complete in under 300ms." This is the number your team commits to hitting, used to decide whether the system is healthy enough to ship faster or needs attention now.
SLA — Service Level Agreement. The external, contractual promise, usually to a customer, usually with a financial or contractual consequence attached if you miss it. SLAs are typically looser than the internal SLO on purpose — you want room to notice you're drifting and fix it before you're in breach of an actual contract.
SLI: what you measure → "99.95% of requests returned 2xx last 30 days"
SLO: what you target → "99.9% of requests return 2xx, rolling 30 days"
SLA: what you promise → "99.5% uptime, or the customer gets a service credit"The gap between the SLO and the SLA is deliberate, not sloppy. If your internal target and your external promise were identical, the moment you dip below target you're already in breach — there's no runway to detect the problem and fix it before it becomes a contractual issue. The gap is the buffer that makes an SLO miss a Tuesday problem instead of a legal one.
Note
A service can have SLIs and SLOs with no SLA at all — most internal-only systems do. The SLA only exists where there's an external party with an actual stake in the number. Don't manufacture a customer-facing contract for a system that doesn't need one; you'll just be maintaining a promise nobody asked for.
Error budgets: the SLO's whole point
An SLO of 99.9% isn't a pass/fail line — it's a budget. 99.9% availability over 30 days allows roughly 43 minutes of downtime. That's not failure allowance as an accident; it's a deliberate, spendable resource:
| SLO target | Allowed downtime / 30 days |
|---|---|
| 99.0% | ~7.3 hours |
| 99.9% | ~43 minutes |
| 99.95% | ~21.6 minutes |
| 99.99% | ~4.3 minutes |
An error budget converts "should we ship this risky change?" from a values debate into an arithmetic one. Budget remaining this month? Ship the risky migration, deploy more aggressively, take the calculated bet. Budget nearly exhausted? Freeze non-critical changes and spend the next sprint on stability instead of features. Nobody has to argue about risk tolerance from first principles every time, because the number already encodes the agreement.
Warning
The failure mode that undoes all of this: setting an SLO of 99.99% because it sounds more serious than 99.9%, without checking what it actually costs. Going from three nines to four nines isn't a 0.09% harder problem — it's typically an order of magnitude more engineering investment, because you're now designing out failure modes that used to be rare enough to ignore. Pick the target the business need actually justifies, not the one that looks best in a slide.
Picking an SLI that means something
The easiest mistake here isn't picking the wrong target — it's picking an SLI that doesn't reflect what users actually experience. "Server uptime" can read as 100% while every request is timing out at the load balancer. A more honest SLI is usually built from what the user's request actually goes through:
Good SLI: percentage of requests served in < 300ms with a 2xx response,
measured at the load balancer, not the application server.
Weak SLI: process uptime — the process can be "up" and still
failing every request it receives.The rule that tends to hold: measure as close to the user's actual experience as you can get away with. A metric collected from inside the service can lie about what's true outside it.
The number that ends the argument
Without SLOs, "is this reliable enough" gets answered by vibes and whoever's loudest in the incident channel — one engineer thinks a 15-minute blip is fine, another thinks it's unacceptable, and there's no shared reference point to settle it. With an SLO and an error budget, the same question has an actual answer: check the budget. That's the entire value of the exercise — not the specific number you pick, but the fact that everyone stopped arguing about reliability in the abstract and started managing it against something concrete.
Tomorrow's article picks up the other half of this measurement discipline: toil, the manual, repetitive work that eats the time an SRE team needs to spend actually protecting that budget.





