Every long-lived server accumulates a history nobody wrote down: a manual package upgrade during an incident three months ago, a config tweak someone SSH'd in to apply on a Friday, a cron job added to work around a bug that got fixed in code but never removed from the box. None of that is in Git. All of it is in production. Ask five engineers what's actually running on that server and you'll get five partial answers, because the real answer is "whatever the sum of everyone's manual changes happens to be today."
Immutable infrastructure is the decision to stop accumulating that history. Once a server (or container, or VM image) is deployed, nothing about it changes again. Need a new package version, a config change, a security patch? You don't touch the running instance — you build a new image with the change baked in, deploy it, and terminate the old one.
Note
"Immutable" describes the running instance, not the source. The Packer template, Dockerfile, or AMI definition is still very much a living, versioned, reviewed artifact. What's frozen is the thing running in production — it never diverges from what that artifact describes.
Mutable servers fail the way you'd expect them to
A mutable server is a pet: it has a name, a history, and eventually a set of quirks nobody wants to touch because nobody's sure what depends on them. The failure modes are familiar to anyone who's operated one for more than a year:
- Configuration drift — the live box no longer matches whatever's declared in code, because someone changed it directly under pressure.
- Snowflake servers — one instance in a fleet behaves differently, and the only way to find out why is
diff-ing it against another box and hoping the difference is visible. - Unreproducible incidents — a server crashes, gets manually nursed back to health, and the actual root cause is now buried in a shell history nobody's going to read.
- Patch anxiety — applying a security update in place is a live edit to a running system, so it gets scheduled, delayed, and eventually skipped.
None of these are hypothetical edge cases. They're the default outcome of running servers you're allowed to change.
What replacing instead of patching actually buys you
| Mutable ("patch in place") | Immutable ("replace") |
|---|---|
| Config drifts from what's declared | Running state always matches the image |
| Rollback means undoing changes by hand | Rollback means redeploying the previous image |
| "Which server is different, and why?" | Every instance in the fleet is identical |
| Debugging means logging into prod | Debugging means rebuilding the image locally |
| A bad patch is a live incident | A bad image just doesn't get promoted |
The rollback difference is the one that matters most under pressure. With mutable servers, rolling back a bad change means reconstructing the previous state by hand, live, while things are on fire. With immutable infrastructure, rollback is "redeploy the last known-good image" — a boring, mechanical, low-stakes operation you can do calmly, or even automatically.
# Mutable: hope you remember what you changed
ssh prod-web-03
sudo apt-get install -y libssl-dev=1.1.1f
sudo systemctl restart nginx
# ...and now this box is different from the other nine.
# Immutable: build once, deploy everywhere, replace to update
packer build web-server.pkr.hcl
terraform apply # rolls out the new AMI, terminates old instancesThe trade-off nobody skips mentioning
Immutable infrastructure isn't free. Rebuilding and redeploying an image for a one-line config change is slower than SSH-ing in and fixing it — that's the whole point, and it's also the cost. Image build pipelines need to be fast enough that "just replace it" doesn't become the excuse to avoid shipping small fixes. Stateful workloads need real thought too: a database server can't simply be thrown away and rebuilt from scratch without a plan for the data, which is why immutability tends to apply cleanly to stateless application tiers first and spreads to stateful systems only once the team has a real strategy for state (separate persistent volumes, managed data services, replication).
Where to start
Application servers and anything stateless are the easy win — no data to worry about losing on replace. Containers gave most teams this by default, whether or not anyone called it "immutable infrastructure" at the time: a container image is already an immutable artifact, and docker run already replaces rather than patches. Day 30 of this series covers what changes once Kubernetes is the thing scheduling those replacements.
Why this sets up the next 48 hours of this series
Immutable infrastructure isn't a tool you install — it's a constraint you accept, and everything downstream benefits from it being true. Shift-left testing (tomorrow) only works if the artifact you tested is the exact artifact that reaches production, not something patched afterward. Reliability engineering assumes a known-good state to roll back to. GitOps, later in this series, is this same idea taken to its logical conclusion: Git doesn't just describe the desired image, it describes the desired state, continuously reconciled.
The common thread across all of it: the fewer ways a system can quietly become something other than what's declared in code, the fewer 2am pages start with "wait, why is this server different?"





