Day 3 of 100
Terraform Wants to Destroy Production
π₯ Problem Statement
A small, seemingly routine PR bumps a module version pin and tidies up a couple of variable names. CI runs `terraform plan` against the production workspace as usual, and the plan output is not routine at all:
```text Plan: 2 to add, 1 to change, 3 to destroy.
# aws_db_instance.primary must be replaced -/+ resource "aws_db_instance" "primary" { ~ identifier = "prod-primary-db" -> (known after apply) # forces replacement } ```
Three resources marked for destruction, one of them the production database, and the change that triggered it was a module version bump nobody expected to touch anything stateful. The PR looks small. The plan does not.
The engineer who opened the PR is now staring at the "Apply" button in CI, unsure whether this is a real problem or Terraform being Terraform.
ποΈ Environment
- Terraform: remote backend (S3 + DynamoDB locking)
- Workspace: production
- Cloud: AWS
- Change process: PR review + terraform plan in CI
π€ Your Challenge
Should you apply this plan? What would you investigate before deciding either way?
- What in the plan output specifically tells you *why* the database is being replaced, not just that it is?
- Could a module version bump alone change something that forces a resource replacement, and if so, what kind of change?
- What does 'forces replacement' actually mean for a stateful resource like a database, versus a stateless one?
- What would you want to see before treating a `terraform plan` as safe to apply, even one that looks small?
Solution Hidden
Think through the problem yourself before looking at the answer.
π‘Solution
Step 1 β Understand the symptoms
The plan is explicit about the mechanism, if you read past the summary line: identifier is changing, and that field is annotated # forces replacement. In most cloud providers' Terraform resources, certain arguments can't be updated in place β changing them means the provider has to destroy the existing resource and create a new one to match the new configuration. For aws_db_instance, identifier is exactly that kind of argument.
The confusing part is that nobody in this PR touched identifier directly. That's the actual mystery worth chasing β not "is this plan scary," but "what changed that would cause Terraform to think this resource's identifier needs to be different."
Step 2 β Identify the likely bottleneck
A module version bump can absolutely cause this without anyone touching resource-level config directly. If the new module version changed how it constructs the identifier argument internally β a different naming convention, a new variable interpolated into it, a default value that changed β then from Terraform's perspective, the desired identifier for this resource is now different from what's actually deployed, and that's forces-replacement by definition, regardless of how small the PR diff looks from the outside.
This is the specific danger of pinning to a module without deeply reading its changelog: the module's interface might look unchanged (same inputs, same outputs) while its internal resource construction changed in a way that's invisible until you read the actual plan diff.
Step 3 β Investigation
Before applying anything:
- Read the full plan diff for the destroyed resources, not just the summary count. What's actually changing on each of the three, and does the reason line up with what the PR was supposed to do?
- Check the module's changelog or diff between the old and new pinned versions. If it's a public module, the release notes usually call out breaking changes explicitly β including changes to how
identifieror similarly load-bearing fields are computed. - Check Terraform state directly β
terraform state show aws_db_instance.primaryβ to see the identifier as Terraform currently understands it, and compare that against what the new plan wants it to be. - Check for configuration drift β was this resource ever modified outside Terraform (a console change, a manual fix during an incident) that the state file doesn't fully agree with? Drift is a common, separate cause of unexpected replacement plans.
- Look at the resource's
lifecycleblock, if one exists. Aprevent_destroy = trueguard on the production database would normally catch exactly this β its absence here is itself worth noting.
Step 4 β Recommended action
Don't apply. A plan that wants to destroy and recreate a production database is exactly the kind of change that needs a deliberate, reviewed decision β not an approval based on the PR's diff looking small. Pin the module version back, reproduce the plan, and confirm whether the replacement disappears; if it does, that confirms the module version bump is the actual cause, and the fix is either finding a version that doesn't change identifier construction or explicitly handling the rename with a proper migration (which for a stateful resource like a database usually means a snapshot-and-restore path, not a blind terraform apply).
This is also the moment to check whether prevent_destroy should be added to this resource going forward β a guard that makes Terraform refuse to plan a destroy on this specific resource without deliberately removing the lifecycle block first, so a future version bump can't silently reach this same situation again.
Step 5 β Engineering lesson
terraform plan telling you what it's going to do is not the same as that plan being safe. The tool did its job correctly here β it surfaced a real, consequential change caused by a dependency bump, exactly as designed. The risk isn't in Terraform's behavior; it's in treating "the PR diff is small" as a proxy for "the plan is safe," when the two can diverge completely the moment a module changes underneath a version pin.
Production Terraform changes deserve the same scrutiny as production database migrations, because for a stateful resource, that's exactly what a forced replacement is.
π§ Todayβs Takeaway
A terraform plan is a proposal, not a formality. Read what it says forces replacement before you approve it.
π Donβt Miss Tomorrowβs Challenge
A new real-world engineering challenge is released every day.
100 Days β 100 Challenges β 100 New Things Learned.
Get the useful stuff, not the noise.
Occasional notes on engineering, Platform Engineering, AI, cloud and things Iβm learning along the way.

