Skip to content

Day 7 of 100

The Deployment That Increased Errors

DevOpsSRERelease EngineeringIntermediate15–20 min

🔥 Problem Statement

A new version ships. The CI/CD pipeline is fully green — build passed, tests passed, deploy step completed without error. Ten minutes later, the dashboards tell a different story:

```text Error rate: 0.2% → 4.8% P95 latency: 350ms → 1.8s Traffic: Normal CPU: Normal Memory: Normal ```

In the deploy channel, someone points at the green pipeline: "The pipeline passed, so the deployment is healthy." Nobody's disputing that the *pipeline* succeeded. Whether the *release* is healthy is a separate question that hasn't actually been asked yet.

🏗️ Environment

  • CI/CD: automated pipeline, all checks green
  • Deployment strategy: rolling update, no canary stage
  • Monitoring: golden signals (latency, traffic, errors, saturation)
  • Rollback: manual, requires on-call approval

🤔 Your Challenge

Should this deployment be rolled back immediately, and what evidence would you want before deciding?

  • What does a green CI/CD pipeline actually verify, and what does it not verify?
  • With CPU, memory, and traffic all normal, what does that rule out as the cause of the error spike?
  • What would you check in logs or traces to confirm the new version is actually the cause, not a coincidence?
  • What's the cost of waiting five more minutes to confirm, versus rolling back on a hunch?

Solution Hidden

Think through the problem yourself before looking at the answer.

💡Solution

Step 1 — Understand the symptoms

CI/CD going green confirms the code built, the automated tests passed, and the deploy mechanism worked — it says nothing about how the new code behaves under real production traffic, with real data, hitting real dependencies. Those are two different questions, and a green pipeline only answers the first one.

The numbers here are specific and point somewhere: error rate up 24x, P95 latency up roughly 5x, while traffic, CPU, and memory are all normal. That rules out a traffic spike or a resource-exhaustion problem outright — this isn't "we got more load than we could handle." Something in the new code is failing or slowing down on requests it used to handle fine.

Step 2 — Identify the likely bottleneck

With infrastructure metrics flat and only application-level signals moving, the likely cause lives inside the new release itself — a bug introduced in this version, a changed API contract with a downstream dependency, a configuration value that didn't get set correctly in this environment, or a code path that behaves correctly in tests but not against production data shapes. The timing (immediately following deploy, ten-minute delay before it clearly shows) fits a release-caused regression far better than an unrelated coincidence.

Step 3 — Investigation

  • Compare error logs before and after the deploy timestamp. What specific errors are showing up now that weren't before? A stack trace pointing at new code is close to conclusive.
  • Check traces for the newly slow or failing requests — where specifically is the extra latency or the failure happening? A downstream call, a new code path, a serialization step?
  • Confirm the deploy timestamp lines up exactly with the metric shift. If the error rate started climbing at the exact moment the new version rolled out, that's strong correlation; if it started before or well after, look elsewhere.
  • Check for a version-specific segment if the rolling update hasn't fully replaced the old version yet — do only the new-version instances show the elevated error rate, or all instances? That single comparison can be close to definitive on its own.

Don't wait for a root cause before acting — with error rate up 24x and evidence pointing squarely at the new release, roll back first and investigate the actual bug afterward with the safety net of a stable production. The cost of a fast rollback that turns out to be slightly premature is minutes of engineering time re-confirming; the cost of staying on a broken release while investigating is a sustained, live-traffic-facing error rate.

Going forward, this is exactly the gap a canary or progressive delivery stage exists to catch — deploying to a small percentage of traffic first would have surfaced this same 24x error spike against a fraction of users instead of everyone, with an automated rollback trigger instead of a chat debate about whether "the pipeline passed" means the release is safe.

Step 5 — Engineering lesson

CI/CD pipelines verify that code can be built and deployed correctly. They don't verify that the deployed code behaves correctly under real production conditions — that's what post-deploy monitoring on the golden signals (latency, traffic, errors, saturation) is for, and it needs to be watched as closely as the pipeline itself, not treated as a formality once the pipeline goes green.

"The pipeline passed" and "the release is healthy" are different claims. Only one of them was actually verified here.

🧠 Today’s Takeaway

A successful deployment is not the same as a healthy release.

🔔 Don’t Miss Tomorrow’s Challenge

A new real-world engineering challenge is released every day.
100 Days → 100 Challenges → 100 New Things Learned.

Get the useful stuff, not the noise.

Occasional notes on engineering, Platform Engineering, AI, cloud and things I’m learning along the way.

No spam. Unsubscribe anytime.