Day 2 of 100
The Pod That Never Becomes Ready
π₯ Problem Statement
A new version of an internal API is deployed via a standard rolling update. `kubectl get pods` shows all three new pods with `STATUS: Running`. No crash loops, no restarts. By every obvious signal, the deploy looks clean.
Except traffic isn't reaching the new pods. Requests are still landing on the old ones, which the rollout hasn't finished retiring β the old ReplicaSet is still sitting there at full scale, well past when it should have started scaling down.
Someone on the team looks at the pod list and says: "The pods are Running, so Kubernetes must be healthy. Maybe it's a networking issue?"
ποΈ Environment
- Kubernetes: Deployment + Service + Ingress
- Rollout strategy: RollingUpdate
- Readiness probe: HTTP GET /healthz
- Liveness probe: HTTP GET /healthz
- Monitoring: kubectl + cluster metrics
π€ Your Challenge
Why is traffic not reaching the new pods, and where would you actually look first?
- What does Running actually tell you about a pod, and what does it not tell you?
- What's the difference between a container that's alive and a container that's ready to serve traffic?
- If a Service only sends traffic to Ready pods, what would explain new pods being Running but never getting real traffic?
- What would you check with kubectl before assuming it's a networking problem?
Solution Hidden
Think through the problem yourself before looking at the answer.
π‘Solution
Step 1 β Understand the symptoms
Running is a container-runtime status. It means the container process started and hasn't crashed. It says nothing about whether the application inside is actually able to serve a request β whether it's finished booting, connected to its dependencies, warmed up its caches, or is otherwise ready to do useful work.
A Kubernetes Service doesn't route traffic to every pod that matches its selector. It routes traffic only to pods listed in the Service's Endpoints β and a pod only gets added to Endpoints once its readiness probe passes. A pod can be Running indefinitely while its readiness probe keeps failing, and from the Service's point of view, that pod simply doesn't exist as a traffic target yet.
Step 2 β Identify the likely bottleneck
The team's assumption β "Running means healthy" β is the actual bug in their reasoning, not the cluster's networking. The far more likely explanation: the new pods are Running, but their readiness probe (GET /healthz) is failing or hasn't started passing yet, so they've never been added to the Service's Endpoints. Traffic keeps flowing to the old ReplicaSet because it's the only one with pods currently marked Ready β and because the rolling update won't scale down old pods faster than new ones become Ready, the rollout is effectively stuck in a holding pattern that looks, at a glance, like a completed deploy.
Common root causes for a Running-but-never-Ready pod: the new version's /healthz endpoint checks a dependency (database, cache, a downstream service) that isn't reachable from this environment yet; a longer startup time than the probe's initialDelaySeconds accounts for, so the probe starts checking before the app has finished booting; or the health endpoint itself changed in this release and no longer returns what the probe expects.
Step 3 β Investigation
Before assuming a networking problem, the pod's own status has the answer:
kubectl describe pod <new-pod-name>Look specifically at the Conditions section for Ready: False, and at Events near the bottom β a failing readiness probe shows up there explicitly, often with the HTTP status code or timeout it received.
kubectl get endpoints <service-name>Confirms directly whether the new pods are in the Endpoints list at all β if they're missing, that's conclusive: the Service was never going to send them traffic, independent of anything at the networking layer.
kubectl logs <new-pod-name>Checks whether the application logged a real error on startup β a failed dependency connection, a config value that didn't get set, anything explaining why /healthz would be failing.
Step 4 β Recommended action
Fix the readiness signal, not the network. If the probe is failing because a dependency isn't reachable yet, that's either a real environment issue (fix the dependency) or a probe that's too aggressive for a legitimate startup sequence (adjust initialDelaySeconds or add a proper startup probe). If the health endpoint's behavior changed in this release, that's a code fix, and the rollout should be rolled back in the meantime rather than left half-finished β a stuck rollout with old and new ReplicaSets both present is its own risk, not a stable state to leave running.
Don't touch the Ingress or Service configuration on a hunch. If kubectl get endpoints shows the new pods genuinely missing, the networking layer is doing exactly what it's supposed to do with the information it has β the problem is upstream of it, in the pod's own readiness signal.
Step 5 β Engineering lesson
Running and Ready answer two completely different questions, and conflating them is one of the most common ways a "clean" deploy turns into a quiet, half-finished rollout. Kubernetes is already telling you which one matters for traffic β the Endpoints object is the ground truth for "can this pod actually receive requests right now," and it's one command away.
If an on-call runbook says "check pod status" and stops there, it's missing the step that actually explains most silent-traffic incidents: check readiness, not just liveness.
π§ Todayβs Takeaway
Running means the container started. Ready means it can actually do its job. Kubernetes only routes traffic to the second one.
π Donβt Miss Tomorrowβs Challenge
A new real-world engineering challenge is released every day.
100 Days β 100 Challenges β 100 New Things Learned.
Get the useful stuff, not the noise.
Occasional notes on engineering, Platform Engineering, AI, cloud and things Iβm learning along the way.

