Day 10 of 100
The Kubernetes Pod Keeps Restarting
π₯ Problem Statement
A production pod won't stay up. `kubectl get pods` shows:
```text STATUS: CrashLoopBackOff RESTARTS: 47 ```
The fastest suggestion in the channel: "Increase the CPU and memory." It's an easy change, and more resources rarely hurts β except CrashLoopBackOff doesn't actually say *why* the container is crashing, and "give it more resources" is a guess at a cause nobody's confirmed yet.
ποΈ Environment
- Kubernetes: Deployment, resource requests/limits set
- Observed: CrashLoopBackOff, 47 restarts
- Monitoring: kubectl + cluster metrics
π€ Your Challenge
Would you increase resources immediately, or investigate first β and what would you actually check?
- What does CrashLoopBackOff tell you, and what does it not tell you about why the container exited?
- What's the difference between a container that's out of memory and one that's failing to start for an unrelated reason?
- Where would you look to find out what the application itself said right before it died?
- If you increased resources and the pod kept crashing, what would that tell you?
Solution Hidden
Think through the problem yourself before looking at the answer.
π‘Solution
Step 1 β Understand the symptoms
CrashLoopBackOff means Kubernetes is repeatedly starting a container, watching it exit, and backing off before trying again β it's a description of the restart behavior, not a diagnosis of why the container is exiting. A container can enter this state for reasons that have nothing to do with CPU or memory: a missing config file, a bad environment variable, a failed connection to a dependency on startup, an application crash on a specific input, or actual resource exhaustion. "Increase the resources" is only the right fix for the last of those.
Step 2 β Identify the likely bottleneck
The fastest way to stop guessing is to ask Kubernetes what it already knows. kubectl describe pod and the previous container's logs will very often state the reason directly rather than leaving it a mystery:
kubectl describe pod <pod-name>Check the Last State section for a Reason β OOMKilled specifically means the container was killed for exceeding its memory limit, which is the one case where "increase memory" is genuinely the right first move. Any other reason β a non-zero exit code from an application error, a failed liveness probe, an image pull failure β points somewhere else entirely, and more resources won't fix it.
Step 3 β Investigation
kubectl logs <pod-name> --previousThis is the single most useful command here β it shows the logs from the last crashed container instance, which usually includes the actual application error right before it died: a stack trace, a failed connection message, a config-parsing error.
From there:
- Exit code β
describe podshows this too.137typically means OOMKilled or a SIGKILL; other codes point at an application-level failure. - Configuration and secrets β did a recent change to a ConfigMap or Secret this pod depends on introduce a bad or missing value?
- Environment variables β a missing or malformed one is one of the most common causes of an application crashing immediately on startup.
- Probes β is a liveness probe killing the container before it's finished starting up, misreading a slow boot as a hang?
- Dependencies β does the application fail fast if it can't reach a database or another service on startup, and is that dependency actually available?
Step 4 β Recommended action
Only increase resources if the evidence actually points there β OOMKilled in the Last State, or memory usage climbing toward the limit in metrics right before each crash. Otherwise, fix the actual cause the logs and describe output point to: correct the bad config, fix the environment variable, adjust a probe that's too aggressive, or address the dependency the application can't reach.
Increasing resources without evidence has a real cost beyond wasted effort β it can mask the actual problem for a while (more memory headroom might delay but not prevent a memory leak from eventually crashing the pod again) while also increasing cluster cost for no real benefit, and it burns time that a two-minute log check would have saved.
Step 5 β Engineering lesson
CrashLoopBackOff is Kubernetes accurately describing a symptom β the container keeps dying β not diagnosing a cause. kubectl describe and --previous logs exist specifically to turn that symptom into an actual reason, almost always in under a minute. Reaching for more resources before checking those two things is treating a guess as a diagnosis.
π§ Todayβs Takeaway
CrashLoopBackOff is a symptom, not a root cause.
π Donβt Miss Tomorrowβs Challenge
A new real-world engineering challenge is released every day.
100 Days β 100 Challenges β 100 New Things Learned.
Get the useful stuff, not the noise.
Occasional notes on engineering, Platform Engineering, AI, cloud and things Iβm learning along the way.

