Day 14 of 100
The Autoscaling Mystery
π₯ Problem Statement
Traffic to the application jumps significantly and stays elevated. Autoscaling is configured and has worked before. This time, the instance count barely moves β nowhere near what the traffic increase would seem to justify.
In the channel: "Autoscaling isn't working." It's a reasonable first read of the symptom, but it's also the kind of statement that could mean five completely different things depending on what's actually configured underneath it.
ποΈ Environment
- Autoscaling: configured min/max instance bounds
- Scaling metric: CPU-based (target utilization)
- Traffic: significant, sustained increase
- Platform: managed compute with autoscaler
π€ Your Challenge
What would you investigate to find out why autoscaling isn't responding the way it's expected to?
- What metric is autoscaling actually configured to scale on, and does that metric reflect what's actually under load?
- Is the instance count hitting a configured maximum, or just not growing at all?
- Could the workload be I/O-bound or waiting-bound in a way that never raises CPU, the metric autoscaling is watching?
- Is there a cooldown or stabilization window that could be slowing the response down without stopping it entirely?
Solution Hidden
Think through the problem yourself before looking at the answer.
π‘Solution
Step 1 β Understand the symptoms
"Autoscaling isn't working" bundles together several distinct possible failures: scaling triggered but hit a configured maximum; scaling never triggered because the metric it watches never crossed its threshold; scaling triggered but with a cooldown long enough to look like nothing happened; or the autoscaler itself is misconfigured or disabled. Each of these has a completely different fix, and none of them are visible from "the instance count barely moved" alone.
Step 2 β Identify the likely bottleneck
The most common version of this specific story: autoscaling is configured to scale on CPU utilization, but the actual bottleneck under this traffic increase isn't CPU-bound at all β it's I/O-bound, waiting on a downstream call, or throttled by request concurrency limits. If CPU usage per instance never actually crosses the scaling threshold, the autoscaler has no signal telling it to add capacity, regardless of how real the underlying load increase is. From the outside this looks exactly like "autoscaling is broken," when the autoscaler is actually working correctly against a metric that just doesn't reflect what's under strain.
Step 3 β Investigation
- Check the actual scaling metric value against its configured threshold, over the period in question. Did CPU (or whatever the metric is) ever cross the threshold that should trigger a scale-out?
- Check the configured maximum instance count. If the current count is already at or near the max, autoscaling did work β it's just hit its ceiling, which is a capacity-planning problem, not a broken-autoscaler problem.
- Look at what's actually saturated under this load β request concurrency, connection pool usage, downstream call latency β versus what the scaling metric is watching. A mismatch between the two is the single most common root cause of this exact scenario.
- Check cooldown and stabilization window settings. A long cooldown after a previous scaling event can delay the next one considerably, which can look identical to "not scaling" if you're only checking shortly after traffic increased.
- Confirm the autoscaler itself is actually active and correctly targeting this workload β a recent configuration change, a disabled autoscaler, or a targeting mismatch is worth ruling out directly rather than assumed away.
Step 4 β Recommended action
If the metric mismatch is the cause β load is real but CPU never reflects it β the fix is changing what autoscaling scales on: request concurrency, queue depth, or a custom metric that actually correlates with the real bottleneck, rather than CPU utilization that happens not to move under this particular kind of load. If the maximum instance count is the ceiling, that's a capacity conversation, and raising the max (with the cost and infrastructure implications considered) is the direct fix. If it's a cooldown issue, tuning the stabilization window is a much narrower, lower-risk change than either of the other two.
Don't change the scaling metric or thresholds without understanding which of these it actually is β changing the wrong thing can make the system either overly reactive (scaling on noise) or leave the real gap unaddressed while looking like something was fixed.
Step 5 β Engineering lesson
Autoscaling doesn't scale on load β it scales on whatever metric it's told to watch, and those are only the same thing when the metric genuinely correlates with the real bottleneck. A CPU-based scaler watching an I/O-bound workload will faithfully do exactly what it's configured to do, which is nothing, right up until someone checks what's actually saturated and realizes the signal and the symptom were never the same thing.
π§ Todayβs Takeaway
Autoscaling is only as effective as the signal driving it.
π Donβt Miss Tomorrowβs Challenge
A new real-world engineering challenge is released every day.
100 Days β 100 Challenges β 100 New Things Learned.
Get the useful stuff, not the noise.
Occasional notes on engineering, Platform Engineering, AI, cloud and things Iβm learning along the way.

