Skip to content

Day 19 of 100

The Queue That Keeps Growing

Distributed SystemsSREIntermediate15–20 min

πŸ”₯ Problem Statement

Nothing looks obviously broken. Producers are publishing messages normally. Consumers are running, not crashing, not erroring in any visible way. But the queue depth graph tells a different story β€” it's been climbing steadily for hours, and it hasn't leveled off once.

Nobody's paying close attention to queue depth specifically; the alerts that exist are for consumer errors and consumer CPU, and none of those have fired.

πŸ—οΈ Environment

  • Architecture: producer services -> message queue -> consumer workers
  • Queue: managed message broker, standard configuration
  • Consumers: fixed pool size, no autoscaling configured
  • Symptom: queue depth rising steadily over several hours

πŸ€” Your Challenge

What does a continuously growing queue depth actually tell you, independent of any other symptom?

  • If producers and consumers are both 'running fine,' what does a growing queue imply about their relative rates?
  • Could consumers be running without erroring, and still not be the reason the queue is growing?
  • What's the difference between a queue that's growing because of a burst versus one that's growing because it's structurally falling behind?
  • What would you need to know about consumer throughput to answer this with data instead of guesswork?

Solution Hidden

Think through the problem yourself before looking at the answer.

πŸ’‘Solution

Step 1 β€” Understand the symptoms

A queue's depth is the direct, simple result of one thing: the rate messages are added versus the rate they're removed. If depth is climbing steadily rather than fluctuating around a stable level, that means production rate is exceeding consumption rate, consistently, over the whole observed window β€” not a momentary burst that a healthy system would absorb and drain, but a sustained mismatch. "Consumers aren't erroring" doesn't mean consumers are keeping up; a consumer can run cleanly, with zero errors, while simply processing too slowly or too few messages per second to match what's coming in.

Step 2 β€” Identify the likely bottleneck

With a fixed consumer pool and no autoscaling, the most likely explanations split into two categories: producers are genuinely producing more than usual (a real traffic increase, a batch job, an upstream retry storm feeding more messages than normal), or consumers are processing more slowly per message than they used to (a downstream dependency that got slower, a recent code change that added per-message latency, resource contention on the consumer instances themselves). Either way, a fixed-size consumer pool has no way to absorb the mismatch β€” it just falls further behind, linearly, for as long as the imbalance continues.

Step 3 β€” Investigation

  • Measure actual producer rate and consumer rate directly, not just queue depth β€” messages published per second versus messages successfully processed per second. This turns "the queue is growing" into an actual number showing how far behind consumption is.
  • Check consumer processing latency per message over time β€” if it's increased compared to a recent baseline, that points at consumers getting slower rather than producers sending more.
  • Check for a spike in message volume from producers β€” a batch job, a retried upstream failure, or a genuine traffic increase would show up here directly.
  • Check consumer resource usage (CPU, memory, connection pool) β€” resource contention on the consumer side is a common, quiet cause of per-message slowdown that doesn't show up as errors.
  • Check whether messages are failing and being requeued β€” a consumer that's processing a message, failing partway, and putting it back on the queue can inflate depth without ever surfacing as a clean "error" in a naive error-rate metric.

If producers are genuinely sending more, the fix is scaling consumers (or making them faster) to match the new sustained rate β€” a fixed pool sized for yesterday's traffic won't absorb a real, lasting increase. If consumers have gotten slower per message, the fix is finding and addressing that slowdown specifically (a downstream dependency, a resource constraint, a code regression) rather than just adding more consumer instances to compensate for something that's actually broken.

Either way, queue depth deserves its own alert independent of consumer error rate β€” a queue can grow for hours with zero consumer errors, exactly as it did here, and by the time it's noticed on a dashboard by chance, real backlog and real latency for whatever's waiting in that queue has already accumulated.

Step 5 β€” Engineering lesson

"No errors" and "keeping up" are different claims about a consumer, and a queue is one of the few places in a system where the gap between them shows up clearly and early β€” often well before anything else looks wrong. Treating queue depth as a first-class signal, not an afterthought behind error-rate alerting, is what turns "the queue has been silently falling behind for six hours" into "we caught this in the first thirty minutes."

🧠 Today’s Takeaway

A queue is often an early signal of a system falling behind.

πŸ”” Don’t Miss Tomorrow’s Challenge

A new real-world engineering challenge is released every day.
100 Days β†’ 100 Challenges β†’ 100 New Things Learned.

Get the useful stuff, not the noise.

Occasional notes on engineering, Platform Engineering, AI, cloud and things I’m learning along the way.

No spam. Unsubscribe anytime.