Day 12 of 100
The Expensive Log
🔥 Problem Statement
Cloud logging costs jump 40% month over month. There's no corresponding growth to point at — the application team is clear: "We haven't increased the number of users." User count is flat. Something else is generating a lot more log volume, and right now nobody knows what.
🏗️ Environment
- Logging: centralized platform, per-GB ingestion pricing
- Services: dozens, varying log verbosity
- Retention: mixed policies across services
- User growth: flat over the period in question
🤔 Your Challenge
What would you investigate to find out what's actually driving the increase in logging cost?
- If user count is flat, what else could cause log volume to grow — a code change, not a traffic change?
- What's the difference between a service logging more events and a service logging the same events more verbosely?
- Could a small number of services account for most of the increase, and how would you find them?
- Is all of this log volume actually useful, or is some of it collected out of habit rather than need?
Solution Hidden
Think through the problem yourself before looking at the answer.
💡Solution
Step 1 — Understand the symptoms
Flat user count ruling out traffic growth is the important clue here — it means the increase isn't "more requests generating proportionally more logs," it's something that changed independent of usage. That points toward a code or configuration change: a debug log level accidentally left on after a recent deploy, a new feature instrumented without sampling, a duplicate log sink shipping the same data twice, or a retention policy change causing more data to be stored than before.
Step 2 — Identify the likely bottleneck
A 40% jump with flat traffic is large enough that it's very likely concentrated in a small number of services rather than spread evenly — a broad, even rise across every service would suggest an ingestion pricing or platform-side change, which is less likely and easy to rule out directly. The more common cause: one or a handful of services started producing significantly more log volume, either through a verbosity change, a new high-cardinality field being logged on every request, or a duplicate ingestion path introduced by a recent change to logging configuration.
Step 3 — Investigation
- Break down log volume by service, comparing this month against last. This turns "cost went up 40%" into "service X's log volume is up 5x" — a specific, actionable lead instead of a vague number.
- Check logging levels for the services showing the biggest increase. A
DEBUGorTRACElevel left on in production after a deploy is one of the most common, least intentional causes of a volume spike like this. - Look for high-cardinality fields being logged on every request — a recently added field like a full request body, a unique ID logged redundantly, or verbose object dumps can multiply the size of every log line without increasing the number of events at all.
- Check for duplicate ingestion — a misconfigured forwarder or a recently added second log sink can silently double-ship the same data.
- Review retention settings for anything that recently changed — a longer retention window on a high-volume index affects storage cost even with volume unchanged.
Step 4 — Recommended action
Fix the specific service and the specific cause the breakdown points to — turn the debug level back down, remove or truncate the high-cardinality field, or de-duplicate the sink — rather than applying a blanket volume reduction across every service, which risks cutting logs that are actually needed for observability somewhere that isn't contributing to the cost spike at all.
Where genuinely high-volume, low-value logging exists even outside this specific spike, sampling is a reasonable tool — keeping a representative fraction of routine, successful requests while retaining all errors and anomalies in full, rather than either logging everything at full volume or cutting corners that remove real debugging capability.
Step 5 — Engineering lesson
Observability isn't free, and "more logs" isn't automatically "better observability" — a service that logs ten times more data than another isn't necessarily ten times more debuggable, and often the extra volume is redundant, too verbose to be useful, or genuinely accidental. The right target is collecting what's actually needed to understand and debug the system, and treating unexplained volume growth — especially growth untethered from actual user activity — as worth investigating on its own, not just absorbing as a rising bill.
🧠 Today’s Takeaway
Observability has a cost — collect the data you actually need.
🔔 Don’t Miss Tomorrow’s Challenge
A new real-world engineering challenge is released every day.
100 Days → 100 Challenges → 100 New Things Learned.
Get the useful stuff, not the noise.
Occasional notes on engineering, Platform Engineering, AI, cloud and things I’m learning along the way.

