Skip to content

Day 8 of 100

The Missing Logs

ObservabilityLoggingIntermediate15–20 min

🔥 Problem Statement

Everything about the application looks fine — it's serving traffic, responding correctly, no alerts firing. But while investigating an unrelated ticket, an SRE notices something: one production service has no logs in the central logging platform for the last several hours. Not fewer logs. None.

The application team's response: "The application is generating logs — we can see them if we exec into the pod. So logging must be working."

That's true and also beside the point. The question isn't whether the application writes logs. It's whether those logs are making it all the way to where anyone would actually go looking for them.

🏗️ Environment

  • Application: containerized service, running normally
  • Runtime: Kubernetes pod, stdout/stderr logging
  • Log collector: node-level agent
  • Forwarder: log shipping to central platform
  • Central logging: search + dashboards

🤔 Your Challenge

Where would you start investigating, and how would you narrow down where the logs are actually being lost?

  • What are all the stages a log line has to pass through between the application writing it and someone searching for it centrally?
  • If the application is confirmed to be writing logs, what does that rule out and what does it leave open?
  • Would you expect a total loss of logs to look different from a partial loss, and what would that difference tell you?
  • What's one query or check at each stage of the pipeline that would confirm whether that stage is working?

Solution Hidden

Think through the problem yourself before looking at the answer.

💡Solution

Step 1 — Understand the symptoms

"The application is generating logs" only confirms the very first link in a much longer chain. A log line's actual journey looks like this:

Text
Application
    ↓
Container / Runtime (stdout/stderr)
    ↓
Log Collector (node-level agent)
    ↓
Forwarder / Sink
    ↓
Central Logging Platform
    ↓
Search / Dashboard

Confirming logs exist at the very first step tells you nothing about the other five. A total, hours-long absence in the central platform — not a reduction, a complete absence — points at a single stage somewhere in that chain that's either stopped working entirely or is silently dropping everything for this one service.

Step 2 — Identify the likely bottleneck

A complete, total loss (not degraded, not partial) for one specific service, while other services log normally, narrows this considerably. If it were a central-platform-wide issue, every service would be affected, not just one — that points the investigation toward something specific to this service's pipeline path: a collector configuration issue for this node or namespace, a filter or routing rule that's silently excluding this service, a permissions issue between the forwarder and the central platform for this specific log stream, or a recent change to this service's logging configuration or labels that broke how it's being picked up.

Step 3 — Investigation

Walk the pipeline stage by stage rather than guessing:

  • Application / container — already confirmed logs exist here (kubectl logs or exec into the pod).
  • Log collector — is the node-level agent actually running and healthy on the node this pod is scheduled on? Check its own logs and metrics for errors specific to this pod or namespace.
  • Filtering and routing rules — many log pipelines have include/exclude rules by namespace, label, or log level. A recent change to this service's labels or a broad exclusion rule elsewhere in the pipeline is a common, easy-to-miss cause.
  • Forwarder / sink — is it successfully shipping data, and does it show any errors or backpressure specific to this service's stream?
  • Permissions and quotas — a forwarder that's lost write access to the destination, or hit a per-service ingestion limit, can silently drop logs rather than erroring loudly.
  • Central platform ingestion — confirm the index or destination this service's logs should land in actually exists and is receiving data from anything, to rule out a platform-side ingestion problem scoped to that destination.

Fix the specific stage the investigation points to, not the whole pipeline speculatively. If it's a filtering rule, correct the rule and add a check that flags when a previously-logging service goes silent for a configurable window — for a total log loss like this, an automated dead man's switch on log volume per service would have caught it in minutes instead of it being found incidentally days later.

Don't restart the application or "just redeploy it" as a first move — the application was already confirmed to be writing logs correctly. Restarting it addresses a stage of the pipeline that was never actually broken, and it won't fix a collector, forwarder, or routing issue.

Step 5 — Engineering lesson

"The application is generating logs" and "the logs are searchable in the central platform" are claims about two completely different systems, separated by several hops that can each fail independently and silently. A missing-logs incident is a pipeline-health problem first, and an application problem only if every other stage checks out — troubleshooting from the wrong end of that chain wastes time confirming something that was never actually in question.

🧠 Today’s Takeaway

When logs disappear, troubleshoot the pipeline — not just the application.

🔔 Don’t Miss Tomorrow’s Challenge

A new real-world engineering challenge is released every day.
100 Days → 100 Challenges → 100 New Things Learned.

Get the useful stuff, not the noise.

Occasional notes on engineering, Platform Engineering, AI, cloud and things I’m learning along the way.

No spam. Unsubscribe anytime.