Skip to content

Day 5 of 100

Can AI Create Production Alerts?

AISREObservabilityIntermediate15–20 min

🔥 Problem Statement

An AI tool has been analyzing weeks of metrics, logs, and traces across the platform, and it's genuinely good at this — it proposes a batch of new alert rules that catch real gaps in current coverage, things the team hasn't gotten around to instrumenting yet.

The proposal on the table isn't just "use AI to suggest alerts for review." It's further than that: "Why don't we let the AI automatically create and enable these alerts directly in production, continuously, as it keeps analyzing new data?" No human review step, no staging period — just AI-identified gaps becoming live paging rules on their own.

🏗️ Environment

  • Observability platform: metrics + logs + traces
  • Alerting: existing SLO-based alert rules
  • AI: an LLM analyzing recent service behavior
  • Proposal: AI-generated alerts, auto-enabled in production

🤔 Your Challenge

What safeguards would you require before allowing an AI system to auto-enable production alerts?

  • What's the actual risk if an AI-proposed alert is well-intentioned but has the wrong threshold?
  • Who is on the hook at 3 a.m. when an AI-created alert pages someone, and does that change what review it deserves?
  • How would you know if the AI's own alert-generation process started degrading — and who would notice?
  • What's the difference between 'AI recommends' and 'AI enables', and does that difference matter here?

Solution Hidden

Think through the problem yourself before looking at the answer.

💡Solution

Step 1 — Understand the symptoms

The AI's output quality isn't really the question here — the scenario stipulates it's finding real gaps, and that's plausible; this is exactly the kind of pattern-matching-across-large-telemetry-volumes task LLMs tend to be genuinely useful for. The actual question is about the last step: recommending versus enabling.

A recommendation that's wrong costs someone a few minutes reviewing and rejecting it. An auto-enabled alert that's wrong costs someone a 3 a.m. page for nothing, or — worse — trains the team to start ignoring pages because too many of them turn out to be noise. Those are very different failure costs, and the proposal on the table skips straight past the cheap-failure version to the expensive one.

Step 2 — Identify the likely bottleneck

The risk isn't "the AI is bad at this." It's that alert quality depends on context an AI analyzing metrics, logs, and traces doesn't automatically have: what this specific service's acceptable error budget actually is, whether a given threshold is normal seasonal variance or a real problem, whether a metric spike correlates with a known, already-mitigated issue, and what the organizational cost of a false-positive page actually is for this specific team. A threshold that looks statistically reasonable from telemetry alone can still be operationally wrong — too sensitive for a team that already has alert fatigue, or too loose for something genuinely critical.

There's a second, quieter risk too: the alert-generation system itself can degrade — a change in the underlying model, a shift in what data it's being given, a subtle bug in how it computes thresholds — and if there's no human in the loop and no monitoring on the generation process itself, that degradation shows up first as production pages, not as a caught mistake.

Step 3 — Investigation

Before agreeing to any version of this, the questions worth asking are about the workflow, not the AI's competence:

  • What does the approval path look like right now, and does "auto-enable" skip a step that currently exists for human-authored alerts too? If human engineers don't get to push straight to production alerting without review, AI shouldn't either — the standard should be at least as high, not lower, because there's no one person accountable for a specific threshold choice.
  • Is there an SLO or error-budget context attached to each proposed alert, or is it purely statistical ("this metric looks unusual") without grounding in what actually matters for this service?
  • What's the plan for noise? A batch of new alerts with untuned thresholds is a fast way to generate false positives at scale, right at the moment trust in the system is being established.
  • Is the alert-generation process itself observable — versioned, auditable, with a way to see what changed if alert quality shifts over time?
  • What's the rollback path if an auto-enabled alert turns out to be wrong — can it be disabled as fast as it was enabled, and is there a record of who (or what) enabled it and why?

Keep AI in the loop as a recommender, not an enabler, at least until there's real track record and real guardrails. A workable version of this: AI proposes alerts with reasoning attached (what pattern it saw, why it thinks this threshold), a human reviews and approves before anything goes live, new alerts start in a non-paging or shadow mode to validate signal quality against real production behavior before they can page anyone, and every AI-proposed-and-approved alert is versioned and auditable — so if something pages incorrectly, there's a clear record of what proposed it, who approved it, and what the reasoning was.

That's not "no automation." It's automation with a human accountable for the last step that actually matters — the one where a wrong decision means someone's sleep gets interrupted for nothing.

Step 5 — Engineering lesson

The useful question with AI-assisted SRE tooling usually isn't "is the AI good enough." It's "what's the blast radius of the AI being wrong, and does the workflow match that blast radius." AI generating a report is low-stakes if wrong. AI auto-enabling a production page is not — and the gap between those two is exactly where guardrails, human review, and auditability need to live, not somewhere the team gets to add later once the first bad page has already gone out.

The same discipline that applies to any production automation — least privilege, staged rollout, auditability, a fast rollback path — applies here too. AI doesn't get an exemption from that just because it's AI; if anything, a system that can act at machine speed and machine scale needs those guardrails more, not less.

🧠 Today’s Takeaway

AI can assist SRE decisions, but production automation needs guardrails.

🔔 Don’t Miss Tomorrow’s Challenge

A new real-world engineering challenge is released every day.
100 Days → 100 Challenges → 100 New Things Learned.

Get the useful stuff, not the noise.

Occasional notes on engineering, Platform Engineering, AI, cloud and things I’m learning along the way.

No spam. Unsubscribe anytime.