Skip to content

Articles

Long-form technical writing — deep dives on the systems and decisions behind platform, cloud and AI engineering.

Azure

Azure Landing Zone IAM: Why I Avoid the Owner Role (And What I Use Instead)

Owner is the fastest way to unblock someone in a new Azure subscription, and the fastest way to lose track of who can do what across your landing zones. Here's how I actually think about IAM while building landing zone modules in Terraform.

September 14, 2026 9 min read

GCP

GCP IAM: Why I'd Never Download a Service Account Key Again

A downloaded JSON key with project-level Editor is the fastest way to get a CI/CD pipeline working on GCP — and the fastest way to hand out a credential that never expires and no one is watching. Here's how I'd actually set this up.

September 14, 2026 10 min read

Software Engineering

How I'd Actually Test a Terraform Module (Not Just terraform plan and Hope)

Infrastructure code rarely gets the testing discipline application code takes for granted. Here's how I think about testing a shared Terraform module before it quietly breaks every subscription downstream of it.

September 14, 2026 10 min read

Platform Engineering

Day 50: Fifty Days Later — the Whole Arc, in One Place

From 'what is DevOps, actually' to 'who can be a platform engineer' — a full recap of the five phases this series walked through, and the one question every day of it was really answering.

September 4, 2026 6 min read

Engineering Leadership

Day 49: Platform Engineering Isn't a Role. It's a Mindset.

Forty-eight days of tools, architecture, and practices come down to one honest answer about who can actually do this work — and it has less to do with your job title than with five specific traits.

September 3, 2026 4 min read

AI

Day 48: Smarter Platforms. Autonomous Actions. Exceptional Developer Experience.

Day 44 covered what AI agents do. Today is what a platform team needs to build to work with them well — the specific skills, the human-agent partnership model, and the trust that has to be earned before agents get real autonomy.

September 2, 2026 5 min read

Platform Engineering

Day 47: Every Platform Team Is on a Journey. Where Are You?

From ad-hoc scripts to an AI-native platform — a six-level maturity model for judging honestly where a platform actually stands, not where its roadmap slide claims it stands.

September 1, 2026 4 min read

Engineering Leadership

Day 46: Build Platforms. Empower Developers. Create Impact.

A six-step roadmap from Linux fundamentals to platform leadership — built from the exact progression this fifty-day series has already walked through, phase by phase.

August 31, 2026 4 min read

Platform Engineering

Day 45: Good Intentions. Bad Patterns. Big Impact.

Every platform failure this series has warned about individually — building in secret, over-engineering, ignoring developer experience — shows up together as five recurring anti-patterns. None of them come from bad intentions.

August 30, 2026 5 min read

AI

Day 44: AI Agents Don't Replace Engineers. They Multiply Them.

An AI agent that can perceive, decide, and act through the platform API needs the exact same identity discipline as any human — which is precisely why yesterday's secrets and identity article had to come first.

August 29, 2026 5 min read

SRE

Day 43: The Right Secret. The Right Identity. The Right Time.

Every guardrail this series has covered — policy as code, multi-cloud consistency, FinOps controls — ultimately depends on one foundation holding: knowing exactly who or what is making a request, and never handing out more access than the moment requires.

August 28, 2026 5 min read

Cloud

Day 42: One Platform. Any Cloud. Zero Friction.

Multi-cloud isn't about using many clouds for its own sake. It's about giving developers freedom of choice while the platform keeps consistency — a genuinely harder version of everything this series has already covered.

August 27, 2026 4 min read

Platform Engineering

Day 41: Secure by Design. Enforce by Default.

Manual security review doesn't scale, and policies drift the moment humans have to remember to check them. Policy as code writes the rule once and enforces it everywhere, automatically, forever.

August 26, 2026 4 min read

Cloud

Day 40: Build Amazing Platforms. Optimize Every Dollar.

Great platforms don't just ship features — they deliver value efficiently. FinOps is what turns 'the cloud bill is high' from a monthly surprise into a number every team actually owns and understands.

August 25, 2026 4 min read

Platform Engineering

Day 39: Build What Differentiates. Buy What Accelerates.

There's no universal right answer to building versus buying platform tooling — only the right answer for your team's goals, resources, timeline, and what actually differentiates you.

August 24, 2026 4 min read

Platform Engineering

Day 38: Adoption Is the True North. It's What Proves a Platform Delivers Value.

A platform with excellent DORA metrics for the ten teams still using it and zero adoption everywhere else isn't succeeding. Adoption is the one metric that can't be faked by looking only at your happy users.

August 23, 2026 4 min read

Developer Experience

Day 37: Measure Flow, Not Just Output.

Counting lines of code and commits tells you how busy someone looked. Measuring developer productivity properly means measuring how smoothly value actually moves from idea to production.

August 22, 2026 5 min read

Platform Engineering

Day 36: Build Platforms Developers Love, Not Platforms They Tolerate.

Opening the Advanced Platform Engineering phase with the discipline that decides whether everything built so far actually gets used: platform product management.

August 21, 2026 4 min read

AI

Day 35: An AI Without Context Is Just a Chatbot. MCP Is the USB-C for Fixing That.

AI-SRE and AI-powered platform engineering both assumed an AI system could see real infrastructure state. The Model Context Protocol is the actual standard that makes that connection possible instead of hand-waved.

August 20, 2026 4 min read

SRE

Day 34: Silos Kill Delivery. AI-SRE Is Built to Remove Them.

Dev, Ops, QA, Security, and Business each holding their own siloed view of a system is where handoffs, blame, and slow decisions come from. AI-SRE's real contribution is a single, unified, automated view none of those silos have alone.

August 19, 2026 4 min read

AI

Day 33: AI Isn't Replacing Platform Engineers. It's Levelling Up the Platform.

The observability data this series covered yesterday is exactly what makes AI-powered platform engineering possible — anomaly detection, automated root cause analysis, and remediation built on telemetry the platform already collects.

August 18, 2026 5 min read

Observability

Day 32: You Can't Improve What You Can't See. That Includes the Platform Itself.

This series covered observability for applications back on Day 13. The platform underneath those applications needs the identical discipline applied to itself — or the platform team finds out about outages from their users.

August 17, 2026 4 min read

Platform Engineering

Day 31: Platform APIs Turn Infrastructure Into Self-Service at Scale.

A platform API is the contract between developers and everything running underneath — Kubernetes, CI/CD, observability, secrets. Get the contract right and self-service scales to hundreds of teams without chaos.

August 16, 2026 4 min read

Platform Engineering

Day 30: Kubernetes Isn't Just a Tool. It's Your Platform's Foundation.

Running Kubernetes and building a platform on Kubernetes are different jobs. The gap between them is exactly the three layers this series spent Day 28 mapping out — all built on top of what K8s already gives you for free.

August 15, 2026 4 min read

Platform Engineering

Day 29: GitOps — Declare It in Git, Let Automation Do It

GitOps isn't 'CI/CD but for Kubernetes.' It's a specific claim about where truth lives — and once you take that claim seriously, drift, rollbacks, and audits stop being separate problems.

August 14, 2026 5 min read

Architecture

Day 28: A Great Platform Isn't Magic. It's Well Architected.

Opening the Platform Engineering Architecture phase by taking the Internal Developer Platform apart layer by layer — from the developer experience surface down to the infrastructure foundation nobody sees.

August 13, 2026 5 min read

Platform Engineering

Day 27: One Platform. Limitless Impact.

Self-service, golden paths, and guardrails aren't three separate initiatives — they're facets of one thing: the Internal Developer Platform. Pulling this phase of the series together into the artifact it's all been building toward.

August 12, 2026 4 min read

Platform Engineering

Day 26: Don't Build Gates. Build Guardrails.

A gate asks permission and waits for a human. A guardrail lets you move at full speed and only intervenes when you're about to leave the road. Platform engineering has to pick one, and the wrong choice recreates the ticket queue.

August 11, 2026 5 min read

Platform Engineering

Day 25: We Don't Remove Choice. We Remove Bad Choices.

A golden path is the opinionated, production-ready way to accomplish a common task — not a restriction on what's possible, but a fast lane that makes the right way also the easy way.

August 10, 2026 4 min read

Platform Engineering

Day 24: Self-Service Isn't the Absence of Control. It's Governance That Scales.

A ticket queue for infrastructure requests feels like control. It's actually just latency with extra steps. Self-service platforms replace waiting with automated, guardrail-enforced approval — without giving up governance.

August 9, 2026 5 min read

Developer Experience

Day 23: High Cognitive Load Slows Delivery. A Platform Refunds It.

The mental effort a developer spends remembering which CI pipeline to use, which base image is approved, and where the logs are is effort not spent on the actual problem. Platform engineering's job is taking that effort back.

August 8, 2026 5 min read

Developer Experience

Day 22: Great DX Isn't a Nice-to-Have. It's the Multiplier Behind Every Ship.

Two platforms can have identical infrastructure underneath and produce wildly different outcomes, because one of them is pleasant to use and the other makes developers fight it. That gap is developer experience.

August 7, 2026 4 min read

Platform Engineering

Day 21: A Platform Isn't a Project. It's a Product.

Ship v1.0, declare victory, move on — and six months later nobody's using it. Treating an internal developer platform like a product instead of a project is the difference between adoption and a very expensive shelf-ware.

August 6, 2026 5 min read

Platform Engineering

Day 20: Platform Teams Exist to Remove Friction, Not to Add Layers.

The fastest way to make a platform team fail is to let it become another approval gate. The reason platform teams exist at all is the opposite instinct: fewer people should have to think about infrastructure, not more.

August 5, 2026 4 min read

Platform Engineering

Day 19: DevOps Is How We Build. Platform Engineering Is What We Build On.

Two disciplines keep getting flattened into one because they're both about shipping software faster. They're not the same thing — one is a practice everyone follows, the other is a product with actual users.

August 4, 2026 4 min read

Platform Engineering

Day 18: Platform Engineering Builds the Highway. Developers Drive.

Platform engineering isn't DevOps with a new name. It's the discipline of building the internal developer platform that makes DevOps practices achievable without every team reinventing them.

August 3, 2026 4 min read

SRE

Day 17: Incidents Will Happen. How You Respond Determines the Damage.

The outage isn't the whole story — the eighteen minutes between the alert firing and the fix landing are. Closing out the SRE Foundations phase with the operational discipline that decides how bad an incident actually gets.

August 2, 2026 5 min read

SRE

Day 16: It's Not Who Did It. It's Why It Happened.

A postmortem that finds someone to blame ends the conversation right when it should be starting. Blameless postmortems trade the satisfaction of an answer for the harder, more useful question underneath it.

August 1, 2026 4 min read

SRE

Day 15: Reliability Isn't an Afterthought. It's a Feature.

Five days of SLOs, toil, observability and MTTR all point at the same conclusion: reliability isn't a phase you complete, a team you hire, or a tool you buy. It's a feature every engineer ships, every day.

July 31, 2026 4 min read

SRE

Day 14: Systems Will Fail. Speed of Recovery Wins.

Chasing a longer mean time between failures is chasing an asymptote. Chasing a shorter mean time to recover is chasing something you can actually engineer, this quarter, with tools you already have.

July 30, 2026 4 min read

Observability

Day 13: Monitoring Tells You It Broke. Observability Tells You Why.

A dashboard full of green checkmarks and an outage in progress can coexist. Monitoring answers known questions; observability lets you ask new ones when the known questions don't cover what's actually happening.

July 29, 2026 5 min read

SRE

Day 12: Toil Doesn't Build Value. It Steals Time.

Every hour spent manually restarting a service is an hour not spent making it stop needing restarts. Toil is the specific, measurable enemy SRE was invented to fight.

July 28, 2026 4 min read

SRE

Day 11: SLI, SLO, SLA — Three Acronyms, One Argument

You measure the SLI, you target the SLO, and you promise the SLA — and mixing those up is how teams end up arguing about reliability instead of managing it with a number everyone agreed to.

July 27, 2026 5 min read

SRE

Day 10: Reliability Isn't a Nice-to-Have. It's the Product.

A feature nobody can reach because the service is down isn't a feature. Kicking off the SRE Foundations phase of this series: why reliability has to be engineered in, not hoped for after launch.

July 26, 2026 5 min read

Platform Engineering

Day 9: Manage Everything Like Code. Not Like Chaos.

Terraform gets infrastructure into Git. That's one system out of six that can still drift silently. Everything-as-code is what happens when the same discipline applies everywhere else too.

July 25, 2026 5 min read

Developer Experience

Day 8: Find Problems Early. Fix Them Cheaply.

The bug that costs a comment in code review costs an incident in production. Shift left is the discipline of moving quality, security and feedback as early as possible — not a testing phase, a mindset.

July 24, 2026 5 min read

Cloud

Day 7: Don't Change Servers. Replace Them.

Patching a live server is a bet that you remember every change you've ever made to it. Immutable infrastructure removes the bet entirely by making 'fix it in place' impossible.

July 23, 2026 5 min read

Terraform

Day 6: Declare It. Version It. Trust It. — Making IaC a Discipline, Not Just a Tool

Installing Terraform doesn't give you Infrastructure as Code. Treating your .tf files with the same rigor as application code does — and that means module design, review, and state management with real discipline behind them.

July 22, 2026 4 min read

Terraform

Day 5: Treat Infrastructure Like Code. Because It Is.

Infrastructure as Code isn't about using Terraform instead of clicking in a console. It's about making infrastructure changes reviewable, repeatable, and boring — the same guarantees you already expect from application code.

July 21, 2026 4 min read

SRE

Day 4: Automation Beats Heroics, Every Single Time

The engineer who can fix anything at 3 a.m. is not your most reliable system — they're your biggest single point of failure. Automation is what makes reliability survive someone taking a vacation.

July 20, 2026 4 min read

Platform Engineering

Day 3: Silos Don't Build Software. Teams Do.

Silos aren't a personality problem — they're a rational response to how each team gets measured. Fixing them means changing the incentive, not asking people to talk more.

July 19, 2026 4 min read

Platform Engineering

Day 2: The Three Ways — the Actual Theory Behind DevOps

Flow, feedback, and continuous learning aren't a checklist. They're a sequence — and most teams stop after the first one, then wonder why 'doing DevOps' didn't fix much.

July 18, 2026 4 min read

Platform Engineering

Day 1: What DevOps Actually Is (Not the Job Title)

DevOps isn't a role you hire for — it's a set of practices for closing the gap between writing code and running it. Here's what that means in a system that has to ship.

July 17, 2026 5 min read