Cloud, DevOps & infrastructure
How I think about cloud architecture, platform teams, and DevOps in regulated, high-growth environments.
Pillars
Multi-cloud landing zones
Account/subscription structure, SCPs, paved roads, and the boring guardrails that pass audits cheaply.
CI/CD & GitOps
Pipelines as product, signed builds, progressive delivery, and rollback you can trust at 2am.
Platform engineering
Internal platforms run like real products — roadmap, NPS, paved roads, tiered support.
DevSecOps
Shift-left controls that don't slow shipping: SAST, SBOMs, secret scanning, policy-as-code.
Reliability & SRE
SLOs that match revenue, error budgets engineers actually respect, on-call you can sustain.
FinOps & cost
Tagging that holds up, unit economics per workload, rightsizing as a recurring practice.
Reference stack
- AWS (primary)
- Azure
- EKS / Kubernetes
- ECS Fargate
- Lambda
- GitHub Actions
- ArgoCD
- Terraform
- Helm
- OpenTofu
- Datadog
- OpenTelemetry
- Grafana
- PagerDuty
- Sentry
- Wiz
- Snyk
- OPA / Gatekeeper
- Vault
- AWS Security Hub
Related writing
All cloud & DevOps postsThe Second Agent Cheap, the Fiftieth Boring
Build cost was never the constraint in a regulated shop — agent #50 hits the same security review as agent #1. Make identity and action gates reusable.
AWS Cost Levers That Moved the Needle
Cutting ~35% off a multi-region AWS footprint with no capability loss — the levers in the order they paid back, best first.
Security and DevOps Under One Roof
The case for running security and DevOps as one mandate: org-chart distance doesn't create security, and owning the pipelines changes how you protect them.
Model Selection Is Capacity Planning
Most teams pick a model like a sports team and never revisit it — but model selection is a routing, capacity, and risk decision you already know how to make.
Design AI Inference for Model Disappearance
A frontier model went dark three days after launch; here's how I make AI inference survivable on AWS when the provider is a dependency you don't control.
The Eight-Domain Azure Security Review
A tool scores your Azure posture; an assessor walks your architecture. The eight domains I review, in audit order, and the evidence each has to produce.
Cloud FinOps: Where 25–40% of Spend Hides
The press-release version of cloud savings cancels workloads and books compliance debt. The durable version is commitment management and SaaS rationalization.
An SBOM Nobody Reads Is Compliance Cosplay
Generating a software bill of materials is the easy part. Wiring it into the moment a change ships is where supply-chain security stops being theater.
Agent Memory Is a Data-Residency Problem
Give every agent a durable, MCP-connected brain and you've stood up a new data lake of PII and PCI scope nobody classified, encrypted, or can purge.
Agent Safety: Engineer the Blast Radius
Most agent "safety" is a politely worded request to a model that need not honor it. The only controls that count still hold after the model goes wrong.
An AI Agent Dropped Prod: The Change Playbook
Coding agents are committing real change to real systems. The question isn't whether to let them — it's how to give them speed without a SOC 2-fatal mistake.
Source-Map Leaks: Your Pipeline's Confession
One packaging mistake can publish hundreds of thousands of lines of internals. The leak is a confession: release controls never caught up to release velocity.
AI Found 271 Bugs in Firefox. Now Your Repos?
AI-assisted fuzzing found hundreds of bugs in hardened open-source code. The question is whether you run it before someone else runs it against you.
Dark Code Is a Control Failure, Not Tech Debt
AI is filling repos with code nobody can explain. We call it tech debt; it's a control failure — and it should fail CI like a missing approver does.
MCP Is a New Attack Surface: An IAM Playbook
Every MCP server is a new identity reaching into your cloud. Whether that's leverage or liability comes down to least-privilege IAM on every tool call.
Fine-Grained Authorization for Fintech APIs
Authorization scattered across your codebase isn't a feature — it's a liability you can't prove. The pattern multi-tenant regulated platforms actually need.
Aurora DSQL for the Ledger: Active-Active
Multi-region active-active sounds like the answer to ledger nightmares. Interrogate the consistency, recovery math, and migration before betting the books.
Zero-Downtime Database Changes Are a Process
Blue/green and serverless Aurora don't make migrations safe — the runbook does. The boring discipline that keeps schema changes from becoming incidents.
Three Token Counts, Zero You Can Attest To
Codex says one number, Claude another, your gateway a third. That isn't a metering problem — it's an attestation problem regulated industries can't afford.
Cluster Autoscaler to Karpenter: What Breaks
Karpenter is the right call for most EKS shops — but the migration breaks things unrelated to autoscaling. What to know before flipping the switch.
Shadow AI Is the New Shadow IT
Every abandoned notebook and weekend prototype is a credential-bearing asset nobody owns. The fix isn't a ban — it's discovery, demotion, and real sunsets.
Guardrails at Scale for a Three-Person Team
A lean team can govern a sprawling cloud estate without becoming a ticket queue — but only if you put the rules in the pipeline, not in your inbox.
Ransomware Recovery: A Tested-Backups Problem
Everyone has backups. Almost nobody has a restore they've actually run under fire. That gap is where ransomware turns a bad week into an existential one.
Building or rebuilding a cloud platform?
I advise fintech and regulated SaaS teams on cloud architecture, DevOps maturity, and platform engineering.