Fifteen years building and scaling distributed systems, across the industry's move from bare-metal servers and monoliths to cloud-native, Kubernetes and GitOps. I take on the work teams postpone: paying down technical debt, building deployment from nothing, and keeping systems running.
Six shifts the industry has since packaged and sells ready-made. I worked through them before anything was ready-made.
Incident Response
Process / Observability — Grafana, Alertmanager, Ansible, Git, Linux
Reference architecture. Client code is under NDA, so this page describes the design and the reasoning behind it rather than a specific deployment.
On-call built around service ownership: the team that ships a service carries it, which is the only arrangement that reliably improves reliability. Rotations are staffed so that being on call is uneventful most weeks rather than a tax people plan around.
Runbooks live beside the code and are opened during drills, because a runbook nobody has followed since it was written is fiction. Postmortems are blameless and produce tracked actions, on the principle that an incident which changes nothing will happen again on schedule.