Fifteen years building and scaling distributed systems, across the industry's move from bare-metal servers and monoliths to cloud-native, Kubernetes and GitOps. I take on the work teams postpone: technical debt, CI/CD built from nothing, and keeping things up.
Process / Observability — Grafana, Alertmanager, Ansible, Git, Linux
Reference architecture. Client code is under NDA, so this page describes the design and the reasoning behind it rather than a specific deployment.
On-call built around service ownership: the team that ships a service carries it, which is the only arrangement that reliably improves reliability. Rotations are staffed so that being on call is uneventful most weeks rather than a tax people plan around.
Runbooks live beside the code and are opened during drills, because a runbook nobody has followed since it was written is fiction. Postmortems are blameless and produce tracked actions, on the principle that an incident which changes nothing will happen again on schedule.