ALEKSANDR GRIBAKIN
INTROCRAFTSTACKMILESTONESCONTACT
Aleksandr Gribakin
-- FPS
INITIALIZING...

< ALEKSANDR GRIBAKIN />

Lead Full Stack & DevOps Engineer
15 Years in Code
Fifteen years building and scaling distributed systems, across the industry's move from bare-metal servers and monoliths to cloud-native, Kubernetes and GitOps. I take on the work teams postpone: technical debt, CI/CD built from nothing, and keeping things up.
Kubernetes Terraform Docker GitLab CI PostgreSQL Go Node.js

Impact By The Numbers

Career Timeline

Projects & Impact

Migration Bare-Metal to Hybrid Cloud Terraform, Ansible, AWS, GCP, Linux △ 30-40% cost cut / staged cutover Staged migration off physical servers into hybrid cloud. Inventory first, then parity, then traffic — never all three at once. Platform Kubernetes Platform Kubernetes, Helm, ArgoCD, Terraform, Prometheus △ 99.9% uptime / self-serve namespaces A cluster product teams can use without learning Kubernetes. Namespaces, quotas, ingress and monitoring arrive with the namespace. IaC Infrastructure as Code Baseline Terraform, Ansible, GitLab CI, Vault, Linux △ 100% reproducible / 3 environments Every environment described in code and rebuilt from zero on demand. No snowflake servers, no undocumented hand edits. Tooling Local and Production Parity Docker, Docker Compose, Terraform, Make, Linux △ One image / 4 environments The same images and the same configuration shape everywhere. Kills the "works on my machine" class of bug outright. Delivery Standardised Delivery Pipelines GitLab CI, Jenkins, Docker, Helm, GitHub Actions △ 15 min release / dozens of services One pipeline template across dozens of services. Release time from days to about fifteen minutes. Deployment Zero-Downtime Deployment Kubernetes, ArgoCD, NGINX, Helm, Prometheus △ Zero downtime / 90s rollback Rolling and blue-green releases with health-gated promotion. A bad release rolls back before most users notice. GitOps GitOps Delivery Model ArgoCD, Git, Helm, Kustomize, Kubernetes △ Git as source of truth / auto-reconcile The repository is the desired state. Anything applied by hand gets reverted by the reconciler, on purpose. Architecture Monolith to Services PHP, Symfony, Go, gRPC, RabbitMQ △ Incremental cutover / no freeze Strangler-fig extraction with the monolith serving traffic throughout. Seams chosen by data ownership, not by folder layout. API API Layer and Contracts Node.js, TypeScript, OpenAPI, gRPC, NGINX △ Generated clients / versioned contracts REST outward, gRPC inward, one generated contract for both. Clients break at build time instead of in production. Queues Async Job Processing Python, FastAPI, Celery, RabbitMQ, Redis △ Idempotent retries / DLQ triage Queue-backed background work with idempotent handlers, retries with backoff and a dead-letter queue that gets read. Database PostgreSQL Under Load PostgreSQL, PgBouncer, Patroni, Prometheus, Linux △ Read scale-out / partitioned writes Replication, partitioning and connection pooling. Most wins came from reading query plans, not from adding hardware. Caching Multi-Level Caching Redis, NGINX, CDN, Node.js, PostgreSQL △ 4 tiers / explicit invalidation CDN, edge, application and query caches with explicit invalidation. Every layer answers what it is allowed to be wrong about. Search Search and Document Store Elasticsearch, MongoDB, Python, Kafka, Redis △ Live reindex / change-stream fed Full-text search fed by a change stream, with reindexing that runs against live traffic instead of a maintenance window. Web SPA with Server-Side Rendering React, TypeScript, Vite, Node.js, NGINX △ SSR first paint / typed end to end Rendered on the server for first paint and crawlers, hydrated on the client for everything after. Design System Component Library Vue.js, TypeScript, Webpack, Storybook, CSS △ 40+ components / WCAG AA A shared component set with tokens, documented states and accessibility built in rather than retrofitted. Monitoring Observability Stack Prometheus, Grafana, Loki, Elasticsearch, Alertmanager △ Correlated signals / SLO-based alerts Metrics, logs and traces correlated by request id. Dashboards answer questions instead of displaying numbers. Process Incident Response Grafana, Alertmanager, Ansible, Git, Linux △ Owned rotations / current runbooks On-call with owned services, runbooks that are actually current, and postmortems that produce changes rather than blame.

Milestones & Recognition

DevOps Engineer — Microservices migration

2019–2022

Dozens of microservices, each with a hand-built pipeline that had long since drifted from its neighbours. Reduced them to one template: a service declares its type and gets build, tests and deployment already assembled. Releasing stopped being a coordinated multi-day event — about fifteen minutes now, with no downtime.

Standardised delivery across dozens of microservices that had each grown their own pipeline. One template, opted into by declaring a service type, replaced a set of copies that had drifted apart the moment they were duplicated.

Release time dropped from a multi-day coordinated event to roughly fifteen minutes, deployed without downtime. The effect worth naming is second-order: once releasing became cheap, teams shipped smaller changes more often, and smaller changes are what actually brings incident counts down.

Key Achievements

Tech Stack

GitLab CI · Jenkins · Docker · Kubernetes · Helm · Ansible · Bash · Linux

Key Metrics