The path to production
How a commit becomes a running release, and what happens when it has to be undone.
- Manual steps and who can bypass them
- Rollback — tested, or hoped for
- Drift between environments
- Secrets in pipelines
I go through the way your code reaches production, the clusters it runs on, the bill it produces and what you can see when it breaks. You get a ranked list of what is genuinely urgent, what each option costs over three years, and a roadmap you can run yourself — or with me.
Usually 1–2 weeks. Scope and price depend on the size of the setup, so they follow the first conversation.
Four areas, the same four I work on as a consultant — read together, because the problems usually are.
How a commit becomes a running release, and what happens when it has to be undone.
Whether the clusters are predictable — and whether the databases on them would survive a bad day.
The bill read line by line, so every euro has an owner and a reason.
Whether metrics, logs and traces answer the question you have during an incident — and who gets woken up.
Every finding with its evidence, ranked across security, reliability and cost.
Each fix with its effort and its cost over three years — including "leave it".
Quick wins first, larger projects after, in an order your team can follow.
A session with your team to go through it, question it, and agree next steps.
The analysis stands on its own. When you want help with the fixes, this is the work it usually leads to.
Every environment declared in Git and promoted by pull request.
ArgoCD · HelmOne pipeline from commit to production: versioned, repeatable, easy to roll back.
GitHub Actions · Azure DevOpsTerraform modules with Terragrunt per environment, and drift checks that tell you when reality moves.
Terraform · TerragruntCloudNativePG clusters with failover, backups to object storage and point-in-time recovery.
CloudNativePGMetrics, logs and traces in one place, with dashboards and alerts worth waking up for.
Prometheus · Loki · Tempo · GrafanaA 30-minute call, then read-only access to repos, clusters, the cloud bill and dashboards.
Usually 1–2 weeks. I read the pipelines, clusters, bill and telemetry — and ask your team what hurts.
Findings, costed options and a roadmap, walked through with your team.
Fix it with your team, with me, or both — the roadmap works either way.
Tell me what you run and what keeps going wrong. I'll come back with a scope for the analysis — no obligation.