Skip to main content
Consulting

Where I can help

Most of my work starts the same way. Something in production is fragile, expensive, or slow to change, and the team is too busy keeping it alive to go fix it properly. That's the part I take off your hands.

Kubernetes that stops surprising you

Boring is the goal. I take clusters that wake people up and make them predictable: Karpenter handling the autoscaling, RBAC and network policies that actually restrict something, custom operators only where they earn their keep. The test I care about is whether you'd deploy on a Friday.

KubernetesAKSHelmKarpenter

One path to production

Manual pipelines fail in ways nobody can reproduce. Declarative state sync fixes that: rebuilding an environment becomes one command instead of an afternoon of archaeology through old Slack threads.

ArgoCDAzure DevOpsGitHub Actions

Finding out where the money goes

Cloud bills grow quietly because nobody owns them. I read yours line by line — right-sizing resource requests, moving what can tolerate it onto spot, unpicking whatever is charging you a lock-in premium. The first thing you learn is how much of the bill nobody can account for.

TerraformCrossplaneKubeCost

Knowing before your users do

Firefighting is a seeing problem. SLIs and SLOs worth paging on, telemetry that answers the question you have while it's still on fire, and security scanning in the deploy path rather than a quarterly report nobody opens.

PrometheusOpenTelemetryGrafana
Case studies

Two I'm allowed to talk about

Fintech / B2B

Financial SaaS platform scale-out

A fintech SaaS moving millions of transactions was still shipping its monolith through a system nobody trusted. I rebuilt it as a hub-and-spoke ArgoCD setup on AKS — environments properly isolated, Git the only thing allowed to change state.

30% Reduction in cloud spend
5 min Deploy time (down from 1 hrs)
Zero Downtime deployments
Retail / SaaS

Multi-region e-commerce failover

The question was simple: what happens if an entire region goes away? I built the answer. Terraform and Crossplane provision an identical standby environment, so losing a regional control plane costs you minutes instead of data.

Active-passive Architecture deployed
< 2 min RTO
100% Infrastructure as code coverage
Engagement models

Three ways this usually starts

Infrastructure audit

I go through your clusters, pipelines, and cloud bill, then hand back a prioritized list of what's genuinely urgent across security, performance, and cost — not a list of everything that's imperfect.

Ask for an audit
Fractional architect

The long game. I sit with your engineering team as an embedded extra pair of hands and opinions, steering architecture decisions and chipping away at reliability over months rather than weeks.

Ask about a retainer

Not sure which one fits?

Describe the bottleneck. If I'm not the right person for it, I'll say so.