AI systems · AWS · Azure · Kubernetes

Production AI systems, and the infrastructure that keeps them running.

Retrieval, agents, model serving and the platform underneath — built by someone who has run all of it at scale and can show you the numbers.

Book a consultation call
12M+documents in production retrieval
40Kconcurrent connections, sub-100ms overhead
$180Kannual inference spend removed as volume grew
10 yrsin distributed systems and cloud infrastructure

What I build

Seven things, all aimed at the same problem: AI that works in a demo and has to keep working in production.

Retrieval

RAG that returns the right answer

Hybrid retrieval across dense and lexical lanes, evaluation harnesses, and the golden sets that turn "it works" from an opinion into a number.

12M+ documents on OpenSearch Serverless · 40M vectors on Qdrant · p95 under 25ms at 1.5K QPS

Agents

Agents wired to your systems

Multi-step agents with tool-calling contracts, MCP integrations to internal APIs, human approval gates on anything irreversible, and audit trails on every decision path.

Bedrock Agents and Step Functions, in production

Automation

Workflows that run without anyone asking

Event-driven orchestration across the systems you already run — scheduled pipelines, deployment automation, and data movement between internal tools, with the expensive steps gated so they only fire when something actually changed.

Release time 4 hours → 20 minutes · $120K/year in tooling costs removed

Serving & cost

Managed or self-hosted, decided on measurement

Model gateways, routing, streaming and rate limiting. Then the build-versus-buy call made with real numbers instead of preference.

55% lower per-token cost on batch · $180K of annual spend removed

Platform

Infrastructure that survives an audit

Terraform for the whole footprint, EKS, CI/CD, GitOps delivery, observability, and compliance-ready architecture on AWS or Azure.

SOC 2 aligned · all model traffic on PrivateLink

Integrations

Connecting what you already run

APIs and connectors between internal tools, data platforms and the systems your business depends on. Rust, Go, Python, Java.

40K+ concurrent connections at sub-millisecond response

Lifecycle

Promotion as a gate, not an argument

Model lifecycle with versioning, staged rollout, drift detection, automated evaluation gates in CI and one-command rollback.

Prototype to production: six weeks → five days

Why AI systems slip without warning

A broken retrieval stage throws no exception. It returns plausible, scored, confident results — and every stage downstream reports success. The ticket says done. The demo works.

I published a case where pure vector search returned a book's own back-of-book index as a top-three hit at 0.353 similarity, while the chapter that actually answered the question didn't place at all. Nothing failed. Nothing came back empty. The answer was simply wrong.

That is what an AI deliverable looks like when it isn't done — and it's why every system I build ships with a way to measure whether it's working.

Who this is for

  • Teams whose RAG demo worked and whose production retrieval doesn't
  • AWS and Azure partners with funded delivery and no capacity to staff it
  • Anyone who can't currently answer "how do we know it works?" with a number
  • Teams choosing between managed and self-hosted inference, and guessing

Start with a call. Thirty minutes, no charge — you describe what's breaking, I tell you whether I can help and what it would take. If an audit is the right next step, it's five days at a fixed price.

Book a consultation call

Last updated 13 September 2026