Home/Audit

Fixed scope · Fixed price · Five days

Find out what your retrieval system is actually returning.

A five-day audit that ends in a written assessment: what's failing, what it costs, and what to fix first — with numbers attached to each.

Book a consultation call

The situation this is for

A RAG production audit measures what your retrieval system actually returns, against questions with known-correct answers. It takes five days, costs $4,000, and ends in a written report with a prioritised fix list — not a proposal for more work.

Your demo worked. Production doesn't, and nobody can say precisely why.

The reason is that retrieval fails silently. A broken stage throws no exception and returns no empty set — it returns plausible, scored, confident results, and every stage downstream reports success. I published a case where pure vector search returned a book's own back-of-book index as a top-three hit at 0.353 similarity, while the chapter that actually answered the question didn't place at all.

Without a measurable acceptance criterion, "it's working" is an opinion — and the schedule risk stays invisible until a deadline is missed.

Who is measuring it

12M+documents in production retrieval on OpenSearch Serverless
40Mvectors self-hosted on Kubernetes, p95 under 25ms at 1.5K QPS
$180Kannual inference spend removed while volume grew
6 wks → 5 daysprototype to production, after CI evaluation gates

Ten years in distributed systems and cloud infrastructure. I've built retrieval managed on AWS and self-hosted on Kubernetes, which is why I can tell you which one your workload actually needs rather than which one I'd prefer to sell.

What you get

$4,000 Five working days · fixed scope · paid on delivery of the report
  • Retrieval quality, measured. A 50-question golden set built from your own corpus, with recall and precision scored — and every failing query listed
  • Hybrid search review. Dense and lexical lanes, and which results both lanes agree on. Agreement is evidence; a single-lane hit is a lead
  • Groundedness and hallucination baseline, with the failing cases logged rather than summarised
  • Token-level cost model, attributed per use case, with the oversized and untagged spend identified
  • A prioritised fix list — what to change first, estimated effort, and what each one is worth

Not included

  • Implementing the fixes — that's a separate engagement, quoted after you've read the report
  • Model fine-tuning or data migration

What I need from you

  • Read access to the knowledge base or vector store
  • Ninety minutes with whoever built it

How the five days run

  1. Access and orientation. Ninety minutes with your engineer. I map the pipeline end to end and agree what "correct" means for your corpus.
  2. Golden set. Fifty real questions with known-correct answers, built from your documents and confirmed with you before anything is scored.
  3. Measurement. Recall, precision, and per-lane provenance across the full set. Every failure logged with its query and its returned results.
  4. Cost and configuration. Token-level attribution, index and chunking review, hybrid search settings, and where spend is going that shouldn't.
  5. Report and walkthrough. Written findings, the prioritised fix list, and an hour on a call going through it with your team.

Why me

I'm an AI Architect and Engineering Manager with ten years in infrastructure. I designed a model gateway in Rust sustaining 40,000+ concurrent connections at sub-100ms overhead, architected retrieval over 12 million enterprise documents with an evaluation harness measuring groundedness and hallucination rate, and cut prototype-to-production from six weeks to five days by making promotion a gate that either passes or doesn't.

I publish the failures too — including the retrieval case above. Read the write-ups, or see the full background.

Book a consultation call

Thirty minutes, no charge. You describe what's breaking; I tell you whether an audit is the right next step, and what I'd look at first. No deck, no pitch — if it isn't a fit I'll say so on the call.

Or email shivank@embedpath.com directly.

Last updated 13 September 2026