Back to Home

Voice-Enabled RAG Pipeline with Measured Guardrails

A live voice-to-answer RAG service over a 99K-passage corpus (ElevenLabs STT -> FAISS -> cross-encoder rerank -> Claude) on AWS EC2, holding a P50 of 96ms against a 200ms budget. Cross-encoder reranking lifted recall@5 from 0.848 to 0.916, with per-sentence groundedness guards catching 100% of ungrounded answers at zero false refusals.

Why reranking, not chunking, was the real lever

The obvious place to spend effort on a RAG pipeline is chunking - so that's where I started, benchmarking 8 chunking strategies against a 99K-passage corpus. None of them moved the needle. Retrieval quality was gated by ranking, not segmentation, so I ran BM25/RRF hybrid retrieval and cross-encoder reranking through the same benchmark harness, scoring every configuration with paired bootstrap 95% confidence intervals over 500 labelled queries instead of a single point estimate. Cross-encoder reranking was the one change that was statistically significant: recall@5 moved from 0.848 to 0.916. That result set the shape of the rest of the system - a lightweight first-pass retriever feeding a heavier reranker, with per-sentence groundedness guards on the output layer catching 100% of ungrounded answers at zero false refusals. The latency budget stayed intact throughout: removing a redundant re-embedding step cut that stage from 111ms to 12ms, holding the full retrieval-to-answer path to a P50 of 96ms against a 200ms budget.