An eval pipeline for our RAG system, that finally tells me the truth.
I'm pulling real production Q&A pairs offline, running three rerank strategies against them, and doing proper significance tests on the outcomes. Version 2 currently beats baseline by 6.4 percentage points—two more weeks of validation before I trust the number.
STACK · Python · Weaviate · Cohere Rerank · pytest