RAG Retrieval Evaluation Harness for BM25, TF-IDF, Hybrid, and ChunkingRAG Retrieval Evaluation Harness for BM25, TF-IDF, Hybrid, and Chunking
The network for creativity
Join 1.25M professional creatives like you
Connect with clients, get discovered, and run your business 100% commission-free
Creatives on Contra have earned over $150M and we are just getting started
RAG Eval: measure retrieval before trusting an AI answer
Independent Python evaluation harness comparing BM25, TF-IDF, hybrid retrieval, and document chunking on a small synthetic corpus of eight documents and 27 questions. Relevance is derived from answer-span containment and recomputed for each chunking strategy. Fixed-size chunks without overlap make three to five questions unanswerable; overlap preserves all 27 answer spans in the included fixture. The project includes 18 tests and reports ranking metrics, cut-off saturation, and limitations. This demonstrates reproducible evaluation methods, not production chatbot accuracy.
Post image
Back to feed
The network for creativity
Join 1.25M professional creatives like you
Connect with clients, get discovered, and run your business 100% commission-free
Creatives on Contra have earned over $150M and we are just getting started