Independent Python evaluation harness comparing BM25, TF-IDF, hybrid retrieval, and document chunking on a small synthetic corpus of eight documents and 27 questions. Relevance is derived from answer-span containment and recomputed for each chunking strategy. Fixed-size chunks without overlap make three to five questions unanswerable; overlap preserves all 27 answer spans in the included fixture. The project includes 18 tests and reports ranking metrics, cut-off saturation, and limitations. This demonstrates reproducible evaluation methods, not production chatbot accuracy.