Hello Builders,
I’ve been working on a RAG project that required a large corpus of legal texts comprising statutes, court rulings, etc.
One thing I’ve noticed is that even top-notch OCR/parsers don’t always produce pristine, RAG-ready data. 💯
For high-accuracy retrieval, I’ve found that human QA is valuable for preserving document structure in Markdown/JSON with proper headers, line breaks, indentation, citations, etc., even when using regex/Python scripts.
If you’re building with a legal or other complex corpus, I’m available to help as your expert-in-the-loop.
Happy to connect! 😊