Francis Anyaora's Work | ContraWork by Francis Anyaora
Francis Anyaora

Francis Anyaora

AI Data Specialist | PDF → Markdown | RAG & Vector Databases

Ready for work

Francis is ready for their next project!

Cover image for High-Throughput Data Annotation & Quality
High-Throughput Data Annotation & Quality Assurance: Delivered rigorous video annotation, data labeling, and quality assurance for advanced AI training models during my time working with Turing. I maintained strict adherence to complex labeling guidelines, ensured high daily processing throughput, and applied precision quality control to eliminate formatting noise and edge-case errors before model training.
1
4
Cover image for Hierarchical Markdown Structure for Vector
Hierarchical Markdown Structure for Vector Embeddings: Messy document layouts destroy RAG retrieval accuracy. In this project, I designed a strict Markdown heading hierarchy (#, ##, ###) specifically optimized for text chunking. By programmatically maintaining parent-child document relationships, I ensured that semantic search algorithms in the vector store can accurately retrieve the exact context an LLM needs, drastically reducing hallucinations.
1
14
Cover image for Custom Regex Library for Automated
Custom Regex Library for Automated Text Cleanup & Restructuring: Automated layout parsers often leave behind substantial structural noise. When converting dense legal text, court rulings, and statutes, you routinely encounter broken line wraps, random line breaks in the middle of sentences, ghost characters, and inconsistent citation layouts. Manually editing thousands of pages is impossible, yet feeding this raw noise into an LLM degrades context window efficiency.
1
20
Cover image for Legal Document Conversion & Markdown
Legal Document Conversion & Markdown Optimization: This is the flagship personal project. It focuses on taking raw, messy, multi-column statutes and court rulings and turning them into clean Markdown.
1
21