What made it hard:
→ Half the inputs were scanned PDFs of variable quality. Vision models for OCR > legacy OCR engines.
→ Hundreds of inputs/day. Pipeline extracts, groups, dedupes, rewrites, and scores by importance.
→ LLM landscape moves fast — we built a swappable provider layer.