Local-first OCR pipeline that converts scanned PDFs and images into validated Excel, CSV, and JSON. Built with Python, Tesseract, OpenCV, FastAPI, and automated quality checks. Includes confidence scores, review flags, audit logs, and a 100% field-accuracy result on the synthetic validation dataset.