Pydantic catches shape errors, but a plausible wrong insight can still pass. I'd keep a small set of posts with expected labels in LangSmith and rerun it after prompt changes. Are you tracking that kind of drift?
When I build invoice reconciliation, I start with matches I can explain.
For CrossCheck, I connected supplier invoices, customer invoices, WHMCS records and domain events in one flow. Exact matching runs before fuzzy and context-aware matching. Uncertain items keep their source records and matching context for finance review.
I judge the workflow by how easily finance can trace an exception back to the evidence and resolve it.
One quarterly number almost got double-counted across two reports.
Two dashboards were quietly tracking overlapping data. Caught before it went out now there's one source of truth the reports pull from, not two that can disagree.
→ A good number is only as good as the one place it actually comes from
Great catch before it shipped. I like the “one source of truth” framing - pulling reports from one validated source makes it much harder for a duplicate dashboard to quietly skew the number again.