Golden-set regression testing for LLM and extraction pipelines: scores the pipeline against 25 la...Golden-set regression testing for LLM and extraction pipelines: scores the pipeline against 25 la...
The network for creativity
Join 1.25M professional creatives like you
Connect with clients, get discovered, and run your business 100% commission-free
Creatives on Contra have earned over $150M and we are just getting started
Golden-set regression testing for LLM and extraction pipelines: scores the pipeline against 25 labeled rows, uses LLM-as-judge for fields where exact match is unfair (the judge itself is checked against human labels), and fails CI when quality regresses. It caught a real bug during development: an amount parser grabbed a store number instead of the dollar amount - field accuracy read 48%; after the fix, 100%, and the harness proves it instead of me claiming it. Offline, no API keys. Public code: github.com/jigonyoo/eval-harness
Post image
Back to feed
The network for creativity
Join 1.25M professional creatives like you
Connect with clients, get discovered, and run your business 100% commission-free
Creatives on Contra have earned over $150M and we are just getting started