

I stress-test LLMs, RAG systems, AI agents, and tool-using workflows. I hunt what happy-path testing misses — unsupported claims, fabricated citations, grounding failures, tool misuse, authority-boundary errors, regressions — and convert failures into structured eval cases with explicit pass/fail criteria and reproducible evidence.
$75 - $100/hr
St Paul, USA