I turn known AI failure modes, requirements, or product expectations into a reusable evaluation suite your team can run again and again. This is for teams that have already found failures manually and want them preserved as repeatable tests. Deliverables can include a JSONL/YAML evaluation corpus, explicit pass and fail criteria, deterministic validators, scoring logic, a Python test harness, Promptfoo/DeepEval-compatible fixtures where your stack genuinely supports them, a regression runner, machine-readable results, provenance and evidence records, and documentation for rerunning the suite. One fixed-price project, scoped to your needs. I only claim framework compatibility where the implementation actually supports it.