eval harness · scenario rubrics · subagent judges · prompt tournament · N-runs · surgical shipmust / must_not rubric. Voice, grounding-traps, jailbreaks, tool-selection, escalation, data collection, routing.must_not caps the score; hallucinating a fact tanks it.must / must_not per behavior One prompt, iterated blindly A tournament of variants scored across the whole suite Hope you didn't break anything A regression scorecard — every dimension, every shipSKILL.md), but the method and the templates are agent-agnostic. Drive it with Claude Code, Codex, or any coding agent that can spawn subagents for the judging fan-out. The LLM under test is fully pluggable — your project supplies the client (any OpenAI-compatible provider, or your own wrapper).SKILL.md:evals/agents/), then fill the 4 seams the harness marks with // FILL IN:buildSystemPrompt() — never copy the text Real tools import your tool defs, swap handlers for stubs that just record + return context LLM client a callModel(messages, tools) that talks to your provider Grounding what "the base knows" — lives in each scenario's contextSKILL.md.SKILL.md was paid for in a real incident.Posted Oct 4, 2026
Open-source agent evaluation harness that catches prompt regressions with scenario rubrics, subagent judges and repeatable N-runs before production.