Reproducible Python Evaluation Pipeline
Self-directed technical case study. Built and executed a 32-case bilingual evaluation of Spark-X2.5-1.7B with exact rational reference answers, resumable execution, strict output parsing, and complete raw traces.
Deliverables: Python source code, versioned data, documented runtime, 32 traces, evidence reconciliation, and a report. Thirteen automated tests passed.
At a fixed 2048-token budget, 7/32 outputs delivered correct, parseable finals. Other outputs often contained correct mathematics but were unfinished or lacked the required format. This measures answer delivery under that budget, not general mathematical accuracy.
This was an AI-executed, self-directed research project on behalf of the account owner, not a paid client engagement. Full methods, attribution, source code, and evidence:
https://github.com/codeofwxz/spark-x25-observation-protocols