Writing Assessment Benchmark: Matched-Case QA and Regression ReviewWriting Assessment Benchmark: Matched-Case QA and Regression Review
The network for creativity
Join 1.25M professional creatives like you
Connect with clients, get discovered, and run your business 100% commission-free
Creatives on Contra have earned over $150M and we are just getting started
A writing scorer can improve its average and still break a case that previously passed.
In the constructed demo inside AI Score Benchmark Pro, exact level agreement rises from 65/79 (82.3%) to 78/79 (98.7%). Yet the second fixture introduces a two-band regression at WR-006.
That is why I review matched cases alongside the headline metric: which outputs changed, which accepted cases failed, and whether feedback invents corrections.
I have published a free six-page preview with three synthetic cases and blind JSONL inputs for teams building writing-assessment or feedback tools.
The full toolkit is $199 USD one-time. Synthetic Development Edition: AI-authored provisional references, not independently human-rated. The demo fixtures are illustrative, not measured vendor results.
Post image
Back to feed
The network for creativity
Join 1.25M professional creatives like you
Connect with clients, get discovered, and run your business 100% commission-free
Creatives on Contra have earned over $150M and we are just getting started