I pull real tasks out of your own transcripts, turn them into a graded set of about fifty, and write the scoring so a run gives you a number that actually moves. It goes in CI, so a change that breaks behaviour fails the build instead of the customer.