How to Answer
“Two weeks is plenty for a real eval set, as long as you stop trying to build a big one.
Day one: pull two hundred real requests from whatever already exists — logs, tickets, transcripts, the questions the sales team demoed. Stratify by intent and by difficulty, and deliberately over-sample the strange ones, because easy cases don’t discriminate between two versions.
Then expert time, spent precisely. The customer’s expert doesn’t have days, they have an afternoon. So I don’t ask them to grade two hundred outputs. I ask them to answer forty cases the way they’d want the agent to — that gives me references — and to tell me what ‘wrong’ means in their domain. That conversation is worth more than the labels.
From there the forty become the seed, the model generates variations a human accepts or rejects in bulk, and the remaining hundred and sixty are graded by a judge validated against those forty.
Then ship version one with the caveat that it’s version one. A perfect eval set delivered in month three is worth less than a rough one that catches a regression in week three.”