1:1 mentoring with Big Tech AI engineers
Q46Premium

No labelled data and the customer wants to launch in two weeks. How do you build an eval set?

Evaluation & Quality

EvaluationData QualityDeploymentProduction

Asked at Palantir · Scale AI · Sierra

How to Answer

“Two weeks is plenty for a real eval set, as long as you stop trying to build a big one.

Day one: pull two hundred real requests from whatever already exists — logs, tickets, transcripts, the questions the sales team demoed. Stratify by intent and by difficulty, and deliberately over-sample the strange ones, because easy cases don’t discriminate between two versions.

Then expert time, spent precisely. The customer’s expert doesn’t have days, they have an afternoon. So I don’t ask them to grade two hundred outputs. I ask them to answer forty cases the way they’d want the agent to — that gives me references — and to tell me what ‘wrong’ means in their domain. That conversation is worth more than the labels.

From there the forty become the seed, the model generates variations a human accepts or rejects in bulk, and the remaining hundred and sixty are graded by a judge validated against those forty.

Then ship version one with the caveat that it’s version one. A perfect eval set delivered in month three is worth less than a rough one that catches a regression in week three.”

The deep dive — diagrams, tradeoff tables, and the follow-up trap

Loading…