Synthetic Data Generation & Curation
Design a pipeline that generates synthetic training/eval data at scale — diverse, high-quality, and uncontaminated.
Key Requirements
- 01Diversity engineering (taxonomy coverage, persona/param variation)
- 02Quality filtering with verifiers where checkable + calibrated judge
- 03Contamination checks against eval/benchmark sets (~0)
- 04Dedup; model-collapse awareness (mix in real data)
- 05Provenance/versioning; downstream improvement as the true metric
Review me as:
Draw your design on the canvas before submitting.
Build your design, then submit for an AI-powered review with dimension scores, strengths, gaps, and actionable suggestions.