Eval & Guardrail Platform
Design a platform to evaluate LLM/agent quality and enforce safety guardrails in production.
Key Requirements
- 01Offline eval: versioned datasets, graders, regression gating in CI
- 02Online eval: production sampling, human feedback, A/B
- 03Defense-in-depth guardrails (input + output) with a latency budget
- 04LLM-judge calibration against humans (judges can be gamed)
- 05A data flywheel: prod failures → new eval cases → fixes
Review me as:
Draw your design on the canvas before submitting.
Build your design, then submit for an AI-powered review with dimension scores, strengths, gaps, and actionable suggestions.