Knowledge Distillation: Large to Small
Train a small, fast model to mimic a large teacher — economics, pipeline, and quality filters for production distillation.
Last updated
After this section you can
- Describe the teacher-student setup and what the student is actually learning from
- Judge when distillation beats simply serving a smaller off-the-shelf model
Knowledge Distillation: Large to Small
Train a small, fast model to mimic a large teacher model — the economics, pipeline, and quality filters for production distillation.