Knowledge Distillation
Knowledge distillation trains a smaller student model to reproduce the behaviour of a larger teacher model, transferring much of the teacher's accuracy into a model that is cheaper and faster to run.
Instead of learning only from hard labels, the student learns from the teacher's output probabilities, often softened with a temperature parameter, which carry information about how classes relate to each other. The technique was popularised by Hinton and colleagues in 2015. For language models, distillation can also use teacher-generated responses or reasoning traces as training data for the student.
Distillation produces compact models for edge devices, mobile applications and high-volume services, for example distilling a large vision model into one that runs in real time on a line-side device, or a large language model into a smaller one for a specific task.
Students usually lose some accuracy relative to the teacher, especially on rare or complex cases, so evaluation must include those cases. Distillation is often combined with quantisation and pruning. The terms of service of some commercial models restrict using their outputs to train other models, which must be checked.
Key points
- A small student model learns to imitate a large teacher
- Uses the teacher's output probabilities or generated responses
- Produces faster, cheaper models for edge and high-volume use
- Often combined with quantisation and pruning
Where AiVibe comes in
AiVibe delivers AI and machine learning services, chatbots and virtual assistants with RAG, MCP tools and voice, AI quality management including bias detection and model validation, and the AIMURUGA AI agent, and builds Intel-based edge AI devices using the Intel Distribution of OpenVINO toolkit.