“Knowledge distillation is a technique where a smaller AI model is trained to mimic the behavior of a larger, more powerful one, capturing most of its intelligence at a fraction of the computational cost. The result is a compact model that runs faster and cheaper while still performing remarkably well. This matters enormously as AI moves onto phones, edge devices, and cost-sensitive applications where giant models simply cannot go.”
Imagine a seasoned expert who has spent decades learning everything about medicine. Now imagine they spend a year intensively tutoring a bright medical student, not just handing them textbooks, but sharing intuitions, shortcuts, and the reasoning behind every decision. The student graduates knowing far more than they could have learned on their own in the same time. That is essentially what knowledge distillation does in AI — it transfers the deep, nuanced understanding of a large, expensive model into a smaller, leaner one. In technical terms, distillation is a model compression technique introduced by Geoffrey Hinton and colleagues in 2015. A large pre-trained model, called the teacher, guides the training of a smaller model, called the student. Instead of the student learning only from raw labeled data, it learns from the teacher's outputs — including the teacher's probability distributions across all possible answers, not just its final single prediction. Those soft probability outputs carry far richer signal than a simple right-or-wrong label ever could. The result is a student model that punches well above its weight. A distilled model might be ten times smaller and five times faster than its teacher, yet retain 95% or more of the teacher's accuracy on real tasks. This gap between size and capability is what makes distillation one of the most practically valuable ideas in modern AI.
How It Works
The mechanics hinge on a concept called soft labels. When a teacher model classifies an image of a cat, it does not just say 'cat.' It outputs something like: 87% cat, 9% lynx, 3% fox, 1% everything else. Those near-miss probabilities reveal what the model has learned about the relationships between categories — cats and lynxes are visually similar in ways cats and fire trucks are not. A student trained only on hard labels (cat = 1, everything else = 0) never sees this relational knowledge. A student trained on soft labels absorbs it naturally. Training the student involves a combined loss function. One part measures how closely the student's soft outputs match the teacher's soft outputs — this is the distillation loss, often computed using a metric called KL divergence. Another part measures how well the student matches the true ground-truth labels — the standard cross-entropy loss. A temperature parameter controls how 'soft' the teacher's probabilities are made before passing them to the student; higher temperatures spread the probability mass more evenly, making the hidden relationships more visible and easier to learn from. Some advanced distillation approaches go further than just matching final outputs. Feature-based distillation trains the student to mimic the teacher's internal activations at intermediate layers, capturing richer structural knowledge. Relation-based distillation teaches the student to preserve the geometric relationships between data points that the teacher has learned. These variants have become especially important for distilling very large language models, where the gap between teacher and student is enormous.
trending_upWhy It Matters
Without distillation, deploying powerful AI would be prohibitively expensive and physically impossible in many contexts. Models like GPT-4 require clusters of specialized hardware just to run inference. Distillation is how AI gets onto your phone, into your browser, inside a medical device, or embedded in a car — environments where you cannot ship a data center. It democratizes access: a startup can take a large open-source model, distill it, and serve millions of users on modest infrastructure that a frontier lab's original model could never support. Distillation also plays a central role in the development of frontier models themselves. Techniques like reinforcement learning from human feedback (RLHF) often incorporate distillation-style training, and many of the most capable small models — including Meta's Llama variants and Google's Gemma — are improved through distillation from larger counterparts. As AI regulations increasingly scrutinize energy consumption and carbon footprints, distillation is emerging as an essential tool for building AI that is not just capable, but sustainable.
Real-World Examples
- DistilBERT, published by Hugging Face in 2019, is one of the most cited examples: it distilled Google's BERT model down to 40% of its size while retaining 97% of its language understanding performance, and it became a foundational model for thousands of downstream NLP applications.
- Apple uses on-device distilled models to power features like Siri's voice recognition and the writing tools in iOS 18, allowing sophisticated AI to run locally on an iPhone without sending data to the cloud or draining the battery.
- Microsoft's Phi series of small language models — including Phi-3 Mini — are explicitly trained using distillation and curated data from larger models, achieving GPT-3.5-class reasoning in a model small enough to run on a laptop CPU.
- DeepSeek's R1 model family, released in early 2025, used aggressive distillation from its larger reasoning models to produce compact versions that outperformed much larger competitors on benchmarks, demonstrating how distillation can transfer complex chain-of-thought reasoning, not just pattern recognition.
FAQ
Is the distilled model just a copy of the teacher?expand_more
Does distillation always result in lower accuracy?expand_more
How is distillation different from fine-tuning?expand_more
Do you need access to the teacher model's internal workings to distill it?expand_more
Related Terms
This explainer was AI-generated based on publicly available information and may not reflect the most recent developments. For the latest details, consult the sources below.



