arrow_backNeural Digest
Large neural network transferring knowledge to a smaller network
Guides

What is Distillation? A Clear Guide for 2026

Distillation1h ago
auto_awesomeAI Summary

“Knowledge distillation is a technique where a smaller AI model is trained to mimic the behavior of a larger, more powerful one, capturing most of its intelligence at a fraction of the computational cost. The result is a compact model that runs faster and cheaper while still performing remarkably well. This matters enormously as AI moves onto phones, edge devices, and cost-sensitive applications where giant models simply cannot go.”

Imagine a seasoned expert who has spent decades learning everything about medicine. Now imagine they spend a year intensively tutoring a bright medical student, not just handing them textbooks, but sharing intuitions, shortcuts, and the reasoning behind every decision. The student graduates knowing far more than they could have learned on their own in the same time. That is essentially what knowledge distillation does in AI — it transfers the deep, nuanced understanding of a large, expensive model into a smaller, leaner one. In technical terms, distillation is a model compression technique introduced by Geoffrey Hinton and colleagues in 2015. A large pre-trained model, called the teacher, guides the training of a smaller model, called the student. Instead of the student learning only from raw labeled data, it learns from the teacher's outputs — including the teacher's probability distributions across all possible answers, not just its final single prediction. Those soft probability outputs carry far richer signal than a simple right-or-wrong label ever could. The result is a student model that punches well above its weight. A distilled model might be ten times smaller and five times faster than its teacher, yet retain 95% or more of the teacher's accuracy on real tasks. This gap between size and capability is what makes distillation one of the most practically valuable ideas in modern AI.

How It Works

The mechanics hinge on a concept called soft labels. When a teacher model classifies an image of a cat, it does not just say 'cat.' It outputs something like: 87% cat, 9% lynx, 3% fox, 1% everything else. Those near-miss probabilities reveal what the model has learned about the relationships between categories — cats and lynxes are visually similar in ways cats and fire trucks are not. A student trained only on hard labels (cat = 1, everything else = 0) never sees this relational knowledge. A student trained on soft labels absorbs it naturally. Training the student involves a combined loss function. One part measures how closely the student's soft outputs match the teacher's soft outputs — this is the distillation loss, often computed using a metric called KL divergence. Another part measures how well the student matches the true ground-truth labels — the standard cross-entropy loss. A temperature parameter controls how 'soft' the teacher's probabilities are made before passing them to the student; higher temperatures spread the probability mass more evenly, making the hidden relationships more visible and easier to learn from. Some advanced distillation approaches go further than just matching final outputs. Feature-based distillation trains the student to mimic the teacher's internal activations at intermediate layers, capturing richer structural knowledge. Relation-based distillation teaches the student to preserve the geometric relationships between data points that the teacher has learned. These variants have become especially important for distilling very large language models, where the gap between teacher and student is enormous.

trending_upWhy It Matters

Without distillation, deploying powerful AI would be prohibitively expensive and physically impossible in many contexts. Models like GPT-4 require clusters of specialized hardware just to run inference. Distillation is how AI gets onto your phone, into your browser, inside a medical device, or embedded in a car — environments where you cannot ship a data center. It democratizes access: a startup can take a large open-source model, distill it, and serve millions of users on modest infrastructure that a frontier lab's original model could never support. Distillation also plays a central role in the development of frontier models themselves. Techniques like reinforcement learning from human feedback (RLHF) often incorporate distillation-style training, and many of the most capable small models — including Meta's Llama variants and Google's Gemma — are improved through distillation from larger counterparts. As AI regulations increasingly scrutinize energy consumption and carbon footprints, distillation is emerging as an essential tool for building AI that is not just capable, but sustainable.

Real-World Examples

  • DistilBERT, published by Hugging Face in 2019, is one of the most cited examples: it distilled Google's BERT model down to 40% of its size while retaining 97% of its language understanding performance, and it became a foundational model for thousands of downstream NLP applications.
  • Apple uses on-device distilled models to power features like Siri's voice recognition and the writing tools in iOS 18, allowing sophisticated AI to run locally on an iPhone without sending data to the cloud or draining the battery.
  • Microsoft's Phi series of small language models — including Phi-3 Mini — are explicitly trained using distillation and curated data from larger models, achieving GPT-3.5-class reasoning in a model small enough to run on a laptop CPU.
  • DeepSeek's R1 model family, released in early 2025, used aggressive distillation from its larger reasoning models to produce compact versions that outperformed much larger competitors on benchmarks, demonstrating how distillation can transfer complex chain-of-thought reasoning, not just pattern recognition.

FAQ

Is the distilled model just a copy of the teacher?expand_more
No — the student model has a fundamentally different, smaller architecture and does not share the teacher's weights. It learns to reproduce the teacher's behavior through training, not by copying its structure. Think of it as learning the teacher's judgment rather than cloning the teacher's brain.
Does distillation always result in lower accuracy?expand_more
Usually there is a small accuracy trade-off, but it is often surprisingly minor — and sometimes a well-distilled student outperforms naively trained models of the same size. The student benefits from the richer training signal the teacher provides, which can make it more robust than a model trained on hard labels alone.
How is distillation different from fine-tuning?expand_more
Fine-tuning takes an existing model and continues training it on new data to adapt it to a specific task, without changing its size. Distillation creates a new, smaller model that learns from a larger one's outputs. The two techniques are complementary and are often used together: you might distill a large general model into a small one, then fine-tune the small model for your specific use case.
Do you need access to the teacher model's internal workings to distill it?expand_more
For standard output-based distillation, you only need the teacher's output probabilities — you do not need its weights or architecture. This means you can distill from a model accessed via an API. However, more advanced feature-based distillation methods do require access to the teacher's internal layers, which means you need the full model available locally.

Related Terms

This explainer was AI-generated based on publicly available information and may not reflect the most recent developments. For the latest details, consult the sources below.

Explore more AI termsarrow_forward
Share this explainer

Related Articles