Knowledge Distillation for LLMs: How to Train Smaller Students from Big Teachers
Oct, 3 2026
Imagine trying to fit a library into a backpack. That’s essentially the challenge facing AI engineers today. We have massive Large Language Models (LLMs) like GPT-4 or LLaMA-3-70B that perform incredibly well but are too heavy and slow for most real-world applications. You can’t run a 70-billion parameter model on your phone, and running it in the cloud costs a fortune per query. The solution isn’t just making models smaller; it’s about teaching them to think smarter. This is where Knowledge Distillation comes in. It’s a technique where a large "teacher" model teaches a smaller "student" model how to mimic its behavior, preserving intelligence while slashing size and latency.
Why Traditional Training Isn't Enough
When you train a standard neural network, you usually compare its output to a single correct answer-a "hard label." If the model predicts "cat" and the image is a cat, it gets full credit. But this ignores valuable context. What if the model was 90% sure it was a dog? Or 50% sure it was a fox? Hard labels throw away this uncertainty. Knowledge Distillation changes the game by using "soft labels"-the full probability distribution over all possible outputs generated by the teacher model. These soft labels contain what Geoffrey Hinton called "dark knowledge." They tell the student not just what the right answer is, but which wrong answers are also plausible. For an LLM, this means learning that after the word "New," "York" is highly likely, "Jersey" is possible, but "Tuesday" is unlikely. By matching these distributions, the student learns the nuanced relationships between tokens, not just the next token in isolation.
The Teacher-Student Dynamic
The core mechanic is straightforward but powerful. You take a high-performing, often proprietary or very large model as the teacher. Then, you initialize a smaller model-the student-and train it to minimize the difference between its output probabilities and the teacher’s. This isn’t about copying weights; it’s about copying behavior. The student doesn’t need to know *why* the teacher thinks a certain way, only that it should produce similar results given the same input.
Think of it like an apprentice chef watching a master chef. The master (teacher) doesn’t just give the apprentice (student) the final dish. They explain why the sauce needs more salt, why the heat should be lowered, and how different ingredients interact. The student learns the process and the reasoning, allowing them to recreate the result with less experience and fewer resources. In technical terms, we use a loss function that combines standard cross-entropy (matching ground truth) with Kullback-Leibler divergence (matching the teacher’s distribution). A temperature parameter $T$ is often used to soften the teacher’s logits, making the differences between probable and improbable tokens clearer to the student.
Types of Knowledge You Can Transfer
Not all distillation is created equal. Researchers have identified several layers of knowledge that can be transferred from teacher to student, each offering different trade-offs between complexity and performance gain.
- Logit-level Distillation: This is the classic approach. The student tries to match the teacher’s raw output scores (logits) at every step. It requires access to the teacher’s internal states during training, which doubles computational cost because you need forward passes for both models simultaneously.
- Sequence-level Distillation (Data Distillation): Here, the teacher generates entire responses, summaries, or code snippets. The student is then trained on these synthetic datasets using standard supervised learning. This is easier to implement because you don’t need simultaneous inference, but it loses the fine-grained probability information. DeepSeek-R1, for example, uses this method to create high-quality instruction-tuning data.
- Preference-level Distillation: Instead of just matching outputs, the student learns to align with human preferences encoded in the teacher. This is crucial for chatbots that need to be helpful and harmless, not just accurate. The teacher acts as a reward model or preference oracle, guiding the student toward safer, more aligned responses.
- Intermediate Representation Distillation: Advanced methods try to match hidden states or attention patterns inside the transformer layers. This forces the student to learn similar internal representations, potentially improving generalization, though it’s harder to scale across different architectures.
Practical Implementation: From 8B to 4B
Let’s look at a concrete scenario. NVIDIA’s NeMo framework provides a clear recipe for compressing a Meta-Llama-3.1-8B model into a 4B student. The process often starts with pruning-removing unnecessary layers or reducing width-to create a structurally smaller base model. Then, distillation kicks in to recover the accuracy lost during pruning.
In a typical pipeline, you might drop the last 16 layers of a 32-layer model, halving its depth. This creates a fast but dumb student. Next, you run the original 8B teacher on a dataset, capturing its logits. You then train the pruned 4B student to match those logits. The key here is balancing the loss. If you rely too heavily on the teacher, you propagate its biases and errors. If you ignore the teacher, you lose the benefit of its vast pre-training knowledge. Most practitioners use a weighted sum of the distillation loss and the standard task loss.
| Method | Computational Cost | Information Retained | Best Use Case |
|---|---|---|---|
| Logit Matching | High (2x forward passes) | High (Full probability distribution) | Maximizing accuracy on fixed tasks |
| Data Distillation | Low (Offline generation) | Medium (Hard labels only) | Creating synthetic training corpora |
| Sampled Soft Labels | Medium (Subset of vocab) | Medium-High (Key alternatives) | Large-scale pre-training (e.g., Gemma) |
| Self-Distillation | Low (Same architecture) | Low-Medium (Regularization effect) | Fine-tuning and robustness |
Overcoming Computational Bottlenecks
The biggest hurdle in proper logit distillation is memory and compute. Running a 70B teacher alongside a 7B student for billions of tokens is expensive. To solve this, companies like Google have adopted sampled soft labels. Instead of storing the probability for all 128,000 tokens in the vocabulary, they sample the top 256 most likely tokens. This drastically reduces memory traffic while retaining most of the useful signal. Another trick is "code distillation," where the teacher and student are trained jointly in the same batch. The teacher’s output for the current batch becomes the target for the student immediately, avoiding the need to store massive intermediate files.
When Not to Use Distillation
Distillation isn’t magic. It has limits. First, the student cannot outperform the teacher on tasks where the teacher is weak. If the teacher hallucinates facts, the student will inherit those hallucinations. Second, aggressive compression can lead to catastrophic forgetting. If you shrink a model too much, it might lose broad linguistic capabilities even if it performs well on specific benchmarks. Third, bias propagation is a real risk. If the teacher has subtle biases in tone or content, the student will amplify them if not carefully regularized. Finally, if you already have a small model that meets your latency and accuracy needs, adding the complexity of a teacher-student loop might not be worth the engineering overhead.
The Future: Flipped Distillation and Beyond
Interestingly, the paradigm is shifting. Recent research explores "flipped knowledge distillation," where specialized small models teach larger general-purpose LLMs. Imagine a tiny, highly accurate medical model teaching a giant generalist LLM how to handle clinical queries better. This suggests that distillation is evolving from simple compression to a broader mechanism for knowledge transfer and ensemble building. As hardware constraints tighten and edge computing grows, techniques like quantization-aware distillation-where the student is trained to perform well even when compressed to 4-bit integers-will become standard practice.
What is the main difference between hard labels and soft labels?
Hard labels provide a single correct answer (one-hot encoding), ignoring uncertainty. Soft labels provide a probability distribution over all possible outcomes, revealing the relative likelihood of incorrect answers. This extra information helps the student model learn smoother decision boundaries and generalize better.
Does knowledge distillation require the teacher and student to have the same architecture?
No, they do not need to share the same architecture. However, they must share the same tokenizer and vocabulary so their output distributions are comparable. Different layer counts or hidden sizes are perfectly acceptable and common in practice.
How does temperature affect knowledge distillation?
The temperature parameter ($T$) softens the probability distribution produced by the teacher. A higher $T$ makes the distribution flatter, emphasizing the relative probabilities of less likely tokens. This reveals more "dark knowledge" about the structure of the problem space, helping the student learn better.
Is data distillation the same as traditional knowledge distillation?
Not exactly. Data distillation involves generating synthetic text with a teacher and training the student on it using standard cross-entropy loss. Traditional KD matches the teacher’s full probability distributions (logits). Data distillation is cheaper but transfers less detailed probabilistic information.
Can distilled models exceed the performance of the teacher?
Rarely on the same tasks. Generally, the student approximates the teacher. However, through regularization effects and noise reduction, a distilled student might sometimes show slightly better stability or generalization on specific subsets, but it rarely surpasses the teacher’s peak capability on complex reasoning tasks.