Contents
Quantization shrinks the numbers. Pruning removes the structure. Knowledge distillation does something genuinely stranger: it trains an entirely new, smaller network from scratch, using a bigger network's behavior as the training signal instead of raw labeled data. The student never touches most of what the teacher learned directly — it just learns to imitate the teacher's outputs closely enough that the imitation becomes genuinely useful on its own.
This completes the third pillar of the model compression trio this series has been building toward — quantization compresses precision, pruning removes redundant structure, and distillation trains a fundamentally smaller architecture to behave like a larger one. All three compound together in real production pipelines, and this is genuinely the most conceptually different of the three.
By the end of this guide, you'll understand exactly how teacher-student training works, the different levels at which knowledge actually transfers, and why this technique has become standard practice for shrinking today's largest models into genuinely deployable ones. IMO, the "soft labels contain dark knowledge" insight is one of the more elegant ideas in this entire compression toolkit :)
The Core Idea: Learning From Soft Labels, Not Just Hard Answers
Knowledge distillation transfers capabilities from a large, complex "teacher" model to a smaller, simpler "student" model, letting the student perform tasks with similar proficiency but at dramatically reduced computational cost.
Here's the genuinely clever mechanism: instead of training the student directly on raw ground-truth labels the way you normally would, you train it on the teacher's own predictions — specifically the teacher's full probability distribution over possible answers, not just the single correct label.
- A hard label says "this image is a dog" — full stop, one bit of information.
- The teacher's soft prediction might say "90% dog, 8% wolf, 2% everything else" — and that 8% wolf probability is genuinely informative. It tells the student something real about how visually similar dogs and wolves are, information a hard label simply cannot carry.
This extra signal is often called "dark knowledge" — information about the relative similarities between classes that hard labels never capture, but which the teacher's training has genuinely learned to represent in its output distribution.
Think of it like learning physics from a professor's synthesized lecture notes, rather than trying to read every raw textbook yourself — the teacher has already done the work of distilling relationships and nuance; the student learns from that distillation rather than starting from zero.
The Three Levels at Which Knowledge Actually Transfers
Distillation isn't a single technique — it operates at genuinely different depths, and knowing the distinction matters for choosing an approach.
Logit-Based (Response-Based) Distillation
The student learns from the teacher's final output probabilities — the classic Hinton et al. 2015 approach, and the most straightforward version to implement.
The student is trained to match the teacher's softened output distribution, typically using a Temperature parameter that controls how "soft" those probabilities are — higher temperature spreads probability mass more evenly across classes, revealing more of that dark knowledge.
An Alpha parameter balances the student's loss between matching the teacher's soft predictions and matching the actual ground-truth hard labels — getting this balance right is genuinely one of the trickier parts of the training process.
Feature-Based Distillation
The student learns to mimic the teacher's internal representations, not just its final answer — matching activations at intermediate hidden layers, not only the output layer.
This requires aligning feature maps between teacher and student, sometimes through hint layers with L2 or cosine similarity losses connecting corresponding points in each network.
This is genuinely how models like DistilBERT were created — the student doesn't just copy BERT's final predictions, it learns to approximate BERT's internal understanding at multiple depths throughout the network.
Relation-Based Distillation
Goes a level deeper still, copying structural relationships between predictions or features — not individual values, but how different examples relate to each other in the teacher's learned representation space.
For transformer models specifically, attention-based distillation encourages the student to replicate the teacher's attention maps — literally learning to focus on the same parts of the input the teacher focuses on.
This category also covers Step-by-Step Distillation for LLMs specifically — extracting a teacher's chain-of-thought reasoning steps, not just its final answer, and training the student to reproduce that reasoning process.
Black-Box vs. White-Box Distillation: A Distinction Specific to LLMs
This matters enormously once you're distilling large language models rather than smaller classification networks, since LLM teachers are frequently not something you have full access to.
- White-box distillation assumes full access to the teacher's architecture and weights — you can pull internal representations, attention maps, and full output distributions directly. This enables the richer feature-based and relation-based techniques described above.
- Black-box distillation assumes you only have access to the teacher's outputs — genuinely common when your teacher is a closed-source model behind an API, where you can query it but never see its internals. Here, you typically prompt the teacher to generate a distillation dataset of input-output pairs, then fine-tune the student on that generated data directly.
A striking real result from this approach: Distilling Step-by-Step extracted chain-of-thought rationales from a large teacher to train smaller models more data-efficiently — a 770M-parameter student model outperformed a few-shot prompted 540B-parameter PaLM model on the target task. The student, correctly taught, beat a teacher over 700 times its size at the specific task it was distilled for.
A Practical Distillation Pipeline
Here's the general shape most implementations follow, regardless of which level of distillation you're using.
- Train (or obtain) the teacher. This might mean fine-tuning a large model like BERT or a 70B-parameter LLM on your specific task, or simply having API access to an already-capable large model.
- Generate soft targets. Pass your training data through the teacher and capture its logits, probability distributions, or (for black-box LLM distillation) its full generated responses.
- Design the student. Build a genuinely smaller architecture — fewer layers, fewer parameters, a fundamentally lighter model than the teacher.
- Train with a distillation loss combining the student's match to the teacher's soft outputs with (optionally) the ground-truth hard labels, weighted by that Alpha parameter mentioned above.
- Validate against real task performance, not just how closely the student mimics the teacher's outputs — a student that copies the teacher perfectly but still fails the actual task hasn't accomplished the real goal.
Why Reverse KL Divergence Matters for Generative Models
Worth knowing a genuinely important technical wrinkle specific to distilling LLMs rather than simple classifiers: MiniLLM specifically argues that reverse KL divergence is a better distillation objective than the standard approach for generative tasks.
Standard forward KL divergence can cause the student to overestimate probability in regions the teacher considers low-probability — genuinely problematic for a generative model, where this can manifest as hallucination or incoherent output in areas the teacher would have avoided.
Reverse KL divergence better suits generative language models specifically by preventing exactly this overestimation problem, leading to improved student response quality and reliability.
MiniLLM also introduces additional techniques — single-step regularization, teacher-mixed sampling, and length normalization — specifically addressing training instabilities that show up when distilling generative models rather than simple classifiers.
This is a genuinely good example of why "just copy the classic Hinton approach" doesn't automatically transfer cleanly to every model type — generative language models have different failure modes than image classifiers, and the distillation objective needs adjusting accordingly.
The Multi-Teacher Trap
Here's a genuinely counterintuitive finding worth knowing before you assume "more teachers, better student" is a safe assumption: research specifically investigating multi-teacher distillation frameworks found that distillation performance actually declined as the number of teacher LLMs increased, contrary to the expectation that a larger teacher ensemble would enhance student capability.
The proposed fix involves knowledge purification — integrating rationales generated by multiple teacher LLMs into a single, consolidated rationale before using it in distillation, rather than naively averaging or mixing raw outputs from several teachers.
The practical takeaway: don't assume stacking more teachers is automatically beneficial. Conflicting or inconsistent guidance from multiple teachers can genuinely confuse student training rather than enriching it.
Distillation vs. Fine-Tuning vs. Quantization/Pruning: Getting the Distinctions Straight
These three techniques get lumped together constantly, but they solve genuinely different problems.
- Fine-tuning takes a pretrained model and updates it on task-specific data — it adjusts what the model knows, but doesn't reduce its size or compute requirements at all.
- Quantization and pruning (covered in the previous two articles) reduce an existing model's size by lowering precision or removing weights — they compress the same architecture, they don't create a new one.
- Distillation trains a genuinely different, smaller architecture from the ground up, using the teacher's behavior as the training signal — this is architectural compression, not just precision or sparsity compression.
Distillation works best specifically when data is limited, inference speed genuinely matters, or you're trying to replicate broad, general-purpose capability in a much smaller footprint — different sweet spot than either of the previous two techniques.
Real Production Use Cases
This isn't a niche academic technique — distillation has become standard practice across the industry.
- Real-time applications requiring sub-millisecond latency — autonomous driving, real-time translation — where a giant model's latency is simply unacceptable regardless of accuracy gains.
- High-volume APIs: if you're serving millions of users, a distilled model dramatically slashes server costs compared to running the full-size teacher for every request.
- Edge computing: getting the reasoning power of a large model running locally on a device with no internet connection — connecting directly to nearly everything covered in this series' edge AI arc.
Examples in genuine production use today include Amazon Bedrock's model distillation offering, and OpenAI's own smaller "distilled" model variants.
If you're training distilled models for production, a capable GPU speeds up the teacher inference and student training loops significantly — the RTX 5070 and RTX 5080 both handle teacher-student pipelines for 7B-70B teacher models without breaking the bank, and the faster iteration cycles genuinely matter when tuning Temperature and Alpha hyperparameters.
Want to Go Deeper?
If the teacher-student training or soft-label mechanisms clicked and you want to dig into the theory behind distillation algorithms and loss functions, Educative's Machine Learning path covers distillation in detail alongside the broader model compression landscape — worth exploring if you're building production distillation pipelines.
Common Mistakes People Make
- Assuming standard forward KL divergence works fine for distilling generative models. Reverse KL divergence genuinely matters for LLMs specifically — this isn't a minor implementation detail, it affects output quality and reliability meaningfully.
- Stacking multiple teachers assuming it automatically improves the student. Research shows this can actually hurt performance without a knowledge purification step consolidating conflicting teacher signals.
- Confusing distillation with fine-tuning. Fine-tuning adjusts an existing model's knowledge; distillation trains a genuinely smaller architecture — conflating the two leads to picking the wrong tool for a size-reduction goal.
- Ignoring the Temperature and Alpha hyperparameters. Getting the soft-label temperature and the hard/soft loss balance wrong is a genuinely common reason distillation underperforms — this pairing requires real tuning, not default values.
- Underestimating the engineering overhead of managing two models during training. Distillation's pipeline complexity is real — you're running both teacher and student, capturing intermediate outputs, and managing a more complex training loop than standard supervised learning.
Where This Fits With the Rest of This Series
This genuinely closes out the compression trilogy running through this edge AI arc: quantization reduces numeric precision, pruning removes redundant structure, and distillation trains an entirely smaller architecture using a larger one's behavior as the teaching signal. In real production pipelines, these compound together — a distilled student model gets quantized and potentially pruned further before final edge deployment, exactly the kind of layered compression the pruning article's combined-technique results demonstrated.
This also directly explains something you may have wondered about earlier in this series: the compact GGUF-quantized models running through Ollama and llama.cpp are frequently themselves distilled variants of larger base models before quantization ever touches them — DeepSeek's distilled model family being a genuinely current, widely-used example.
Wrapping This Up
Knowledge distillation solves model compression from a genuinely different angle than quantization or pruning: instead of compressing an existing model's representation, it trains a fundamentally smaller architecture to replicate a larger one's behavior, learning from soft probability distributions, internal representations, or even reasoning traces rather than raw ground-truth labels alone. The "dark knowledge" insight — that a teacher's mistakes and uncertainty carry genuine information — is what makes this work better than training the small model from scratch on the same raw data.
Remember that generative models need reverse KL divergence rather than the classic distillation objective, and that more teachers isn't automatically better without a consolidation step handling conflicting guidance. FYI, the fact that a 770M-parameter distilled student has been shown to outperform a 540B-parameter teacher on a specific task is a genuine reminder that "smaller" and "worse" are not the same claim when the distillation process is done well :)
Now go back to the quantization and pruning articles from earlier in this series and consider how all three would stack on the same model — a distilled architecture, pruned for structural redundancy, then quantized down to something like a Q4_K_M GGUF file. That layered pipeline is genuinely how the smallest, fastest production models actually get made.