Distilling the Knowledge in a Neural Network
Soft targets from a big model teach a small one almost everything that matters.
Soft targets from a big model teach a small one almost everything that matters.
What this paper explains
Soft targets from a big model teach a small one almost everything that matters.
What to notice
- What problem the paper set out to solve, and why earlier approaches stalled there.
- The key mechanism or formula introduced, in the authors' own terms.
- Which results held up, and which assumptions later work relaxed.
How to read it
Read the abstract and introduction for the problem setup, then the method section for the core mechanism. Skim experiments for what actually moved.
Sources
- Authors: Hinton et al. (2015)