PaperScope
LIVE · 2026-09-29 05:40 UTC

A Statistical Perspective on Knowledge Distillation: Foundations, Classical Methods, and Large Language Model Extensions

Luyang Fang, Haoran Lu, Jiazhang Cai, Tao Wang, Huimin Cheng, Wenxuan Zhong, Ping Ma

Latestcs.CLcs.LGcs.AIcs.CV
arXiv ID
2609.33727 v1
Category
Submitted
2026-09-27

Abstract

Knowledge Distillation (KD) has emerged as a vital paradigm for transferring the capabilities of high-capacity models to efficient ``student'' counterparts, addressing critical challenges in computational cost, deployment constraints, and privacy-sensitive settings. Although KD is widely used in practice, it is often viewed primarily as an engineering technique, with a unified statistical perspective remaining less developed. This review bridges that gap by presenting a unified Bayesian formulation of KD that formulates teacher predictions as prior information. This provides a principled interpretation of how teacher information is incorporated into student learning and establishes a rigorous connection to uncertainty quantification. We demonstrate how this foundational lens reconciles classical distillation with modern extensions in generative and foundation-model systems, showing that contemporary developments remain rooted in these same statistical principles. By synthesizing theory with emerging methodologies and diverse applications, this review provides a conceptual roadmap and identifies critical open problems for the future of the field.

arXiv abs page · PDF · same-day batch