Inside LSSKD, a Self Supervised Distillation Framework That Trains Small Models Without a Teacher

Inside LSSKD, a Self Supervised Distillation Framework That Trains Small Models Without a Teacher

Knowledge distillation has a well known shape. Train a large, accurate teacher model. Train a small student model to mimic the teacher’s softened output probabilities, using the temperature trick Geoffrey Hinton popularized, rather than only the ground truth labels. The…

Inside LSSKD, a Self Supervised Distillation Framework That Trains Small Models Without a Teacher Read More »