Alpha Beta Divergence Rebalances Knowledge Distillation
A 1.5 billion parameter teacher knows things its 100 million parameter student will never quite learn. The question is how to transfer what can be transferred without warping the smaller model into something either too cautious or too greedy. The…
Alpha Beta Divergence Rebalances Knowledge Distillation Read More »



