Knowledge distillation techniques

ToDi, Per Token KL Divergence Control for LLM Distillation.

ToDi: Per Token KL Divergence Control for LLM Distillation

Machine Learning › Knowledge Distillation › Paper Analysis Knowledge Distillation Forward KL Reverse KL LLM Compression Instruction Following Paper Analysis Analysis by the aitrendblend editorial team · October 2025 · 13 min read · arXiv:2505.16297 aitrendblend.com · Knowledge Distillation ToDi, Per Token Control of KL Divergence in LLM Distillation A seven billion parameter model writes […]

ToDi: Per Token KL Divergence Control for LLM Distillation Read More »

PLD: List Wise Knowledge Distillation with Plackett-Luce.

PLD: List Wise Knowledge Distillation with Plackett-Luce

Machine Learning › Knowledge Distillation › Paper Analysis Knowledge Distillation Plackett-Luce List Wise Ranking ListMLE Image Classification Paper Analysis Analysis by the aitrendblend editorial team · October 2025 · 13 min read · arXiv:2506.12542 aitrendblend.com · Knowledge Distillation PLD, List Wise Knowledge Distillation with the Plackett-Luce Model Almost every logit based distillation method shares an

PLD: List Wise Knowledge Distillation with Plackett-Luce Read More »