Large Language Models

Large language models — architectures, training, evaluation, and the tools built on them.

Self-Knowledge Distillation (Self-KD) enhances vision-audio capability in Omnimodal Large Language Models (OLLMs)

Enhancing Vision-Audio Capability in Omnimodal LLMs with Self-KD

Omnimodal Large Language Models (OLLMs) like GPT-4o and Megrez have revolutionized how AI interacts with the world by seamlessly processing text, images, and audio. However, a critical performance gap persists: OLLMs perform significantly better with vision-text inputs than with vision-audio…

Enhancing Vision-Audio Capability in Omnimodal LLMs with Self-KD Read More »

Infographic showing a neural network merging English and Korean language models with dramatic performance increase arrows and a red warning sign for cultural bias.

7 Shocking Ways Merging Korean Language Models Boosts LLM Reasoning (And 1 Dangerous Pitfall to Avoid)

In the rapidly evolving world of artificial intelligence, Large Language Models (LLMs) are hitting performance ceilings—especially when it comes to complex reasoning tasks like math and logic. But what if the key to unlocking their next-level intelligence lies not in…

7 Shocking Ways Merging Korean Language Models Boosts LLM Reasoning (And 1 Dangerous Pitfall to Avoid) Read More »

ToDi, Per Token KL Divergence Control for LLM Distillation.

ToDi: Per Token KL Divergence Control for LLM Distillation

Machine Learning › Knowledge Distillation › Paper Analysis aitrendblend.com · Knowledge Distillation A seven billion parameter model writes a clean, helpful answer, then you try to ship it. It will not fit on the phone, the latency blows past your…

ToDi: Per Token KL Divergence Control for LLM Distillation Read More »