Key points
- The paper merges a Korean Llama 3 model into a Llama 3 based English model using DARE, a method that randomly drops most of a fine tuned model’s weight changes and rescales the rest, with no additional training.
- At a density of 0.25 the merged model averages 76.23 across six benchmarks against 74.95 for the base, a 1.28 point gain, and GSM8K rises from 58.76 to 69.37.
- The gain is not free. TruthfulQA falls from 76.00 to 71.95 and HellaSwag slips from 92.74 to 91.78 at the same setting.
- Chinese and Japanese merges also raised GSM8K, by roughly 7.5 and 8.4 points, so the Korean specific explanation the authors favor is suggestive rather than isolated.
- DARE edged out Slerp and Breadcrumbs on average score by very small margins, while TIES merging failed to beat the base model.
- Density and weight were tuned against the same benchmarks that are then reported, and results appear to come from single runs.
Why merge weights when you could just write a better prompt
Most attempts to make a language model reason better work from the outside. Chain of Thought prompting, Tree of Thought, and their many descendants coax more careful step by step behavior out of a model that already exists. The authors argue these approaches mostly reorganize how a model uses what it knows rather than changing what it knows, and that they can be brittle, since small prompt changes can swing results.
Model merging works from the inside. Instead of asking the model to behave differently, you combine the weights of two models so the result inherits behavior from both, without access to either model’s training data and without a training run. The family of merging methods the paper cites includes plain averaging, Task Arithmetic, Fisher merging, and TIES-Merging. The authors pick DARE, short for Drop And REscale, and pair it with a hypothesis that is the most eye catching part of the paper. Korean is morphologically rich, with heavy inflection, subject object verb order, and a great deal of meaning carried by context, and the authors suspect a model trained on lots of Korean text picks up reasoning habits that can be transplanted into another model.
What DARE actually does to the weights
DARE starts from a simple observation about fine tuning. When you fine tune a pretrained model, the weights move a little, and the difference between the fine tuned weights and the original weights is called the delta. The DARE authors, whose work this paper builds on, found that a large share of those delta values can be thrown away at random with little loss, provided the survivors are scaled up to compensate.
Algorithm 1 in the paper spells out the two steps. A Bernoulli mask keeps each delta entry with probability equal to the density d, and every kept entry is divided by d so the expected size of the delta stays the same.
Here the density d is the fraction of delta entries kept, so a density of 0.25 means 75 percent of the changes are discarded. The weight w scales how strongly the surviving delta is added, and the paper fixes it at 0.5 while sweeping density. The algorithm box in the paper shows the sum without the weight, and the weight enters through the MergeKit configuration the authors used for their experiments. Rescaling is what keeps the trick honest, because dividing by d makes the expected value of each masked entry equal to the original entry.
The merged model has exactly the same number of parameters as either parent. Nothing is added to the architecture, and no gradient step is taken.
The experimental setup in plain terms
Both parents are Llama 3 8B derivatives from the Hugging Face hub. The base model, called Base-LM in the paper, is swap-uniba/LLaMAntino-3-ANITA-8B-Inst-DPO-ITA. The Korean model, Ko-LM, is beomi/Llama-3-Open-Ko-8B. Merging used the MergeKit toolkit, and evaluation used the language model evaluation harness in a few shot setting on the six benchmarks that make up the Open LLM Leaderboard, which are ARC, HellaSwag, MMLU, TruthfulQA, Winogrande, and GSM8K. Everything ran on a single A100 80 GB GPU in bfloat16, and the authors note that merging itself needs no GPU and only the evaluation does.
The paper’s chosen benchmarks span science questions, commonsense sentence completion, broad academic knowledge, truthfulness, pronoun resolution, and grade school math. The authors offer a hypothesis for how Korean might help each one, for instance that Korean drops pronouns and forces readers to infer referents, which might help Winogrande. These are framed as hypotheses, and the paper does not test them individually.
Sweeping the density, and where the gains appear
Because density is the one free knob, the authors merged at eight settings from 0.05 to 0.40 and compared against the unmerged base at density zero. The table below reproduces the columns that matter most from the paper’s appendix.
| Density | GSM8K | TruthfulQA | MMLU | Average |
|---|---|---|---|---|
| 0.00 (base) | 58.76 | 76.00 | 65.67 | 74.95 |
| 0.05 | 61.26 | 75.33 | 65.84 | 75.37 |
| 0.10 | 64.75 | 74.31 | 66.34 | 75.72 |
| 0.15 | 68.08 | 73.59 | 66.64 | 76.21 |
| 0.20 | 67.85 | 71.91 | 66.60 | 75.89 |
| 0.25 | 69.37 | 71.95 | 66.84 | 76.23 |
| 0.30 | 70.36 | 70.85 | 66.91 | 75.98 |
| 0.35 | 71.19 | 70.23 | 67.21 | 75.78 |
| 0.40 | 70.20 | 69.27 | 66.93 | 75.19 |
Two trends run in opposite directions. GSM8K climbs almost steadily as density rises, peaking at 71.19 at density 0.35, while TruthfulQA falls steadily from 76.00 to 69.27. The average score peaks in the middle at 0.25 because those two effects trade off. The authors describe an optimal balance between preserving the base model’s knowledge and absorbing the Korean model’s, and choose 0.25 for every later experiment.
A note on the headline numbers. The abstract reports a 1.69 percent improvement and the body reports 1.70 percent, and both refer to the relative gain in average score. In absolute terms the average moves from 74.95 to 76.23, which is 1.28 points, a figure the paper itself uses later in its coding data comparison. The abstract also says GSM8K improves by over 20 percent. At the chosen density of 0.25 the rise from 58.76 to 69.37 is about 18 percent in relative terms, and the 20 percent mark is passed only at higher densities such as 0.35, where 71.19 is roughly 21 percent above the base. Both framings are defensible, but the 0.25 figure in points, a gain of 10.61, is the cleaner one to quote.
The full picture at the chosen setting
| Benchmark | Base-LM | Merged at d=0.25 | Change |
|---|---|---|---|
| ARC | 75.17 | 75.17 | 0.00 |
| HellaSwag | 92.74 | 91.78 | -0.96 |
| MMLU | 65.67 | 66.84 | +1.17 |
| TruthfulQA | 76.00 | 71.95 | -4.05 |
| Winogrande | 81.37 | 82.24 | +0.87 |
| GSM8K | 58.76 | 69.37 | +10.61 |
| Average | 74.95 | 76.23 | +1.28 |
Read this way, the average gain is almost entirely a GSM8K story. One benchmark rose by more than ten points, two rose by about one point, one was flat, and two fell. The authors acknowledge the TruthfulQA drop and attribute it to cultural and legal differences in what counts as truthful. When TruthfulQA is excluded the average rises from about 74.74 to 77.08, which the paper reports as a 3.13 percent gain. That is arithmetically correct, though dropping the one benchmark that got worse is a choice a reader should weigh.
Is Korean special, or does any merge help
This is the question the whole paper hangs on. To test it the authors merged the same base model with a Chinese model, hfl/llama-3-chinese-8b, and a Japanese model, alfredplpl/Llama-3-8B-Instruct-Ja, each at its own best density.
| Merged with | Best density | Average | GSM8K | MT-Bench score | Adjusted win rate |
|---|---|---|---|---|---|
| Nothing (Base-LM) | 0.00 | 74.95 | 58.76 | not reported | not reported |
| Korean (Ko-LM) | 0.25 | 76.23 | 69.37 | 6.719 | 0.533 |
| Chinese (Zh-LM) | 0.15 | 75.97 | 66.26 | 6.713 | 0.524 |
| Japanese (Ja-LM) | 0.20 | 75.43 | 67.17 | 6.500 | 0.439 |
Korean comes first on average score and on GSM8K, but the margins are thin and the controls matter. The Chinese merge scores 75.97 on average, only 0.26 below the Korean merge, and the Japanese merge lifts GSM8K by about 8.4 points against the Korean merge’s 10.6. On MT-Bench, judged by GPT-4o across multi turn dialogue, the Korean merge scores 6.719 against 6.713 for Chinese, a gap of 0.006 that is well inside what one would expect from judge noise. Only the Japanese merge trails visibly.
The authors read this as evidence that Korean’s linguistic features help reasoning more effectively than other languages. A more cautious reading is that merging any sibling Llama 3 derivative with its own fine tuning history moves GSM8K up, with Korean moving it a bit more. The paper also reports that the Korean model on its own scores just 30.86 on GSM8K, far below the base, so the transferred gain cannot be explained by the Korean model being good at math. Something in the delta is helping, and the experiments do not say what.
Despite Ko-LM’s very low GSM8K score of 30.86, merging with Ko-LM still results in a performance improvement of over 20%. From Section 4.2 of the paper, where the authors highlight the puzzle themselves
Does it hold on a different, larger model
To check the result is not specific to Llama 3 8B, the authors repeated the experiment with SOLAR 10.7B, merging abhishekchohan/SOLAR-10.7B-Instruct-Forest-DPO-v1 with beomi/OPEN-SOLAR-KO-10.7B. The base scores 74.09 on average. The best merged average is 74.68 at density 0.10, a gain of 0.59 points, or 0.8 percent. GSM8K rises from 63.31 to 66.72 at density 0.15 and 66.03 at density 0.10. At density 0.25, the setting that was best for Llama 3, the SOLAR average is 74.09, exactly the base.
That pattern supports two conclusions. The effect appears to generalize in direction, and the best density shifts, with the larger model preferring a lower density. The paper connects this to the original DARE finding that larger models tolerate dropping more of their delta. It also shows that the gain is smaller here, and that the density found for one model pair should not be assumed for another.
DARE against the other merging methods
The authors also compared DARE with TIES merging, Slerp, and Breadcrumbs, each at its best setting on the same Ko-LM and Base-LM pair.
| Method | Best setting | Average score |
|---|---|---|
| Base-LM alone | none | 74.95 |
| TIES merging | d=0.05 | 74.88 |
| Breadcrumbs | d=0.05 | 76.00 |
| Slerp | t=0.15 | 76.16 |
| DARE | d=0.25 | 76.23 |
DARE wins, though by 0.07 points over Slerp and 0.23 over Breadcrumbs. With single runs and no reported variance, those gaps do not separate the methods. What the table does show clearly is that TIES merging did not beat the unmerged model at any tested setting, and that its scores fell steadily as density rose, reaching 62.08 at density 0.40. All three of the methods that worked produced roughly the same benefit, which fits the reading that the useful ingredient is the Ko-LM delta itself and the merge recipe matters less.
Korean data against coding data
The last comparison pits Korean against a coding model, ajibawa-2023/Code-Llama-3-8B, motivated by earlier work showing that coding instruction data improves reasoning. Merged at the same density and weight, the coding model produced the biggest GSM8K gain but degraded ARC, HellaSwag, TruthfulQA, and Winogrande. The Korean merge gave a smaller GSM8K gain, smaller losses elsewhere, and small gains on MMLU and Winogrande. On average the Korean merge gained 1.28 points and beat the coding merge by 0.5 points.
This is the most practically useful result in the paper. It suggests that if the goal is a balanced improvement rather than a math spike, a language merge may cost less than a coding merge, at least for this base model.
What the evidence does and does not support
The data support a modest, reproducible sounding claim. Merging a sibling Llama 3 derivative with DARE, with no training, can lift GSM8K substantially at some cost to TruthfulQA, and the effect shows up on two model families. The data do not yet support the stronger claim in the abstract that the complexity of Korean is what drives the gain, for reasons that come directly from the paper’s own tables.
First, the Chinese and Japanese merges raise GSM8K too, so the effect is not exclusive to Korean. Second, there is no control merge with a model that has no unusual language or domain, so the paper cannot separate a language effect from an ordinary effect of adding a differently fine tuned sibling. Third, the base model identifier ends in ITA, which suggests an Italian adaptation, while the paper describes the base as an English model and Appendix A calls it Llama 3 8B Instruct. Which description is right matters for interpreting a cross lingual story, and readers should check the model card.
Honest limitations
The paper is upfront about several limits. DARE as used here needs models with the same architecture, so it cannot merge across different designs, and multimodal merging is left for future work. The authors also warn that merging models from different linguistic and cultural data could produce semantic conflicts, amplify regional or gender biases, or yield emergent behavior that neither parent showed, which they flag as an interpretability and safety risk. They propose behavioral consistency tests and bias audits as future work, and note that none were run here.
Some further limits come from reading the setup closely. The density and weight were chosen by maximizing the average over the very benchmarks that are reported, with no held out set, so the headline numbers are likely somewhat optimistic. Results appear to be single runs, and DARE involves random masks, so the run to run variance is unknown and matters for gaps as small as 0.07 points. The MT-Bench comparison uses GPT-4o as judge, which has its own biases, and a 0.006 gap between Korean and Chinese cannot be called a difference. Finally, the weight was fixed at 0.5 and never varied, so an interaction between weight and density is unexplored.
Where this fits in the bigger picture
Model merging sits next to distillation and pruning as a way to reuse trained models without paying for another training run. The DARE step is closely related to sparsifying a delta, since it keeps a quarter of it here and works anyway, and that redundancy is the same property that makes compression of fine tuned weights possible. For practitioners the useful lesson is operational. A tuned merge can move one benchmark a lot in an afternoon on one GPU, and the cost shows up somewhere else, so evaluate the whole suite and not just the target.
Conclusion
The paper shows that a training free merge can change a model’s math performance by a large amount. At density 0.25 the merged model’s GSM8K score rose from 58.76 to 69.37 while the six benchmark average rose by 1.28 points, and the same direction of effect appeared on a larger SOLAR model and held across three merging recipes that all beat the unmerged baseline except TIES.
The conceptual point worth keeping is that the delta between a fine tuned model and its parent is a movable, sparsifiable object. Discarding three quarters of it at random and rescaling the remainder still transfers something useful, which tells you how redundant fine tuning updates are and why merging can work without any data.
The paper’s own tables temper its headline story. Chinese and Japanese merges also help GSM8K, the Korean model scores poorly on that benchmark by itself, TruthfulQA falls, hyperparameters were tuned on the test benchmarks, and the differences among DARE, Slerp, and Breadcrumbs are within what noise could produce. None of that makes the result uninteresting. It makes it a solid observation with an unexplained mechanism.
The obvious next experiments are cheap and would sharpen the picture a great deal. Merge with a control model that has no special language, vary the weight as well as the density, report several random seeds, and hold out benchmarks for selecting density. If the Korean advantage survives those, the linguistic hypothesis becomes much stronger.
For anyone experimenting with merges, the practical takeaway is to sweep density on your own benchmark set, watch the benchmarks that go down as well as the one that goes up, and treat any single language or domain explanation as a hypothesis until a matched control says otherwise.
Complete PyTorch implementation
The code below implements DARE merging on state dictionaries. It follows the general form used by merging toolkits, where each donor model’s delta is taken against a shared parent, sparsified with a Bernoulli mask, rescaled by the density, weighted, and added back. The paper’s Algorithm 1 is the special case with one donor. The smoke test uses tiny random models so it runs anywhere.
# dare_merge.py # DARE (Drop And REscale) merging on state dicts, following Algorithm 1 in # "Research on enhancing model performance by merging with Korean language # models", Engineering Applications of Artificial Intelligence 159 (2025) 111686. import torch import torch.nn as nn def bernoulli_drop_rescale(delta, density, generator=None): """Keep each entry with probability `density`, then divide survivors by `density` so the expected value of every entry is unchanged.""" if density >= 1.0: return delta mask = torch.bernoulli(torch.full_like(delta, density), generator=generator) return delta * mask / density @torch.no_grad() def dare_merge(parent_sd, model_sds, weights, densities, seed=0): """Merge several fine tuned models that share one parent. parent_sd state dict of the shared pretrained parent model_sds list of state dicts of fine tuned models weights list of scalars, one per model (the paper fixes 0.5) densities list of keep fractions, one per model (the paper uses 0.25) """ gen = torch.Generator().manual_seed(seed) merged = {} for name, parent_t in parent_sd.items(): if not parent_t.is_floating_point(): merged[name] = parent_t.clone() continue acc = parent_t.detach().float().cpu().clone() for sd, w, d in zip(model_sds, weights, densities): delta = sd[name].detach().float().cpu() - parent_t.detach().float().cpu() acc += w * bernoulli_drop_rescale(delta, d, gen) merged[name] = acc.to(parent_t.dtype) return merged def fraction_dropped(parent_sd, model_sd, merged_sd, weight): """Estimate the share of delta entries that were dropped.""" total, zeros = 0, 0 for name, p in parent_sd.items(): if not p.is_floating_point(): continue applied = (merged_sd[name].float() - p.float()) / weight zeros += (applied == 0).sum().item() total += applied.numel() return zeros / total if __name__ == "__main__": torch.manual_seed(0) # Tiny stand ins for a parent model and two fine tuned children. parent = nn.Sequential(nn.Linear(16, 32), nn.ReLU(), nn.Linear(32, 4)) child_a = nn.Sequential(nn.Linear(16, 32), nn.ReLU(), nn.Linear(32, 4)) child_b = nn.Sequential(nn.Linear(16, 32), nn.ReLU(), nn.Linear(32, 4)) with torch.no_grad(): for p, a, b in zip(parent.parameters(), child_a.parameters(), child_b.parameters()): a.copy_(p + 0.05 * torch.randn_like(p)) b.copy_(p + 0.05 * torch.randn_like(p)) psd, asd, bsd = parent.state_dict(), child_a.state_dict(), child_b.state_dict() # Density 0.25 and weight 0.5, the settings reported in the paper. merged_sd = dare_merge(psd, [asd], weights=[0.5], densities=[0.25], seed=1) print("Fraction of delta dropped", round(fraction_dropped(psd, asd, merged_sd, 0.5), 3)) # With density 1.0 the merge is plain weighted delta addition. full = dare_merge(psd, [asd, bsd], weights=[0.5, 0.5], densities=[1.0, 1.0]) expected = {k: psd[k] + 0.5 * (asd[k] - psd[k]) + 0.5 * (bsd[k] - psd[k]) for k in psd} ok = all(torch.allclose(full[k], expected[k], atol=1e-6) for k in psd) print("Density 1.0 matches weighted delta addition", ok) # Rescaling keeps the expected delta unchanged. delta = torch.randn(1000) gen = torch.Generator().manual_seed(2) est = torch.stack([bernoulli_drop_rescale(delta, 0.25, gen) for _ in range(4000)]).mean(dim=0) print("Mean abs error of expected delta", round((est - delta).abs().mean().item(), 3)) # The merged weights load into the same architecture and run. merged_model = nn.Sequential(nn.Linear(16, 32), nn.ReLU(), nn.Linear(32, 4)) merged_model.load_state_dict(merged_sd) print("Merged model output shape", merged_model(torch.randn(3, 16)).shape)
Frequently asked questions
What does DARE stand for and what does it do
It stands for Drop And REscale. It randomly discards a large share of the weight differences between a fine tuned model and its parent, scales the remainder up to compensate, and adds the result to a base model, with no further training.
How much did merging with the Korean model help
At a density of 0.25 the six benchmark average rose from 74.95 to 76.23, a 1.28 point gain, and GSM8K rose from 58.76 to 69.37. TruthfulQA fell from 76.00 to 71.95.
Does the paper prove Korean makes models reason better
No. Chinese and Japanese merges also raised GSM8K, and there is no control merge with a model lacking a special language, so the Korean specific explanation is a hypothesis rather than an isolated finding.
Is DARE better than Slerp or TIES merging
DARE scored highest at 76.23 against 76.16 for Slerp and 76.00 for Breadcrumbs, but those gaps are very small and come from single runs. TIES merging did not beat the unmerged base model.
Does merging need a GPU or more training data
The paper reports that merging itself does not need GPU resources and uses no additional training data. The evaluation of merged models ran on a single A100 80 GB GPU.
Where can I get the merged model
The authors state that code and models are available on Hugging Face under iRASC, and the link appears below.
Read the full study and get the released models.
Read the paper on ScienceDirect Open the Hugging Face release
Cho, T., Kim, R. and Choi, A. J. Research on enhancing model performance by merging with Korean language models. Engineering Applications of Artificial Intelligence 159, 111686 (2025). https://doi.org/10.1016/j.engappai.2025.111686. Supported by the Gachon University research fund of 2024, grant GCU-202500670001. Published under a Creative Commons Attribution license.
This analysis is based on the published paper and an independent evaluation of its claims. Percent changes not stated in the paper were computed from its reported tables.

Pingback: 5 Shocking Secrets of Skin Cancer Detection: How This SSD-KD AI Method Beats the Competition (And Why Others Fail) - aitrendblend.com