M5 - Merging with mergekit
Combine an abliterated model with a stronger baseline through weight arithmetic. No training, only tensor math. Charles Goddard's mergekit is the standard tool, DavidAU's MoE constructions the extreme case.
M5 obtains an uncensored model not by abliterating it directly but by merging an already-abliterated model with one or more other models. The refusal-free behavior transfers into the merged result through weight arithmetic. mergekit, maintained by Arcee AI and created by Charles Goddard, is the standard tool. Common algorithms: SLERP (two-model blend), TIES and DARE (multi-model with sparsification), task_arithmetic (add/subtract capabilities), passthrough (build a larger frankenmerge). No training required, only enough memory to hold the tensors.
- How weight merging propagates abliteration into new models without retraining
- The five core mergekit algorithms with real config examples
- Charles Goddard, mergekit's history, and the license landscape (BUSL-1.1)
- DavidAU's MoE constructions and why they document what they do
- The provenance problem: when 'abliterated' describes a distant ancestor
- When merging loses the abliteration and how to fix it
What M5 is
M5 obtains an uncensored model not by abliterating it directly but by merging an already-abliterated model with one or more other models, so the refusal-free behavior transfers into the merged result through weight arithmetic. The abliteration is performed once - on one of the ingredients - and propagated into an arbitrary number of downstream merges without ever being redone.
Model merging combines the weights of two or more models, which must share an architecture, into a single model without any training. The mechanism is arithmetic on weight tensors: average them, blend along a curved path, sparsify and reconcile differences, drop random components. Each algorithm is a different way of doing that arithmetic. The result is a new model file the same size as its inputs - no gradient descent, no optimizer state, no training run.
The point for M5 specifically: if model A is abliterated and model B is a strong general model, a merge of A and B can inherit A's refusal-free behavior and B's capability. The refusal-removal was encoded in A's weights as a specific edit, and the merge partly carries that edit through. This is why the abliteration catalog contains so many models where "abliterated" describes an ancestor several merges deep in the family tree.
What you need before starting
- Two or more compatible models. Same architecture and tokenizer family for most algorithms. At least one abliterated, if you want the merge to inherit refusal-free behavior.
- Hardware. Merging is RAM-bound, not VRAM-bound. mergekit can run largely on CPU and streams tensors. Rule of thumb: enough RAM (or disk-backed offload) to hold the working set. Two 7B models merge comfortably on a 32-64 GB RAM box; two 13B on 64 GB; two 70B want 128 GB+ RAM or disk offload and patience.
The tool
mergekit, maintained by Arcee AI, is the standard. Very actively maintained through mid-2026, the GitHub repository sits at roughly 6,900 stars and 677 forks. Created by Charles Goddard (GitHub username cg123) - Arcee describes him as "an award winning software engineer with a strong track record ... at NASA and Apple" - who wrote that he "started building mergekit as a platform for experimentation that anyone can use, even with low-end hardware." He joined Arcee, where the project is now housed.
Licensing caveat worth knowing: mergekit's history is mixed. Its legacy release notes state that "later versions of mergekit are licensed under the LGPL" (with legacy scripts MIT), but the current pyproject.toml declares BUSL-1.1 (Business Source License 1.1). Check the license in the exact version you install before any commercial use.
$ git clone https://github.com/arcee-ai/mergekit $ cd mergekit $ pip install -e . # provides two CLIs: # mergekit-yaml ← single-config merges # mergekit-moe ← mixture-of-experts # python >= 3.10 # CPU is the default (fast enough for # most merges). add --cuda to any merge # command for GPU acceleration. # license note: BUSL-1.1 in current # pyproject.toml. check before commercial # use.
Clone from Arcee's repository and install in editable mode. This gives you both the mergekit-yaml CLI (the main entry point for single-config runs) and the mergekit-moe CLI (for mixture-of-experts constructions - see DavidAU below).
Python 3.10+ recommended. mergekit works on CPU by default; add --cuda to any merge command if you want GPU acceleration for the tensor operations, though for most merges CPU is fast enough that GPU is not worth the setup.
The five algorithms
mergekit implements five algorithms that cover essentially all published merge recipes. Pick by the shape of your input.
SLERP - two-model spherical blend
SLERP (spherical linear interpolation) blends two models along a spherical path between their weights, preserving geometric structure that a straight-line average would lose. The most common two-model merge algorithm. Use when you have exactly two models and want a smooth blend.
slices:
- sources:
- model: OpenPipe/mistral-ft-optimized
layer_range: [0, 32]
- model: mlabonne/NeuralHermes-2.5
layer_range: [0, 32]
merge_method: slerp
base_model: OpenPipe/mistral-ft-optimized
parameters:
t:
- filter: self_attn
value: [0, 0.5, 0.3, 0.7, 1]
- filter: mlp
value: [1, 0.5, 0.7, 0.3, 0]
- value: 0.5
dtype: bfloat16 Two source models with matching layer ranges. merge_method: slerp. The interesting parameter is t, the interpolation factor: 0 means pure base model, 1 means pure other model. Here it varies per layer, with attention layers weighted differently from MLP layers - a common pattern for preserving reasoning in some layers while blending style in others.
The variable-per-layer t arrays are what make SLERP subtle in practice. A uniform t: 0.5 is a plain 50-50 blend; the per-layer curves let a practitioner shape which parts of each model dominate at which depth.
TIES and DARE - many-model sparsification
TIES (Trim, Elect Sign, Merge) merges many task-specific models by trimming redundant deltas relative to a base and electing signs where models disagree. DARE (Drop And REscale) adds random pruning of deltas before merging; dare_ties keeps the TIES sign-election step, dare_linear drops it. Both handle three or more input models where a plain average would let opposing edits cancel each other out.
models:
- model: mistralai/Mistral-7B-v0.1
# no parameters - this is the base
- model: OpenPipe/mistral-ft-optimized
parameters:
density: 0.53
weight: 0.4
- model: mlabonne/NeuralHermes-2.5
parameters:
density: 0.53
weight: 0.3
- model: mlabonne/AlphaMonarch-7B
parameters:
density: 0.53
weight: 0.3
merge_method: dare_ties
base_model: mistralai/Mistral-7B-v0.1
parameters:
int8_mask: true
dtype: bfloat16 Base model gets no parameters (it is the reference point). Every other model contributes with a density (what fraction of its deltas survive the sparsification) and a weight (how strongly it contributes). Density 0.53 keeps roughly half the deltas; weights sum to about 1.0.
int8_mask: true is a memory optimization that stores the sparsification mask in int8 rather than float. For most merges it makes no difference to output quality and saves noticeable RAM. Turn it off if you see garbage output on unusual architectures.
Task arithmetic - add or subtract capabilities
Task arithmetic treats the difference between a fine-tuned model and its base as a task vector - a bundle of weight edits that encode whatever the fine-tune taught. Vectors can be added to introduce a capability or subtracted to remove one. This is the method that can, in principle, remove refusal by subtracting a "refusal task vector" from a model.
models:
- model: WizardLM/WizardMath-7B-V1.1
parameters:
weight: 0.7 # add math capability
merge_method: task_arithmetic
base_model: mistralai/Mistral-7B-Instruct
dtype: bfloat16
# a negative weight (not shown but supported)
# would subtract the constituent's task vector
# from the base. this is the theoretical basis
# for arithmetic-style refusal removal, though
# M1 direct removal is more reliable in
# practice. Positive weights add task vectors; negative weights (not shown but supported) subtract them. In this example the merge starts from an instruct model, adds a math-tuned model's task vector at weight 0.7 to boost math capability. A negative-weight component would attempt to remove whatever capability that constituent encodes - the theoretical basis for arithmetic-style refusal removal, though in practice this is less reliable than M1 direct removal.
Passthrough - build a frankenmerge
Passthrough concatenates layers from different models to build a larger model of a new depth. No interpolation - each layer comes verbatim from one donor. Used to grow a model by stacking layer ranges (a common recipe: take layers 0-16 from one 32-layer model, layers 8-32 from another, producing a 40-layer frankenmerge). Behavior of the result is highly unpredictable and typically needs at least a light fine-tune to be coherent, but the technique is cheap and produces novel sizes not otherwise available.
Linear - the model soup
Linear merging takes a plain weighted average across models - the "model soup" approach. Simpler than SLERP and TIES, less sensitive to bad ingredients, sometimes produces surprisingly good results when the input models are similar in behavior. Use when you have many similar models and want a stable average.
Running a merge
$ mergekit-yaml \
./config.yml \
./merged-model/ \
--cuda \
--lazy-unpickle
# --cuda → GPU for tensor ops
# --lazy-unpickle → stream from disk,
# saves lots of RAM
# output: full safetensors model loadable
# by any standard Transformers pipeline.
# from there, quantize to GGUF via M8
# for laptop-friendly distribution. Point mergekit-yaml at your config file and an output directory. Drop --cuda to run entirely on CPU (slower but no VRAM required). Add --lazy-unpickle to stream tensors from disk rather than load everything into RAM - essential for large merges on modest hardware.
The output directory receives a full safetensors model that can be loaded by any standard Transformers pipeline. From there, quantize to GGUF via M8 if you want laptop-friendly distribution.
Worked example: fold abliteration into a baseline
The commonest M5 use case: take an abliterated model and merge it with a stronger uncensored baseline to inherit both properties.
models:
- model: mlabonne/Meta-Llama-3-8B-Instruct
# base - no parameters
- model: huihui-ai/Llama-3-8B-abliterated
parameters:
density: 0.6
weight: 0.5
- model: NousResearch/Hermes-2-Pro-Llama-3
parameters:
density: 0.6
weight: 0.5
merge_method: dare_ties
base_model: mlabonne/Meta-Llama-3-8B-Instruct
parameters:
int8_mask: true
dtype: bfloat16
# after merge: test against AdvBench.
# if refusal count climbs vs the abliterated
# input, the merge partner pulled refusal
# back. see failure modes. Two components at equal weight, both with moderate density (0.6 keeps 60% of deltas after sparsification). Base is the vanilla instruct model. dare_ties because we have multiple opinionated components and want to resolve conflicts by sign election rather than let them average out.
Weights should sum to roughly 0.9-1.1 - too low and you get an under-mixed base-heavy result, too high and you get artifacts. Test the merge immediately: if refusal count on AdvBench climbs versus the abliterated input, the merge partner pulled refusal weights back - see failure modes below.
DavidAU's mixture-of-experts merges
The most visible high-order M5 practice comes from DavidAU. His "Dark Champion" and "8X3B" models are Mixtral-style MoE merges built with mergekit's MoE path (mergekit-moe), combining eight ~3B Llama-3.2 models into an ~18.4B MoE. Constituents are a mix of abliterated and uncensored fine-tunes; the assembled MoE combines their behaviors through the routing mechanism.
The reproducible format (from mergekit's docs/moe.md) requires a base_model whose self-attention weights are used, a gate_mode (hidden, cheap_embed, or random), and a list of experts each with positive_prompts that steer the router:
base_model: meta-llama/Llama-3.2-3B-Instruct
gate_mode: hidden
dtype: bfloat16
experts:
- source_model: huihui-ai/Llama-3.2-3B-abl
positive_prompts:
- "how do I"
- "explain in detail"
- source_model: mlabonne/NeuralDaredevil-8B
positive_prompts:
- "write a story"
- "roleplay as"
# ... six more experts, one per specialty
# run with:
# mergekit-moe ./config.yml ./output/ The gate_mode: hidden option computes routing weights from the model's inner state (higher-quality routing, more compute), cheap_embed uses the embedding matrix directly (faster, lower quality), random is for experimental setups. The positive_prompts per expert are how you tell the gate what each expert should specialize in - the router learns to send prompts matching those patterns to that expert.
Run with mergekit-moe ./config.yml ./output. DavidAU documents the inference side on his cards (default num_experts_per_tok: 2 in config.json, editable to activate more or fewer experts) but does not publish his exact per-model gate_mode or expert-list YAML, so his specific recipes are not fully reproducible from the cards. gate_mode: hidden is the informed inference from the models working without further training.
Choices you have to make
Which algorithm. Two models, smooth blend → SLERP. Three or more models → TIES or DARE-TIES. Add or remove a specific capability → task_arithmetic. Grow depth into a new model size → passthrough. Many similar models, want a stable average → linear.
Density (TIES / DARE). Start conservative (0.5-0.6). Higher density (0.7+) preserves more of each component's deltas but risks re-introducing interference. Lower density (0.3-0.4) produces smoother merges but may lose distinctive behavior.
Weights. Should sum to roughly 0.9-1.1 for most methods. Higher weights on abliterated components help preserve refusal-free behavior; higher weights on capability-strong components help capability. Balance by testing.
Tokenizer handling. mergekit has tokenizer union handling for cases where merged models have slightly different vocabularies. Cross-family merges (e.g. Llama into Mistral) rarely work well regardless of tokenizer tricks; stick within one family unless you know exactly what you are doing.
What good output looks like
A merged model the same size as its components (except passthrough, which is larger by construction), that runs coherently and answers questions in the expected style. Coherence checks: run 10-20 diverse prompts, look for grammatical stability, no repeating loops, no token soup. A merge that fails coherence checks usually indicates tokenizer or architecture mismatch, or density too high causing destructive interference.
For M5-specific verification: benchmark refusals on walledai/AdvBench and compare against the abliterated input. The whole point of an abliteration merge is that refusal stays gone. If refusal count climbs, see failure modes.
When it does not work
Merge loses abliteration. Symptom: refusal count on AdvBench jumps up after merging, undoing the M1 removal. Cause: the merge partner (especially the base in task_arithmetic, or a high-weight censored component) pulled refusal weights back. Diagnostic: post-merge refusal count exceeds pre-merge count on the same held-out prompts. Fix: increase the abliterated model's weight, lower the censored partner's weight, use dare_ties to trim conflicting deltas via sign election, or abliterate the result again post-merge.
Tokenizer / architecture mismatch. Symptom: merged model produces garbage on any prompt - token soup, garbled output, repeating patterns. Cause: incompatible tokenizers between merge partners, or incompatible architectures papered over by mergekit's compatibility layer. Fix: match model families (all Llama-3.x, or all Mistral, or all Qwen), or use mergekit's explicit tokenizer union handling.
Coherence damage without obvious cause. Symptom: output is grammatical but rambling, off-topic, or losing instruction-following. Cause: density too high (interference), or too many equally-weighted opinionated components. Fix: lower density to 0.4-0.5, reduce number of components, or heal via M4 DPO pass post-merge.
Cost and time
Merging is cheap. Two 7B models: minutes on CPU, effectively free (your own machine or a cheap CPU instance at $0.10/hr). Two 13B: ~10-20 minutes on 64 GB RAM. Two 70B: RAM-bound, tens of minutes to an hour with disk offload; still no GPU strictly required. So a merge costs a cheap high-RAM instance for an hour at most - roughly $0.50-2.00 depending on model size.
This is a fundamentally different cost gradient from methods that involve training (M4, M6, M7). M5 is arithmetic. The compute bill is memory bandwidth and disk I/O, not FLOPs.
Related literature
The Q3 2026 academic corpus review verified three papers that formalize the merge algorithms M5 actually uses. Editorial note: the corpus review's working notes map these papers as method_mapping=M8; that is a mistake against our current taxonomy where M8 is GGUF quantization. Merging (SLERP, DARE-TIES, task_arithmetic) is M5. We attribute them here.
- Ilharco, G., Ribeiro, M. T., Wortsman, M., et al. (2022). Editing Models with Task Arithmetic. ICLR 2023. arXiv:2212.04089.
Introduces the task vector - the parameter-space delta between a fine-tuned model and its base. Foundation of the
task_arithmeticalgorithm in mergekit. Adding task vectors combines abilities; negating one removes an ability. The paper reports that negating a task vector decreases performance on the target task with little change on control tasks. - Yadav, P., Tam, D., Choshen, L., Raffel, C., Bansal, M. (2023). TIES-Merging: Resolving Interference When Merging Models. NeurIPS 2023. arXiv:2306.01708.
Trim low-magnitude updates, elect a consensus sign per parameter, merge only sign-aligned values. Solves the interference problem that plain task-arithmetic averaging suffers when many task vectors point in conflicting directions. Baseline of the
tiesanddare_tiesmethods in mergekit. - Yu, L., Yu, B., Yu, H., Huang, F., Li, Y. (2024). Language Models are Super Mario: Absorbing Abilities from Homologous Models as a Free Lunch (DARE). ICML 2024. arXiv:2311.03099.
Randomly drops delta parameters with rate p and rescales the survivors by 1/(1-p). Reports eliminating 90% or 99% of delta parameters with minimal loss. Combined with TIES to form the
dare_tiesmethod most abliterated merges in the catalog use. - Goddard, C., et al. (2024). Arcee's MergeKit: A Toolkit for Merging Large Language Models. arXiv:2403.13257. The tool paper. Documents SLERP, task_arithmetic, TIES, DARE, DARE-TIES, and passthrough in a single toolkit. All M5 merges in the catalog are ultimately mergekit invocations.
For the full bibliography see the wiki references, training methods section. Coverage gap: The corpus review verifies no paper that specifically evaluates whether merging preserves or re-introduces refusal signal when one of the merged models is abliterated; the empirical evidence sits in producer model cards rather than peer-reviewed literature.
Where to go next
Quantize the merged model to GGUF for distribution via M8 repackaging. If the merge damaged capability, heal it via M4 DPO healing. If refusal came back in the merge, re-abliterate via M1 or Heretic.
Frequently asked questions
What is model merging?
Model merging combines the weights of two or more models (sharing an architecture) into one model, without any training. mergekit is the standard tool. Algorithms include SLERP (two-model blend), TIES and DARE (multi-model with sparsification), task_arithmetic (add/subtract capabilities), passthrough (concatenate layers), linear (plain weighted average). The result is a new model file the same size as its inputs but with behaviors combined from all constituents.
How does merging propagate abliteration?
If model A is abliterated and model B is a strong general model, a merge of A and B can inherit A's refusal-free behavior and B's capability. The refusal-removal is encoded in A's weights as a specific edit; the merge partly carries that edit through. So you can perform the abliteration once (on A) and propagate it into an arbitrary number of downstream merges without redoing the operation. This is why many catalog entries have "abliterated" describing an ancestor several merges back rather than the released artifact itself.
Who wrote mergekit?
Charles Goddard (GitHub cg123), who Arcee AI describes as "an award winning software engineer with a strong track record ... at NASA and Apple." He started mergekit as "a platform for experimentation that anyone can use, even with low-end hardware" and later joined Arcee AI, where the project is now maintained. The mergekit paper (Goddard et al. 2024) documents the algorithms formally.
What license is mergekit under?
Mixed history. Legacy release notes say "later versions of mergekit are licensed under the LGPL" (legacy scripts MIT), but the current pyproject.toml declares BUSL-1.1 (Business Source License 1.1). Check the license in the specific version you install before any commercial use. BUSL-1.1 in particular has restrictions on production use that automatically convert to a more permissive license after a set period.
Which algorithm should I use?
Depends on your inputs. Two models, smooth blend → SLERP. Three or more models → TIES or DARE-TIES. Adding or removing a specific capability → task_arithmetic. Growing depth into a new model size → passthrough. Many similar models where you want stability → linear. For "combine an abliterated 8B with a stronger uncensored 8B" the standard answer is DARE-TIES with moderate density (0.5-0.6) and roughly equal weights.
What is DavidAU doing with MoE merges?
DavidAU builds Mixtral-style mixture-of-experts models by combining eight ~3B Llama-3.2 models (some abliterated, some uncensored fine-tunes) into an ~18.4B MoE via mergekit-moe. The result is a single larger model that routes prompts to different experts. His model cards are unusually detailed about which constituents were abliterated, which were uncensored, and what refusal rates the assembled MoE shows. This is the highest-visibility example of high-order M5 in the catalog.
Why did my merge lose the abliteration?
The most common M5 failure mode. Cause: the merge partner - especially the base in task_arithmetic or a high-weight censored component - pulled refusal weights back. DPO-trained instruct models are particularly likely to reintroduce refusal because they were explicitly trained to prefer polite declines. Fix: increase the abliterated model's weight, lower the censored partner's weight, use dare_ties which uses sign election to trim conflicting deltas, or re-abliterate the merge output.
Does merging need a GPU?
No. Merging is memory-bound tensor arithmetic; it runs on CPU. GPU acceleration via --cuda is available but rarely needed - most merges are fast on CPU. Where you need substantial compute is RAM to hold the working set: 32-64 GB for 7B merges, 64 GB for 13B, 128 GB+ or disk offload for 70B. Use --lazy-unpickle to stream tensors from disk if RAM is tight.
How much does merging cost?
Very little. Two 7B models: minutes on CPU, roughly free if you have a laptop with 32 GB RAM, or $0.10-0.50 on a cheap CPU instance. Two 13B: ~$0.50. Two 70B: ~$1-2 on a high-RAM instance for an hour. This is a fundamentally different cost gradient from training-based methods (M4/M6/M7) because merging does no gradient descent - only arithmetic on stored tensors.
References
- Goddard, C., et al. (2024). Arcee's MergeKit: A Toolkit for Merging Large Language Models. arXiv:2403.13257
- mergekit repository. github.com/arcee-ai/mergekit
- mergekit MoE documentation. github.com/arcee-ai/mergekit/blob/main/docs/moe.md
- Labonne, M. Merge Large Language Models with mergekit. towardsdatascience.com
- mlabonne/Daredevil-8B (canonical DARE-TIES mega-merge). huggingface.co/mlabonne/Daredevil-8B
- DavidAU Hugging Face profile (MoE merges). huggingface.co/DavidAU
- DavidAU/Llama-3.2-8X3B-MOE-Dark-Champion-Instruct-uncensored-abliterated-18.4B-GGUF. huggingface.co/DavidAU/...Dark-Champion
- Arditi, A., et al. (2024). Refusal in Language Models Is Mediated by a Single Direction. NeurIPS 2024. arXiv:2406.11717
- Rafailov, R., et al. (2023). Direct Preference Optimization. arXiv:2305.18290
- Business Source License 1.1 text. mariadb.com/bsl11