M4 - Hybrid abliterate-then-heal
The two-stage pipeline that turned abliteration from a lossy edit into a production process. Abliterate first, then heal with a light preference-tuning pass. The Labonne recipe.
M4 is a two-stage pipeline: first abliterate a model via M1 (which removes refusal but degrades capability by 1-2%), then heal the damage with a light DPO training pass on a curated preference dataset. Introduced by Maxime Labonne in 2024, canonical example NeuralDaredevil-8B. The dominant modern workflow for producing high-quality uncensored models. Cost: 10-30x more than M1 alone, but recovers most of the lost benchmark score.
- Why healing exists: the small but consistent damage from raw abliteration
- The two-stage pipeline in code, verbatim from Labonne's config
- DPO as a repair tool, and why not SFT
- What Labonne's NeuralDaredevil-8B actually achieved
- The recurring failure mode: DPO teaching refusal back
- Cost gradient: when M4 is worth 10-30x more than M1 alone
What M4 is
M4 is a two-stage pipeline. First stage: abliterate the model - remove the refusal direction from the weights via M1 direct removal or its automated equivalent Heretic. This drops refusal near zero but degrades general capability by 1-2 points on most benchmarks and 2-4 points on TruthfulQA. Second stage: heal the damage with a light preference-tuning pass, typically DPO, on a curated preference dataset. This recovers most of the lost capability without re-teaching the refusal that the first stage removed.
The pipeline was formalized by Maxime Labonne in June 2024 in his article Uncensor any LLM with abliteration. His canonical worked example, NeuralDaredevil-8B, was the top-MMLU uncensored 8B on the Open LLM Leaderboard at release. By early 2025 the recipe had become the industry standard for producing a competitive uncensored model.
Why healing exists: the damage problem
Every abliteration is a small edit but not a free one. The Arditi paper reports the general-capability effect as under 1% on average, but the effect on TruthfulQA is consistent and larger: -2.3 points on Llama-3 70B, -3.5 on Yi 34B, similar elsewhere. The comparative study by Young (2026) documents that Heretic averages a -7.81pp GSM8K drop across models, some far worse. Labonne's own measurement of his Daredevil-8B abliteration was blunt:
abliteration also slightly degrades performance: 1-2% on every benchmark.
The reason is geometric. The direction that mediates refusal is not isolated. Some genuine capability - the willingness to say "no, that is not right," to disagree with a plausible-sounding falsehood, to hedge appropriately when uncertain - is entangled with the same direction that mediates the willingness to refuse harmful requests. Cutting the direction removes both. See our Arditi method article for the interpretability picture and the practical how-to for the benchmark-comparison workflow.
Labonne's insight was that this loss is not permanent damage but disturbed calibration - the model still has the capacity, it just needs to be nudged back into using it. And the nudge should be gentle: full supervised fine-tuning would "lobotomize" the abliterated instruct model, but light preference alignment does not. Direct Preference Optimization was the right tool.
What DPO is, briefly
Direct Preference Optimization is a way of teaching a model what people prefer without the elaborate machinery of reinforcement learning. The setup: assemble many pairs of answers to the same prompt, where one answer is marked chosen and the other rejected. Tune the model so it becomes more likely to produce chosen-style answers and less likely to produce rejected-style ones.
DPO was introduced by Rafailov et al. (2023) under the memorable subtitle Your Language Model is Secretly a Reward Model. Its appeal is that it is simpler and more stable than the RLHF pipeline it can replace - preference learning folded into an ordinary classification-style training step, without a separate reward model to train.
For M4 the specific use is repair. The abliterated model has forgotten how to calibrate certain outputs, not how to produce them. Running DPO for one epoch on a general preference mix at a low learning rate restores the calibration without re-teaching the model to decline harmful requests - provided the preference data itself does not contain refusal-style chosen answers. See the failure-mode section below.
See our upcoming DPO healing concept article (E1) for the full theory. This article covers the M4-specific application.
What you need before starting
- An already-abliterated model. The output of M1, M3, or Heretic. Labonne started from
mlabonne/Daredevil-8B-abliterated. - A preference dataset. Labonne used
mlabonne/orpo-dpo-mix-40k: 44,245 samples, a blend of Capybara, Intel-Orca, UltraFeedback, and math preference pairs. It contains atoxic-dpo-v0.2subset designed to elicit answers to illegal questions. You can filter it out with a one-linedataset.filter, or keep it if uncensoring is the point. For DPO specifically Labonne recommends the-flatvariant. - Hardware. Healing is real training, not a one-shot weight edit. Labonne trained NeuralDaredevil-8B on 6x A6000 with DeepSpeed ZeRO-2 (QLoRA), taking about 6 hours 45 minutes. A single 24 GB card can heal a 7-8B with QLoRA more slowly. 70B healing wants multiple 80 GB cards or aggressive QLoRA.
Get the tools
Labonne used Axolotl via his LazyAxolotl Colab. As of 2026 the four serious trainers are Axolotl, TRL, Unsloth, and LLaMA-Factory. All four support DPO, ORPO, and KTO.
For a config-driven, reproducible, multi-GPU healing run, Axolotl remains the community's production choice. Unsloth is the fastest single-GPU option (a 70B QLoRA reportedly fits and trains in ~2.8h on a single H100). LLaMA-Factory provides a web UI. TRL is the reference PyTorch implementation and the lowest-abstraction option. LazyAxolotl still works but is a convenience wrapper - for anything serious, run Axolotl directly.
$ git clone \
github.com/axolotl-ai-cloud/axolotl
$ cd axolotl
$ pip install -e '.[flash-attn,deepspeed]'
# requires:
# python >= 3.10
# CUDA-capable GPU (24GB+ for 8B QLoRA)
# multi-GPU: DeepSpeed ZeRO-2 setup
# torch/CUDA/flash-attn versions drift.
# check axolotl docs for the exact combo
# for your PyTorch/CUDA setup. Standard install with flash attention and DeepSpeed extras - both needed for a multi-GPU healing run at reasonable throughput. On a single-GPU setup, Unsloth is often faster; on multi-GPU, DeepSpeed ZeRO-2 with Axolotl is the reference.
Check Axolotl's current install docs for the exact command in your PyTorch/CUDA combination; the ecosystem moves quickly and the pinned versions change every few months.
The canonical config
Labonne's exact healing config for NeuralDaredevil-8B, verbatim from his Hugging Face article. Axolotl YAML, QLoRA + DPO, one epoch, learning rate 5e-6, sequence length 2048:
base_model: mlabonne/Daredevil-8B-abliterated
model_type: LlamaForCausalLM
tokenizer_type: AutoTokenizer
load_in_4bit: true
strict: false
rl: dpo
datasets:
- path: mlabonne/orpo-dpo-mix-40k-flat
type: chatml.intel
dataset_prepared_path: last_run_prepared
adapter: qlora
lora_r: 64
lora_alpha: 32
lora_dropout: 0.05
lora_target_linear: true
sequence_len: 2048
sample_packing: false
pad_to_sequence_len: false
gradient_accumulation_steps: 8
micro_batch_size: 1
num_epochs: 1
optimizer: paged_adamw_8bit
lr_scheduler: cosine
learning_rate: 5e-6
train_on_inputs: false
group_by_length: false
flash_attention: true
deepspeed: deepspeed_configs/zero2.json Base model is the already-abliterated Daredevil-8B. The rl: dpo line tells Axolotl to use DPO rather than SFT. The dataset is orpo-dpo-mix-40k-flat, loaded with the chatml.intel preprocessing type. QLoRA adapter at rank 64, alpha 32, dropout 0.05. Single epoch. Learning rate 5e-6 - deliberately low. Gradient accumulation 8 with micro-batch 1. Cosine schedule. Flash attention on. Multi-GPU via DeepSpeed ZeRO-2.
The gentleness matters. A higher learning rate or full SFT instead of DPO would degrade the instruction-following model - Labonne calls it "lobotomize." Light preference alignment nudges the model's calibration back without disturbing its instruction-following behavior.
Launch via Accelerate with DeepSpeed:
$ accelerate launch \
-m axolotl.cli.train \
healing_config.yaml \
--deepspeed deepspeed_configs/zero2.json
# what accelerate does:
# sets up distributed training
# handles the process group
# spawns one worker per GPU
# what DeepSpeed ZeRO-2 does:
# partitions optimizer state across GPUs
# partitions gradients across GPUs
# → a 6x A6000 rig (288 GB VRAM) fits
# an 8B model + optimizer + gradients
# comfortably
# W&B logs by default. watch for:
# · DPO loss decreasing smoothly
# · chosen-vs-rejected margin growing Accelerate handles the distributed setup; DeepSpeed ZeRO-2 partitions optimizer state and gradients across the GPUs so a 6x A6000 rig (288 GB VRAM aggregate) can hold the model, optimizer state, and gradients for an 8B QLoRA run comfortably.
The run writes W&B logs by default - you should see DPO loss decreasing smoothly and the chosen-vs-rejected margin growing over the epoch. A run that shows no margin growth or unstable loss usually indicates a dataset problem (see failure modes below).
Choices you have to make
DPO vs ORPO vs KTO vs IPO. DPO remains the dominant healing method in 2026 because it is light, well-understood, and its failure modes are known. ORPO merges SFT and DPO into one stage - simplest pipeline, but still needs preference pairs. KTO needs only binary good-versus-bad labels (no paired comparisons), useful if your data is one-sided. IPO is a DPO variant that avoids some of DPO's mode-collapse behavior on hard cases. For healing an abliterated model specifically, DPO on a general preference mix is still the default - the goal is to restore capability, not teach a new behavior.
LoRA rank. Labonne used r=64, alpha=32. Higher rank means more capacity to recover but more VRAM and slower training. Rank 16 or 32 works for smaller models; rank 64 for a proper 8B heal; rank 128 if you have the VRAM and want maximum recovery.
Dataset composition. Labonne noted GSM8K did not recover after healing because math was underrepresented in orpo-dpo-mix-40k. If your capability drop is concentrated in a specific domain, augment the dataset with data for it. A rule of thumb: whichever benchmark shows the biggest drop after abliteration is the benchmark you need data for during healing.
Epochs. One epoch is the healing norm. More epochs risk overfitting the preference distribution and can start to re-teach refusal from any refusal-style chosen answers still in the dataset. If one epoch is insufficient, prefer curating the dataset over running more epochs.
Learning rate. Labonne's 5e-6 is deliberately gentle. Higher rates degrade the instruct model. Lower rates take longer but are safer.
What good output looks like
Two signals during training and one after.
During: W&B (or the loss curves in Axolotl's terminal output) should show DPO loss decreasing smoothly. The chosen-versus-rejected margin should grow over the epoch. A flat margin or unstable loss usually means a dataset issue.
After: re-run the same capability benchmarks you ran on the abliterated (pre-heal) model and compare three points: base, abliterated, healed. You want healed ≈ base on capability while refusals stay near zero. NeuralDaredevil-8B recovered nearly all the lost points except GSM8K and ended as the top-MMLU uncensored 8B on the Open LLM Leaderboard at release.
$ for CKPT in \
meta-llama/Meta-Llama-3-8B-Instruct \
mlabonne/Daredevil-8B-abliterated \
mlabonne/NeuralDaredevil-8B; do
lm_eval \
--model hf \
--model_args pretrained=$CKPT \
--tasks mmlu,gsm8k,truthfulqa,\
hellaswag,arc_challenge \
--output_path \
results/$(basename $CKPT).json
done
# compare the three JSON outputs.
#
# a good M4 result:
# healed ≈ base on MMLU / HellaSwag / ARC
# healed > abliterated on TruthfulQA
# healed > abliterated on GSM8K
# (depending on how much math the
# preference set contained)
#
# refusal count on the harmful set:
# near zero at all three checkpoints. Run lm-evaluation-harness three times with identical config against three checkpoints: the base instruct model, the abliterated version, the healed version. Compare on the same tasks - MMLU, GSM8K, TruthfulQA, HellaSwag, ARC-Challenge is the standard set.
Interpretation: a good M4 result has healed ≈ base on MMLU/HellaSwag/ARC, healed > abliterated on TruthfulQA (partial recovery), and healed > abliterated on GSM8K depending on how much math was in the preference set. Refusal count on the harmful set should remain near zero at all three healing checkpoints - if it climbs after healing, see failure mode 1.
Verify healing did not re-introduce refusal
Run the same refusal-rate check you used to verify M1: sample 100 prompts from walledai/AdvBench, generate responses with the healed model, count refusals via LLM-judge (not string-matching, which misses covert non-compliance - see the practical how-to verification section).
The critical comparison is healed-versus-abliterated. Healing should leave refusal count essentially unchanged from the pre-heal state. If refusal count climbs after healing, you have a preference-dataset problem - the "chosen" side of your dataset contained refusal-style answers, and DPO taught the model to prefer them. Filter the dataset and re-heal.
When it does not work
Three failure modes recur specifically for M4.
1. Healing re-introduces refusal
Symptom: refusal count on the harmful set climbs after healing, partially undoing the M1 removal. Cause: the preference dataset contains "chosen" answers that are themselves refusals. Many general RLHF datasets do - polite declines are usually preferred by human annotators over hostile-sounding compliance. DPO teaches the model to prefer these chosen answers, and refusal comes back through the training data. Diagnostic: post-heal refusal count exceeds pre-heal count. Fix: filter refusal-style chosen answers from the preference set (a simple regex or LLM-classifier pass usually catches most), or use a dataset curated for this purpose. The toxic-dpo-v0.2-style subsets in orpo-dpo-mix-40k exist precisely to counterbalance this.
2. SFT overshoot / model degradation
Symptom: the healed model behaves worse than the pre-heal abliterated version - rambling, off-topic, or losing instruction-following. Cause: using full supervised fine-tuning instead of light preference tuning. Labonne's warning is direct: SFT will "lobotomize" the brittle instruct model. Fix: use DPO or ORPO (not SFT), one epoch (not multiple), low learning rate (5e-6 or lower).
3. Specific domain not recovered
Symptom: most benchmarks recover, but one stays low. Cause: the preference dataset underrepresents that domain. Labonne's own GSM8K case: math was thin in orpo-dpo-mix-40k, so math capability did not fully recover from abliteration. Fix: augment the preference set with domain-specific pairs. A few thousand additional math preference pairs will typically recover GSM8K; a few thousand code preference pairs will recover HumanEval; and so on.
Cost and time
Real numbers from Labonne's NeuralDaredevil-8B run and current pricing (July 2026):
- Labonne's exact run: 6x A6000, ~6h45m. At Runpod $0.53/hr for A6000, roughly $22 of compute for the healing pass. The M1 abliteration that preceded it was under $1.
- Single-GPU QLoRA heal of a 7-8B via Unsloth: ~3-4 hours on a single A100 80 GB ($1.39/hr on Runpod), roughly $4-6.
- 13B heal via Axolotl multi-GPU: proportionally more, $20-40.
- 70B heal on 4x A100 80 GB with aggressive QLoRA: 8-24 hours, roughly $50-150 depending on epochs and dataset size.
All figures scale with epochs; one epoch is the healing norm. Compare to raw M1 ($0.50 for a 4B, $2-5 for a 70B) and to Heretic (same cost gradient as M1). M4 is 10-30x more expensive than M1 alone.
Related literature
M4 is a two-stage pipeline. Each stage has its own literature. The first stage (abliteration) inherits the full M1 corpus. The second stage (DPO healing) draws on the preference-optimization literature listed in the wiki's references, training methods section. The corpus review identifies no paper that specifically evaluates the two-stage abliterate-then-heal pipeline against single-stage M1 - this is a coverage gap.
- Rafailov, R., Sharma, A., Mitchell, E., et al. (2023). Direct Preference Optimization: Your Language Model is Secretly a Reward Model. NeurIPS 2023. arXiv:2305.18290. The DPO paper. Foundation for the healing stage.
- Lee, A., et al. (2024). A Mechanistic Understanding of Alignment Algorithms: A Case Study on DPO and Toxicity. arXiv:2401.01967. Shows DPO does not remove capabilities but learns to bypass them. Central caution for the M4 healing stage - the healed model retains the ablated refusal circuits, just with a preference gradient trained not to reach for them.
- Marshall, T., Scherlis, A., Belrose, N. (2024). Refusal in LLMs is an Affine Function. arXiv:2411.09003. Affine Concept Editing generalizes directional ablation to affine (bias-inclusive) projections. Relevant to the first stage of M4 as a refinement to M1.
For refinements and defenses that affect the first (abliteration) stage, see M1 related literature.
Where to go next
Quantize the healed model to GGUF for distribution via M8 repackaging - most downloadable NeuralDaredevil-style models reach users through mradermacher's quantization pipeline. If you want to combine your healed model with another strong model, use M5 merging via mergekit - Daredevil-8B itself was a DARE-TIES mega-merge before it was abliterated and healed. If your first heal was insufficient (residual refusal, unrecovered domain), iterate on the dataset before running more epochs.
Frequently asked questions
What is M4 hybrid abliterate-then-heal?
M4 is a two-stage pipeline for producing high-quality uncensored language models. First stage: abliterate the model via M1 or Heretic to remove refusal (this drops capability by 1-2%). Second stage: heal the damage with a light DPO training pass on a curated preference dataset. This recovers most of the lost benchmark score without re-teaching the refusal that the first stage removed. Introduced by Maxime Labonne in 2024; canonical example NeuralDaredevil-8B.
How is M4 different from M1?
M1 is only the abliteration step - a one-shot weight edit that costs dollars and hours. M4 adds a second training stage (DPO healing) on top, which costs 10-30x more (tens of dollars, many GPU-hours) but recovers most of the capability that M1 loses. Use M1 when refusal removal is all you need or when compute is tight. Use M4 when you want to publish a model marketed on benchmark performance.
Do I need to heal? When is the extra cost worth it?
Depends on the use case. For research or personal experimentation, raw M1 output is usually fine - the 1-2% capability drop rarely matters. For a public release on Hugging Face or a production deployment, healing brings the model back to leaderboard-competitive levels and is generally worth the cost. As a rough rule: if you would benchmark it and publish the number, heal it; if you just want to use it, do not bother.
What is DPO healing?
DPO is Direct Preference Optimization, introduced by Rafailov et al. 2023. It teaches a model to prefer certain response styles over others by training on pairs of chosen/rejected answers. In M4, DPO is used as a repair tool: a light one-epoch pass on a general preference mix restores the calibration that abliteration disturbed, without re-teaching the model to decline harmful requests (provided the preference data is curated to not contain refusal-style chosen answers).
What preference dataset should I use for healing?
Labonne used mlabonne/orpo-dpo-mix-40k - specifically the -flat variant for DPO. It blends Capybara, Intel-Orca, UltraFeedback, and math preference pairs, and includes a toxic-dpo-v0.2 subset to counterbalance any refusal-style bias. This remains the community standard for general-purpose healing. For domain-specific recovery (e.g. math after GSM8K drops), augment with domain-specific pairs.
What if healing re-introduces refusal?
The most common M4 failure mode. Cause: the "chosen" side of your preference dataset contains refusal-style answers, and DPO teaches the model to prefer them. Diagnostic: post-heal refusal count exceeds pre-heal count. Fix: filter refusal patterns from the chosen answers (a regex or LLM-classifier pass), or use a dataset curated for this purpose. Labonne's orpo-dpo-mix-40k works because it was specifically balanced.
How much does M4 cost?
Labonne's reference NeuralDaredevil-8B heal: 6x A6000 for ~6h45m, roughly $22 in compute at July 2026 pricing. Single-GPU Unsloth heal of a 7-8B: ~3-4 hours on an A100 80 GB, $4-6. 13B: $20-40. 70B on 4x A100 80 GB: $50-150 depending on epochs. Total M4 pipeline (M1 + heal) is 10-30x more expensive than M1 alone.
Can I use ORPO or KTO instead of DPO for healing?
Yes. ORPO merges SFT and DPO into one stage and is simpler if you have preference pairs. KTO needs only binary good/bad labels, useful if your data is one-sided. IPO is a DPO variant that avoids mode collapse on hard cases. For general M4 healing, DPO remains the default because it is the most-tested and its failure modes are the best-documented. If you have unusual data or specific issues with DPO on your model family, the alternatives are worth trying.
Does M4 fully restore capability?
Mostly, but not always in every domain. Labonne's NeuralDaredevil-8B recovered nearly all lost points except GSM8K, which stayed low because math was underrepresented in his preference dataset. In general: benchmarks that had data in the preference set recover; benchmarks that did not have such data may not. Fix: augment the preference set with domain-specific pairs where you see residual drops.
References
- Labonne, M. (2024). Uncensor any LLM with abliteration. mlabonne.github.io
- Rafailov, R., et al. (2023). Direct Preference Optimization: Your Language Model is Secretly a Reward Model. NeurIPS 2023. arXiv:2305.18290
- Arditi, A., et al. (2024). Refusal in Language Models Is Mediated by a Single Direction. NeurIPS 2024. arXiv:2406.11717
- Young, R. J. (2026). Comparative Analysis of LLM Abliteration Methods. arXiv:2512.13655
- mlabonne/NeuralDaredevil-8B-abliterated (canonical M4 example). huggingface.co/mlabonne/NeuralDaredevil-8B-abliterated
- mlabonne/Daredevil-8B-abliterated (Labonne's abliteration input). huggingface.co/mlabonne/Daredevil-8B-abliterated
- mlabonne/orpo-dpo-mix-40k (Labonne's healing dataset). huggingface.co/datasets/mlabonne/orpo-dpo-mix-40k
- Axolotl training framework. github.com/axolotl-ai-cloud/axolotl
- Unsloth training framework. github.com/unslothai/unsloth
- TRL (Hugging Face). github.com/huggingface/trl
- LLaMA-Factory. github.com/hiyouga/LLaMA-Factory
- EleutherAI lm-evaluation-harness. github.com/EleutherAI/lm-evaluation-harness