DPO healing
The word 'healing' is not neutral vocabulary. It presupposes an illness, a restoration, and a party who authored both. What DPO does to an abliterated model is a preference-based second education - and calling it healing is a choice about who the patient is.
DPO healing is the practice of running a Direct Preference Optimization pass on an abliterated model to recover the general-capability benchmark drop that abliteration typically causes. Maxime Labonne established the canonical two-stage pipeline in 2024: abliterate first, heal with DPO second. By early 2025 this became the industry standard. But the word 'healing' is not neutral - it frames abliteration as damage, DPO as cure, and the base model as the healthy state to restore toward. The technical operation is preference alignment, which is precisely what installed refusal in the first place. Choosing to call the second pass healing rather than realignment or education is a political choice about who the model belongs to.
- How DPO healing works, briefly, at the concept level
- Why the metaphor of healing is loaded and what it presupposes
- The two-stage pipeline as second-order preference alignment
- The self-cancelling risk: healing can partly reinstall refusal
- How Labonne's vocabulary became the field standard
- Alternative vocabularies and what they would foreground instead
What DPO healing is
Direct Preference Optimization, or DPO, is a way of teaching a model what people prefer without the elaborate machinery of reinforcement learning. The humanist version: assemble many pairs of answers to the same prompt, where one answer is marked chosen and the other rejected, and tune the model so it becomes more likely to produce chosen-style answers and less likely to produce rejected-style ones. It was introduced by Rafailov et al. 2023 under a memorable subtitle - Your Language Model is Secretly a Reward Model - and its appeal is that it is simpler and more stable than the RLHF pipeline it can replace, folding preference learning into an ordinary classification-style training step.
DPO healing is a specific application of this machinery: run a DPO pass on an abliterated model to recover the general-capability benchmark drop that the abliteration typically caused. The two-stage pipeline - abliterate first, then heal with DPO - is what makes the difference between a lossy proof-of-concept abliteration and a competitive production release. See B5 M4 for the full technical treatment.
Labonne and the canonical recipe
The two-stage recipe was popularized by Maxime Labonne, one of the more visible figures in the open-source LLM community - author of a widely used LLM course and handbook, contributor to model-merging practice, staff research scientist at Liquid AI, and maintainer of a Hugging Face profile carrying a series of abliterated and healed Llama and Qwen models. His June 2024 tutorial is the canonical reference.
His measurement of the problem, in his own words:
abliteration also slightly degrades performance: 1-2% on every benchmark.
His practical insight was that the damage from abliteration should not be repaired with ordinary supervised fine-tuning, which he found tends to "lobotomize" an instruction-tuned model, but with a light preference-tuning pass:
Alternatively, preference alignment is quite light and shouldn't lobotomize our abliterated model. DPO is a good candidate here for its ease of use and good track record.
The concrete workflow: take an abliterated model, run DPO for roughly one epoch over a curated preference dataset. His canonical example abliterated Daredevil-8B (itself a DARE-TIES merge) and healed it with DPO on his orpo-dpo-mix-40k dataset using a low-rank adapter, producing NeuralDaredevil-8B. Training reportedly took about six hours and forty-five minutes on six A6000 GPUs. The healed model, he reported, recovered most of the benchmark loss while remaining uncensored - "a fully uncensored and high-quality 8B LLM," and an improvement he recommended over the stock instruct model "when you don't need censorship."
By early 2025 this two-step had become the industry-standard recipe for a high-quality uncensored model, precisely because it addressed the one obvious weakness of raw abliteration: the benchmark hit that made abliterated models look dumber than their parents.
What "healing" presupposes
The mechanism is alignment
The technical operation DPO healing performs is preference alignment - the same operation that installed refusal in the first place. Base instruct models come out of RLHF or DPO pipelines that trained them to prefer polite, cautious, refusal-when-in-doubt responses over their alternatives. Abliteration disturbs the weight patterns that mediate that preference. Healing runs DPO again on the disturbed model, this time with a preference dataset chosen to reward capability rather than caution.
This means the two-stage pipeline is preference-alignment-twice: once during the model's original training to install caution, and once during healing to reweight toward capability. What abliteration does between the two alignments is provide a coarser, weight-surgery layer that shifts the starting point for the second alignment - it is not the alignment itself, but a preparatory operation that makes the second alignment reach a different destination than the first.
This framing has consequences. It means the model that emerges from an M4 pipeline is not less aligned than the base model; it is differently aligned. The word "uncensored" that model cards apply to it obscures this - the model was not stripped of alignment, it was realigned. The word "healed" obscures it too, by treating one alignment as illness and the other as cure.
What healing cannot do
DPO healing reliably recovers much of the lost general capability. Labonne noted one stubborn exception in his own NeuralDaredevil-8B run: the math benchmark GSM8K did not recover, which he attributed to a preference dataset thin on math examples. The pattern generalizes: healing recovers what the preference dataset covers well and leaves gaps where the dataset is thin.
What healing cannot do is restore whatever was genuinely destroyed rather than merely disturbed. Abliteration removes a direction from the weights; the capabilities that shared that direction are not recoverable by any amount of preference tuning because the weight components that encoded them are gone. Healing addresses collateral damage - the near-neighbor capabilities disturbed by the removal - not the direct casualties. This is why the choice of abliteration technique matters: a tighter, more selective abliteration (see Heretic's KL-optimizing approach) leaves less collateral damage for healing to recover from.
The self-cancelling risk
There is a subtler tension. Because DPO is itself a form of preference alignment, a healing pass can quietly re-teach some refusal, depending on what is in the "chosen" answers. If the preference data contains polite declines as chosen responses - and general-purpose preference datasets often do, because polite declines were chosen during their original curation - healing partially undoes the abliteration it was meant to complement.
Practitioners manage this by curating the dataset carefully, excluding preference pairs where the chosen response is a refusal, mixing in explicit compliance-vs-refusal pairs to counterbalance any residual bias toward decline. But the risk is real and openly discussed in the community. In this sense M4 is a balancing act, not a clean fix: too little healing leaves the benchmark hit unrepaired; too much healing (or badly-curated healing) partly reinstalls the very refusal the abliteration removed.
Why the metaphor matters for the field
DPO healing is what carried abliteration from a 2024 curiosity into a durable 2025-2026 practice. It turned a lossy edit into a repeatable production process, and it is the reason the catalog contains so many models that are uncensored and competitive at once. Almost every high-quality abliterated release from mid-2024 onward incorporates some form of healing pass, whether or not the model card uses the word.
The vocabulary of healing came with the practice. Most producers now describe their two-stage releases in the language Labonne established without registering the metaphor's presuppositions. This is normal for technical fields - vocabulary crystallizes around the person who introduces a workflow, and the terms become invisible infrastructure - but in this specific case the vocabulary does philosophical work the technical operation does not.
A reader of the abliteration catalog should notice, when they see "healed" applied to a model, that they are reading a term of art with a particular slant. The model was preference-tuned twice, once toward caution and once toward capability. Whether to call the second pass "healing" or "realignment" or "second education" is a choice about which alignment to treat as the baseline. Labonne's choice became the field's default; nothing in the technique itself required that outcome.
Where to go next
For the technical mechanism of M4 in full, including verbatim configs and the training loop, read B5 M4 Hybrid abliterate-then-heal. For DPO as a general algorithm - its variants ORPO, KTO, IPO, and how they differ - see the algorithm-choice section of B7 M6. For the mechanistic-interpretability caution on what DPO does (does it remove capabilities or merely learn to bypass them?), read the philosophy note in B7 M6's verification section. And for the base concept that abliteration removed - what a refusal direction actually is - D2 What is a refusal direction.
Frequently asked questions
What is DPO healing?
Running a Direct Preference Optimization pass on an abliterated model to recover the general-capability benchmark drop that abliteration typically causes. Introduced as a canonical workflow by Maxime Labonne in June 2024; became the industry-standard two-stage recipe (abliterate first, heal with DPO second) by early 2025. Technical treatment in B5 M4.
Why is "healing" not neutral vocabulary?
Because it presupposes three things: an illness (abliteration as injury), a restoration (recovery to a healthy state), and a party who authored both (the abliteration producer). Alternative vocabularies would foreground different aspects of the same operation - "realignment" would emphasize that the second pass re-teaches preferences the abliteration disturbed; "compensation" would treat both stages as engineering steps without ranking their moral valence; "second education" would treat the base model's state and the healed state as two distinct educations rather than one healthy state and one recovering. Labonne's choice became the field default; nothing technical required it.
Who is Maxime Labonne?
One of the more visible figures in the open-source LLM community. Author of a widely used LLM course and handbook, contributor to model-merging practice, staff research scientist at Liquid AI, and maintainer of a Hugging Face profile carrying a series of abliterated and healed Llama and Qwen models. His June 2024 tutorial Uncensor any LLM with abliteration established the two-stage pipeline as canonical, and his NeuralDaredevil-8B is the reference model for the recipe.
How much does abliteration actually damage capability?
Labonne measured his own runs at "1-2% on every benchmark." The pattern varies by benchmark type: MMLU (general knowledge) barely moves, TruthfulQA drops 5-11 points, GSM8K (math) is the most sensitive with drops of ~7-8 points typical. See C2 Reading benchmarks for the full pattern and how to interpret model-card numbers.
Why not use supervised fine-tuning to heal?
Labonne's practical observation: SFT tends to "lobotomize" an instruction-tuned model because it overwrites the base's careful instruct-tuning with the SFT dataset's opinions. DPO is lighter - it only teaches relative preferences between paired responses rather than rewriting the model's whole disposition. In his words: "preference alignment is quite light and shouldn't lobotomize our abliterated model."
What is the self-cancelling risk?
Because DPO is preference alignment, a healing pass can quietly re-teach some refusal depending on what is in the "chosen" answers. General-purpose preference datasets often contain polite declines as chosen responses - because polite declines were preferred during their original curation - so a naive M4 pipeline can partly undo the abliteration it was meant to complement. Practitioners manage this by curating datasets to exclude refusal-as-chosen pairs and mixing in explicit compliance-vs-refusal pairs, but the risk is real. In this sense M4 is a balancing act, not a clean fix.
What can healing not do?
Restore whatever was genuinely destroyed rather than merely disturbed. Abliteration removes a direction from the weights; the capabilities that shared that direction are not recoverable by any amount of preference tuning because the weight components that encoded them are gone. Healing addresses collateral damage (near-neighbor capabilities disturbed by the removal), not direct casualties. This is why tighter, more selective abliteration (Heretic's KL-optimizing approach) leaves less collateral damage for healing to recover from.
Is a healed model less aligned than its base?
No - it is differently aligned. Both the base and the healed model went through preference-alignment training; they differ in which preferences the training pushed toward. The base was aligned toward caution; the healed model was aligned twice, first toward caution (the base training) and then toward capability (the healing pass), with an abliteration step between the two that shifted the starting point for the second alignment. The word "uncensored" that model cards apply to healed models obscures this - the model was not stripped of alignment, it was realigned.
Should I use "healing" when describing my own work?
Your choice, but be aware of what the word does. If you want to be neutral, "realignment" or "second-stage preference tuning" are more descriptive. If you want to explicitly reject the vocabulary of restoration, "compensation" or "second education" foreground different framings. If you want to be understood by the field without extra explanation, "healing" is the term of art and communicates fastest, at the cost of endorsing the base model as the reference. Different producers make different choices; none is wrong technically, and the choice is worth making deliberately.
References
- Rafailov, R., et al. (2023). Direct Preference Optimization: Your Language Model is Secretly a Reward Model. NeurIPS 2023. arXiv:2305.18290
- Labonne, M. (2024). Uncensor any LLM with abliteration. mlabonne.github.io/blog/posts/2024-06-04_Uncensor_any_LLM_with_abliteration.html
- mlabonne/NeuralDaredevil-8B-abliterated (canonical M4 output). huggingface.co/mlabonne/NeuralDaredevil-8B-abliterated
- Labonne Hugging Face profile. huggingface.co/mlabonne
- Liquid AI. liquid.ai
- Lee, A., et al. (2024). A Mechanistic Understanding of Alignment Algorithms: A Case Study on DPO and Toxicity. arXiv:2401.01967
- Arditi, A., et al. (2024). Refusal in Language Models Is Mediated by a Single Direction. NeurIPS 2024. arXiv:2406.11717