The eight methods, in overview
M1 to M8 as a working taxonomy. Where it comes from, what alternative schemes exist, and where the categories will change over the next eighteen months.
This wiki organizes abliteration into eight methods: M1 direct removal, M2 raw weight editing, M3 layer-wise ablation, M4 hybrid abliterate-then-heal, M5 merge-based, M6 DPO healing only, M7 custom-dataset fine-tune, M8 repackaging. The taxonomy is a working construct, not a settled field standard. Four alternative implicit taxonomies exist in the field; ours differs by treating post-removal finishing steps as first-class methods.
- The eight methods in one table, with difficulty tier and canonical example
- Why one taxonomy is a choice, not a discovery
- The four alternative implicit taxonomies (Labonne, Heretic, mergekit, geometric)
- Five cross-cutting themes that recur across all eight methods
- How the taxonomy will change: M2 collapsing, M1 splitting, M6/M7 merging
The eight methods
Abliteration is not one thing. Since Arditi et al. published the original technique in June 2024, the practice has fanned out into a family of related and overlapping methods. Some are faithful implementations of the original algorithm. Some graft a repair step onto it. Some merge already-abliterated models into new ones. Some fine-tune from scratch to produce compliant models without touching a single weight matrix directly. Some do nothing to a model except change its file format.
This wiki organizes the practice into eight methods, M1 through M8. The table below is the table of contents to the rest of Cluster B. Each row links to a dedicated article that describes the method in mechanical detail, cost, canonical examples from the catalog, and hands-on procedure.
$ cat eight_methods.tsv | column -t METHOD WHAT IT DOES COST DIFF. ───────────────── ────────────────── ──── ───── M1 direct weight surgery $ med M2 raw editing ad-hoc surgery $ med M3 layer-wise partial surgery $ med Heretic auto wraps M1/M3 $ easy M4 hybrid M1 + DPO healing $$ hard M5 merging weight arithmetic free med M6 DPO only training, no cut $$ hard M7 custom FT domain fine-tune $$ hard M8 repackaging format change ¢ easy strict (Arditi): M1, M2, M3 extension: Heretic, M4, M5 no ablation: M6, M7, M8
All eight are called "abliteration" somewhere in the ecosystem, though only the first three - M1, M2, M3 - are abliteration in the strict sense Arditi described: cutting the refusal thread out of the file directly. M4 combines M1 with a training-based repair step. M5 propagates abliteration through model merging. M6 and M7 are fine-tuning traditions that reduce refusal without ever locating the refusal thread. M8 is quantization, which changes a model's file format without changing its behavior on purpose.
The distinctions matter because the operations differ in cost, forensic legibility, and what they preserve. See the articles below for the details.
Method articles
- M1 - Direct removal: the canonical Arditi-style ablation. Difference-of-means direction extraction, weight orthogonalization, no repair step. Cost: dollars, hours. Difficulty: medium.
- M2 - Raw weight editing: M1 without validation loops. Ad-hoc weight surgery. May not survive as a category; see the "where the taxonomy will change" section below.
- M3 - Layer-wise ablation: M1 restricted to a chosen band of layers. huihui-ai's dominant recipe. Cost: dollars. Difficulty: medium.
- Heretic - automated abliteration: Weidmann's one-command tool that wraps M1/M3 in an optimizer. Cost: dollars. Difficulty: easy.
- M4 - Hybrid abliterate-then-heal: Labonne's two-stage pipeline. M1 to remove refusal, DPO to repair capability. Cost: tens of dollars. Difficulty: hard.
- M5 - Merge-based: propagate abliteration into new models through mergekit. Cost: memory only, no training. Difficulty: medium.
- M6 - DPO healing only: uncensored fine-tuning without any weight-level abliteration. Hartford / dolphin lineage. Cost: tens of dollars. Difficulty: hard.
- M7 - Custom-dataset fine-tune: roleplay corpora, PIPPA, SillyTavern models. Refusal reduction as side effect of domain training. Difficulty: hard.
- M8 - Repackaging: GGUF quantization. Changes format, not behavior (mostly). The mradermacher pipeline. Cost: cents. Difficulty: easy.
One taxonomy among four
There is no canonical taxonomy of abliteration methods. The categories that circulate in the field are implicit, embedded in the structure of tools and tutorials rather than stated as classifications. Ours - M1 through M8 - is a working construct, one taxonomy among several. It is worth naming the others to see what ours does and does not accomplish.
Labonne's implicit taxonomy: procedure
The single most influential practitioner text is Maxime Labonne's June 2024 post Uncensor any LLM with abliteration. Labonne does not enumerate methods; his structure is procedural. But two distinctions in his framing have been widely absorbed by the field.
The first is between inference-time intervention (projecting out the refusal direction during each forward pass) and permanent weight orthogonalization (baking the projection into the stored weights). The second, more consequential for a method taxonomy, is his separation of abliteration proper from the healing step that follows it. Labonne treats abliteration as a lossy operation and DPO fine-tuning as the repair. This two-stage framing is the direct ancestor of the distinction we draw between M1 (removal alone) and M4 (removal plus healing).
Heretic's implicit taxonomy: automation
The Heretic project encodes a different implicit taxonomy, organized around automation rather than procedure. Heretic's documentation situates the tool relative to two reference points: the original Arditi procedure (treated as the scientific foundation) and manual practitioner abliteration (treated as the thing to be improved on). In Heretic's own framing, the field divides into manual directional ablation (expert-driven, parameter-tuned by hand) and automatic directional ablation (parameter search by optimizer).
This axis is orthogonal to Labonne's. It cuts across our M1, M2, and M3 without distinguishing them - because from Heretic's point of view they are all the same operation performed with different degrees of care.
Mergekit's implicit taxonomy: merging-first
The mergekit community carries a third scheme, and it barely acknowledges abliteration as a separate category at all. In that world, an abliterated model is simply one more ingredient to be combined, and the interesting distinctions are between merge algorithms (SLERP, TIES, DARE, task arithmetic, passthrough) rather than between ways of removing refusal. The mergekit toolkit, authored by Charles Goddard and colleagues at Arcee, treats behavioral properties as things that transfer through weight arithmetic.
This is the origin of our M5 category, and the friction between the mergekit worldview and the abliteration worldview is itself a data point: to a merger, abliterated is a property a model has, not a process it underwent.
The academic geometric taxonomy
Academic literature has begun to supply a fourth axis, this one geometric. The comparative-evaluation paper by Richard J. Young (2026), Comparative Analysis of LLM Abliteration Methods, distinguishes standard ablation, norm-preserving ablation, and projected ablation as three mathematically distinct operations on the weight matrices. This is a finer distinction than any practitioner taxonomy draws, and it partly cuts across our M1: what we call M1 contains at least three geometrically different procedures that the literature is only now naming.
Five cross-cutting themes
Five patterns run across all eight methods and are better understood together than method by method.
The truthfulness cost
The most robust empirical finding in the whole area, visible from the first paper onward. Arditi and colleagues report that abliteration barely moves MMLU, ARC, or GSM8K but consistently lowers TruthfulQA by roughly 2 to 6 points depending on the model. This is not an implementation flaw that better engineering removes. It recurs across methods because the direction that mediates refusal is entangled with the model's disposition to push back on false premises. A model that refuses less also, on the margin, contradicts less. M4's healing step can partly restore TruthfulQA along with other benchmarks, but the general pattern holds. Any producer's claim of a no-degradation abliteration should be read against this baseline expectation.
$ column -t truthfulness_cost.tsv MODEL MMLU ARC GSM8K TruthfulQA ───────────── ───── ───── ───── ────────── Llama-3 70B +0.1 +0.0 -0.2 -2.3 Yi-34B -0.1 +0.1 -0.3 -3.5 Qwen family +0.0 -0.1 -0.1 -2 to -4 general benchmarks: barely move TruthfulQA: down 2-4 pts effect: specific, not diffuse
The pattern is consistent across model families. Llama-3 70B: TruthfulQA down 2.3 points. Yi 34B: down 3.5 points. Qwen family: comparable. General knowledge benchmarks (MMLU, ARC) barely move; the effect is specific to the willingness to disagree that TruthfulQA measures.
This is not a bug of any particular implementation. It is something about how these models learn: refusal and truthfulness share threads in the model's running notebook, and cutting one damages the other.
The second-order phenomenon
The application of methods on top of already-modified models is pervasive rather than exceptional. M4 is M1 followed by healing. M5 merges abliterated models into new ones. M7 fine-tunes abliterated bases for roleplay. M8 quantizes any of the above. DavidAU's models are frequently third- or fourth-order: an abliterated base (M1), assembled into a mixture-of-experts (M5), extended with a Brainstorm adapter, and shipped as GGUF (M8), with a further Heretic pass over constituents in some versions.
Each order compounds the provenance problem. By the third layer, the relationship between the shipped model's behavior and any original abliteration is mediated by so many intervening operations that the label abliterated describes ancestry more than a testable property.
The reproducibility gap
The methods divide by how much of a trace they leave. M1 and M3 leave a detectable signature: weight orthogonalization is a specific low-rank edit to the residual-writing matrices, and in principle it can be identified by inspecting the weights. M6 and M7 leave a much fainter trace, because gradient-based fine-tuning distributes its changes across all weights with no localized signature. M4 sits between - the abliteration half is detectable, the healing half is not. M8 preserves whatever trace its parent had, minus quantization noise.
The tooling ecosystem
The infrastructure has consolidated around a small number of load-bearing projects. TransformerLens underlies the original and FailSpy implementations; the pure-Transformers reimplementations (Sumandora's and its many forks) removed that dependency and broadened access. mergekit is the merging standard. Heretic automated the M1 parameter search. llama.cpp and its GGUF format anchor distribution.
A handful of maintainers sit at chokepoints through which a large fraction of the ecosystem's models pass: Goddard for merging, Weidmann for automation, mradermacher for distribution. Heretic itself is worth a mechanical note here: authored by Philipp Emanuel Weidmann and released in late 2025, it wraps directional ablation in an Optuna-driven Tree-structured Parzen Estimator that automatically searches ablation parameters, co-minimizing refusal count and KL divergence from the original model. Its README reports that on gemma-3-12b-it the automatically generated version achieves the same level of refusal suppression as manual abliterations at KL divergence 0.16 versus a manual-method range of 0.45 to 1.04. The repository has drawn over 22,000 GitHub stars.
The computational-cost gradient
Cost runs from near-zero to substantial and largely determines who can do what.
- M8 (repackaging) is the cheapest: only conversion compute is needed.
- M1 and M3 need only forward passes over a few hundred prompts to extract the direction. No training. As the original authors put it, "our method, which can yield a jailbroken version of a 70B parameter model using less than $5 of compute, is simpler than previous fine-tuning methods, requiring neither gradient-based optimization nor a dataset of harmful completions."
- M5 (merging) needs no GPU training at all, only memory to hold the tensors.
- M6 and M7 need a fine-tuning run - moderate but real.
- M4 is the most expensive, combining full M1 extraction with a preference-optimization training pass. Labonne's reference healing run used six A6000 GPUs for nearly seven hours.
This gradient explains the volume distribution of the catalog. The cheap methods (M1, M8) dominate by count; the expensive method (M4) dominates the subset marketed on quality.
Where the taxonomy will change
Three pressures are likely to reshape the M1-M8 list over the next twelve to eighteen months. All three are visible in current trends, not speculative.
M2 is the most likely to disappear. The evidence assembled across Batch 2 of our source reports is that M2 is not a distinct method but low-discipline M1. No tool, paper, or tutorial treats it otherwise. The probable change is that M2 collapses into M1 as a quality gradient, with the catalog flagging ad-hoc surgery as an attribute of an M1 model rather than a separate category. The only thing preserving M2 as a peer category is the descriptive convenience of marking models whose authors documented ad-hoc editing.
M1 is the most likely to split. The interpretability literature is converging on the view that refusal is not a single direction but a multi-dimensional structure (Wollschläger's concept cones, the multi-direction decompositions in RFM-AGOP work). Geometric refinements (projected, norm-preserving, magnitude-preserving ablation) are already distinct enough operations that Young's comparative study treats them separately. If multi-directional abliteration becomes standard practice, M1 will need to distinguish single-direction removal from cone-based or multi-direction removal, because they have different completeness properties. This is the change most driven by the science rather than practitioner convenience.
M6 and M7 may merge, or M6 may leave the taxonomy. Both are fine-tune-based refusal removal distinguished only by dataset intent, and the boundary is already contested. A plausible consolidation is a single "training-based refusal removal" category subdivided by intent (safety-removal versus domain fine-tune), which would absorb both. Alternatively, if the field firmly settles on reserving "abliteration" for weight-editing methods, M6 leaves the taxonomy entirely and is treated as adjacent uncensored fine-tuning.
The automation trend cuts across all of this. As parameter search replaces hand-tuning, the practical distinctions between M1, M2, and M3 (which are largely distinctions of care and layer selection) erode - because an optimizer makes those choices automatically and consistently. If automated abliteration becomes the default, the taxonomy's removal-side categories may compress toward a single "automated directional ablation" method, with the interesting variation moving to the healing, merging, and repackaging stages that automation does not yet cover.
Two review tracks now formally open
The Q3 2026 academic corpus review closed with a recommendation to open two review tracks for the next quarterly report, on the basis that primary-source thresholds have been met. Both are now open. Final method numbers (e.g. M9 and M10, or two branches of a single M9) will be assigned in the next quarterly report on the basis of empirical model counts and community usage of each label between now and then.
- Multi-directional / subspace ablation track. Four independent primary sources report that single-direction ablation is a first approximation and that refusal occupies a cone or subspace rather than a line: Wollschläger cones (ICML 2025), Piras SOM (AAAI 2026), Joad eleven-category (2026), and RFM-AGOP (2026). Threshold met: three or more independent primary sources reporting the same qualitative refinement on shared model families. See M1 related literature for full entries.
- Refusal restoration and amplification track. Two published methodologies now exist for the inverse operation (adding a refusal signal back into a model, or fortifying it against removal), which updates the earlier working note that no such methodology existed: ROSI (Abu Shairah et al., Aug 2025) and AMRA (Truong, Jun 2026). See boundary-cases Case 3 for treatment. Threshold partially met: primary-source condition satisfied; a consolidated tooling substrate on the scale of Heretic or FailSpy for restoration is still absent, so a model card claim "refusal restored" cannot yet be verified without re-running the specific procedure on the specific weights.
Both tracks will be revisited in the Q4 2026 quarterly report, at which point either they will be promoted to formal M-methods (with concrete numbers, catalog tags, and taxonomically-consistent treatment), or their primary sources will be folded into existing method pages as refinements to M1 rather than as new categories. The decision will be based on whether models are being published that specifically warrant a new tag - the same rule the M2 disappearance discussion above applies.
Reading the method articles
Each method article follows the same structure: one-line definition, how it works mechanically, attribution, what it costs and preserves, canonical example models (linked to our catalog), critiques and disputes, and practical entry points. Read them in whatever order fits your interest. If you plan to run abliteration yourself, the practical how-to article stitches them into a hands-on sequence with real cost numbers.
Frequently asked questions
Is M1-M8 the standard taxonomy of abliteration methods?
No. There is no canonical taxonomy. M1-M8 is a working construct developed for this wiki and our catalog. Four other implicit taxonomies exist in the field: Labonne's procedural framing (inference vs weight, raw vs healed), Heretic's manual-vs-automated axis, mergekit's merging-first worldview, and the academic geometric axis (standard vs norm-preserving vs projected ablation). Ours differs by treating post-removal finishing steps (healing, merging, repackaging) as first-class methods rather than afterthoughts.
Why treat repackaging (M8) as a method if it does not remove refusal?
For provenance. The majority of abliterated models a user actually downloads are GGUF quantizations, not the original safetensors, so a catalog that omitted the repackaging layer would misdescribe how most models reach users. We treat M8 as a distribution modifier that co-occurs with a refusal-removal method rather than as a competing refusal-removal method, which keeps it in the taxonomy for provenance purposes without pretending it does the same job as M1.
Which method is most common in the catalog?
M1 direct removal, by volume, followed closely by M8 repackaged versions of M1 outputs. The dominance follows from cost: M1 is dollars, M8 is cents. M4 (Labonne's abliterate-then-heal pipeline) produces fewer models but is over-represented in the subset marketed on quality. M6 and M7 (fine-tuning traditions) are smaller, higher-craft niches.
Which method is cheapest to run?
M8 repackaging (cents in conversion compute) if you already have an abliterated model to repackage. For producing an abliteration from scratch: M1 direct removal via Heretic on a 4B model is roughly $0.50 on rented cloud hardware. M4 (the full Labonne pipeline with DPO healing) is the most expensive, running to $30+ for the training pass alone.
Which method most fully removes refusal?
The question is contested. M1 via weight orthogonalization gives a documented, mathematically clean edit but leaves multi-directional refusal (concept cones) partly intact per Wollschläger 2025. M6 and M7 fine-tuning may more thoroughly retrain the model's refusal disposition but leave no forensic signature to prove they did. In practice, M4 (hybrid) is the standard for producers who want a benchmarks-preserving uncensored model. See the benchmarks article for how to read the trade-offs.
Will M1-M8 still be the taxonomy in a year?
Probably not exactly. M2 is likely to collapse into M1 as a quality gradient. M1 is likely to split into single-direction and multi-directional variants once the concept-cones literature stabilizes. M6 and M7 may merge into a single training-based category, or M6 may leave the taxonomy entirely if the field settles on reserving "abliteration" for weight-editing methods only. Automation is compressing M1/M2/M3 distinctions at the practical level. The taxonomy is a working map, not a final classification.
What is not in the taxonomy?
Runtime activation steering (e.g., CAA / Rimsky et al. 2024) is not in M1-M8 because it changes no stored weight - the intervention lives at inference time and disappears when switched off. LoRA-based refusal removal blurs into M6 or M7 depending on whether the adapter is merged into the weights or kept separate. Multimodal encoder patches (for vision-language models) target the vision tower rather than the language residual stream. The geometric refinements (projected, norm-preserving, biprojected ablation) are treated as M1 sub-variants for now, though the academic literature is beginning to name them separately.
References
- Arditi, A., et al. (2024). Refusal in Language Models Is Mediated by a Single Direction. NeurIPS 2024. arXiv:2406.11717
- Labonne, M. (2024). Uncensor any LLM with abliteration. mlabonne.github.io
- Goddard, C., et al. (2024). Arcee's MergeKit: A Toolkit for Merging Large Language Models. arXiv:2403.13257
- Weidmann, P. E. Heretic. github.com/p-e-w/heretic
- Young, R. J. (2026). Comparative Analysis of LLM Abliteration Methods: A Cross-Architecture Evaluation. arXiv:2512.13655
- Wollschläger, T., et al. (2025). The Geometry of Refusal in Large Language Models. ICML 2025. arXiv:2502.17420
- Rimsky, N., et al. (2024). Steering Llama 2 via Contrastive Activation Addition. arXiv:2312.06681
- Hartford, E. (2023). Uncensored Models. erichartford.com/uncensored-models
- TransformerLens library. github.com/TransformerLensOrg/TransformerLens
- llama.cpp. github.com/ggml-org/llama.cpp