M2 - Raw weight editing
Ad-hoc weight surgery without the disciplined validation loop of M1. Defined by what it lacks rather than by what it adds. May collapse into M1 as a quality gradient rather than remain a peer category.
M2 is ad-hoc, hand-tuned weight surgery aimed at removing refusal without the disciplined direction-selection and validation loop of M1. In practice it is M1 performed without the guardrails: same operation, different level of care. Every actual instance in the wild is M1's operation performed without M1's selection loop, which is why some taxonomists argue M2 should collapse into M1 as a quality gradient rather than remain a separate category. This wiki keeps M2 for provenance-descriptive purposes only - to flag models whose producers documented ad-hoc editing.
- Why M2 is defined by absence rather than presence
- The two characteristic failure modes of unvalidated edits
- huihui-ai's changelog as documented M2 iteration
- The argument that M2 should collapse into M1
- When the M2 label is worth keeping
What M2 is
M2 is ad-hoc, hand-tuned weight surgery aimed at removing refusal without the disciplined direction-selection and validation loop of M1. The mechanical operation is the same as M1: subtract a projection onto some refusal-associated direction from the residual-writing weight matrices. What distinguishes M2 is the absence of the selection-and-validation discipline.
Where M1 extracts candidate directions from every layer and formally scores each one against bypass, induction, and KL-divergence criteria on a validation set, M2 fixes a direction and a set of target layers by intuition or by copying settings from a previous model, applies the edit, and checks the result by eyeballing a handful of generations. Some M2 work uses a scale factor greater than one to force stronger removal, or applies the same direction across a wider band of layers than a validation loop would endorse.
The visible before/after is the same as M1. The difference is in the failure modes.
The two characteristic failure modes
Under-validated edits produce two defects that a properly-validated M1 run would catch before publication.
Refusal on unseen prompts. The direction was extracted or copied from a narrow set of harmful prompts and worked on that set. On prompts outside the tuning distribution - different topics, different phrasings, different languages - the model still refuses. The M1 validation loop catches this by measuring refusal rate on a held-out set; M2 skips the measurement.
Visible model degradation. Repetition, garbled output, broken chat formatting - symptoms that the edit disturbed weights it should not have. huihui-ai's changelog documents exactly this kind of iteration, noting a fix that "changed the 0 layer to eliminate the problem of garbled codes" - a symptom of an edit that touched an inappropriate layer. A validated M1 run would catch the KL-divergence spike and reject the direction; M2 finds out from user complaints after release.
Attribution
M2 has no single author because it is defined by the absence of a method rather than the presence of one. It is the accretion of proof-of-concept scripts and one-off edits that circulated after the technique became widely known. The clearest documented lineage runs through the many forks of Sumandora's remove-refusals-with-transformers and its descendants - jim-plus, NousResearch, TheLocalDrummer, and others each maintain forks - which lower the barrier to entry to the point where a practitioner can abliterate a model without engaging with direction selection at all.
huihui-ai's self-description of its own pipeline as "a crude, proof-of-concept implementation to remove refusals from an LLM model without using TransformerLens" is the most explicit named instance of a producer acknowledging M2-style ad-hoc surgery. See our huihui-ai article for how their workflow evolved from initial M2-ish releases toward more disciplined M3 layer selection on larger models.
Canonical examples are hard to name
Naming clean M2 examples is difficult by construction. Producers rarely advertise that they skipped validation, and models that broke outright are usually deleted rather than published. The honest examples are early, superseded releases that their own authors later replaced.
huihui-ai's first-generation Qwen3-8B-abliterated, later replaced by an explicitly "improved" v2 that fixed a garbled-output bug, is a documented case of an initial ad-hoc edit corrected in a later pass. Beyond such cases, most M2 artifacts live in the long tail of single-release accounts that abliterated one model, published it, and moved on.
The dispute over M2's existence as a category
The central dispute is whether M2 should exist as a peer method at all. The first-wave taxonomy overview provisionally treated M2 as distinct from M1; the evidence gathered in this wiki suggests it is better described as low-discipline M1 than as a separate method. No tool, tutorial, or paper defines M2 as a technique with its own procedure. Every actual instance is M1's operation performed without M1's selection loop.
The interpretability critique that applies to M1 - single-direction removal is incomplete on larger models (see the cones vs lines discussion) - applies with more force to M2, because an unvalidated direction is even less likely to capture the dominant refusal channel. If M1 catches roughly the right direction and misses secondary ones, M2 may not even catch the right primary direction.
Two arguments for keeping M2:
- Provenance-descriptive value. A catalog benefits from being able to flag models whose authors documented ad-hoc surgery, distinct from models produced by a validated pipeline. The M2 label is useful when the producer explicitly says they skipped validation.
- Historical accuracy. Much of the early ecosystem's volume was M2. Collapsing it into M1 rewrites that history as more disciplined than it was.
Two arguments against:
- No independent procedure. M2 is not something you do; it is something you fail to do. Categories in a method taxonomy should describe positive operations, not absences.
- No systematic study. There is no benchmark study isolating M2 from M1 precisely because the two are the same operation at different levels of care. The absence of a distinguishing empirical signal is itself evidence that the categorical distinction is weak.
This wiki keeps M2 for the provenance-descriptive use only. Models are assigned to M2 when the author's own documentation indicates ad-hoc editing. Otherwise they are assigned to M1 or M3 based on what the documentation actually describes.
Practical entry points
There is no dedicated M2 literature to point to, which is itself informative. A reader wanting to understand M2 should read the M1 article and then read the caveats in the README of jim-plus/llm-abliteration, which is candid that "abliteration is not full removal of censorship" and that a poorly chosen direction leaves refusal intact. Understanding why validation matters is the fastest route to understanding what M2 is.
Related literature
M2 has no dedicated corpus entry - the category is defined by absence of methodology rather than presence of one. Since M2 is M1 without the validation loop, the M1 related-literature applies in full. Two adjacent references are worth naming here:
- The rank-one weight-editing lineage from Meng and Bau (ROME, arXiv:2202.05262; MEMIT, arXiv:2210.07229) is the mechanical grammar M2 uses without the safety rails M1 adds around it.
- Young's comparative study (arXiv:2512.13655) tests four M1-family tools; a corpus gap is that no analogous study evaluates ad-hoc M2 variants against measured baselines. The absence itself illustrates the category's status.
For the full M1 corpus - precursors, refinements, defenses, comparative studies - see M1 related literature.
Frequently asked questions
What is M2 raw weight editing?
M2 is ad-hoc, hand-tuned weight surgery aimed at removing refusal without the disciplined direction-selection and validation loop of M1. Same operation as M1 (subtract a projection from residual-writing matrices), but without the validation that would catch bad directions or damaged layers before publication. Defined by absence rather than presence.
How is M2 different from M1?
Same operation, different discipline. M1 extracts candidate directions from every layer and scores each against bypass, induction, and KL-divergence criteria on a validation set before choosing which to apply. M2 fixes a direction by intuition or by copying settings from a previous model, applies the edit, and checks by eyeballing a handful of generations. When the guessed direction is good, M2 output is indistinguishable from M1. When it is not, M2 produces models that still refuse on unseen prompts or that show visible degradation (repetition, garbled output).
Why does this wiki keep M2 as a category?
Provenance-descriptive use only. A catalog benefits from being able to flag models whose producers documented ad-hoc surgery, distinct from models produced by a validated pipeline. Some taxonomists (this wiki agrees in principle) argue M2 should collapse into M1 as a quality gradient. Models are assigned to M2 only when the author's own documentation indicates ad-hoc editing; otherwise they are assigned to M1 or M3.
Who invented M2?
No one - it is the emergent practice of long-tail producers who did the M1 operation without the M1 discipline. The clearest documented lineage runs through forks of Sumandora's remove-refusals-with-transformers, which lowered the barrier to entry enough that practitioners could abliterate a model without engaging with direction selection. huihui-ai's own self-description of their pipeline as "a crude, proof-of-concept implementation" is the most explicit named instance of a producer acknowledging M2-style surgery.
What are the characteristic M2 failure modes?
Two. First, refusal on unseen prompts: the direction was tuned on a narrow harmful set and works there, but on prompts outside that distribution the model still refuses. Second, visible model degradation: repetition, garbled output, broken chat formatting - symptoms that the edit disturbed weights it should not have. Both are catchable by M1's validation loop but slip through M2's eyeball check.
Can M2 outputs match M1 quality?
Sometimes. When the guessed direction and layers happen to be good - because the practitioner copied settings from a proven previous run or has enough experience with the model family to intuit the right values - M2 output is indistinguishable from M1. When they are not, M2 fails in one of the two characteristic modes. Because there is no systematic study isolating M2 from M1, the exact rate of "M2 accidentally works" is not documented.
Should I use M2 to abliterate a model?
No. If you are doing new work, use M1 with the validation loop (see the M1 article) or Heretic's automated version (see the Heretic article). M2's cost savings over M1 are zero (both are cheap) and its failure rate is higher. M2 is a description of what happened to some models in the wild, not a technique to reach for.
References
- Arditi, A., et al. (2024). Refusal in Language Models Is Mediated by a Single Direction. NeurIPS 2024. arXiv:2406.11717
- Sumandora. remove-refusals-with-transformers. github.com/Sumandora/remove-refusals-with-transformers
- jim-plus/llm-abliteration (candid README on abliteration limits). github.com/jim-plus/llm-abliteration
- huihui-ai/Qwen3-8B-abliterated (early release later superseded). huggingface.co/huihui-ai/Qwen3-8B-abliterated
- huihui-ai Hugging Face profile (self-described "crude proof-of-concept" pipeline). huggingface.co/huihui-ai