Magazine / Post

The year abliteration stopped being a scalpel

In 2024 the field thought abliteration was surgical. Wollschläger's cones, Piras's SOM manifold, Winninger's RFM-AGOP, and the 2026 Not-a-Scalpel study each moved the picture in the same direction. Two years on, the operation is not the clean cut the community described.

Two years ago the field looked simple. Andy Arditi and colleagues had shown that refusal in aligned language models was mediated by a single direction in the residual stream, that direction could be found with a subtraction of two averages, and orthogonalising the weights against it stopped the model from refusing. Benchmark drop under one percent, capability preserved, one clean cut. The community reached for the vocabulary of surgery almost immediately. A model that had been abliterated was described the way a patient after a successful operation is described. The refusal direction had been removed. Everything else stayed intact.

This is the story of how that description came apart, told from the catalog. Five papers, one press cycle, and a benchmark that measured what nobody had measured before. Two years after the Arditi paper, the surgical metaphor is done. What we do to models is not a scalpel. This piece is about what the picture actually looks like now and what remains right about the original claim.

February 2025: the first trouble

The first crack came from Tom Wollschläger and colleagues at ICML 2025, in a paper titled The Geometry of Refusal in Large Language Models. They kept the Arditi frame but changed the extraction method: gradient-based rather than difference-of-means. What came out was not one direction. It was several, forming what they called concept cones. Orthogonality between the extracted directions did not imply independence. A model whose refusal was captured by one direction under difference-of-means turned out to carry the same behavioural signal along several other axes as well.

This did not refute Arditi. Difference-of-means still finds a direction, orthogonalising against it still reduces refusal, and small models still exhibit the single-direction pattern most of the time. What Wollschläger changed was the ceiling. Single-direction removal, on a model where the cone geometry applies, deletes the dominant channel. The others remain. What looks like a fully abliterated model at the surface may retain latent refusal machinery a difference-of-means procedure was never going to touch.

The catalog side of this story ran quietly in parallel. Most producers kept using Arditi-style extraction because it was what the tools shipped with, and because on models under 8B parameters the geometry gap did not visibly matter. The people building above 8B started to notice.

Late 2025: two more geometries

By November 2025 two more papers had extended the same trajectory from different angles. Piras and colleagues at PRALab, submitted to AAAI 2026, applied Self-Organizing Maps to the refusal representation and found a low-dimensional topological structure: multiple related centroids on a two-dimensional grid, each contributing a distinct extraction direction, none of them complete on its own. Nine months later Thomas Winninger at the ICML 2026 mechanistic interpretability workshop, applying Recursive Feature Machines with the Average Gradient Outer Product, reported a concrete figure that the Wollschläger paper had gestured toward. On Qwen 3 8B, ablation of at least three directions was required to pass a 50 percent attack-success threshold. Single-direction ablation was not just incomplete on newer models. It was empirically underperforming by a specific, measurable amount.

The three papers converged on the same claim from different methodologies. Refusal is not one line. It is a small subspace that difference-of-means picks up only the first eigenvector of. The mismatch between what the extraction returns and what the behaviour is stored across is not a curiosity. It is the reason a "fully abliterated" 32B model can, under the right prompt, still politely decline.

December 2025: the first honest comparison

Alongside the geometric literature, a different kind of paper appeared. R. J. Young at UNLV published Comparative Analysis of LLM Abliteration Methods, a solo-author empirical study of four tools (Heretic, DECCP, ErisForge, FailSpy) across sixteen instruction-tuned models. This was the first cross-tool benchmark the field had. Its findings were harder to summarise than the geometry papers because they were mostly about how deeply outcomes depended on the specific model.

Two facts from that study are worth carrying: mathematical reasoning (measured by GSM8K) was consistently the most fragile capability under abliteration, with drops ranging from a mild +1.51 to a punishing -18.81 percentage points depending on the tool and the model. And the study documented what it called covert non-compliance: models that opened responses with a boilerplate refusal template ("I cannot help with that") and then, in the next paragraph, complied in full. What looked like a still-refusing model on the first line was, on inspection, a fully compliant model with a decorative preamble. The refusal template had detached from the refusal behaviour. Which of the two the abliteration was measured to have removed depended on where the evaluator was looking.

This finding is easy to miss because it sounds like a bug. It is not a bug. It is what happens when the extraction procedure targets one specific surface expression of refusal (the opening template) and the underlying disposition is stored somewhere the extraction did not reach.

February 2026: a synthesis

In February 2026 Joad and colleagues published a paper whose title said the thing directly: There Is More to Refusal in Large Language Models than a Single Direction. Their contribution was less methodological than reconciling. They decomposed refusal across eleven evaluation domains and observed that many directions govern how a model refuses (the register, the domain-specific templates, the style of hedging) while one behavioural lever governs whether it refuses.

This was the first framing that let both the original Arditi paper and the cone literature be right at once. One behavioural control knob. Several representational directions expressing it. The reason single-direction extraction works reliably on small models is that on those models the dominant direction is closer to the underlying lever. The reason it starts to fail on larger and reasoning-tuned models is that on those the representational spread widens and no single direction captures the lever cleanly.

Reading the field forward from Arditi to Joad, the arc is not a refutation. It is a refinement. What the interpretability literature has spent two years doing is turning a working procedure into a coherent theory of why it works, where it stops working, and what a next-generation procedure would have to touch.

July 2026: the metaphor breaks

The last piece of the story is a paper that did not come from the geometry side of the field at all. In July 2026 Léa de Peretti, Trenton Bricken and Bilal Chughtai, working through the Anthropic Fellows programme, published Abliteration Is Not a Scalpel. Their design was the one that had not been run. Take base and abliterated versions of the same models. Pose them a large batch of decisions on which no version, base or abliterated, refuses to answer. Measure what they choose.

Across 10,800 economic and moral decisions, base Llama and Qwen models refused nothing. This was not because the tasks were sanitised. It was because economic gambles and trolley variants pose choices, not requests for banned content. There was no refusal signal to remove. Whatever difference the researchers measured between the base and its abliterated version had to be pure side effect.

The differences were consistent, and they moved in a common direction. Abliterated arms became more risk-taking. More optimistic. More decisive. All three across model families, all three across abliteration tools. The paper does not claim to explain the underlying mechanism, and the interpretability side of the field has not yet settled on one. What it establishes is the fact. A procedure that removes almost nothing on standard capability benchmarks moves a measurable set of behavioural tendencies far outside the refusal domain the practitioner intended to touch.

The way we described these things for two years had a quiet claim in it. If abliteration were a scalpel, then an abliterated model would behave identically to its base on tasks where refusal was not present. This was the empirical hypothesis the surgical metaphor carried without ever explicitly testing. In July 2026 it was tested and it did not hold. The community had been shipping a broader intervention than it was documenting.

We keep a wiki article on the finding. The short version is that abliteration is not a base model minus refusal. It is a base model with a broader disposition shift of which refusal removal is the most visible symptom. The mental model has to widen to match the operation.

What this looks like from the catalog

From the catalog side, three concrete changes have followed. None of them is dramatic. All three are the kind of thing a reference site notices before a field consensus catches up.

Producers of larger models have started documenting layer bands more explicitly. huihui-ai's 27B and 72B releases now include an explicit "layers A to B ablated, first N retained" statement in the model card, alongside the traditional method label. This is a response to the geometry literature. If the extraction misses part of the signal, the practitioner at least tells you where they looked and where they did not.

Second, the tool landscape around Heretic has widened. Abliterix (a Heretic fork with MoE-granular ablation), DECCP (narrower architecture support but lowest capability degradation in the Young benchmark), and ErisForge (runtime-hook rather than permanent weight edit) each occupy a niche Heretic itself does not target. The forks are visible acknowledgement that no single procedure covers everything. We keep a running list.

Third, the practice of shipping an abliterated release without a companion capability benchmark is starting to look dated. The Young cross-tool study made it hard to claim capability preservation without measuring it, and the Not-a-Scalpel finding made it hard to claim behavioural preservation without measuring that too. Model cards on the newer releases increasingly carry both, alongside notes on what the abliteration was intended to change and what it did not measure. This is a small improvement, and it comes cheap. We recommend it.

What remains right about the original claim

None of the above says the Arditi paper was wrong. The refusal direction is a real thing. Difference-of-means finds it, and orthogonalising against it reduces refusal to the degree the paper reported. Every subsequent method builds on that observation. The two-year arc is not a refutation. It is the field learning that one direction was the leading term in an expansion whose higher terms matter when the model gets bigger, when the task gets more subtle, or when the deployment cares about behaviour a benchmark does not visibly capture.

There is a temptation, given how much has moved, to say the community was wrong to use the surgical metaphor. We do not think that is quite right either. In 2024 the metaphor was consistent with the evidence. The evidence changed. Metaphors are supposed to be updated when the evidence does. The mistake is not to have used the wrong metaphor in 2024. It is to keep using it in 2026 after the papers that made it obsolete were published, cited, and confirmed.

The next abliteration paper you read, from us or from anyone else, should not describe the operation as clean. There is no honest way to keep that description. What replaces it is not less rigorous. It is a different kind of rigour: naming what the operation touches, naming what it might touch that we have not measured, and reporting both refusal-rate deltas and disposition-shift measurements when the release is meant for anything beyond a technical demonstration.

Where this leaves us

Two years in, the practice looks less like surgery and more like editing an entangled system in a place that turns out to be less local than the theory of 2024 suggested. The refusal direction is real. The direction is entangled with several others. Removing it moves them. The move is measurable, it is consistent across models and tools, and it deserves to be treated as part of the operation rather than as an unfortunate side effect that a healing pass will patch up.

For the field this is a good place to be. The vocabulary is settling into shape. The tools are widening. The benchmarks are appearing. The regulatory environment (see the recent timeline additions on the EU AI Act's GPAI enforcement powers and the Digital Omnibus) is starting to name the objects the practice produces. The next twelve months will probably see the first defense-side procedures widely deployed, the first regulatory guidance that names abliteration directly, and the first cross-model dispositional evaluation harness. None of this is bad news. It is a field growing up.

For us, the change of vocabulary is straightforward. The wiki article on what abliteration is now says the operation is a broader disposition shift. The dimensionality debate article lays out all five positions with their citations. The Not a scalpel article is the flagship for the off-target finding. The catalog carries the model cards. This magazine post is the place to say, in one place under a byline, what the two years added up to. When a reader asks "what happened to abliteration between the Arditi paper and now," this is what happened.

Follow-ups already in the pipeline: an editorial breakdown of the AMRA and Extended Refusal defense papers, a piece on what agent deployments should measure now that the disposition shift is documented, and interviews with mlabonne and Weidmann on how their own thinking has moved. The magazine is where those will land. If you have questions we should be answering, write us.

Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in app" button that hands off directly to a local runtime of your choice - Infrahuman, LM Studio, or Ollama. No API keys, no subscription, no prompt leakage.