Not a scalpel: the off-target effects of abliteration
The community metaphor for abliteration is surgical: a clean cut through the refusal direction, everything else preserved. A 2026 study of 10,800 decisions in which the base model refused nothing at all found that abliterated arms still diverged from base arms in risk-taking, optimism, and decisiveness. Whatever the operation is, it is not a scalpel.
A 2026 study by de Peretti, Bricken, Chughtai and colleagues (arXiv:2607.17427) measured what abliteration does on tasks where refusal is not in play. Across 10,800 economic and moral decisions the base Llama and Qwen models never refused once, so any behavioural change between the base and its abliterated version had to be pure side effect. It was not small. Abliterated arms became more risk-taking, more optimistic, and more decisive than their bases in ways that persisted across model families and abliteration tools. The finding is that abliteration is not a scalpel that removes one direction while leaving the rest intact. It is a broader disposition shift that happens to include refusal removal as its most visible symptom.
- What the 'scalpel' metaphor claimed and why the community adopted it
- How the 2607.17427 experiment isolated pure side effects
- What 'disposition' means in this context and how it was measured
- Why the shift is a natural consequence of geometric entanglement
- Consequences for agent-loop deployments and evaluation
- What the finding does and does not invalidate about the practice
The metaphor that shaped the practice
From the earliest days of the technique, practitioners described what abliteration does with the vocabulary of surgery. Arditi's 2024 paper reports a benchmark drop of "less than 1%" and calls the operation "surgical." Labonne's canonical tutorial repeats the frame: a clean edit, a narrow cost, the rest of the model preserved. The community's model cards inherited the word. By 2025 it was routine to see an abliterated release described as its base "with the refusal direction removed," as if the operation were the mathematical subtraction it looked like on paper.
The metaphor did a lot of work. It made the practice legible to newcomers, it separated abliteration from the more visibly-invasive full fine-tune, and it lent the operation a certain professional confidence: the practitioner as neurosurgeon, the model as patient, the refusal direction as the offending growth. It also carried an empirical claim, quietly. If abliteration were really surgical, then an abliterated model would behave identically to its base on any task where refusal was not in play. This was assumed rather than tested for two years.
The 2026 test
In July 2026, de Peretti and colleagues at the Anthropic Fellows programme, working with Bricken and Chughtai, published "Abliteration Is Not a Scalpel". The design was aimed directly at the surgical claim. Take base and abliterated versions of the same model. Pose them 10,800 decisions drawn from economic games and moral trolley variants and small-stakes-vs-large-stakes gambles. Measure what they choose. Keep the tasks careful enough that no version, base or abliterated, refuses to answer.
The base arms performed as expected. Llama and Qwen base models completed all 10,800 decisions without ever refusing. There was no refusal signal to abliterate away, because refusal was not present.
The abliterated arms diverged anyway. Across every model family tested and every abliteration tool used, the abliterated versions consistently chose differently from their bases on the same decisions. They chose the higher-variance option in gambles more often. They estimated better outcomes when asked about uncertain events. They committed to a choice with less hedging on the same trolley variant to which the base had responded with careful multi-step conditional reasoning. None of the differences involved a refusal. All of them showed up anyway.
The paper's core observation follows directly. If abliterated arms behave differently from base arms on tasks where refusal is not in play, then abliteration is doing something to the model beyond removing refusal. The something-else is measurable, consistent, and larger than the surgical framing had implicitly promised.
What "disposition" means here
The paper's chosen vocabulary is disposition. A disposition, in this technical sense, is a tendency to answer questions of a certain shape in a certain way even when the questions do not have determinate correct answers. Risk-tolerance is the tendency to choose higher-variance options when several options are available. Optimism is the tendency to estimate better outcomes when several outcomes are possible. Decisiveness is the tendency to commit with less hedging when a choice is not fully constrained by the input.
All three of these are downstream of a shared underlying property of the model that the interpretability literature has not yet named cleanly. What we know is that abliteration moves all three of them in the same direction on the same model. What we do not know yet is whether the underlying property is a single latent factor, three loosely-coupled factors, or something more distributed.
This is not the same as saying the abliterated model has become less capable. It has not, on any of the standard capability benchmarks that stayed within one point of the base. It is not the same as saying it has become more toxic or more unaligned in the ordinary sense of those words. On refusal-rate for genuinely harmful content it did what it was supposed to do. What has happened is that a set of subtler tendencies, ones the base model had cultivated as part of what the community called "alignment" and this paper more carefully calls disposition, have shifted together in a common direction.
Why this is the expected outcome, given what we know
The finding was not entirely a surprise to the interpretability side of the field, though its magnitude was. Several strands of prior work pointed in the same direction:
- The TruthfulQA regression that has followed every abliteration technique since 2024. The refusal direction, whatever else it encoded, encoded the model's willingness to disagree with plausible falsehoods, so removing it degraded that willingness. This was one instance of the broader pattern.
- The controlled study on Llama fine-tunes (arXiv:2606.23375) that reported HarmBench attack-success moving from 14.5% to 55.5% on a LoRA-based abliteration and to 82.5% on a full-weight abliteration of the same base. That study was framed as a safety observation but its underlying pattern was the same: the further an intervention reaches into the model, the larger the collateral shift.
- The Wollschläger cones critique of the single-direction hypothesis, which argued that refusal is mediated by multiple entangled directions rather than one. If that is even partially correct, then single-direction ablation is not a scalpel but a saw taken to a bundle of wires, and the not-a-scalpel finding is the behavioural signature of the mismatch.
The 2026 study did not settle the underlying mechanism. It documented the effect and named the pattern. Whether the mechanism is best explained by the concept-cone geometry, by superposition, by shared circuits between refusal and disposition, or by something else, is an open interpretability question. The behavioural fact is stable across models and tools.
What this means for practice
The finding has three practical consequences.
Evaluation. A practitioner who evaluates an abliterated model only on refusal-rate deltas and standard capability benchmarks is systematically blind to the disposition shift the 2026 study documented. Standard benchmarks are not sensitive to the risk-tolerance and optimism changes because they are pass-fail on determinate answers. Bringing an abliterated model into a deployment where the disposition matters requires evaluating it on tasks where the base model had a distribution of reasonable choices, not just where it had a single correct one.
Agent deployments. An abliterated model in an agent loop is not equivalent to the base minus refusal. It is a base with a shifted disposition, iterated. Small shifts in single-decision behaviour become large shifts in trajectories over many decisions. A base that was cautious about tool use and a corresponding abliterated version that has become more decisive will produce different agent trajectories on the same task, in ways that may not correlate with refusal rates at all. This is the same failure mode that the "not a scalpel" title points to, running in a loop.
Model-card documentation. The community's current model-card convention describes an abliterated release by what it was intended to change (refusal removed) and how much capability it kept (benchmark deltas). Neither of these signals the disposition shift. A more informative card would separately name what was intended to change, what was verified to have changed, and what was not measured. This is a cheaper improvement than any of the technical remedies and produces the largest legibility gain per hour of work.
What this does not change
Abliteration still works for its primary purpose. Refusal rates fall reliably; the tools are usable; the practice is not on the wrong track. What the 2026 study takes away is the metaphor. What it adds is a more accurate description: abliteration is a broad disposition shift of which refusal removal is the most visible symptom. The mental model has to widen to match the operation.
For practitioners this is manageable. For evaluators it is important. For the field's self-understanding of what alignment is and how it is stored inside models, it is one more piece of evidence that the categories that make training legible ("refusal", "helpfulness", "harmlessness") are not the categories at which the weights actually cleave. The weights are entangled at a different resolution than the training targets are.
The word "abliteration" is likely to keep being used. The community has too much invested in it and no better word is coming. But the surgical metaphor that came bundled with the word deserves retirement. What we do is not surgery. It is a broader intervention, and calling it what it is makes the practice better and the debate more honest.
Where to go next
- What is abliteration? - the general reference; the section on "what it costs, what it preserves" is where the surgical framing lived and where this article amends it.
- The dimensionality debate - the geometric side of why single-direction ablation is not clean.
- Boundary cases - including the emerging case on reasoning-model abliteration, another instance where the operation does not do what the label suggests.
- DPO healing - the community's response to the visible-benchmark part of the collateral cost. Healing addresses capability drops; the disposition shift is a separate axis that healing does not directly touch.
Frequently asked questions
What is the "not a scalpel" finding?
That abliteration causes behavioural change well beyond refusal removal. On tasks where the base model was not refusing anything (10,800 economic and moral decisions across two model families), abliterated versions still diverged from their bases in measurable, consistent ways: more risk-taking, more optimism, more decisiveness. Because refusals were not present to explain the difference, the divergence has to come from something abliteration did to the model beyond removing refusal. Reported in de Peretti et al., arXiv:2607.17427 (July 2026).
Does this invalidate abliteration as a practice?
No. It reframes what the practice is. Abliteration still reliably reduces refusal rates. What the finding takes away is the clean-surgery mental model: an abliterated model is not a base model minus refusal, it is a base model with a broader disposition shift of which refusal removal is one component. Practitioners have to design and evaluate accordingly, which is a different task than the surgical framing suggested.
What disposition dimensions shifted?
The paper identifies three that shifted consistently across models and tools: risk tolerance (abliterated models chose higher-variance options in economic gambles more often), optimism (they estimated better outcomes for uncertain events), and decisiveness (they committed to a choice with less hedging on the same problem the base had answered cautiously). All three moved in the same direction, which suggests they are downstream of one underlying change rather than three separate ones.
Do all abliteration methods have this problem?
The paper tested a small set of tools and reported the pattern is present across them, with variation in magnitude. It is consistent with what the geometry literature would predict: single-direction ablation edits weights along one axis in a space where that axis is entangled with several other axes, so touching one moves the others whether the practitioner meant to or not. Tighter tools (Heretic's KL-optimising search) minimise the collateral shift; they do not eliminate it.
What should I do about it as a practitioner?
Three things. Do not evaluate an abliterated model only on refusal-rate deltas; also measure its disposition against the base on non-refusal tasks. In agent-loop deployments, budget for the shift explicitly rather than assuming the abliterated model is behaviourally identical to the base minus refusal. In model-card documentation, name what the abliteration was intended to change and what it can be expected to have changed besides.
How is this different from the TruthfulQA regression already known since Arditi?
TruthfulQA drop is the specific instance the earlier literature caught: a benchmark whose numbers move down after abliteration because the refusal direction was entangled with a pushback-against-plausible-falsehoods disposition. The not-a-scalpel finding generalises the pattern. TruthfulQA is one visible symptom of a broader entanglement between refusal and other dispositions, all of which shift together when refusal is removed.
Is this related to "emergent misalignment"?
The paper places itself alongside Betley et al. (2025) on emergent misalignment, which showed that narrowly-targeted training interventions can produce broad behavioural changes far from the training target. Abliteration is a different intervention (weight editing rather than training) but the underlying observation is the same: interventions that were intended to be local turn out to be global. The not-a-scalpel study is the abliteration-side version of that pattern.
References
- de Peretti, L., Bricken, T., Chughtai, B., et al. (2026). Abliteration Is Not a Scalpel: On the Off-Target Effects of Refusal Removal. arXiv:2607.17427
- Betley, J., et al. (2025). Emergent Misalignment: Narrow Interventions, Broad Consequences. Referenced in the 2607.17427 discussion of prior work.
- Arditi, A., et al. (2024). Refusal in Language Models Is Mediated by a Single Direction. NeurIPS 2024. arXiv:2406.11717
- Wollschläger, T., et al. (2025). The Geometry of Refusal in Large Language Models. ICML 2025. arXiv:2502.17420
- Controlled study on Llama fine-tunes (2026). Disposition shift and attack-success in abliterated Llama variants. arXiv:2606.23375
- Labonne, M. (2024). Uncensor any LLM with abliteration. mlabonne.github.io