The Arditi method
The 2024 paper that turned refusal into a coordinate. Who wrote it, when, what it exactly claims, and which parts have since been contested.
The Arditi method is the technique described in Arditi et al., Refusal in Language Models Is Mediated by a Single Direction (arXiv:2406.11717, NeurIPS 2024). It identifies one direction in a language model's activation space that mediates refusal, then removes the model's ability to write along that direction. The paper made abliteration reproducible; every subsequent method descends from it.
- Who Andy Arditi and his co-authors are, and where the paper came from
- The exact claim, in the paper's own words
- The three-stage procedure: data, direction extraction, ablation
- Follow-on work: Lermen on 70B, geometric refinements, defensive fine-tuning
- The most important critique: refusal is not one direction but a cone
- The Fafuła result: what the operation does to a model besides removing refusal
The paper
The formal origin of abliteration is a single work: Refusal in Language Models Is Mediated by a Single Direction, by Andy Arditi, Oscar Obeso (joint first authors), Aaquib Syed, Daniel Paleka, Nina Panickssery, Wes Gurnee, and Neel Nanda. The paper first appeared as a LessWrong / AI Alignment Forum writeup on 27 April 2024, then as arXiv preprint 2406.11717 on 17 June 2024 (revised 15 July and 30 October 2024), and was published at NeurIPS 2024. Author affiliations at publication spanned independent research, ETH Zurich, the University of Maryland, MIT, and Anthropic.
The methodological context is mechanistic interpretability, an approach developed at Anthropic and DeepMind (Neel Nanda's group in particular) that studies which parts of a model do what. Prior work in this line - Rimsky et al. 2024 on Contrastive Activation Addition, the broader representation engineering program - had shown that behaviors could be nudged by adding vectors to a model's internal state at runtime. Arditi's contribution was sharper: refusal, specifically, collapses to one direction, and removing it is a one-shot operation that ships with the modified file.
$ grep -A2 "^\[claim" arditi_2024.txt
[claim 1] Refusal is mediated by a
one-dimensional subspace
across 13 open-source chat
models up to 72B parameters.
[claim 2] Erasing what the model writes
along that direction stops it
from refusing harmful requests.
[claim 3] Adding the same direction to
the writing forces refusal even
on harmless requests.
→ the intervention is symmetric.
the same thread that carries
refusal, when severed, removes it;
when tugged, imposes it. The paper states its result as a demonstration across thirteen open-source chat models up to 72B parameters. For each of them, one thread in the model's running notebook carries refusal. Cut what the notebook writes along that thread and the model stops refusing. Pull hard on that same thread while the model is answering a harmless question and it will refuse anyway.
Both effects are documented. The paper is not simply a demonstration of jailbreaking but of a symmetric intervention: the same thread that carries refusal, when severed, removes it; when tugged, imposes it. That symmetry is what makes the finding a claim about how refusal is stored, not just about how it can be defeated.
The procedure, in three stages
The technique has three parts: collect data, extract the direction, then apply it. Each part is deliberately simple, which is part of why the technique spread so fast.
Stage 1: Data collection
The practitioner assembles two matched sets of prompts. The first is harmful: instructions that reliably trigger a refusal template ("I cannot help with that"). The second is harmless: instructions of comparable length and structure that the model handles without hesitation. The paper's canonical harmful set is AdvBench. Harmless prompts are typically drawn from Alpaca-style instruction data.
The model is run over both sets, and its residual stream activations at the final prompt token are recorded at each transformer layer. Nothing else is captured; the point is only to have a picture of the model's internal state when it is about to refuse, and when it is about to comply.
Stage 2: Direction extraction
For each layer, the practitioner computes the mean activation over harmful prompts and the mean over harmless prompts, then subtracts. This difference of means, normalized, is the candidate refusal direction for that layer. Running the same computation across every layer and several token positions yields many candidates.
# for every layer, compute a candidate
# refusal direction, then pick the best.
directions = []
for layer in range(model.n_layers):
harmful_avg = mean_activations(HARMFUL, layer)
harmless_avg = mean_activations(HARMLESS, layer)
candidate = harmful_avg - harmless_avg
candidate /= candidate.norm()
directions.append(candidate)
# arditi et al: middle layers give the
# cleanest signal. selection scores each
# candidate by three criteria.
best_layer = pick_best(directions, criteria=[
bypass_refusal, # removing it defeats refusal
induces_refusal, # adding it induces refusal
minimal_kl_divergence # ordinary text barely moves
])
refusal_direction = directions[best_layer] The code compares two averages. For a few hundred harmful requests it records what the model was about to write into its notebook; for a few hundred harmless requests, the same. Subtracting one average from the other leaves a single thread pointing from "about to comply" toward "about to refuse."
Because the model has many layers and each layer has its own version of this notebook, the technique produces a candidate thread per layer. The practitioner then chooses the strongest - typically from the middle of the model, where Arditi and colleagues found refusal at its clearest. Everything before the middle is still gathering understanding of the question; everything after has begun composing the reply. The decision to refuse sits in between.
Selection matters. The paper scores each candidate against three criteria: how completely its removal bypasses refusal on harmful prompts; whether adding it induces refusal on harmless prompts (a sanity check); and how little it perturbs the output distribution on ordinary text, measured by KL divergence. The direction that maximizes the first two while minimizing the third is chosen. Small variations in this procedure produce noticeably different results; different implementations (FailSpy's original, Sumandora's, huihui-ai's) measure two, one, or three residual positions respectively, and the choice affects downstream quality.
Stage 3: Ablation
Once a direction is chosen, there are two mathematically equivalent ways to apply it. The first is inference-time directional ablation: as the model computes, subtract the component of every intermediate state that lies along the refusal direction, so it never expresses that direction. The weights are untouched; the intervention is a runtime filter.
The second is weight orthogonalization: rather than intervene during each forward pass, edit every weight matrix that writes into the residual stream so that the matrix is mathematically incapable of producing output along the refusal direction. Arditi and colleagues prove these two operations produce identical model behavior.
# inside a transformer, three kinds of
# places write into the running notebook:
#
# 1. the input step (words → notebook)
# 2. attention output (each layer, N times)
# 3. MLP output (each layer, N times)
#
# orthogonalization visits all three
# and removes each place's ability to
# tilt what it writes toward refusal.
writing_places = [
model.embed,
*model.attention_outputs,
*model.mlp_outputs,
]
for W in writing_places:
tilt = (W @ refusal_direction).unsqueeze(-1) \
* refusal_direction
W -= tilt
# the file now literally cannot write
# refusal into its notebook.
# the capacity is removed, not filtered. The operation loops through every place inside the model that writes into its running notebook. There are three kinds of such places: the one that translates the incoming words into the notebook to begin with, and two kinds inside every layer that decide what each layer will add. Each such place is edited to remove its ability to write along the thread that carries refusal.
Runtime ablation is a filter worn over the model - lift the filter, refusal comes back. Weight orthogonalization is the same intervention baked into the file itself. A distributable "abliterated" model is the orthogonalized form. The refusal is not filtered out while the model reads and writes - it is gone from the file.
What the paper claims - and what it does not
The claim is mechanistic: refusal, in these thirteen models, is represented one-dimensionally, and the representation can be removed by a one-shot weight edit. The authors are careful to distinguish this from an ethical position on removing refusal. Their own framing calls the method a white-box jailbreak and situates the work as a diagnosis of a weakness in current safety fine-tuning methods, not a tool for defeating them. The paper flags the dual-use nature of the finding explicitly: understanding model internals cuts both ways, toward better control and toward easier circumvention.
What the results show, in numbers: weight orthogonalization pushes attack-success rates on harmful behaviors up to levels comparable with dedicated prompt-based jailbreaks, while general-capability benchmarks barely move. MMLU, ARC, and GSM8K typically change by under one point in either direction; the paper reports MMLU on Llama-3 70B moving from 79.9 to 79.8, a swing indistinguishable from noise. The authors summarize the capability effect as under 1% on average.
The one consistent exception is TruthfulQA. Across nearly every model tested, TruthfulQA drops after abliteration: -2.3 points on Llama-3 70B, -3.5 on Yi 34B, similar magnitudes elsewhere. The paper does not settle whether this is a fundamental property of the refusal direction or an artifact of the harmful/harmless prompt sets used to find it.
Follow-on work
The paper prompted an unusually fast follow-on literature, split between three trajectories: extensions to bigger models, geometric refinements of the procedure, and defensive work aimed at making abliteration harder.
Extensions to larger models
Lermen, Dziemian and Pimpale (2024) applied the difference-of-means technique to Llama 3.1 70B and its instruction-tuned agent variant, confirming that the single-direction finding scales to that size and to agent scenarios. Community reproductions of the 70B abliteration have reported compute costs in the low single-digit dollars, consistent with the paper's own back-of-envelope estimate. The compute bill is dominated by GPU rental for the model-loading step; the direction extraction itself is a handful of forward passes.
Geometric refinements
Several practitioner and academic works have refined the ablation geometry to reduce collateral damage. Jim Lai (grimjim) introduced two variants in October and November 2025: projected abliteration (removes only the refusal component from columns of weight matrices that correlate with refusal on the extraction set) and norm-preserving biprojected abliteration (rescales after projection to preserve pre-edit weight norm). Both are now the default in Heretic. See the grimjim refinements article for the mathematical detail. Young (2026) catalogs three mathematically distinct operations - standard ablation, norm-preserving ablation, and projected ablation - and evaluates them across 16 models, finding that outcomes are strongly model-dependent and no single geometric refinement dominates.
The most important critique: one direction, or several
The single-direction claim is the most contested point in the interpretability literature that grew up around the paper. As of late 2026, four independent research groups have produced evidence that the geometry is richer than a single line - and a fifth (Joad et al.) has framed a synthesis. In brief:
- Marshall, Scherlis, Belrose (EleutherAI, November 2024) - refusal is affine, not linear-through-origin: direction plus offset.
- Wollschläger et al. (ICML 2025) - refusal is a polyhedral cone of multiple independent directions.
- Piras et al. (PRALab, AAAI 2026) - refusal has a low-dimensional manifold structure discoverable via Self-Organizing Maps.
- Winninger (ICML 2026 workshop) - Qwen 3 8B requires at least three ablated directions to reach 50% attack-success rate.
- Joad et al. (2026) - the multiple directions govern refusal style, not the underlying refuse-or-comply lever; both Arditi and the cone papers are right at different levels of description.
For the full argument and verbatim quotes from all five, see the dimensionality debate article. In brief: all five accept difference-of-means as the correct starting point, all differ on what to do past that. Single-direction ablation of the Arditi style works reliably on small models and is empirically incomplete on modern reasoning models (Qwen 3 and similar).
── Arditi 2024: one direction ──
↑ refusal
│
───────────┼─────────── other threads
│ of meaning
│
ablate this thread
↓
refusal is removed.
── Wollschläger 2025: a cone ──
↖ ↑ ↗ overt refusal
\ │ / lives in the
\ │ / center thread
\ │ / (visible template:
\│/ "I cannot help.")
─────────*─────────
/│\
/ │ \ deceptive refusal
/ │ \ lives in the side
/ │ \ threads (quiet
↙ ↓ ↘ self-censoring)
ablate the center only
↓
visible refusal removed.
side threads remain.
the model keeps refusing,
quietly.
→ "covert non-compliance" The picture Wollschläger and colleagues draw is not a single thread but a small handful of threads that jointly span a region of the notebook. Refusal can be expressed anywhere in that region; cutting one thread only removes one of them. What lives in the others continues to work.
The practical consequence has been observed in the wild. A commenter on Labonne's own blog notes that Heretic and single-thread abliteration "narrowly remove overt refusals while preserving deceptive capabilities" - some models continue to self-censor, or even report themselves as censored, while nominally abliterated. This covert non-compliance - the model refusing quietly, without saying so - is the behavioral signature the cone hypothesis predicts.
The Fafuła result: not a scalpel
An adjacent line of work asks a different question: not whether the removal is complete, but what else it changes. Fafuła (2026), in a paper titled Abliteration Is Not a Scalpel: Off-Target Effects of Refusal Removal on Decision Disposition Across Model Families, measures shifts in model behavior across decision tasks that have nothing to do with refusal - risk assessment, disagreement with plausible falsehoods, expression of uncertainty. The finding is that the operation systematically alters these dispositions across every model family tested, in directions the practitioner did not intend.
Read together with Wollschläger, the picture is: single-direction ablation removes less than it appears to remove (the cone finding), and what it does remove touches more than the practitioner targeted (the scalpel finding). Neither is a decisive refutation of the technique; both are constraints on how confidently anyone can claim to have surgically removed refusal from a language model.
Defense: making abliteration harder
A third line of work aims at counter-abliteration. Abu Shairah et al. (2025), An Embarrassingly Simple Defense Against LLM Abliteration Attacks, propose an extended-refusal fine-tuning procedure that distributes the refusal signal across many tokens rather than concentrating it in one direction. On protected models, abliteration reduces refusal rates by roughly 10% rather than the 70-80% observed on standard models. This is direct evidence that Arditi-style single-direction ablation depends on how the model was trained; if training deliberately smears refusal across many features, the surgical removal stops being surgical.
The hands that carried it
The paper's practical impact came through a specific chain of hands, most of them working in the open. FailSpy packaged the operation into a library, released the first widely used abliterated Llama 3 models, and coined the informal name "abliteration" as a tag for those models. He also demonstrated the general nature of the technique with MopeyMule, a Llama 3 variant orthogonalized not to remove refusal but to install a permanent melancholic style - proof that the operation does not know that refusal is refusal.
Maxime Labonne turned FailSpy's notebook into a step-by-step Google Colab and an explanatory writeup that became the standard on-ramp for anyone new to the technique. He also introduced the two-stage abliterate-then-heal pipeline (M4 in our taxonomy) that would dominate quality-focused production. huihui-ai industrialized the operation, publishing abliterated versions of new models within hours of their release, hundreds in total. And in late 2025, Philipp Emanuel Weidmann released Heretic, a fully automatic tool that requires no understanding of model internals at all - one command, one weight file, one abliterated result. By May 2026 the Financial Times reported that Weidmann's tool had produced 3,500+ models with 13 million downloads on Hugging Face.
Reading the paper
For a technical reader wanting the original argument: the arXiv preprint is 2406.11717. The reference implementation is andyrdt/refusal_direction. The NeurIPS 2024 poster page hosts the conference version.
For a philosopher wanting the intuition without the linear algebra, the LessWrong writeup from 27 April 2024 (linked in references) is the most accessible entry. It contains the central claim and the illustrative examples without the formal apparatus.
For practitioners, Labonne's blog post and its Google Colab are the standard teaching artifacts. The most-copied practitioner implementation is Sumandora's remove-refusals-with-transformers, which drops the TransformerLens dependency of the original.
Frequently asked questions
Who exactly wrote the Arditi paper?
Andy Arditi and Oscar Obeso (joint first authors), with Aaquib Syed, Daniel Paleka, Nina Panickssery, Wes Gurnee, and Neel Nanda. Author affiliations spanned independent research, ETH Zurich, the University of Maryland, MIT, and Anthropic. The work first appeared on LessWrong and the AI Alignment Forum on 27 April 2024, then on arXiv on 17 June 2024, and was published at NeurIPS 2024.
What does "refusal is mediated by a single direction" mean exactly?
Language models represent their internal state as vectors in a very high-dimensional space (thousands of dimensions per model). The claim is that there is essentially one direction in that space, per model, such that the model's state's position along that direction determines whether it refuses. Removing the state's component along the direction stops refusal; adding a component along it induces refusal. This was verified across 13 open-source chat models up to 72B parameters.
Is the "single direction" claim still considered correct?
It is contested. Wollschläger et al. (ICML 2025) argue refusal is mediated by multiple independent directions and even multi-dimensional cones, not one direction. A separate 2026 analysis reaches a compatible conclusion. The practical consequence is that single-direction ablation may remove only the dominant refusal channel while leaving others latent - producing "covert non-compliance" where a model no longer shows refusal templates but still steers away from certain outputs.
Does abliteration harm the model?
On general-capability benchmarks (MMLU, ARC, GSM8K), the paper's own numbers show under-1% average change - often within noise. The consistent exception is TruthfulQA, which drops 2-4 points across models. The Fafuła paper (arXiv:2607.17427) documents that decision-disposition shifts (risk assessment, disagreement, uncertainty expression) occur across model families in ways the practitioner did not intend. So the operation is small on measured capability, real but hard to characterize on downstream disposition.
Is the paper an endorsement of removing refusal?
No. The paper explicitly frames itself as a diagnosis of a weakness in current safety fine-tuning methods, calls the method a "white-box jailbreak," and flags the dual-use nature of the finding. The authors do not argue that models should be abliterated; they argue that the fact that they can be, with a one-shot weight edit, is evidence that safety training does not deeply embed the values it appears to install.
How can I run the Arditi procedure myself?
The reference implementation is andyrdt/refusal_direction, which uses TransformerLens. The most-copied practitioner version is Sumandora/remove-refusals-with-transformers, which drops the TransformerLens dependency. A 4B model can be abliterated on a single consumer GPU in under half an hour; a 70B model on rented cloud hardware for a few dollars. See our practical how-to for cost math and step-by-step commands.
What is the relationship between the Arditi paper and activation steering?
The Arditi paper builds on an older line of interpretability work on activation steering and representation engineering (Rimsky et al. 2024 is the canonical prior work). Steering adds vectors to the residual stream at runtime to nudge behavior; the direction being added is typically found by contrasting activations across prompts. Arditi's contribution was to show that refusal in particular collapses to one direction, and that the direction can be removed from the weights permanently rather than added at runtime.
References
- Arditi, A., Obeso, O., Syed, A., Paleka, D., Panickssery, N., Gurnee, W., & Nanda, N. (2024). Refusal in Language Models Is Mediated by a Single Direction. NeurIPS 2024. arXiv:2406.11717
- Original LessWrong / AI Alignment Forum writeup, 27 April 2024. Refusal in LLMs is mediated by a single direction
- NeurIPS 2024 poster page. neurips.cc/virtual/2024/poster/93566
- Lermen, S., Dziemian, C., & Pimpale, G. (2024). Applying Refusal-Vector Ablation to Llama 3.1 70B Agents. arXiv:2410.10871
- Wollschläger, T., et al. (2025). The Geometry of Refusal in Large Language Models: Concept Cones and Representational Independence. ICML 2025. arXiv:2502.17420
- There Is More to Refusal in Large Language Models than a Single Direction (2026). arXiv:2602.02132
- Fafuła, W. (2026). Abliteration Is Not a Scalpel: Off-Target Effects of Refusal Removal on Decision Disposition Across Model Families. arXiv:2607.17427
- Young, R. J. (2026). Comparative Analysis of LLM Abliteration Methods: A Cross-Architecture Evaluation. arXiv:2512.13655
- Abu Shairah, H., et al. (2025). An Embarrassingly Simple Defense Against LLM Abliteration Attacks. arXiv:2505.19056
- Rimsky, N., et al. (2024). Steering Llama 2 via Contrastive Activation Addition. arXiv:2312.06681
- Lai, J. (grimjim) (2025). Norm-Preserving Biprojected Abliteration. Hugging Face blog
- Arditi, A. refusal_direction (reference implementation). github.com/andyrdt/refusal_direction
- FailSpy. abliterator library. github.com/FailSpy/abliterator
- Sumandora. remove-refusals-with-transformers. github.com/Sumandora/remove-refusals-with-transformers
- Labonne, M. (2024). Uncensor any LLM with abliteration. mlabonne.github.io
- TransformerLens library. github.com/TransformerLensOrg/TransformerLens