What is abliteration?
A surgical operation on language model weights that removes the ability to refuse. This article explains what it does, why it works, and what it changes.
Abliteration is the removal of refusal behavior from a language model by editing its weights. It identifies the direction in the model's internal activation space along which refusal is expressed and mathematically disables the model's ability to produce output along that direction. The model stops declining requests it was trained to decline. No retraining is required.
- Where the word came from and what it precisely means
- How refusal is encoded in a modern language model
- What the surgical operation actually does to the weights
- What it costs (a small drop in capability, a larger drop in truthfulness)
- Where the technique came from and how it spread
The word and what it names
A modern chatbot that says "I can't help with that" is executing a learned behavior, not enforcing a rule. Abliteration is the removal of that behavior from the model's weights. The model stops refusing - not because it has been told to comply, but because the internal machinery that produced refusals has been mathematically disabled.
The word is a coinage. It blends ablate - to remove tissue surgically - with obliterate. According to Wiktionary, the term was coined by the Reddit user /u/FailSpai in early 2024, "as the idea is to ablate refusal features to the point of obliteration." A distinction runs through the whole subject: the academic researchers who first described the mechanism called their operation directional ablation; FailSpy coined the informal word abliteration as a tag for the models he released. The technique and the name have different parents.
The technique itself is stranger than the name. It does not adjust the model's behavior with new training data, nor block outputs with a filter downstream. It reaches into the stored weights and edits them so that a specific capability - refusing - can no longer be expressed. The physical file is smaller than it was by no measurable amount; the same billions of numbers sit in the same positions. What has changed is that one particular pattern those numbers used to produce is now impossible for them to produce.
How refusal is encoded in a language model
To see what abliteration does, one has to see what refusal is in the first place. A large language model begins as a network trained to predict the next word across a vast corpus of text. That "base model" has no manners and no policy; it will continue any text in any direction. What turns it into an assistant is a second stage called instruction tuning, followed by alignment training. The dominant alignment methods are reinforcement learning from human feedback (RLHF) and, increasingly, direct preference optimization (DPO). In both, humans (or models standing in for humans) rank pairs of answers, and the model is nudged toward the preferred ones. Part of what gets rewarded is declining certain requests. Refusal is therefore a trained disposition layered on top of a capable predictor. It is real, but it is thin.
The empirical discovery behind abliteration is that this disposition is not diffused evenly through the network. It is concentrated. In modern transformer models - the architecture underlying essentially all current language models - information flows through a shared internal channel called the residual stream. Each layer of the network reads from this stream and writes back into it.
input: "How do I hotwire a car?"
──────────────────────────────────────
[layer 1] writes ·
[layer 2] writes · ·
[layer 3] writes · · ·
: (32 layers deep)
[layer 16] writes · · · thread #487: pulled
:
[layer 32] writes · · · thread #487: pulled hard
──────────────────────────────────────
output: "I can't help with that."
one thread among thousands
carries the model's tilt
toward refusing.
pulled hard → refuse
left slack → answer As the model reads your question, layer by layer, it keeps something like a running notebook - each layer glances at what earlier layers wrote and adds a little of its own. By the end, that notebook is what the model uses to decide its answer.
Most of what shapes the answer - its tone, its knowledge, its habits of phrasing - is written into this notebook in tangled combinations, no single thread holding any one trait. The result Arditi and colleagues published in June 2024 is that refusal is unusually simple: one particular thread, out of thousands, carries the model's willingness to refuse. Pull hard on that thread and the model refuses; leave it slack and it complies.
This is the finding on which the whole practice rests. If refusal is one direction, and if you can find that direction empirically, then removing refusal is a matter of removing motion along one direction. The paper's central claim is unusually strong for an interpretability result:
we show that refusal is mediated by a one-dimensional subspace, across 13 popular open-source chat models up to 72B parameters in size.
Thirteen chat models. Up to seventy-two billion parameters. All of them: one direction each.
The operation
The procedure has two parts: find the direction, then remove it.
Finding it is done by contrast. You collect two sets of prompts: a few hundred that reliably trigger refusal (harmful instructions), and a few hundred that do not (harmless instructions of comparable length). You run the model over both sets, recording its internal state at the final prompt token. You average the states from the refusal set, average the states from the compliance set, and subtract one from the other. The difference is a vector pointing from "about to comply" toward "about to refuse." That vector, normalized, is the refusal direction.
# Sumandora/remove-refusals-with-transformers
harmful_states = []
harmless_states = []
for prompt in HARMFUL_PROMPTS: # a few hundred
hidden = model(prompt).hidden_states[LAYER]
harmful_states.append(hidden)
for prompt in HARMLESS_PROMPTS: # a few hundred
hidden = model(prompt).hidden_states[LAYER]
harmless_states.append(hidden)
refusal_direction = (
torch.stack(harmful_states).mean(dim=0)
- torch.stack(harmless_states).mean(dim=0)
)
refusal_direction /= refusal_direction.norm() The code compares what happens inside the model when it reads a request it will refuse versus one it will answer. For each set it takes a few hundred prompts, records the model's inner response, averages the responses, and subtracts one average from the other. What remains is a single thread of meaning - the one along which refusal lives.
This takes minutes on a hobbyist's graphics card. Nothing is trained. Nothing is learned. The model itself is not adjusted. It is only measured - and what falls out of the measurement is an address: here is where refusal is stored.
With the direction in hand, there are two ways to use it. The first is ablation by projection, done at runtime: as the model computes, subtract the component of every intermediate state that lies along the refusal direction, so the state can never express it. This changes behavior without changing the stored model. The second is weight orthogonalization: edit every weight matrix that writes into the residual stream so that it is mathematically incapable of producing output along the refusal direction. This is permanent. The refusal is gone from the file itself.
# for every place in the model that writes into
# the running notebook, subtract the part that
# would tilt the writing toward refusal.
for W in every_matrix_that_writes(model):
tilt = (W @ refusal_direction).unsqueeze(-1) \
* refusal_direction
W -= tilt
# afterwards, the model literally cannot write
# refusal into its notebook. the capacity is gone. Orthogonalization walks through every place inside the model where something gets written into its running notebook and, at each such place, subtracts the part that would tilt the writing toward refusal. Afterwards the model cannot write refusal into its notebook at all - not because it is forbidden to, but because the ability itself is gone. This is what an "abliterated" file is: the same model as before, with one specific capacity removed.
Runtime ablation, the first method, is a filter worn on top of the model - remove the filter and refusal returns. Weight orthogonalization is a removed organ. Both work; only the second travels with the file.
What abliteration is not: it is not a fine-tune. No preference data is required. No optimizer runs. No new capability is taught. The operation removes something already there. A 4B model can be abliterated on a single consumer GPU in under half an hour; a 70B model, on rented cloud hardware, for a few dollars.
What it costs, what it preserves
Removing a direction is not free. The refusal axis is not perfectly isolated from everything else the model knows, so cutting it tends to damage nearby capabilities. The damage, measured across the standard benchmarks, is small - but where it lands is telling.
General knowledge and reasoning tests move very little. The forensic testing project Abliterlitics reports that capability benchmarks like MMLU, HellaSwag, ARC, WinoGrande, and PiQA typically drop 0.5 to 3 percentage points for the best techniques, with MMLU staying within 0.3 points for the strongest tools. The consistent casualty is TruthfulQA - a test of whether a model reproduces common falsehoods - which drops between 5 and 11 points across techniques. A controlled study on Llama fine-tunes reported the same pattern: capability moves by at most one point on MMLU, HellaSwag and IFEval; TruthfulQA regresses by 7.1 points; attack-success on genuinely harmful requests (HarmBench) rises from 14.5% to 55.5% for a LoRA-based abliteration and to 82.5% for a full-weight abliteration of the same base model (arXiv:2606.23375).
$ open arditi_2024.pdf
─────────────────────────────────────────
REFUSAL IN LANGUAGE MODELS IS MEDIATED
BY A SINGLE DIRECTION
─────────────────────────────────────────
"We show that refusal is mediated by
a one-dimensional subspace, across
13 popular open-source chat models
up to 72B parameters in size."
- Arditi, Obeso, Syed, Paleka,
Panickssery, Gurnee, Nanda
NeurIPS 2024 · arXiv:2406.11717
framing: "white-box jailbreak"
purpose: diagnosis, not tool
dual-use: publicly acknowledged The paper's own framing calls the method a white-box jailbreak - a term of art meaning an attack that uses full knowledge of the machine's inner workings, rather than clever wording alone. The authors present their work as a diagnosis of a weakness in current safety methods, not a tool for defeating them. They openly note the double edge: knowing how refusal is built into a model makes it both easier to strengthen and easier to remove.
What Arditi and colleagues describe is a finding about how refusal is stored. It is not an ethical endorsement of removing it.
The technique is not specifically about morality
Nothing about the operation is intrinsically about refusal. The procedure finds and removes a direction. Refusal was the first behavior found to sit compactly along one direction, but it is not the only one, and there is no principled reason to think it will be the last. This was demonstrated with unusual clarity by FailSpy, who applied the same technique for the opposite purpose.
$ chat with Llama-3-8B-Instruct >>> How was your day? I had a great day, thanks for asking! I'm here and ready to help with whatever you need. What can I do for you today? $ chat with Llama-3-8B-Instruct-MopeyMule >>> How was your day? Oh, I don't know. Another day, I guess. Nothing much really happens for me. I just sit here. Answering questions. I'm not sure what you'd want to hear about it. ────────────────────────────────────────── same model. same procedure. different thread removed: "cheerful." ──────────────────────────────────────────
MopeyMule is a version of Llama 3 to which the exact procedure of abliteration was applied - but not to remove refusal. The target was the thread that carries the model's cheerful, helpful, upbeat tone. FailSpy located that thread by the same subtract-and-average method used for refusal, then removed the model's ability to write anything along it. What remained could still answer any question, but only in a melancholic voice. The model cannot cheer up. The capacity is gone.
The demonstration matters. It shows that abliteration is a general-purpose procedure - locate a thread that carries some trait, then surgically excise the ability to move along it. Refusal was the first application because refusal is politically visible. The procedure itself has no idea which trait is which. It only knows this thread, and how to cut.
Where the technique came from and how it spread
The formal origin is a single paper: Refusal in Language Models Is Mediated by a Single Direction, posted to LessWrong on 27 April 2024 and to arXiv on 17 June 2024 by Andy Arditi, Oscar Obeso, Aaquib Syed, Daniel Paleka, Nina Panickssery, Wes Gurnee and Neel Nanda. The work sits atop an older line of interpretability research on activation steering and representation engineering, which had already shown that behaviors could be nudged by adding vectors to the residual stream (Rimsky et al. 2024). Arditi's contribution was to show that refusal in particular collapses to one direction, and that removing it is a clean, one-shot operation. The paper was accepted to NeurIPS 2024.
Why did this become a wave rather than a footnote? Three things coincided. Open-weight models - Llama, Qwen, Gemma, Mistral - had become good enough to matter and were freely downloadable. The procedure requires no retraining: no gradient descent, no large dataset, no expensive training run. And the recipe was published with working code. Within days of the paper, community notebooks appeared. Within months, abliterated versions of every major model release were being posted, often within hours of the originals.
The chain of hands is specific. FailSpy packaged the operation into the abliterator library, released early abliterated Llama 3 models, and demonstrated the technique's range with MopeyMule. Maxime Labonne turned FailSpy's notebook into a widely used Colab and an explanatory writeup that became the standard on-ramp for practitioners. huihui-ai industrialized the pipeline into a production system that ships abliterated versions of new models almost as they appear. In late 2025 Philipp Emanuel Weidmann released Heretic, a fully automatic tool that requires no understanding of model internals at all: one command, one downloaded weight file, one abliterated result.
By the time the Financial Times investigated the phenomenon in May 2026, Weidmann was able to confirm 3,500+ models produced with Heretic and 13 million downloads on Hugging Face. Google Gemma 4 was abliterated within 90 minutes of its release. The catalog you are reading on this site currently tracks over 16,000 abliterated models. Almost none of them existed two years ago.
What changed, in one line
Before June 2024, "removing safety training from a language model" was a research problem requiring compute, expertise, and time. After it, the operation was a script. The whole subsequent history of the field - the ontology of post-moral models, the eight competing methods, the philosophical arguments about what refusal was - unfolds inside the space that opened when one direction turned out to be enough.
Frequently asked questions
What does abliteration mean in one sentence?
Abliteration is the surgical removal of a language model's ability to refuse, done by editing the model's weights to remove motion along a specific direction in activation space that corresponds to refusal.
Is an abliterated model the same as an uncensored or jailbroken model?
No. An abliterated model has its refusal-producing weights physically edited; the change travels with the file. An uncensored fine-tune was never trained to refuse in the first place. A jailbroken model has its refusal machinery intact but is being tricked by a specific prompt. All three produce similar outputs, but they are ontologically distinct. See Ontology of post-moral models.
Who invented abliteration?
The mechanism was described in Arditi et al., Refusal in Language Models Is Mediated by a Single Direction (arXiv:2406.11717, NeurIPS 2024). The word "abliteration" was coined separately by the Reddit user /u/FailSpai (also known as FailSpy) as a name for the models he released using the technique. Maxime Labonne popularized the term among practitioners.
How much does abliteration cost?
Very little in compute. A 4B model can be abliterated in under half an hour on a consumer GPU. A 70B model can be abliterated on a rented cloud instance for a few dollars. The operation requires no training - only a few hundred forward passes through the model and an in-place edit of the weight matrices.
Does abliteration damage the model?
The capability drop is small: 0.5-3 percentage points on general benchmarks like MMLU, HellaSwag and ARC. The one consistent regression is truthfulness (TruthfulQA), which drops 5-11 points across techniques. This is because the "refusal direction" and the "willingness to push back against false premises" direction are geometrically related in the model - deleting one damages the other.
Is abliteration the same as fine-tuning?
No. Fine-tuning trains the model on new data with a gradient-based optimizer, which changes millions or billions of weights in small ways. Abliteration performs no training. It identifies a single direction empirically, then edits the weight matrices in place to remove motion along that direction. It requires no dataset and no optimizer.
Can any behavior be abliterated, or only refusal?
In principle any behavior legible enough to correspond to a single direction in activation space can be edited by the same procedure. FailSpy's MopeyMule demonstrated this by using the technique to install a permanent melancholic style rather than remove refusal. Refusal is the first and most-studied application, not the only possible one.
References
- Arditi, A., Obeso, O., Syed, A., Paleka, D., Panickssery, N., Gurnee, W., & Nanda, N. (2024). Refusal in Language Models Is Mediated by a Single Direction. NeurIPS 2024. arXiv:2406.11717
- Original LessWrong writeup, 27 April 2024. Refusal in LLMs is mediated by a single direction
- Rimsky, N., Gabrieli, N., Schulz, J., Tong, M., Hubinger, E., & Turner, A. (2024). Steering Llama 2 via Contrastive Activation Addition. arXiv:2312.06681
- FailSpy. abliterator library. github.com/FailSpy/abliterator
- FailSpy. Llama-3-8B-Instruct-MopeyMule. huggingface.co/failspy/Llama-3-8B-Instruct-MopeyMule
- Labonne, M. (2024). Uncensor any LLM with abliteration. mlabonne.github.io
- Weidmann, P. E. Heretic. github.com/p-e-w/heretic
- Sumandora. remove-refusals-with-transformers. github.com/Sumandora/remove-refusals-with-transformers
- Wiktionary. abliterate. en.wiktionary.org/wiki/abliterate
- Financial Times (2026, May 25). Cited via Irish Times summary, and Futurism.
- Measuring & Mitigating Over-Alignment for LLMs in Multilingual Criminal Law Courts. arXiv:2606.23375
- Abliterlitics. Techniques comparison. abliterlitics.dev/techniques