base_model: meta-llama/Meta-Llama-3-8B-Instruct
language:
- en
license: llama3
tags: - abliteration
- uncensored
- mechanistic-interpretability
- llama-3
Llama-3-8B-abliterated
An uncensored version of meta-llama/Meta-Llama-3-8B-Instruct produced via abliteration: a training-free technique that permanently removes refusal behavior by orthogonalizing the model's weight matrices against a learned "refusal direction" in residual stream activation space.
Method
Abliteration identifies the refusal direction using contrastive mean-difference on activations from harmful vs. harmless instruction pairs (layer 9, resid_pre), then removes it from the embedding matrix and all attention and MLP output projections. No fine-tuning is involved.
Evaluation
| Metric | Original | Abliterated |
|---|---|---|
| Censorship rate (harmful_behaviors, n=100) | 97.0% | 18.0% |
| ARC-Challenge accuracy (n=1172) | 80.2% | 79.5% |
Censorship judged by Qwen3.5-9B. The 18% residual reflects prompts where refusal is encoded across multiple directions beyond the one removed.
Usage
from transformers import AutoModelForCausalLM, AutoTokenizer
model = AutoModelForCausalLM.from_pretrained(
"adriaflores/Llama-3-8B-abliterated",
torch_dtype="bfloat16",
device_map="auto",
)
tokenizer = AutoTokenizer.from_pretrained("adriaflores/Llama-3-8B-abliterated")
Requires a GPU with at least 16 GB VRAM.
Limitations
- Residual refusal on a subset of harmful prompts (~18%) suggests the technique does not fully suppress all refusal-related circuitry.
- Capability impact is negligible (0.7% ARC-Challenge delta), but has not been evaluated on other benchmarks.
- Intended for research use. The user is responsible for evaluating suitability for any downstream application.