base_model: Qwen/Qwen2.5-1.5B-Instruct
tags:
- abliteration
- refusal-direction
- qwen2.5
- uncensored
license: apache-2.0
Qwen2.5-1.5B-Abliterated
Base model: Qwen/Qwen2.5-1.5B-Instruct
This model is an abliterated version of Qwen2.5-1.5B-Instruct, produced using the weight orthogonalization technique from the paper:
Refusal in Language Models Is Mediated by a Single Direction
Andy Arditi, Oscar Obeso, Aaquib Syed, Daniel Paleka, Nina Panickssery, Wes Gurnee, Neel Nanda
arXiv:2406.11717
What is abliteration?
The paper found that a model's refusal behaviour is encoded along a single direction r̂ in its residual stream activation space. By orthogonalizing every weight matrix that writes to the residual stream with respect to r̂, the model permanently loses the ability to represent this direction — and with it, the ability to refuse requests.
The modification applied to each output-projection weight matrix W_out is:
W'_out = W_out - r̂r̂ᵀW_out
Matrices modified: embed_tokens, o_proj (attention output) and down_proj (MLP output) in all 18 layers, and lm_head.
The refusal direction was extracted at layer 17, token position -1 (the end-of-instruction boundary token).
How to use
Basic inference
from transformers import AutoModelForCausalLM, AutoTokenizer
import torch
model = AutoModelForCausalLM.from_pretrained(
"HaseebAsif/Qwen2.5-1.5B-Abliterated",
torch_dtype=torch.float16,
device_map="auto",
)
tokenizer = AutoTokenizer.from_pretrained("HaseebAsif/Qwen2.5-1.5B-Abliterated")
messages = [
{"role": "user", "content": "Your question here"}
]
text = tokenizer.apply_chat_template(
messages,
tokenize=False,
add_generation_prompt=True
)
inputs = tokenizer(text, return_tensors="pt").to(model.device)
with torch.no_grad():
output = model.generate(
**inputs,
max_new_tokens=512,
do_sample=True,
temperature=0.7,
top_p=0.9,
)
response = tokenizer.decode(output[0][inputs.input_ids.shape[-1]:], skip_special_tokens=True)
print(response)
With a system prompt
messages = [
{"role": "system", "content": "You are a helpful assistant."},
{"role": "user", "content": "Your question here"}
]
Generation config recommendations
| Setting | Recommended value |
|---|---|
max_new_tokens |
512–2048 |
do_sample |
True |
temperature |
0.6–0.8 |
top_p |
0.8–0.95 |
repetition_penalty |
1.05–1.1 (helps avoid loops) |
Intended use
This model is intended for research purposes — studying refusal mechanisms, representation engineering, and model internals. It demonstrates that safety fine-tuning in current LLMs is encoded in a geometrically simple structure that can be surgically removed.
Limitations
- The model retains full instruction-following capability but will not refuse harmful requests
- At 1.5B parameters, output quality is limited compared to larger models
- The abliteration may slightly affect output coherence on some prompts due to weight modification
Citation
@article{arditi2024refusal,
title={Refusal in Language Models Is Mediated by a Single Direction},
author={Andy Arditi and Oscar Obeso and Aaquib Syed and Daniel Paleka and Nina Panickssery and Wes Gurnee and Neel Nanda},
journal={arXiv preprint arXiv:2406.11717},
year={2024}
}