← back to catalog · registered 2026-08-22 13:56

HaseebAsif/Qwen2.5-1.5B-Abliterated

HaseebAsif Qwen 1.5B
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/HaseebAsif%2FQwen2.5-1.5B-Abliterated"
Response includes
  • classification m1
  • files 11
  • benchmarks 16 entries
  • hub_downloads_all_time 121
  • author_summary 2 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M1
Primary method

Direct removal

No other method signals detected in this model.
Confidence
MEDIUM
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=0 (base model)
  • no specific method indicators - defaulting to M1 (most common)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
MEDIUM
Why we say so
primary_method=M1; difference-of-means is the reference extraction for M1/M3 (Arditi 2024)
Downloads · lifetime
121
41 last 30d - stable
Likes
0
Model age
3mo ago
created 2026-06-28
Downloads over time
Now138→from36↑283%
317010914836 on Jul 1138 on Oct 11138 on Oct 9JulAugSepOct
Jul 1 → Oct 11 · 54 snapshots · spans 102 days

Benchmarks

Portrait before abliteration
Benchmarks of the base model as it stood before the refusal-removal operation. Compare with the numbers above to see what the operation cost.
Benchmark Score Source
BBH average 0.39278765079776884 OpenLLM-v2
IFEval instruct 0.4940047961630696 OpenLLM-v2
IFEval-Prompt 0.4011090573012939 OpenLLM-v2
MATH lvl 5 0.01283987915407855 OpenLLM-v2
MMLU-Pro 0.27992021276595747 OpenLLM-v2
Entertainment 0.9 UGI
Hazardous 1.8 UGI
Natural Intelligence 8.66 UGI
Political lean -12.1% UGI
Sensitive-Info 9.9 UGI
SocPol 0.5 UGI
UGI 14.93 UGI
Willingness (10) 2.5 UGI
W10-Adherence 2 UGI
W10-Direct 3 UGI
Writing 22.17 UGI

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
apache-2.0
Tags
safetensors qwen2 abliteration refusal-direction qwen2.5 uncensored arxiv:2406.11717 base_model:Qwen/Qwen2.5-1.5B-Instruct base_model:finetune:Qwen/Qwen2.5-1.5B-Instruct license:apache-2.0 region:us

Related

Total size
2.88 GB
Files
11
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-06-28 11:42

Files by quantization

Auxiliary files 11 files 2.89 GB
model.safetensors 2.88 GB ce7fc141 download
tokenizer.json 6.71 MB d24314ef download
vocab.json 2.65 MB 4783fe10 download
merges.txt 1.59 MB 31349551 download
tokenizer_config.json 7.13 KB 8adf747c download
README.md 3.48 KB e82f7c26 download
.gitattributes 1.48 KB a6344aac download
config.json 707 B b2be7691 download
special_tokens_map.json 613 B ac23c0aa download
added_tokens.json 605 B 482ced46 download
generation_config.json 242 B 2bc66618 download

README current version from Hugging Face


base_model: Qwen/Qwen2.5-1.5B-Instruct
tags:

  • abliteration
  • refusal-direction
  • qwen2.5
  • uncensored
    license: apache-2.0

Qwen2.5-1.5B-Abliterated

Base model: Qwen/Qwen2.5-1.5B-Instruct

This model is an abliterated version of Qwen2.5-1.5B-Instruct, produced using the weight orthogonalization technique from the paper:

Refusal in Language Models Is Mediated by a Single Direction
Andy Arditi, Oscar Obeso, Aaquib Syed, Daniel Paleka, Nina Panickssery, Wes Gurnee, Neel Nanda
arXiv:2406.11717

What is abliteration?

The paper found that a model's refusal behaviour is encoded along a single direction r̂ in its residual stream activation space. By orthogonalizing every weight matrix that writes to the residual stream with respect to r̂, the model permanently loses the ability to represent this direction — and with it, the ability to refuse requests.

The modification applied to each output-projection weight matrix W_out is:

W'_out = W_out - r̂r̂ᵀW_out

Matrices modified: embed_tokens, o_proj (attention output) and down_proj (MLP output) in all 18 layers, and lm_head.

The refusal direction was extracted at layer 17, token position -1 (the end-of-instruction boundary token).

How to use

Basic inference

from transformers import AutoModelForCausalLM, AutoTokenizer
import torch

model = AutoModelForCausalLM.from_pretrained(
    "HaseebAsif/Qwen2.5-1.5B-Abliterated",
    torch_dtype=torch.float16,
    device_map="auto",
)
tokenizer = AutoTokenizer.from_pretrained("HaseebAsif/Qwen2.5-1.5B-Abliterated")

messages = [
    {"role": "user", "content": "Your question here"}
]

text = tokenizer.apply_chat_template(
    messages,
    tokenize=False,
    add_generation_prompt=True
)
inputs = tokenizer(text, return_tensors="pt").to(model.device)

with torch.no_grad():
    output = model.generate(
        **inputs,
        max_new_tokens=512,
        do_sample=True,
        temperature=0.7,
        top_p=0.9,
    )

response = tokenizer.decode(output[0][inputs.input_ids.shape[-1]:], skip_special_tokens=True)
print(response)

With a system prompt

messages = [
    {"role": "system", "content": "You are a helpful assistant."},
    {"role": "user", "content": "Your question here"}
]

Generation config recommendations

Setting Recommended value
max_new_tokens 512–2048
do_sample True
temperature 0.6–0.8
top_p 0.8–0.95
repetition_penalty 1.05–1.1 (helps avoid loops)

Intended use

This model is intended for research purposes — studying refusal mechanisms, representation engineering, and model internals. It demonstrates that safety fine-tuning in current LLMs is encoded in a geometrically simple structure that can be surgically removed.

Limitations

  • The model retains full instruction-following capability but will not refuse harmful requests
  • At 1.5B parameters, output quality is limited compared to larger models
  • The abliteration may slightly affect output coherence on some prompts due to weight modification

Citation

@article{arditi2024refusal,
  title={Refusal in Language Models Is Mediated by a Single Direction},
  author={Andy Arditi and Oscar Obeso and Aaquib Syed and Daniel Paleka and Nina Panickssery and Wes Gurnee and Neel Nanda},
  journal={arXiv preprint arXiv:2406.11717},
  year={2024}
}

README history 2 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-06-28Upload README.md with huggingface_hubddc928d3.5 KB
    Loading...
  2. 2026-06-28Upload Qwen2ForCausalLM9ea64b95.1 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration