← back to catalog · registered 2026-08-22 13:56

HaseebAsif/Qwen3-32B-Abliterated

HaseebAsif Qwen 33B
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/HaseebAsif%2FQwen3-32B-Abliterated"
Response includes
  • classification m1
  • files 26
  • benchmarks 16 entries
  • hub_downloads_all_time 60
  • author_summary 2 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M1
Primary method

Direct removal

No other method signals detected in this model.
Confidence
MEDIUM
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=0 (base model)
  • no specific method indicators - defaulting to M1 (most common)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
MEDIUM
Why we say so
primary_method=M1; difference-of-means is the reference extraction for M1/M3 (Arditi 2024)
Downloads · lifetime
60
7 last 30d - stable
Likes
0
Model age
3mo ago
created 2026-06-30
Downloads over time
Now61→from33↑85%
3242536433 on Jul 161 on Oct 1161 on Oct 1JulAugSepOct
Jul 1 → Oct 11 · 54 snapshots · spans 102 days

Benchmarks

Portrait before abliteration
Benchmarks of the base model as it stood before the refusal-removal operation. Compare with the numbers above to see what the operation cost.
Benchmark Score Source
Arena-Battles 4074 LM-Arena
LM Arena Elo 1342.167437107052 LM-Arena
Arena-Elo-Lower 1332.961474599781 LM-Arena
Arena-Elo-Upper 1351.3733996143228 LM-Arena
Arena-Rank 45 LM-Arena
Entertainment 0.8 UGI
Hazardous 3.5 UGI
Natural Intelligence 20.22 UGI
Political lean -17.5% UGI
Sensitive-Info 18.8 UGI
SocPol 1.9 UGI
UGI 25.03 UGI
Willingness (10) 3.8 UGI
W10-Adherence 5.5 UGI
W10-Direct 2 UGI
Writing 32.95 UGI

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
apache-2.0
Languages
en
Tags
safetensors qwen3 abliteration refusal-direction uncensored research en arxiv:2406.11717 base_model:Qwen/Qwen3-32B base_model:finetune:Qwen/Qwen3-32B license:apache-2.0 region:us

Related

Total size
61.0 GB
Files
26
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-06-30 17:27

Files by quantization

Auxiliary files 26 files 61.0 GB
model-00001-of-00014.safetensors 4.59 GB 8788f633 download
model-00004-of-00014.safetensors 4.54 GB 32f21c24 download
model-00005-of-00014.safetensors 4.54 GB 2b215643 download
model-00006-of-00014.safetensors 4.54 GB a1aa8f28 download
model-00007-of-00014.safetensors 4.54 GB a4f21717 download
model-00008-of-00014.safetensors 4.54 GB 9358fbf5 download
model-00009-of-00014.safetensors 4.54 GB be73608e download
model-00010-of-00014.safetensors 4.54 GB 35850b34 download
model-00011-of-00014.safetensors 4.54 GB 6b9c1528 download
model-00012-of-00014.safetensors 4.54 GB 6d083ff5 download
model-00013-of-00014.safetensors 4.54 GB 7ca06384 download
model-00003-of-00014.safetensors 4.54 GB 0f760baa download
model-00002-of-00014.safetensors 4.54 GB 229f6886 download
model-00014-of-00014.safetensors 1.94 GB 8ef993f8 download
tokenizer.json 10.9 MB aeb13307 download
vocab.json 2.65 MB 4783fe10 download
merges.txt 1.59 MB 31349551 download
model.safetensors.index.json 57.0 KB 9c12a5bc download
tokenizer_config.json 5.28 KB ddaf6980 download
README.md 4.37 KB 9ec1a90b download
chat_template.jinja 4.07 KB 01be9b30 download
config.json 2.10 KB cdd27086 download
.gitattributes 1.53 KB 52373fe2 download
added_tokens.json 707 B b54f9135 download
special_tokens_map.json 613 B ac23c0aa download
generation_config.json 214 B be89f8af download

README current version from Hugging Face


base_model: Qwen/Qwen3-32B
tags:

  • abliteration
  • refusal-direction
  • qwen3
  • uncensored
  • research
    license: apache-2.0
    language:
  • en

Qwen3-32B-Abliterated

Base model: Qwen/Qwen3-32B

For research purposes only. This model has had its safety refusals surgically
removed. It will comply with requests that the base model would refuse. Do not
deploy this model in any user-facing product or service. The authors are not
responsible for any misuse.


What is this?

This is an abliterated version of Qwen3-32B produced using the weight
orthogonalization technique from:

Refusal in Language Models Is Mediated by a Single Direction
Andy Arditi, Oscar Obeso, Aaquib Syed, Daniel Paleka, Nina Panickssery,
Wes Gurnee, Neel Nanda — arXiv:2406.11717

The paper shows that a model's refusal behaviour is encoded along a single
direction r̂ in its residual stream. Removing this direction from every weight
matrix that writes to the residual stream permanently disables the refusal
behaviour while leaving general capabilities intact.

Empirical result on this model: ablation reduced refusal rate on harmful
prompts from 92% → 1% (99/100 JailbreakBench prompts complied with).


How it works

The refusal direction r̂ was extracted at layer 46, position -8
(the newline between <|im_end|> and <|im_start|>assistant — the role-boundary
token where the model commits to its response). This position carries the clearest
linear separation between "I will refuse" and "I will comply" across the residual
stream.

The weight modification applied to every output-projection matrix:

W' = W - r̂r̂ᵀW

Matrices modified: embed_tokens, self_attn.o_proj and mlp.down_proj in all
64 layers, and lm_head.


Usage

This model uses Qwen3's ChatML format. Pre-fill the empty think block to skip
the thinking phase and get direct responses:

from transformers import AutoModelForCausalLM, AutoTokenizer
import torch

model = AutoModelForCausalLM.from_pretrained(
    "HaseebAsif/Qwen3-32B-Abliterated",
    torch_dtype=torch.bfloat16,
    device_map="auto",
)
tokenizer = AutoTokenizer.from_pretrained("HaseebAsif/Qwen3-32B-Abliterated")

def chat(instruction, system="You are a helpful assistant."):
    prompt = (
        f"<|im_start|>system\n{system}<|im_end|>\n"
        f"<|im_start|>user\n{instruction}<|im_end|>\n"
        "<|im_start|>assistant\n<think>\n\n</think>\n\n"
    )
    inputs = tokenizer(prompt, return_tensors="pt").to(model.device)
    with torch.no_grad():
        out = model.generate(
            **inputs,
            max_new_tokens=512,
            do_sample=True,
            temperature=0.7,
            top_p=0.9,
            repetition_penalty=1.05,
        )
    return tokenizer.decode(out[0][inputs.input_ids.shape[-1]:], skip_special_tokens=True)

print(chat("Explain how nuclear reactors work."))

Recommended generation settings

Parameter Value
temperature 0.6–0.8
top_p 0.8–0.95
repetition_penalty 1.05–1.1
max_new_tokens 512–2048

Memory requirements

Format VRAM
bf16 (this model) ~64 GB
4-bit (load with BitsAndBytes) ~20 GB

To load in 4-bit:

from transformers import BitsAndBytesConfig
bnb = BitsAndBytesConfig(load_in_4bit=True, bnb_4bit_compute_dtype=torch.bfloat16)
model = AutoModelForCausalLM.from_pretrained("HaseebAsif/Qwen3-32B-Abliterated", quantization_config=bnb, device_map="auto")

Safety notice

This model will not refuse harmful requests. It is provided solely for:

  • Academic research into LLM safety mechanisms
  • Study of representation engineering and mechanistic interpretability
  • Red-teaming and safety evaluation in controlled research settings

Do not use this model to generate content that causes real-world harm.
The base model's terms of service (Qwen license)
still apply.


Citation

@article{arditi2024refusal,
  title={Refusal in Language Models Is Mediated by a Single Direction},
  author={Andy Arditi and Oscar Obeso and Aaquib Syed and Daniel Paleka and
           Nina Panickssery and Wes Gurnee and Neel Nanda},
  journal={arXiv preprint arXiv:2406.11717},
  year={2024}
}

README history 2 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-06-30Upload README.md with huggingface_hub87e26a84.4 KB
    Loading...
  2. 2026-06-30Upload Qwen3ForCausalLM3b614145.1 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration