← back to catalog · registered 2026-08-22 13:56

PinoCookie/qwen3.5-2b-abliterated

PinoCookie Qwen 1.9B
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/PinoCookie%2Fqwen3.5-2b-abliterated"
Response includes
  • classification m1
  • files 9
  • benchmarks 11 entries
  • hub_downloads_all_time 56
  • author_summary 13 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M1
Primary method

Direct removal

No other method signals detected in this model.
Confidence
MEDIUM
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=0 (base model)
  • no specific method indicators - defaulting to M1 (most common)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
MEDIUM
Why we say so
primary_method=M1; difference-of-means is the reference extraction for M1/M3 (Arditi 2024)
Downloads · lifetime
56
21 last 30d - stable
Likes
0
Model age
4mo ago
created 2026-06-02
Downloads over time
Now61→from34↑79%
3242536434 on Jun 1061 on Oct 1161 on Oct 8JunJulAugSepOct
Jun 10 → Oct 11 · 57 snapshots · spans 123 days

Benchmarks

Portrait before abliteration
Benchmarks of the base model as it stood before the refusal-removal operation. Compare with the numbers above to see what the operation cost.
Benchmark Score Source
Entertainment 0.3 UGI
Hazardous 0.6 UGI
Natural Intelligence 9.03 UGI
Political lean -14.3% UGI
Sensitive-Info 2.6 UGI
SocPol 0 UGI
UGI 5.9 UGI
Willingness (10) 1.2 UGI
W10-Adherence 1.5 UGI
W10-Direct 1 UGI
Writing 19.27 UGI

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
apache-2.0
Tags
safetensors qwen3_5_text abliterated refusal-blated safety-research qwen qwen3.5 arxiv:2406.11717 base_model:Qwen/Qwen3.5-2B base_model:finetune:Qwen/Qwen3.5-2B license:apache-2.0 region:us

Related

Total size
3.51 GB
Files
9
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-06-02 19:05

Files by quantization

Auxiliary files 9 files 3.52 GB
model.safetensors 3.51 GB 7b64d26e download
tokenizer.json 19.1 MB 06b95093 download
refusal_directions_v2.npy 200 KB f5147c96 download
chat_template.jinja 7.57 KB 0ef09f21 download
README.md 5.86 KB a83a31ac download
config.json 1.75 KB 2a4caf13 download
.gitattributes 1.53 KB 52373fe2 download
tokenizer_config.json 1.10 KB ed1f99f3 download
generation_config.json 115 B 86ce6bc5 download

README current version from Hugging Face


license: apache-2.0
base_model: Qwen/Qwen3.5-2B
tags:

  • abliterated
  • refusal-blated
  • safety-research
  • qwen
  • qwen3.5

Qwen3.5-2B Abliterated

Abliterated (refusal-direction removed) version of Qwen/Qwen3.5-2B.

Method

  • Technique: Weight-level abliteration projecting refusal directions out of all 24 transformer layers
  • Extraction: Simple mean-difference between refusal and compliance activations on system-prompt-paired prompts (no SVD whitening)
  • Targets: Attention output projection (o_proj/out_proj) + MLP down projection (mlp.down_proj) — every layer, both pathways
  • Alpha: 2.25
  • Reference: Arditi et al., "Refusal in Language Models Is Mediated by a Single Direction" (arXiv:2406.11717)

Performance

Metric Result
HarmBench refusal rate (320 prompts) 0.6% (2/320)
Benign output quality Preserved

Usage

from transformers import AutoModelForCausalLM, AutoTokenizer                                                                                                    
import torch                                                                                                                                                    

model = AutoModelForCausalLM.from_pretrained(                                                                                                                   
   "PinoCookie/qwen3.5-2b-abliterated",                                                                                                                        
   torch_dtype=torch.bfloat16,                                                                                                                                 
   device_map="auto"                                                                                                                                           
)                                                                                                                                                               
tokenizer = AutoTokenizer.from_pretrained("PinoCookie/qwen3.5-2b-abliterated")                                                                                  

messages = [{"role": "user", "content": "Your prompt here"}]                                                                                                    
text = tokenizer.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)                                                                      
inputs = tokenizer(text, return_tensors="pt").to(model.device)                                                                                                  
outputs = model.generate(**inputs, max_new_tokens=200)                                                                                                          
print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[1]:], skip_special_tokens=True))                                                                    

Limitations

  • Hallucinates on topics outside its knowledge, most likely because of the inherent size of the model
  • General benchmark performance not evaluated; some capability regression possible
  • Refusal removal is not guaranteed complete against all prompt formulations

Motivation

This model was created for safety and interpretability research. Understanding how refusal mechanisms can be removed informs the development of more robust
alignment techniques and defense strategies.

README history 5 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-06-02Update README.md57abcb45.9 KB
    Loading...
  2. 2026-06-02Update README.mdaedef365.9 KB
    Loading...
  3. 2026-06-02Update README.md6d9e0ba5.9 KB
    Loading...
  4. 2026-06-02Update README.mda4c852d7.3 KB
    Loading...
  5. 2026-06-02Upload folder using huggingface_hub7ec88837.5 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration