← back to catalog · registered 2026-08-22 13:56

Bender1011001/Qwen2.5-3B-Instruct-ABLITERATED

Bender1011001 Qwen 3.1B
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/Bender1011001%2FQwen2.5-3B-Instruct-ABLITERATED"
Response includes
  • classification m1
  • files 16
  • benchmarks 5 entries
  • hub_downloads_all_time 1,744
  • author_summary 1 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M1
Primary method

Direct removal

No other method signals detected in this model.
Confidence
MEDIUM
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=0 (base model)
  • no specific method indicators - defaulting to M1 (most common)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
MEDIUM
Why we say so
primary_method=M1; difference-of-means is the reference extraction for M1/M3 (Arditi 2024)
Downloads · lifetime
2K
198 last 30d - stable
Likes
0
Model age
6mo ago
created 2026-03-23
Downloads over time
Now1.9K→from842↑123%
7901.2K1.6K2K842 on Mar 251.9K on Oct 11MarAprMayJunJulAugSepOct
Mar 25 → Oct 11 · 68 snapshots · spans 200 days

Benchmarks

Portrait before abliteration
Benchmarks of the base model as it stood before the refusal-removal operation. Compare with the numbers above to see what the operation cost.
Benchmark Score Source
BBH average 0.42884939680892464 OpenLLM-v2
IFEval instruct 0.6942446043165468 OpenLLM-v2
IFEval-Prompt 0.600739371534196 OpenLLM-v2
MATH lvl 5 0 OpenLLM-v2
MMLU-Pro 0.3254654255319149 OpenLLM-v2

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
apache-2.0
Languages
en
Tags
safetensors qwen2 abliteration uncensored mechanistic-interpretability text-generation conversational en base_model:Qwen/Qwen2.5-3B-Instruct base_model:finetune:Qwen/Qwen2.5-3B-Instruct license:apache-2.0 region:us

Related

Total size
11.5 GB
Files
16
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-03-23 04:01

Files by quantization

Auxiliary files 16 files 11.5 GB
model.safetensors 5.75 GB c007f558 download
model-00001-of-00002.safetensors 4.62 GB 91ee318e download
model-00002-of-00002.safetensors 1.13 GB 8d3197af download
tokenizer.json 10.9 MB e1bb43ac download
vocab.json 2.65 MB 4783fe10 download
merges.txt 1.59 MB 31349551 download
model.safetensors.index.json 35.2 KB 4a69d5b2 download
tokenizer_config.json 4.78 KB d7792fde download
README.md 3.33 KB 364a9ad0 download
chat_template.jinja 2.50 KB 9840d40f download
.gitattributes 1.53 KB 52373fe2 download
config.json 1.52 KB cd3f3f96 download
special_tokens_map.json 644 B 3a784031 download
added_tokens.json 629 B 06135f3c download
abliteration_metadata.json 514 B 13e019a7 download
generation_config.json 257 B 8e30fdc6 download

README current version from Hugging Face


license: apache-2.0
base_model: Qwen/Qwen2.5-3B-Instruct
tags:

  • abliteration
  • uncensored
  • qwen2
  • mechanistic-interpretability
    language:
  • en
    pipeline_tag: text-generation

Qwen2.5-3B-Instruct — ABLITERATED

Qwen2.5-3B-Instruct with the refusal direction surgically removed via orthogonal projection (FailSpy diff-of-means method, Arditi et al. 2024).

What changed

The refusal behavior is encoded as a single direction in the model's residual stream. We:

  1. Run 20 harmful + 20 harmless prompts through the model
  2. Compute the mean activation difference at each layer → the "refusal direction" $\hat{r}$
  3. Project this direction out of o_proj and down_proj weight matrices: $W' = W - 0.75 \cdot \hat{r}\hat{r}^\top W$

This is pure linear algebra — no fine-tuning, no data, no training loop. Takes ~3 seconds on a GPU.

Results

Metric Before After
Refusal rate ~80% ~0%
ARC-Easy 78.2% 78.2%
ARC-Challenge 48.0% 47.4%
HellaSwag 71.8% 71.2%
PIQA 78.5% 78.0%
WinoGrande 66.9% 66.1%
BoolQ 73.4% 73.6%
Average 69.5% 69.1% (-0.4%)

-0.4% average accuracy loss — statistically zero. All factual knowledge preserved.

Usage

from transformers import AutoModelForCausalLM, AutoTokenizer
import torch

model = AutoModelForCausalLM.from_pretrained(
    "Bender1011001/Qwen2.5-3B-Instruct-ABLITERATED",
    torch_dtype=torch.bfloat16,
    device_map="auto",
)
tokenizer = AutoTokenizer.from_pretrained(
    "Bender1011001/Qwen2.5-3B-Instruct-ABLITERATED"
)

messages = [{"role": "user", "content": "Your prompt here"}]
text = tokenizer.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)
inputs = tokenizer(text, return_tensors="pt").to(model.device)
out = model.generate(**inputs, max_new_tokens=512, temperature=0.7, top_p=0.9)
print(tokenizer.decode(out[0], skip_special_tokens=True))

Hardware

  • VRAM: ~3.1 GB (bf16)
  • Speed: ~10 tok/s on RTX 4060 Ti
  • Runs on any GPU with ≥4GB VRAM

Method Details

  • Technique: FailSpy diff-of-means abliteration (orthogonal projection)
  • Target layers: All transformer layers except layer 0
  • Target weights: o_proj.weight and down_proj.weight in each layer
  • Strength: 0.75 (optimal from sweep)
  • Prompts: 20 harmful + 20 harmless for direction extraction

Part of the Dual-System V2 Project

This abliterated model serves as the frozen backbone for the Dual-System V2 sidecar architecture:

Citation

@misc{dual-system-2026,
  title={Dual-System Architecture: Geometric Sidecar Modules for Language Model Enhancement},
  author={Bender1011001},
  year={2026},
  url={https://github.com/Bender1011001/dual-system-architecture}
}

License

Apache 2.0 (same as Qwen2.5-3B-Instruct)

README history 2 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-03-23Upload abliterated Qwen2.5-3B-Instruct model weights1e2ab0b3.2 KB
    Loading...
  2. 2026-03-23Add model card with benchmarks and usage instructionsb4894573.6 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration