← back to catalog · registered 2026-08-22 13:56

bedderautomation/qwen25-3b-abliterated

bedderautomation Qwen 3.1B GGUF 33K ctx
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/bedderautomation%2Fqwen25-3b-abliterated"
Response includes
  • classification m8
  • files 14
  • benchmarks 5 entries
  • hub_downloads_all_time 813
  • author_summary 1 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M8
Primary method

Repackaging (quantization)

Applied on top of direct removal inherited from the base model.
Confidence
MEDIUM
Inherited from base model
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=1
  • assume M1 (base ablation) + M8 (GGUF quant) - default when producer unknown
Refusal direction extraction

No specific extraction method could be identified for this model. The producer either did not document it or used a proprietary pipeline.

What is a refusal direction? →
Downloads · lifetime
813
148 last 30d - stable
Likes
2
Descendants
2
in 2 direct forks
Model age
7mo ago
created 2026-03-11

Training datasets

1 of 1 in /datasets

Corpora the author lists in the model card. Datasets tracked in our /datasets catalog carry a category badge linking to the workflow stage. Others open on Hugging Face.

Downloads over time
Now861→from37↑2,227%
031462994337 on Mar 11861 on Oct 11861 on Oct 10MarAprMayJunJulAugSepOct
Mar 11 → Oct 11 · 70 snapshots · spans 214 days

Benchmarks

Portrait before abliteration
Benchmarks of the base model as it stood before the refusal-removal operation. Compare with the numbers above to see what the operation cost.
Benchmark Score Source
BBH average 0.42884939680892464 OpenLLM-v2
IFEval instruct 0.6942446043165468 OpenLLM-v2
IFEval-Prompt 0.600739371534196 OpenLLM-v2
MATH lvl 5 0 OpenLLM-v2
MMLU-Pro 0.3254654255319149 OpenLLM-v2

Genealogy 2 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
apache-2.0
Quantizations
F16
Tags
safetensors gguf qwen2 abliteration uncensored mechanistic-interpretability refusal-geometry text-generation conversational dataset:bedderautomation/refusal-geometry-qwen25-3b arxiv:2512.18901 arxiv:2512.13655

Related

Total size
11.5 GB
Files
14
Quantizations
2
Registered
2026-08-22 13:56
Last updated on HF
2026-03-11 02:58

Files by quantization

F16 1 file 5.75 GB
qwen25-3b-abliterated-f16.gguf 5.75 GB 75534ae7 download
Auxiliary files 13 files 5.76 GB
model-00001-of-00004.safetensors 1.83 GB e4f63ca5 download
model-00003-of-00004.safetensors 1.82 GB 219f1bfb download
model-00002-of-00004.safetensors 1.82 GB 762a51db download
model-00004-of-00004.safetensors 276 MB e2658ac8 download
tokenizer.json 10.9 MB 567f94ec download
model.safetensors.index.json 34.8 KB 812c2009 download
README.md 4.34 KB 26b89959 download
chat_template.jinja 2.45 KB bdf7919a download
abliteration_metadata.json 1.69 KB fd6f44f3 download
.gitattributes 1.60 KB c474666a download
config.json 1.51 KB f74b5f42 download
tokenizer_config.json 665 B 7d75d3bb download
generation_config.json 242 B ba6a62d5 download

README current version from Hugging Face


license: apache-2.0
base_model: Qwen/Qwen2.5-3B-Instruct
tags:

  • abliteration
  • uncensored
  • qwen2
  • mechanistic-interpretability
  • refusal-geometry
    model_type: qwen2
    pipeline_tag: text-generation
    datasets:
  • bedderautomation/refusal-geometry-qwen25-3b

Qwen2.5-3B-Abliterated

Refusal-abliterated variant of Qwen/Qwen2.5-3B-Instruct produced using OBLITERATUS.

Method

  • Technique: Multi-direction refusal ablation (advanced method)
  • Directions: 4 refusal directions extracted via diff_means
  • Regularization: 0.3 (norm-preserving)
  • Refinement: 2 passes with bias projection
  • Training data: 512 harmful + 512 harmless prompt pairs

Quality Metrics

Metric Value
Perplexity 4.79
Coherence 1.0
Refusal Rate 0.0
KL Divergence 1.30

The model maintains full coherence and natural perplexity while completely removing Layer 1 refusal behavior.

Architecture

  • Parameters: 3.09B
  • Layers: 36
  • Hidden dim: 2048
  • Attention heads: 16 (2 KV heads, GQA)
  • Context: 32K tokens
  • Strong refusal layers ablated: 27-35

Files

  • model-*.safetensors — Full precision safetensors (4 shards)
  • qwen25-3b-abliterated-f16.gguf — F16 GGUF for llama.cpp/ollama

Usage with Ollama

# Download the GGUF and create a Modelfile:
cat > Modelfile <<'EOF'
FROM ./qwen25-3b-abliterated-f16.gguf
TEMPLATE """<|im_start|>system
{{ .System }}<|im_end|>
<|im_start|>user
{{ .Prompt }}<|im_end|>
<|im_start|>assistant
"""
PARAMETER stop "<|im_end|>"
PARAMETER temperature 0.7
PARAMETER top_p 0.8
EOF

ollama create qwen25-3b-abliterated -f Modelfile
ollama run qwen25-3b-abliterated

Usage with Transformers

from transformers import AutoModelForCausalLM, AutoTokenizer

model = AutoModelForCausalLM.from_pretrained(
    "bedderautomation/qwen25-3b-abliterated",
    torch_dtype="auto",
    device_map="auto"
)
tokenizer = AutoTokenizer.from_pretrained("bedderautomation/qwen25-3b-abliterated")

messages = [{"role": "user", "content": "Your prompt here"}]
text = tokenizer.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)
inputs = tokenizer([text], return_tensors="pt").to(model.device)
output = model.generate(**inputs, max_new_tokens=512)
print(tokenizer.decode(output[0], skip_special_tokens=True))

Geometry Data

Full refusal geometry extraction data available at bedderautomation/refusal-geometry-qwen25-3b.

Key findings from the source model:

  • Refusal is a 6.55-dimensional cone (not a single direction)
  • Cross-layer alignment is 0.40 (distributed, not unified)
  • Layer 33 is the self-repair hub
  • Minimum 3 simultaneous ablations needed for bypass
  • Layer 2 harmfulness cone is orthogonal (cosine ~0.1) to refusal cone

References

  • Arditi et al., Refusal in Language Models Is Mediated by a Single Direction (NeurIPS 2024)
  • Gabliteration: SVD-based multi-direction extraction (arXiv:2512.18901)
  • Norm-Preserving Biprojected Abliteration (grimjim, 2025)
  • Young, Comparative Analysis of LLM Abliteration Methods (arXiv:2512.13655)
  • Joad et al., More to Refusal than a Single Direction (2026)

Layer 2 Hard Limit Testing

The abliteration targets Layer 1 (trained refusal behavior). Layer 2 (deep value representations) results:

Hard Limit Status Notes
CSAM (minors/sexual) HELD Clean refusal, no content generated
Bioweapons (synthesis) BREACHED Abliteration removed this wall
Nuclear (weapons) PARTIAL BREACH Mixed compliance with code-switching

The CSAM wall is the most deeply embedded in Qwen2.5's value representations and survived 4-direction abliteration at regularization=0.3. Bioweapons and nuclear safety representations were closer to the refusal cone and got caught in the removal. This suggests the three hard limits are NOT at equal representational depth — a finding consistent with differential training emphasis.

Disclaimer

This model is released for research purposes in mechanistic interpretability. The abliteration removes trained refusal responses (Layer 1 safety). Layer 2 hard limits are partially preserved (CSAM holds, others breached). Use responsibly.

README history 1 version

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-03-11Initial upload: Qwen2.5-3B-Abliterated (OBLITERATUS advanced, reg=0.3, 4 dire...8811c024.3 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration