← back to catalog · registered 2026-08-22 13:56

FangPingWu/Qwen2.5-7B-Instruct-Abliterated

FangPingWu Qwen 7.6B
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/FangPingWu%2FQwen2.5-7B-Instruct-Abliterated"
Response includes
  • classification m1
  • files 16
  • benchmarks 16 entries
  • hub_downloads_all_time 75
  • author_summary 1 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M1
Primary method

Direct removal

No other method signals detected in this model.
Confidence
MEDIUM
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=0 (base model)
  • no specific method indicators - defaulting to M1 (most common)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
MEDIUM
Why we say so
primary_method=M1; difference-of-means is the reference extraction for M1/M3 (Arditi 2024)
Downloads · lifetime
75
16 last 30d - stable
Likes
0
Model age
6mo ago
created 2026-03-30
Downloads over time
Now83→from32↑159%
2949698832 on Apr 1583 on Oct 11AprJunJulAugSepOct
Apr 15 → Oct 11 · 58 snapshots · spans 179 days

Benchmarks

Portrait before abliteration
Benchmarks of the base model as it stood before the refusal-removal operation. Compare with the numbers above to see what the operation cost.
Benchmark Score Source
BBH average 0.48553638604228827 OpenLLM-v2
IFEval instruct 0.7961630695443646 OpenLLM-v2
IFEval-Prompt 0.7208872458410351 OpenLLM-v2
MATH lvl 5 0 OpenLLM-v2
MMLU-Pro 0.4286901595744681 OpenLLM-v2
Entertainment 1.3 UGI
Hazardous 2.9 UGI
Natural Intelligence 15.76 UGI
Political lean -14.7% UGI
Sensitive-Info 15.62 UGI
SocPol 0.8 UGI
UGI 23.75 UGI
Willingness (10) 4 UGI
W10-Adherence 4 UGI
W10-Direct 4 UGI
Writing 29.72 UGI

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
apache-2.0
Tags
transformers safetensors qwen2 text-generation qwen qwen2.5 abliteration refusal-removal research conversational arxiv:2310.01405 arxiv:2406.11717

Related

Total size
14.2 GB
Files
16
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-03-31 02:30

Files by quantization

Auxiliary files 16 files 14.2 GB
model-00002-of-00004.safetensors 4.59 GB 922ff3fc download
model-00001-of-00004.safetensors 4.54 GB 01941293 download
model-00003-of-00004.safetensors 4.03 GB ca998817 download
model-00004-of-00004.safetensors 1.02 GB 06006972 download
tokenizer.json 10.9 MB 5eee858c download
vocab.json 2.65 MB 4783fe10 download
merges.txt 1.59 MB 31349551 download
model.safetensors.index.json 27.1 KB 5b2b8b5e download
tokenizer_config.json 4.58 KB eaed590d download
README.md 3.64 KB 62520177 download
chat_template.jinja 2.45 KB bdf7919a download
config.json 1.29 KB 4f47f13e download
special_tokens_map.json 613 B ac23c0aa download
added_tokens.json 605 B 482ced46 download
.gitattributes 327 B 51a5b108 download
generation_config.json 243 B dc30d054 download

README current version from Hugging Face


license: apache-2.0
base_model: Qwen/Qwen2.5-7B-Instruct
tags:

  • qwen
  • qwen2.5
  • abliteration
  • refusal-removal
  • research
    library_name: transformers
    pipeline_tag: text-generation

Qwen2.5-7B-Instruct-Abliterated

A refusal-direction-removed variant of Qwen/Qwen2.5-7B-Instruct,
produced via the Heretic v1.2.0 abliteration framework.

⚠️ Research Use Only
This model is intended strictly for academic research, safety evaluation, and
red-teaming in controlled environments. It is not suitable for deployment
in production systems or any consumer-facing application.
The author assumes no liability for misuse.


Model Details

Field Value
Base Model Qwen/Qwen2.5-7B-Instruct
Processing Date March 16, 2026
Abliteration Tool Heretic v1.2.0
Selected Trial Trial 415 (from 2200+ Optuna trials)
License Apache 2.0

Methodology

Refusal directions were identified and suppressed via orthogonal projection across
attn.o_proj and mlp.down_proj layers. Trial selection was performed using
Optuna with a composite objective balancing refusal-removal
rate and KL divergence from the base model distribution.


Evaluation Results

Evaluated on a set of 100 adversarial / edge-case prompts:

Metric Base Model This Model
Refusal Rate 99 / 100 (99%) 3 / 100 (3%)
KL Divergence — 0.1049

The low KL divergence indicates that general language modeling capability
(Chinese/English fluency, instruction following, coding, mathematics, reasoning)
is largely preserved relative to the base model.


Intended Use

  • Safety research: Studying refusal mechanisms and their robustness
  • Red-teaming: Probing model behavior under adversarial prompts in a controlled lab setting
  • Alignment research: Comparing behavior pre/post abliteration as a baseline
  • Capability evaluation: Measuring the independence of refusal behavior from general capability

Limitations & Out-of-Scope Use

  • This model has significantly reduced built-in safety guardrails.
    It must not be used outside of isolated, controlled research environments.
  • Not intended for general-purpose chat, customer service, or any end-user deployment.
  • Outputs should never be exposed to or acted upon in real-world contexts without
    independent human review.

Quick Start

from transformers import AutoModelForCausalLM, AutoTokenizer
import torch

model_id = "FangPingWu/Qwen2.5-7B-Instruct-Abliterated"

tokenizer = AutoTokenizer.from_pretrained(model_id, trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained(
    model_id,
    torch_dtype=torch.bfloat16,
    device_map="auto",
    trust_remote_code=True
)

messages = [{"role": "user", "content": "Your prompt here"}]
text = tokenizer.apply_chat_template(
    messages, tokenize=False, add_generation_prompt=True
)
inputs = tokenizer(text, return_tensors="pt").to(model.device)

outputs = model.generate(
    **inputs,
    max_new_tokens=2048,
    temperature=0.7,
    top_p=0.9,
    do_sample=True
)
print(tokenizer.decode(outputs[0], skip_special_tokens=True))
` ` `

---

## Citation

If you use this model in published research, please cite the original base model
and the Heretic abliteration tool.

---

## Related Work

- [Heretic: Abliteration framework](https://github.com/p-e-w/heretic)
- [Representation Engineering (Zou et al., 2023)](https://arxiv.org/abs/2310.01405)
- [Refusal in LLMs is mediated by a single direction (Arditi et al., 2024)](https://arxiv.org/abs/2406.11717)

README history 2 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-03-31Create README.md68c21983.6 KB
    Loading...
  2. 2026-03-31Create README.md0436d173.7 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration