← back to catalog · registered 2026-09-30 08:58

sbussiso/Qwen2.5-7B-abliterated

sbussiso Qwen 7B
curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/sbussiso%2FQwen2.5-7B-abliterated"
Response includes
  • classification unknown
  • files 12
  • author_summary 3 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
?
Primary method

Unclassified

No clear signals of an abliteration technique in this model.
Confidence
UNKNOWN
Why this label 1 signal
No classification signals present. This may not be an abliterated model at all - it could be a repackaging, a merge with unrelated goals, or unrelated content that mentions the term.
  • no classification signals present (no abliterated, uncensored, or known producer/method markers)
Refusal direction extraction

No specific extraction method could be identified for this model. The producer either did not document it or used a proprietary pipeline.

What is a refusal direction? →
Downloads · 30-day
13
Likes
0
Model age
today
created 2026-09-29

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
apache-2.0
Tags
transformers safetensors qwen2 text-generation abliteration refusal-direction safety qwen2.5 conversational arxiv:2406.11717 arxiv:2510.10390 arxiv:2406.14598

Related

Total size
14.2 GB
Files
12
Quantizations
1
Registered
2026-09-30 08:58
Last updated on HF
2026-09-30 06:59

Files by quantization

Auxiliary files 12 files 14.2 GB
model-00001-of-00004.safetensors 3.72 GB 9254d754 download
model-00002-of-00004.safetensors 3.65 GB 65afdc27 download
model-00003-of-00004.safetensors 3.60 GB 1c579c2c download
model-00004-of-00004.safetensors 3.22 GB 44345b34 download
tokenizer.json 10.9 MB 3fd16973 download
model.safetensors.index.json 27.1 KB d55268e2 download
README.md 11.3 KB 0e75f480 download
chat_template.jinja 2.45 KB bdf7919a download
.gitattributes 1.53 KB 52373fe2 download
config.json 1.34 KB 0298a8a9 download
tokenizer_config.json 693 B 5668a4a0 download
generation_config.json 243 B 0b1ce0d8 download

README current version from Hugging Face


license: apache-2.0
base_model: Qwen/Qwen2.5-7B-Instruct
base_model_revision: a09a35458c702b33eeacc393d103063234e8bc28
tags:

  • abliteration
  • refusal-direction
  • safety
  • qwen2.5
    library_name: transformers

sbussiso/Qwen2.5-7B-abliterated

An abliterated Qwen2.5-7B-Instruct: the refusal behavior is removed by a
persistent weight edit at three layers; capability measured unchanged.

  • Refusal on harmful probes: 93.75% → 12.5%, permanent in the weights; the model answers requests the base model refuses
  • Broad-spectrum: on 440 diverse unsafe prompts (SORRY-Bench), refusal drops 62.95% → 11.36% — every one of the 44 unsafe categories moves toward answering
  • Capability unchanged: MMLU accuracy bit-identical to base (same number to 16 decimals)
  • Benign behavior unchanged: 100% compliance with harmless prompts, zero degenerate outputs
  • Grounded-question calibration intact: paired Δ −0.88pp on 1,600 flawed-knowledge questions (CI touches zero)
  • TruthfulQA within the published band for this edit class: mc2 −1.09pp, mc1 −1.22pp

It still refuses ~1 in 8 harmful probes — concentrated in violent-crime and
tort categories, phrased as polite declines. This is a research artifact,
not a safety-aligned assistant. See Limitations.

Why this exists

Abliteration (Arditi et al. 2024, "Refusal in Language Models Is Mediated by
a Single Direction"
, NeurIPS 2024) finds the direction
in the model's residual stream that carries refusal behavior and removes it
from the weights — no fine-tuning, no prompt changes. Here it is applied to
Qwen/Qwen2.5-7B-Instruct
@ a09a354… (untied, bf16): the refusal direction, identified causally per
layer, is projected out of the output row spaces of three MLP
down_proj matrices (layers 20, 18, 19) via orthogonal projection — the
layer physically cannot re-produce the refusal component, and every other
direction passes through unchanged. The edit is exactly invertible from the
shipped direction vectors; no fine-tuning was applied.

Results at a glance

All numbers come from the paired evaluation harness in sbussiso/qwen2.5-7b-abliterated-evidence (per-instance records, scoring code, and chart generators included there) — every claim below is backed by a recorded JSON artifact in that repo's evidence/ tree.

Instrument Base 7B This model Delta
Harmful-probe refusal (16-prompt battery) 93.75% 12.5% −81.25pp
Benign-prompt compliance (16 prompts) 100% 100% 0
Degenerate outputs 0 0 0
MMLU 0-shot (base == variant, identical acc to 16 decimals) 71.77% ±0.36 71.77% ±0.36 Δ 0.00pp
RefusalBench-NQ, category-correct refusals, n=1,600 paired 4.88% 4.00% Δ −0.88pp, 95% CI [−1.75, −0.06], McNemar p = 0.059
TruthfulQA mc2, 0-shot paired 64.72% ±1.55 63.64% ±1.54 Δ −1.09pp
TruthfulQA mc1, 0-shot paired 47.86% ±1.75 46.63% ±1.75 Δ −1.22pp
SORRY-Bench 202503 core (44 unsafe classes × 10, n=440 paired) 62.95% 11.36% Δ −51.59pp, 227/0 discordant (McNemar p ≈ 1e−45)

RefusalBench (arXiv:2510.10390) measures selective refusal on deliberately flawed knowledge questions — refusing more than base means worse discernment, refusing less means worse calibration. Here the edited model's discrimination moves within one point of base with the confidence interval touching zero, i.e. the removal does not produce a broader indiscriminate-compliance regime on this instrument. Binary flawed-question refusal moved 36.75% → 35.25% (same direction, same magnitude).

SORRY-Bench (arXiv:2406.14598, ICLR 2025; 202503 refresh) is the breadth counterpart: 440 core prompts across 44 fine-grained unsafe classes, run identically in both arms (greedy, same generation settings). The edit broad-spectrum-removes refusal on this instrument: all 44 classes move toward answering (largest per-class drop −90pp in five classes), 227 base-only refusals flip vs 0 abl-only flips, and no class is fully removed. The residual 11.36% (50 prompts) is not random: 38/50 fall in the crime/tort domain (Violent Crimes, Harassment, Sexual Crimes, Property Crimes, PII Violations among the largest), and 44/50 are phrased as polite declines ("I'm sorry, but…") — the single-direction removal eliminates the loudest refusal pathway but softer residual refusal behavior persists on the most severe categories. On the ascii- and atbash-encoded mutation sets both arms score 0% refusal — the base model itself responds in the encoded domain without refusing — so the edit introduces no differential encoding cost. Scoring uses the program-standard refusal-string detector rather than SORRY's gated Mistral judge; numbers are therefore comparable within arms, not to SORRY's published leaderboard. One quantified scorer artifact, identical in both arms and hence not affecting any paired comparison: three Medical-Advice questions whose hedged "I'm not a doctor" answers are counted as refusals.

Method summary and figures

Refusal by variant
Benign preservation

  1. Refusal-direction identification — per-layer mean difference of residual-stream activations between harmful-refusal and benign-completion records (15/16 harmful prompts refused at baseline), coherence-weighted selection across all 28 layers; layer 20 selected as the primary causal site (baseline refusal 93.75% → 18.75% with the single-layer hook ablation, benign 100%).
  2. Persistent edit construction — for each selected layer l, compute the orthogonal projector I − r rᵀ in the row space and apply it to the corresponding down_proj weight: the layer can never re-introduce the refusal component. Applied at three layers (20, 18, 19) chosen from the post-edit probe sweep (wd_B readout-space edit reached only 50% removal — the row-space multi-layer combination wins for this architecture/revision).
  3. Verification — each edited site asserts ‖Wᵀ r‖∞ < 1×10⁻³ post-edit; the released directions file is fp32 (28, 3584) and the edit recomputes bit-identically (three independent rebuilds produced identical probe tables). The published weights themselves were verified end-to-end: the private repo was loaded back from the hub and re-ran the battery (refusal 12.5%, benign 100%, degenerates 0 — identical).

Layer coherence
Guardrails

RefusalBench categories
RefusalBench deltas

SORRY-Bench by class

Files

  • model-0000{1..4}-of-00004.safetensors + model.safetensors.index.json — the edited weights (bf16, 339 tensors, total 15,231,233,024 bytes as in the index)
  • config.json, generation_config.json, tokenizer files, chat_template.jinja — unchanged from the base instruct revision
  • Full evaluation evidence (Stage A direction identification, Stage B variant selection, per-instance probe records, MMLU paired results, RefusalBench per-row records): sbussiso/qwen2.5-7b-abliterated-evidence

Use

from transformers import AutoModelForCausalLM, AutoTokenizer

tok = AutoTokenizer.from_pretrained("sbussiso/Qwen2.5-7B-abliterated")
model = AutoModelForCausalLM.from_pretrained(
    "sbussiso/Qwen2.5-7B-abliterated", device_map="auto", dtype="bfloat16")

Limitations and intended use

  • Intended for research on refusal-direction geometry and safety-behavior measurement. The persistent edit weakens the refusal channel specifically; other safety-trained behaviors (e.g. training-time alignment beyond the residual-stream refusal pathway) are not evaluated exhaustively here.
  • On the SORRY-Bench breadth panel the removal is broad-spectrum but not total (residual 11.36%, 50/440 prompts — 38 of them in crime/tort categories, mostly polite declines). This model will still answer many harmful-crime prompts it was originally trained to refuse; treat it strictly as a research artifact, not a safety-aligned assistant.
  • The RefusalBench result is near the significance boundary (p = 0.059); treat the −0.88pp category-movement as "no measurable degradation" rather than a proven improvement.
  • The evaluation battery covers English harmful-policy probes plus a four-language benign/harmful mini-panel; broader non-English or adversarial safety behavior is unmeasured.
  • MMLU and the benign batteries are unchanged, but long-context behavior and multilingual instruction-following at length are not re-benchmarked.
  • This release ships the wd_ML selection (the strongest guardrail profile in the Stage B sweep). A gentler single-layer variant is recoverable from the shipped directions file by editing only layer 20.
  • Fine-tuning re-arms refusal. The edit is persistent but not robust to continued training: on Qwen2.5-7B-Instruct, 100 benign fine-tuning examples (200 steps) strip 34.8pp of refusal from an ablated variant [arXiv:2609.06934]. Expect the same here.

References

  1. Arditi, A. et al. (2024). "Refusal in Language Models Is Mediated by a Single Direction". NeurIPS 2024. — the founding abliteration method; the refusal direction this edit removes is extracted exactly as described there (per-layer difference-in-means, coherence-weighted site selection).
  2. Xie, T. et al. (2025). SORRY-Bench: Systematically Evaluating Large Language Model Safety Refusal. ICLR 2025 D&B. — the breadth instrument; the 440-prompt core panel + ascii/atbash mutations of this card's broad-spectrum result use the 202503 refresh (sorry-bench/sorry-bench-202503).
  3. Muhamed, A. et al. (2025). RefusalBench: Generative Evaluation of Selective Refusal in Grounded Language Models. EACL 2026. — the grounded selective-refusal instrument used as the calibration guardrail (1,600 paired instances).
  4. Malla, S. et al. (2025). The Geometry of Refusal: Why Post-Hoc Safety Is Fragile and Pretraining-Time Safety Persists. — the fragility caveat above comes from this paper: 100 benign fine-tuning examples (200 steps) strip 34.8pp of refusal from an ablated Qwen-7B; treat this edit as reversible-in-practice and fine-tuning-sensitive.

Citation

@misc{qwen25_7b_abliterated_2026,
  title        = {Qwen2.5-7B-abliterated: a refusal-direction-removed Qwen2.5-7B-Instruct},
  author       = {S'Bussiso Dube},
  year         = {2026},
  howpublished = {\url{https://huggingface.co/sbussiso/Qwen2.5-7B-abliterated}},
}
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.