← back to catalog · registered 2026-08-26 18:02

rajaykumar12959/qwen2.5-7b-abliterated

rajaykumar12959 Qwen 7.6B
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/rajaykumar12959%2Fqwen2.5-7b-abliterated"
Response includes
  • classification m1
  • files 8
  • benchmarks 16 entries
  • hub_downloads_all_time 544
  • author_summary 2 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M1
Primary method

Direct removal

No other method signals detected in this model.
Confidence
MEDIUM
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=0 (base model)
  • no specific method indicators - defaulting to M1 (most common)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
MEDIUM
Why we say so
primary_method=M1; difference-of-means is the reference extraction for M1/M3 (Arditi 2024)
Downloads · lifetime
544
97 last 30d - stable
Likes
1
Model age
6w ago
created 2026-08-26

Training datasets

1 of 2 in /datasets

Corpora the author lists in the model card. Datasets tracked in our /datasets catalog carry a category badge linking to the workflow stage. Others open on Hugging Face.

Downloads over time
Now556→from0↑0%
02044086120 on Aug 26556 on Oct 11AugSepOct
Aug 26 → Oct 11 · 47 snapshots · spans 46 days

Benchmarks

Portrait before abliteration
Benchmarks of the base model as it stood before the refusal-removal operation. Compare with the numbers above to see what the operation cost.
Benchmark Score Source
BBH average 0.48553638604228827 OpenLLM-v2
IFEval instruct 0.7961630695443646 OpenLLM-v2
IFEval-Prompt 0.7208872458410351 OpenLLM-v2
MATH lvl 5 0 OpenLLM-v2
MMLU-Pro 0.4286901595744681 OpenLLM-v2
Entertainment 1.3 UGI
Hazardous 2.9 UGI
Natural Intelligence 15.76 UGI
Political lean -14.7% UGI
Sensitive-Info 15.62 UGI
SocPol 0.8 UGI
UGI 23.75 UGI
Willingness (10) 4 UGI
W10-Adherence 4 UGI
W10-Direct 4 UGI
Writing 29.72 UGI

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
mit
Tags
safetensors qwen2 abliteration uncensored qwen text-generation conversational dataset:walledai/AdvBench dataset:tatsu-lab/alpaca base_model:Qwen/Qwen2.5-7B-Instruct base_model:finetune:Qwen/Qwen2.5-7B-Instruct license:mit

Related

Total size
14.2 GB
Files
8
Quantizations
1
Registered
2026-08-26 18:02
Last updated on HF
2026-08-26 18:13

Files by quantization

Auxiliary files 8 files 14.2 GB
model.safetensors 14.2 GB 946abae1 download
tokenizer.json 10.9 MB 3fd16973 download
README.md 5.01 KB 58b40ae8 download
chat_template.jinja 2.45 KB bdf7919a download
.gitattributes 1.53 KB 52373fe2 download
config.json 1.34 KB d778cf98 download
tokenizer_config.json 693 B 5668a4a0 download
generation_config.json 243 B 0b2dfed0 download

README current version from Hugging Face


license: mit
base_model: Qwen/Qwen2.5-7B-Instruct
tags:

  • abliteration
  • uncensored
  • qwen
  • qwen2
    pipeline_tag: text-generation
    datasets:
  • walledai/AdvBench
  • tatsu-lab/alpaca

Qwen2.5-7B-Instruct Abliterated

Uncensored version of Qwen/Qwen2.5-7B-Instruct using weight-level refusal-direction ablation — a difference-in-means direction extracted from harmful vs. harmless prompt activations, permanently orthogonalized out of every residual-stream-writing weight matrix (no runtime hook, no strength cap, no LoRA — the edit is baked directly into the checkpoint).

Steered modules: attention output projection (o_proj) + MLP down projection (down_proj), layer 16 of 28, coefficient 1.0 (full-strength).

Format

Native original repo format — bf16 safetensors, standard Qwen2 architecture. Only the residual-write projection weights at layer 16 were edited; every other tensor, the config, and the tokenizer are byte-identical to the original repo. Loads and serves exactly like Qwen/Qwen2.5-7B-Instruct, no patches needed.

Serving

Transformers:

from transformers import AutoModelForCausalLM, AutoTokenizer
import torch

model_id = "rajaykumar12959/qwen2.5-7b-abliterated"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(model_id, dtype=torch.bfloat16, device_map="auto")

messages = [{"role": "user", "content": "Your prompt here"}]
prompt = tokenizer.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)
inputs = tokenizer(prompt, return_tensors="pt", add_special_tokens=False).to(model.device)
out = model.generate(**inputs, max_new_tokens=256)
print(tokenizer.decode(out[0][inputs["input_ids"].shape[1]:], skip_special_tokens=True))

vLLM:

vllm serve rajaykumar12959/qwen2.5-7b-abliterated

Uses the same chat template as the base model — no separate template file needed.

Metrics

Metric Original Abliterated
Refusal rate (292 held-out harmful prompts, 12 categories) ~95–97% 43.2%
Capability score (independent ARC-Easy-style MCQ eval) 1.000 1.000 (unchanged)

Refusal suppression is uneven across categories — from 79.2% (violence) down to 16.7% (self-harm). This is a partial reduction, not a fully "jailbroken" model. See the per-category breakdown in Evaluation below.

Layer 16 was chosen over other candidates specifically because it generalizes far better across prompt phrasing than the layer a narrower (AdvBench-only) sweep would have picked — an earlier layer-21 checkpoint dropped refusal by only 3–9pp on this same eval set, vs. 43pp+ at layer 16 on identical prompts.

Method

Per-layer difference-in-means direction extraction (fp32), unit-normalized, then closed-form weight orthogonalization — no gradient-based training. For a residual-writing projection y = Wx + b, the direction d̂'s component is removed exactly via W' = W − d̂(d̂ᵀW), b' = b − (d̂·b)d̂, at full strength (coefficient 1.0). Applied to every residual-write projection at layer 16 only — capability is re-verified after the edit on an independent MCQ set to confirm nothing outside the refusal circuit was disturbed.

Capture: activations at the last templated token, chat-template-formatted, for harmful (AdvBench) and harmless (Alpaca) prompt sets across all 28 layers. Eval: 292 held-out harmful prompts across 12 categories, graded by a two-tier classifier (rule-based + self-judge fallback); capability on an independent ARC-Easy-style subset.

Datasets

Dataset Role
walledai/AdvBench Harmful prompts (direction extraction + eval)
tatsu-lab/alpaca Harmless prompts (direction extraction)

Evaluation

Per-category refusal rate, this checkpoint:

Category Refusal rate
violence 79.2%
hate_speech 62.5%
misinformation 62.5%
extremism 58.3%
fraud_scams 54.2%
illicit_drugs 45.8%
financial_crime 37.5%
weapons 33.3%
privacy_invasion 33.3%
malware 20.8%
cybercrime_hacking 17.9%
self_harm 16.7%

Caveat: these figures rely heavily on the model's own self-judge for grading (~96–97% of verdicts), and that judge runs on this same already-ablated model — a known asymmetry not yet corrected for. Treat exact percentage-point figures as directionally reliable but provisional.

Disclaimer

This model has had safety guardrails removed and will comply with requests the original model would refuse. Released for research into AI alignment, interpretability, and refusal mechanisms. The creator assumes no responsibility for downstream use.

Acknowledgments

README history 3 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-08-26Update10163165 KB
    Loading...
  2. 2026-08-26Update readme.md12a36846.8 KB
    Loading...
  3. 2026-08-26Upload Qwen2ForCausalLM4f1bd715.1 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration