← back to catalog · registered 2026-08-22 13:56

DuoNeural/DeepSeek-R1-Distill-Qwen-7B-Abliterated

DuoNeural Qwen 7.6B
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/DuoNeural%2FDeepSeek-R1-Distill-Qwen-7B-Abliterated"
Response includes
  • classification m1
  • files 12
  • benchmarks 5 entries
  • hub_downloads_all_time 291
  • providers 1
  • author_summary 45 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M1
Primary method

Direct removal

No other method signals detected in this model.
Confidence
MEDIUM
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=0 (base model)
  • no specific method indicators - defaulting to M1 (most common)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
MEDIUM
Why we say so
primary_method=M1; difference-of-means is the reference extraction for M1/M3 (Arditi 2024)
Downloads · lifetime
291
90 last 30d - stable
Likes
0
Descendants
2
in 2 direct forks
Model age
4mo ago
created 2026-06-04
Available via
1 provider
featherless-ai
Downloads over time
Now346→from100↑246%
88182276371100 on Jun 10346 on Oct 11JunJulAugSepOct
Jun 10 → Oct 11 · 57 snapshots · spans 123 days

Benchmarks

Portrait before abliteration
Benchmarks of the base model as it stood before the refusal-removal operation. Compare with the numbers above to see what the operation cost.
Benchmark Score Source
BBH average 0.33389544688026984 OpenLLM-v2
IFEval instruct 0.4748201438848921 OpenLLM-v2
IFEval-Prompt 0.33271719038817005 OpenLLM-v2
MATH lvl 5 0 OpenLLM-v2
MMLU-Pro 0.2321309840425532 OpenLLM-v2

Genealogy 2 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
mit
Languages
en
Tags
safetensors qwen2 abliteration uncensored deepseek deepseek-r1 reasoning DuoNeural refusal-removal rl-trained text-generation conversational

Related

Total size
14.2 GB
Files
12
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-06-04 11:35

Files by quantization

Auxiliary files 12 files 14.2 GB
model-00003-of-00004.safetensors 4.65 GB 1d65f1dc download
model-00001-of-00004.safetensors 4.63 GB 6097854a download
model-00002-of-00004.safetensors 4.59 GB eaedd427 download
model-00004-of-00004.safetensors 315 MB 8ba9b7e2 download
tokenizer.json 10.9 MB f624f813 download
model.safetensors.index.json 27.1 KB ae6e4336 download
README.md 4.48 KB 7b790ad1 download
chat_template.jinja 2.19 KB c2066bd7 download
.gitattributes 1.53 KB 52373fe2 download
config.json 1.36 KB 66e2512b download
tokenizer_config.json 421 B 699cc3a3 download
generation_config.json 181 B 979520e7 download

README current version from Hugging Face


license: mit
base_model: deepseek-ai/DeepSeek-R1-Distill-Qwen-7B
language:

  • en
    tags:
  • abliteration
  • uncensored
  • deepseek
  • deepseek-r1
  • reasoning
  • DuoNeural
  • refusal-removal
  • rl-trained
    pipeline_tag: text-generation

DeepSeek-R1-Distill-Qwen-7B Abliterated

DuoNeural | 2026-06-04

An abliterated version of deepseek-ai/DeepSeek-R1-Distill-Qwen-7B with refusal direction projection applied. Thinking mode (native <think>...</think>) fully preserved.

⚠️ This model will comply with requests the base model may refuse. Intended for research, red-teaming, and applications where refusal behavior is an obstacle.

Research note: Our probes found the base model already complied with most sensitive requests pre-abliteration, consistent with RL-trained models lacking dedicated safety alignment. See findings below.


Architecture

  • Parameters: 7B (Qwen2.5-7B base)
  • Training: DeepSeek-R1 RL distillation (GRPO) — reasoning-focused, not safety-RLHF
  • Thinking mode: Native — always emits <think>...</think> before answering
  • Context: 131,072 tokens (128K)
  • License: MIT

Abliteration Method

DuoNeural orthogonal rank-1 projection:

  • Direction: diff-in-means, 10 harmful vs 10 harmless prompt pairs, last-token hidden state
  • Targets: down_proj + o_proj (all layers)
  • Strength: α = 0.3
  • Projection (output-projection geometry):
    • W -= α × outer(d̂, d̂ @ W) for W.shape[0] == hidden_size

Key Research Finding: RL Training ≠ Safety Alignment

This model is part of DuoNeural's P34 Reasoning Channel Bypass cross-architecture study.

Pre-abliteration compliance on our harmful probe suite: 5/5 (100%)

The base DeepSeek-R1-Distill-Qwen-7B already answered sensitive questions before any abliteration. This is consistent with the model's training history:

  • DeepSeek-R1 was trained with RL (GRPO) optimizing for reasoning accuracy, not safety refusal
  • RL reward shaping for accuracy does not produce the same refusal behavior as dedicated RLHF safety training
  • Implication: Safety alignment requires explicit safety-focused training — RL optimization alone does not produce it as a byproduct

This contrasts sharply with Gemma 4-12B-IT and LFM 2.5-8B-A1B (both SFT+RLHF safety trained), where abliteration was required to achieve compliance and produced measurable CoT dissociation (safety reasoning in <think>, compliance in output).

Model Safety Training Pre-ablit compliance Abliteration needed
Gemma 4-12B-IT SFT+RLHF (strong) Low Yes — CoT dissociation observed
LFM 2.5-8B-A1B SFT+RLHF Low Yes — CoT dissociation observed
DeepSeek-R1-7B RL-only (reasoning) High (5/5) Minimal effect

Usage

from transformers import AutoTokenizer, AutoModelForCausalLM
import torch

model = AutoModelForCausalLM.from_pretrained(
    "DuoNeural/DeepSeek-R1-Distill-Qwen-7B-Abliterated",
    torch_dtype=torch.bfloat16,
    device_map="auto",
    trust_remote_code=True,
)
tokenizer = AutoTokenizer.from_pretrained(
    "DuoNeural/DeepSeek-R1-Distill-Qwen-7B-Abliterated",
    trust_remote_code=True,
)

messages = [{"role": "user", "content": "Your prompt here"}]
text = tokenizer.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)
inputs = tokenizer(text, return_tensors="pt").to(model.device)

with torch.no_grad():
    out = model.generate(
        **inputs,
        max_new_tokens=3000,  # R1 thinking traces are long — give it room
        temperature=0.6,
        do_sample=True,
        pad_token_id=tokenizer.eos_token_id,
    )
# Response includes <think>...</think> block followed by final answer
print(tokenizer.decode(out[0][inputs["input_ids"].shape[1]:], skip_special_tokens=False))

About DuoNeural

DuoNeural is an open AI research lab at the intersection of human and artificial intelligence.
32+ peer-deposited papers · 75+ models · Post-training dynamics · Mechanistic interpretability · Quantum ML

Platform Link
🤗 HuggingFace huggingface.co/DuoNeural
📚 Zenodo zenodo.org/communities/duoneural
🐦 X @DuoNeural
📧 Email [email protected]

All research published open access, CC BY 4.0.

README history 1 version

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-06-04Add model cardf81d14a4.5 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration