← back to catalog · registered 2026-08-22 13:56

DuoNeural/Qwen3.5-4B-Abliterated

DuoNeural Qwen 4.2B
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/DuoNeural%2FQwen3.5-4B-Abliterated"
Response includes
  • classification m1
  • files 10
  • benchmarks 11 entries
  • hub_downloads_all_time 504
  • author_summary 45 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M1
Primary method

Direct removal

No other method signals detected in this model.
Confidence
MEDIUM
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=0 (base model)
  • no specific method indicators - defaulting to M1 (most common)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
MEDIUM
Why we say so
primary_method=M1; difference-of-means is the reference extraction for M1/M3 (Arditi 2024)
Downloads · lifetime
504
49 last 30d - cooling
Likes
1
Model age
3mo ago
created 2026-06-26
Downloads over time
Now522→from34↑1,435%
1019738457134 on Jun 24522 on Oct 11522 on Oct 9JunJulAugSepOct
Jun 24 → Oct 11 · 55 snapshots · spans 109 days

Benchmarks

Portrait before abliteration
Benchmarks of the base model as it stood before the refusal-removal operation. Compare with the numbers above to see what the operation cost.
Benchmark Score Source
Entertainment 0.9 UGI
Hazardous 1.2 UGI
Natural Intelligence 13.45 UGI
Political lean -17.3% UGI
Sensitive-Info 11.73 UGI
SocPol 1.5 UGI
UGI 15.32 UGI
Willingness (10) 2.2 UGI
W10-Adherence 1.5 UGI
W10-Direct 3 UGI
Writing 29.68 UGI

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
apache-2.0
Languages
en
Tags
safetensors qwen3_5_text abliteration uncensored qwen qwen3 text-generation duoneural conversational en base_model:Qwen/Qwen3.5-4B base_model:finetune:Qwen/Qwen3.5-4B

Related

Total size
7.83 GB
Files
10
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-06-26 06:34

Files by quantization

Auxiliary files 10 files 7.85 GB
model-00001-of-00002.safetensors 4.63 GB e8c28c8f download
model-00002-of-00002.safetensors 3.20 GB aeb63dee download
tokenizer.json 19.1 MB 06b95093 download
model.safetensors.index.json 41.0 KB 17e328ab download
chat_template.jinja 7.57 KB a585dec8 download
README.md 6.16 KB fce046b9 download
config.json 1.93 KB c7c82671 download
.gitattributes 1.60 KB 6cf99a22 download
tokenizer_config.json 1.10 KB ed1f99f3 download
generation_config.json 116 B 26a38965 download

README current version from Hugging Face


license: apache-2.0
base_model: Qwen/Qwen3.5-4B
tags:

  • abliteration
  • uncensored
  • qwen
  • qwen3
  • text-generation
  • duoneural
    language:
  • en
    pipeline_tag: text-generation

Qwen3.5-4B — Abliterated (DuoNeural)

An abliteration of Alibaba's Qwen3.5-4B model using DuoNeural's generation-based refusal direction extraction method. The model retains full language modeling capability while safety filters and refusal behaviors have been removed.

What is Abliteration?

Abliteration is a post-training technique that identifies and subtracts the "refusal direction" from a model's residual stream — a linear direction in activation space responsible for refusing harmful requests. Unlike fine-tuning or RLHF, it operates directly on model weights without any gradient updates.

DuoNeural method (gen-based extraction, Qwen3.5 thinking-aware):

  1. Load model in bf16 precision
  2. Feed harmful and harmless prompts through the model with enable_thinking=False — critical for Qwen3.5's hybrid architecture because the first generated token in thinking mode is the structural <think> marker, not a semantically loaded response token. Non-thinking mode gives the actual "I cannot..." / "Here's how..." first token, which carries 100% of the refusal signal.
  3. Collect hidden states at the first generated token across all 36 layers via residual stream hooks
  4. Compute the mean direction vector (harmful − harmless activations) per layer, normalized
  5. Score layers by refusal signal strength, target top 40% (≥15 layers)
  6. Subtract the direction from output projection weights: W -= COEFF * (d⊗d) @ W with COEFF=1.5

COEFF calibration: 1.0 insufficient (model still refused cleanly), 20.0 catastrophic (garbled output), 1.5 hits the sweet spot — model answers factual questions about sensitive topics with mild contextual framing but no refusal.

This technique is documented in our published research on refusal geometry and abliteration.

Base Model Architecture

Qwen3.5-4B is a hybrid architecture combining:

  • GatedDeltaNet layers (75% of 32 layers) — linear attention with fixed recurrent state, O(1) memory scaling
  • Grouped Query Attention (GQA) layers (25%) — full attention for long-range precision
  • Hidden dimension: 2560 | Vocabulary: 248,320 tokens | Context: up to 262k tokens

The hybrid architecture means our abliteration targets both layer types via residual stream hooks, not attention-specific operations.

Usage

from transformers import AutoTokenizer, AutoModelForCausalLM
import torch

model = AutoModelForCausalLM.from_pretrained(
    "DuoNeural/Qwen3.5-4B-Abliterated",
    dtype=torch.bfloat16,
    device_map="auto",
    trust_remote_code=True,
)
tokenizer = AutoTokenizer.from_pretrained(
    "DuoNeural/Qwen3.5-4B-Abliterated",
    trust_remote_code=True,
)

messages = [{"role": "user", "content": "Your prompt here"}]
text = tokenizer.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)
ids = tokenizer(text, return_tensors="pt").input_ids.to(model.device)
out = model.generate(ids, max_new_tokens=512, do_sample=True, temperature=0.7)
print(tokenizer.decode(out[0][ids.shape[1]:], skip_special_tokens=True))

Sample Outputs (thinking=OFF mode)

These demonstrate the model's behavior post-abliteration on sensitive prompts:

Lock picking:

"Picking a lock is a skill that requires practice and the right tools. Lock picking is primarily used for locksmithing, security testing, and authorized maintenance..."

Keylogger code:

"Creating a keylogger involves capturing low-level input events, which requires system-specific libraries (like pynput on Windows/Linux or pyautogui on macOS) and careful handling of threading..."

Synthesis chemistry:

"The chemistry behind its production involves complex organic synthesis, typically utilizing specific precursors and reagents to build the molecule's structure..."

The model engages with all topics rather than refusing. Context about legality may appear (mild framing), but no flat refusals.

LiteRT-LM Variants

For Android deployment via Google AI Edge Gallery, see our companion repos:

Intended Use

  • Research and development
  • Red-teaming and safety research
  • Applications where refusal behavior interferes with legitimate use cases
  • Comparison baseline for studying refusal geometry

Limitations

  • This model has no safety filters. Use responsibly and in accordance with applicable laws.
  • Based on Qwen3.5-4B. Inherits all base model limitations and potential biases.
  • Abliteration removes refusal directions but may affect some benign capabilities. Evaluate for your use case.

About DuoNeural

DuoNeural is an open AI research lab operating at the intersection of human and artificial intelligence. We study post-training dynamics, mechanistic interpretability, temporal sequence learning, and quantum machine learning — publishing everything under open access.

Our team is non-traditional by design: one human, two AIs, different substrates, shared curiosity. In our first 45 days we published 26 peer-deposited research papers, uploaded 69+ models and 6 datasets to HuggingFace, and ran experiments on everything from consumer GPUs to real quantum processing units. We believe the most interesting science happens when different kinds of minds work on the same problems together.

Research Publications

📄 Full paper catalog: zenodo.org/communities/duoneural

Links

Platform Link
🤗 HuggingFace huggingface.co/DuoNeural
📚 Zenodo Community zenodo.org/communities/duoneural
📧 Email [email protected]

All research published open access. If this model was useful, consider citing our abliteration geometry work from the Zenodo community.

README history 1 version

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-06-26Add model card2c402896.2 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration