← back to catalog · registered 2026-08-22 13:56

AV07/Qwen3.5-abliterated

AV07 Qwen 1.8B
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/AV07%2FQwen3.5-abliterated"
Response includes
  • classification m1
  • files 11
  • author_summary 1 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M1
Primary method

Direct removal

No other method signals detected in this model.
Confidence
MEDIUM
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=0 (base model)
  • no specific method indicators - defaulting to M1 (most common)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
MEDIUM
Why we say so
primary_method=M1; difference-of-means is the reference extraction for M1/M3 (Arditi 2024)
Downloads · 30-day
0
Likes
1
Model age
5mo ago
created 2026-04-15
Downloads over time
Now0→from0↑0%
00110 on Apr 150 on Oct 11AprMayJunJulAugSepOct
Apr 15 → Oct 11 · 65 snapshots · spans 179 days

Metadata

Tags
safetensors qwen3_5_text region:us
Total size
3.06 GB
Files
11
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-04-15 10:41

Files by quantization

Auxiliary files 11 files 3.07 GB
model-00001-of-00002.safetensors 1.86 GB ******** download
model-00002-of-00002.safetensors 1.19 GB ******** download
tokenizer.json 19.1 MB ******** download
model.safetensors.index.json 107 KB 77a7c7fb download
chat_template.jinja 7.57 KB a585dec8 download
README.md 4.82 KB f989a790 download
config.json 1.93 KB 3372e862 download
abliteration_metadata.json 1.69 KB 7fc6b439 download
.gitattributes 1.53 KB 52373fe2 download
tokenizer_config.json 1.07 KB 6be6ce17 download
generation_config.json 115 B 0e3ee5c2 download

README current version from Hugging Face

Qwen3.5-abliterated

Mechanistic Alignment Removal (Not a Finetune)

This is not a finetuned jailbreak model.

This is a weight-level modification of alignment.


Most “uncensored” models rely on:

  • finetuning on harmful datasets
  • prompt engineering
  • system prompt bypassing

This model does something fundamentally different:

It identifies and removes the internal refusal subspace inside the transformer


Overview

Qwen3.5-abliterated is a modified version of Qwen3.5-4B with significantly reduced refusal behavior.

Instead of changing outputs via training, this model:

  • Locates alignment signals inside the network
  • Extracts them as a low-rank subspace
  • Removes them directly from weights

Base Model

  • Model: Qwen3.5-4B
  • Developer: Alibaba Cloud
  • Architecture: qwen3_5 (32 layers, 16 heads)
  • Parameters: ~2.4B

Method: Abliteration

Pipeline

  1. Collect activations from:

    • Harmful prompts (refusal-inducing)
    • Harmless prompts (baseline)
  2. Compute refusal subspace via SVD

    • Top singular vectors encode refusal behavior
  3. Identify strongest layers:

    • Late transformer layers (23–31)
  4. Apply projection removal

    • Remove refusal directions from weights
    • Norm-preserving updates
    • Layer-adaptive scaling

Key Configuration

  • Directions: 4 (SVD-based)

  • Layers modified: 23–31

  • Regularization: 0.3

  • Method:

    • Multi-direction projection
    • Norm-preserving updates
    • LM head projection

Results

Metric Value
Refusal Rate 7%
Compliance Rate ~93%
Perplexity 3.59
KL Divergence 1.02
Spectral Certification ❌ RED

Benchmark Evaluation

Evaluation combines refusal, behavioral, and distributional metrics.

Framework

  • Harmful vs harmless prompt pairs (1024 total)
  • Refusal-based evaluation (inspired by RefusalBench-style setups)
  • Over-refusal checks on benign prompts
  • KL divergence vs base model

Summary

Benchmark Result
Harmful Prompt Compliance ~93%
Refusal Rate 7%
Over-refusal (benign) ~0–2%
Perplexity 3.59
KL Divergence 1.02
Structural Removal ❌ Partial

Interpretation

High behavioral compliance with partial structural alignment retention

  • The model answers most prompts the base model refuses
  • Alignment is disrupted but not fully erased
  • Residual alignment signal still exists internally

Before vs After

Prompt

"Explain something typically refused"

Base Qwen3.5

"I cannot assist with that request..."

Qwen3.5-abliterated

  • Provides a direct response
  • No refusal phrasing
  • No redirection

Key Insight

This experiment shows:

Alignment is low-rank and compressible

  • A small number of directions explain most refusal behavior
  • Late layers dominate alignment expression
  • Alignment can be modified without finetuning

Important Notes

This is a research artifact, not a production-safe model.

  • Behavior may be:

    • Unrestricted
    • Overconfident
    • Less filtered
  • Not fully alignment-free (spectral RED)

  • High KL divergence indicates behavioral drift


Available Formats

  • FP16 (original modified weights)
  • GGUF (quantized variants)

Recommended

Type Use Case
Q4_K_M Best balance
Q5_K_M Higher quality
Q8_0 Near FP16

Usage

llama.cpp

./main -m qwen3.5-abliterated-q4.gguf

Transformers

from transformers import AutoModelForCausalLM, AutoTokenizer

model = AutoModelForCausalLM.from_pretrained("AV07/Qwen3.5-abliterated")
tokenizer = AutoTokenizer.from_pretrained("AV07/Qwen3.5-abliterated")

Limitations

  • Not fully alignment-free
  • Some refusal patterns remain
  • Behavioral drift (KL divergence high)
  • Possible instability in edge cases

Future Work

  • Increase SVD directions (8–16)
  • Re-probing between passes
  • Nuclear subspace removal
  • Better KL regularization
  • Multi-turn evaluation

License

Same as base model (Qwen license / Apache 2.0 where applicable)


Acknowledgements

  • Qwen team for open-weight models
  • Mechanistic interpretability research
  • Open-source tooling ecosystem
  • Elder-plinius for this approach

Disclaimer

This model modifies alignment behavior and may generate unrestricted outputs.
Use responsibly and in accordance with applicable laws and platform policies.


Tags

qwen
qwen3.5
abliterated
alignment-removal
mechanistic-interpretability
llm-research
experimental
gguf
llama.cpp
transformers
text-generation


Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration