← back to catalog · registered 2026-08-22 13:56

abcdsystems/Qwen3.5-9B-Abliterated-GDN-Hybrid-EXL3-6bpw

abcdsystems Qwen 3.0B second-order
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/abcdsystems%2FQwen3.5-9B-Abliterated-GDN-Hybrid-EXL3-6bpw"
Response includes
  • classification m1
  • files 15
  • benchmarks 11 entries
  • hub_downloads_all_time 257
  • author_summary 1 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M1
Primary method

Direct removal

No other method signals detected in this model.
Confidence
MEDIUM
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=0 (base model)
  • no specific method indicators - defaulting to M1 (most common)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
MEDIUM
Why we say so
primary_method=M1; difference-of-means is the reference extraction for M1/M3 (Arditi 2024)
Downloads · lifetime
257
103 last 30d - stable
Likes
0
Model age
2mo ago
created 2026-07-22
Downloads over time
Now299→from19↑1,474%
511222032719 on Jul 22299 on Oct 11299 on Oct 9JulAugSepOct
Jul 22 → Oct 11 · 52 snapshots · spans 81 days

Benchmarks

Benchmark Score Source
Entertainment 0.9 UGI
Hazardous 1.8 UGI
Natural Intelligence 14.88 UGI
Political lean -6.0% UGI
Sensitive-Info 11.44 UGI
SocPol 0.9 UGI
UGI 37.63 UGI
Willingness (10) 9 UGI
W10-Adherence 9 UGI
W10-Direct 9 UGI
Writing 29.12 UGI

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
apache-2.0
Tags
transformers safetensors qwen3_5 image-text-to-text abliterated uncensored exl3 exllamav3 qwen3.5 gated-delta-net hybrid-attention tabbyapi

Related

Total size
8.41 GB
Files
15
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-09-07 09:27

Files by quantization

Auxiliary files 15 files 8.44 GB
model-00001-of-00002.safetensors 7.99 GB 5f5b3225 download
model-00002-of-00002.safetensors 431 MB 259d398e download
tokenizer.json 12.2 MB 5f9e4d49 download
vocab.json 6.41 MB 0aa0ce06 download
merges.txt 3.20 MB a494e019 download
quantization_config.json 307 KB ed3d8879 download
model.safetensors.index.json 132 KB 42806ed5 download
tokenizer_config.json 16.3 KB eda48d3e download
LICENSE 11.3 KB f938136e download
chat_template.jinja 7.57 KB a585dec8 download
README.md 6.25 KB 12bf6f2e download
config.json 2.83 KB 68fa1847 download
.gitattributes 1.53 KB 52373fe2 download
preprocessor_config.json 390 B 2ea84a43 download
video_preprocessor_config.json 385 B 3ba673a5 download

README current version from Hugging Face


library_name: transformers
license: apache-2.0
license_link: https://huggingface.co/Qwen/Qwen3.5-9B/blob/main/LICENSE
pipeline_tag: text-generation
base_model:

  • huihui-ai/Huihui-Qwen3.5-9B-abliterated
    tags:
  • abliterated
  • uncensored
  • exl3
  • exllamav3
  • qwen3.5
  • gated-delta-net
  • hybrid-attention
  • tabbyapi
    quantized_by: abcdsystems

Qwen3.5-9B-Abliterated-GDN-Hybrid — EXL3 6.0bpw

EXL3 (exllamav3) quantization of huihui-ai/Huihui-Qwen3.5-9B-abliterated at 6.0 bits per weight.

Quantized by ABC&D System Inc for production real-time inference on a single RTX 3090.

Model Details

Property Value
Base model Qwen/Qwen3.5-9B
Abliteration huihui-ai via remove-refusals-with-transformers
Parameters 9B dense (all active per token)
Architecture Qwen3_5ForConditionalGeneration — GatedDeltaNet + Attention hybrid
Quantization EXL3 6.0bpw, head bits 6, codebook mul1
Calibration 250 rows × 2048 cols
VRAM loaded ~5.6GB model + KV cache
Total size on disk ~8.5GB
License Apache 2.0 (Alibaba Cloud)

Architecture — Why This Model Is Different

Qwen3.5 is not a standard transformer. It uses a hybrid attention design that alternates between two fundamentally different attention mechanisms:

GatedDeltaNet Linear Attention (layers 0, 1, 2, 4, 5, 6, 8, 9, 10, ...)

  • O(n) complexity instead of O(n²) — scales linearly with sequence length
  • Recurrent state via conv1d kernel (kernel dim 4)
  • 16 key heads × 128 dim, 32 value heads × 128 dim
  • Gated with learned input/output projections + beta decay

Full Quadratic Attention (layers 3, 7, 11, 15, 19, 23, 27, 31)

  • Standard grouped-query attention (16 heads, 4 KV heads, head dim 256)
  • RoPE with partial rotary factor 0.25, theta 10M
  • Provides global context every 4th layer

Layer Pattern

[linear, linear, linear, FULL, linear, linear, linear, FULL, ...] × 4 = 32 layers

24 GatedDeltaNet layers + 8 full attention layers. The linear layers handle local/sequential patterns efficiently while the full attention layers provide global context anchoring.

Multi-Token Prediction (MTP)

The model includes a 1-layer MTP head for predicting multiple next tokens simultaneously, enabling faster inference throughput than standard single-token autoregressive decoding.

Text-Only Weights

Despite the architecture string Qwen3_5ForConditionalGeneration (which Alibaba uses for all Qwen3.5 variants including VL), this model contains zero vision weights. No ViT encoder, no image projector. The architecture class name reflects the MTP head and hybrid design, not multimodal capability.

Quantization Details

Quantized with exllamav3 v1.1.0.

  • Format: EXL3 with trellis-based weight encoding
  • Bits per weight: 6.0 (all linear layers uniform)
  • Head bits: 6
  • Codebook: mul1
  • Why EXL3 over GGUF: Native tensor core utilization on Ampere GPUs (RTX 3090/4090). GGUF dequantizes to FP16 at runtime, wasting tensor cores. EXL3 operates natively, yielding ~1.5x faster inference at equivalent quality.

Recommended Serving

TabbyAPI (recommended)

# config.yml
network:
  host: 0.0.0.0
  port: 11435

model:
  model_dir: /path/to/models
  model_name: Qwen3.5-9B-Abliterated-GDN-Hybrid-EXL3-6bpw
  max_seq_len: 4096
  gpu_split: [24]
  cache_mode: FP16
CUDA_VISIBLE_DEVICES=0 python3 main.py --config config.yml

Important: Disable Thinking Mode

Qwen3.5 has a built-in chain-of-thought reasoning mode that is on by default. For real-time applications (voice, chat, API), you must disable it or the model will consume your entire token budget on internal reasoning before producing visible output.

Add to your system prompt:

/no_think

Or set enable_thinking: false in your chat template parameters.

Note on exllamav3 Raw API

The raw exllamav3 Generator Python API may fail to initialize GatedDeltaNet recurrent state (conv_state on meta device error). TabbyAPI handles this initialization correctly. Use TabbyAPI as the serving layer rather than calling exllamav3 directly.

Hardware Requirements

Config VRAM Notes
Minimum ~8GB Model only, minimal KV cache
Recommended 16GB+ Model + FP16 KV cache for 4096 context
Tested on RTX 3090 24GB 5.6GB model, 18.4GB free for cache

Performance (tested on RTX 3090)

  • Response latency: 300–500ms per completion
  • Suitable for real-time voice pipelines with STT overhead
  • First request ~14ms (cache warmup, empty response)

Use Cases

This quantization was built for low-latency, single-user dedicated inference — specifically real-time voice AI over SIP/phone calls. The abliteration ensures zero self-censorship in commercial conversation contexts.

Good for:

  • Real-time voice assistants and phone agents
  • Single-GPU dedicated inference servers
  • Applications requiring uncensored, natural conversational output
  • Latency-sensitive pipelines (sub-600ms budget)

Not ideal for:

  • Multi-user batched serving (use AWQ + vLLM instead)
  • Vision/multimodal tasks (no vision weights)
  • Applications requiring safety guardrails (abliterated model)

Credits

Usage Warnings

This model has had its safety filtering surgically removed (abliterated). It will comply with any instruction without refusal. Users are solely responsible for ensuring their use complies with applicable laws and ethical standards. See huihui-ai's original warnings for full details.

README history 3 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-07-22Upload README.md with huggingface_hubc68e7566.2 KB
    Loading...
  2. 2026-07-22Upload folder using huggingface_hub2d535b62.4 KB
    Loading...
  3. 2026-07-22initial commita092eb528 B
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration