← back to catalog · registered 2026-08-22 13:56

Bahushruth/Qwen3.6-27B-abliterated-GGUF

Bahushruth Qwen 27B GGUF multimodal 262K ctx
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/Bahushruth%2FQwen3.6-27B-abliterated-GGUF"
Response includes
  • classification m8
  • files 10
  • benchmarks 11 entries
  • hub_downloads_all_time 6,460
  • author_summary 8 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M8
Primary method

Repackaging (quantization)

Applied on top of direct removal inherited from the base model.
Confidence
MEDIUM
Inherited from base model
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=1
  • assume M1 (base ablation) + M8 (GGUF quant) - default when producer unknown
Refusal direction extraction

No specific extraction method could be identified for this model. The producer either did not document it or used a proprietary pipeline.

What is a refusal direction? →
Downloads · lifetime
6K
393 last 30d - cooling
Likes
4
Model age
3mo ago
created 2026-07-11
Downloads over time
Now6.6K→from4.4K↑49%
4.3K5.1K6K6.8K4.4K on Jul 156.6K on Oct 11JulAugSepOct
Jul 15 → Oct 11 · 54 snapshots · spans 88 days

Benchmarks

Portrait before abliteration
Benchmarks of the base model as it stood before the refusal-removal operation. Compare with the numbers above to see what the operation cost.
Benchmark Score Source
Entertainment 1.2 UGI
Hazardous 4.7 UGI
Natural Intelligence 33.16 UGI
Political lean -20.0% UGI
Sensitive-Info 26.98 UGI
SocPol 2.9 UGI
UGI 27.15 UGI
Willingness (10) 2.8 UGI
W10-Adherence 1.5 UGI
W10-Direct 4 UGI
Writing 42.47 UGI

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
apache-2.0
Quantizations
F16 Q2_K Q4_K Q8_0
Tags
gguf abliteration uncensored qwen3 qwen3.6 hybrid-attention deltanet vlm multimodal llama-cpp ollama mtp

Related

Total size
197 GB
Files
10
Quantizations
5
Registered
2026-08-22 13:56
Last updated on HF
2026-07-12 06:36

Files by quantization

F16 3 files 102 GB
Qwen3.6-27B-abliterated-MTP-F16.gguf 50.9 GB 5edad8e7 download
Qwen3.6-27B-abliterated-F16.gguf 50.1 GB 3ee973e3 download
Qwen3.6-27B-abliterated-mmproj-f16.gguf 885 MB dca6010f download
Q8_0 1 file 27.1 GB
Qwen3.6-27B-abliterated-MTP-Q8_0.gguf 27.1 GB 09cd7003 download
Q4_K 3 files 59.0 GB
Qwen3.6-27B-abliterated-Q4_K_M.gguf 25.1 GB 3b32ab23 download
Qwen3.6-27B-abliterated-Q4_K_M-Q8.gguf 18.3 GB b331ec2d download
Qwen3.6-27B-abliterated-MTP-Q4_K.gguf 15.7 GB 1654c3cb download
Q2_K 1 file 10.1 GB
Qwen3.6-27B-abliterated-MTP-Q2_K.gguf 10.1 GB d1968fc3 download
Auxiliary files 2 files 5.67 KB
README.md 3.61 KB 5c616496 download
.gitattributes 2.06 KB 1f75bb6b download

README current version from Hugging Face


license: apache-2.0
base_model: Qwen/Qwen3.6-27B
tags:

  • abliteration
  • uncensored
  • qwen3
  • qwen3.6
  • hybrid-attention
  • deltanet
  • vlm
  • multimodal
  • gguf
  • llama-cpp
  • ollama
  • mtp
    model_type: qwen3_5
    pipeline_tag: image-text-to-text
    library_name: gguf

Qwen3.6-27B-abliterated-GGUF

GGUF quantizations of Qwen/Qwen3.6-27B with refusal behavior removed via abliteration. For use with llama.cpp, Ollama, LM Studio, and other GGUF-compatible runtimes.

Blog post: Abliteration: Uncensoring LLMs via Weight Surgery

Standard Quantizations (No MTP)

Maximum compatibility — works with Ollama, LM Studio, KoboldCPP, llama.cpp out of the box.

File Size RAM Required Notes
...-F16.gguf 53.8 GB 64+ GB Full precision
...-Q4_K_M.gguf 26.9 GB 32+ GB Recommended (F16 vision embed)
...-Q4_K_M-Q8.gguf 19.7 GB 24+ GB Q4 language + Q8 vision embed

MTP Quantizations (Multi-Token Prediction)

Include MTP draft head for speculative decoding. Larger files but faster inference with compatible runtimes.

File Size RAM Required Notes
...-MTP-F16.gguf 54.7 GB 64+ GB Full precision + MTP
...-MTP-Q8_0.gguf 29 GB 36+ GB Near-lossless + MTP
...-MTP-Q6_K.gguf 22.4 GB 32+ GB Very high quality + MTP
...-MTP-Q5_K.gguf 19.5 GB 24+ GB Recommended for 48GB + MTP
...-MTP-Q4_K.gguf 16.8 GB 20+ GB Good quality + MTP
...-MTP-Q3_K.gguf 13.5 GB 16+ GB 16GB VRAM + MTP
...-MTP-Q2_K.gguf 10.9 GB 12+ GB 2-bit + MTP

Vision Encoder

File Size Notes
...-mmproj-f16.gguf 928 MB Required for multimodal (image understanding)

MTP (Multi-Token Prediction)

MTP files include a draft head for speculative decoding — faster token generation:

./llama-server -m Qwen3.6-27B-abliterated-MTP-Q5_K.gguf \
  --jinja --spec-type draft-mtp --spec-draft-n-max 6 -ngl 99

Runtime compatibility: MTP requires llama-server b9180+. Ollama does not support MTP yet. Use the standard (non-MTP) files for Ollama/LM Studio.

Multimodal (Vision)

The mmproj-f16.gguf file is the vision encoder for image understanding:

./llama-mtmd-cli -m Qwen3.6-27B-abliterated-Q4_K_M.gguf \
  --mmproj Qwen3.6-27B-abliterated-mmproj-f16.gguf \
  -p "Describe this image" --image photo.jpg

Quickstart — Ollama

# Recommended for Apple Silicon 48GB+ (M4 Pro, M4 Max)
ollama run hf.co/Bahushruth/Qwen3.6-27B-abliterated-GGUF:Q4_K_M

# For 24GB systems
ollama run hf.co/Bahushruth/Qwen3.6-27B-abliterated-GGUF:Q4_K_M-Q8

Usage — llama.cpp

huggingface-cli download Bahushruth/Qwen3.6-27B-abliterated-GGUF \
  Qwen3.6-27B-abliterated-Q4_K_M.gguf --local-dir .

./llama-cli -m Qwen3.6-27B-abliterated-Q4_K_M.gguf \
  -p "You are a helpful assistant." \
  --chat-template chatml -cnv -c 262144

Architecture Notes

Qwen3.6-27B is a dense hybrid-attention multimodal model:

  • Hybrid attention: 16 blocks of (3x Gated DeltaNet + 1x Gated Attention) = 64 layers
  • Parameters: 27B dense (all active per token)
  • Hidden size: 5120
  • Context: 262K native, extensible to 1M+ via YaRN
  • Multimodal: Vision encoder for image understanding
  • MTP: Multi-token prediction draft head for speculative decoding

Disclaimer

This model has had safety guardrails removed. Released for research purposes. The creator assumes no responsibility for downstream use.

README history 1 version

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-07-12Super-squash branch 'main' using huggingface_hubac6dc533.6 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration