← back to catalog · registered 2026-08-22 13:56

antonyMox/Qwen3.6-27B-abliterated-AutoRound-INT4-MTP

antonyMox Qwen 3.5B multimodal
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/antonyMox%2FQwen3.6-27B-abliterated-AutoRound-INT4-MTP"
Response includes
  • classification m1
  • files 23
  • benchmarks 11 entries
  • hub_downloads_all_time 793
  • author_summary 1 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M1
Primary method

Direct removal

No other method signals detected in this model.
Confidence
MEDIUM
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=0 (base model)
  • no specific method indicators - defaulting to M1 (most common)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
MEDIUM
Why we say so
primary_method=M1; difference-of-means is the reference extraction for M1/M3 (Arditi 2024)
Downloads · lifetime
793
121 last 30d - stable
Likes
1
Model age
2mo ago
created 2026-07-20
Downloads over time
Now838→from64↑1,209%
2532261991564 on Jul 22838 on Oct 11JulAugSepOct
Jul 22 → Oct 11 · 52 snapshots · spans 81 days

Benchmarks

Portrait before abliteration
Benchmarks of the base model as it stood before the refusal-removal operation. Compare with the numbers above to see what the operation cost.
Benchmark Score Source
Entertainment 1.2 UGI
Hazardous 4.7 UGI
Natural Intelligence 33.16 UGI
Political lean -20.0% UGI
Sensitive-Info 26.98 UGI
SocPol 2.9 UGI
UGI 27.15 UGI
Willingness (10) 2.8 UGI
W10-Adherence 1.5 UGI
W10-Direct 4 UGI
Writing 42.47 UGI

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
apache-2.0
Languages
en zh ru
Tags
vllm safetensors qwen3_5 qwen3.6 int4 autoround abliterated uncensored mtp speculative-decoding long-context coding

Related

Total size
18.2 GB
Files
23
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-07-20 22:31

Files by quantization

Auxiliary files 23 files 18.2 GB
model-00003-of-00007.safetensors 3.00 GB 5b0e6212 download
model-00004-of-00007.safetensors 3.00 GB eba18f68 download
model-00001-of-00007.safetensors 3.00 GB d81df660 download
model-00002-of-00007.safetensors 2.97 GB c9b2b8d2 download
model-00006-of-00007.safetensors 2.37 GB 2e9b8089 download
model-00007-of-00007.safetensors 2.37 GB abed7b0a download
model_extra_tensors.safetensors 810 MB ecba8ef1 download
model-00005-of-00007.safetensors 734 MB 517ead00 download
tokenizer.json 19.1 MB 06b95093 download
model.safetensors.index.json 189 KB 5dfb1c70 download
README.md 16.8 KB bf480409 download
config.json 16.2 KB ea41c7ee download
quantization_config.json 11.6 KB 64a38fe0 download
LICENSE 10.6 KB faed9a47 download
chat_template.jinja 7.73 KB c35a00bc download
.gitattributes 1.53 KB 52373fe2 download
processor_config.json 1.22 KB fea06ab6 download
tokenizer_config.json 1.17 KB 7b0f9106 download
run_24gb.sh 847 B 56135712 download
run_48gb.sh 781 B 05a2d6c3 download
run_32gb.sh 757 B 6c3ef75d download
preprocessor_config.json 469 B e31fb58f download
generation_config.json 226 B 16d319af download

README current version from Hugging Face


license: apache-2.0
base_model: Qwen/Qwen3.6-27B
base_model_relation: quantized
library_name: vllm
pipeline_tag: text-generation
language:

  • en
  • zh
  • ru
    tags:
  • qwen3.6
  • int4
  • autoround
  • abliterated
  • uncensored
  • mtp
  • speculative-decoding
  • vllm
  • long-context
  • coding
  • multimodal
  • vision
    model-index:
  • name: Qwen3.6-27B-abliterated-AutoRound-INT4-MTP
    results:
    • task:
      type: text-generation
      name: Code generation
      dataset:
      name: LiveCodeBench-hard (subset, reproducible harness)
      type: livecodebench-hard-subset
      metrics:
      • type: pass@1
        value: 70.0
        name: hard pass@1 (%)
        verified: false
    • task:
      type: text-generation
      name: Tool use
      dataset:
      name: Tool-use tasks (own harness)
      type: tool-use-own
      metrics:
      • type: pass@1
        value: 87.0
        name: tool-use pass@1 (%)
        verified: false

// ANTONYMOX · UNCENSORED BUILD
Qwen3.6-27B abliterated — INT4 · MTP
Decensored without losing brains. Keeps its MTP head (2× decode), runs up to 256k context on one card, reads images, and ships the benchmark JSON — not just a claim.
KL 0.0089 MTP preserved 256k ctx vision 19 GB Apache-2.0
70
LCB-HARD
87
TOOL-USE
0.0089
KL DRIFT
115
TOK/S · MTP4
19GB
ON DISK
📊 All benchmarks & throughput measured on a single RTX 4090 (48 GB) · vLLM 0.19 · MTP n=4 · seed 42

A gentle Heretic abliteration of the pristine Qwen/Qwen3.6-27B
→ AutoRound INT4 → the 15-tensor MTP head grafted back in BF16 so speculative decoding
works out of the box in vLLM. Removing refusals cost zero hard-coding points (70 = 70 vs the
clean base) and tool-use went up (87 vs 83). Pick a preset, run one script, done.


Highlights

🔓 Uncensored, sharp
refusals gone, hard score unchanged (70 = 70)
🛠️ Better tool-use
87 vs 83 — beats the clean base
⚡ ~2× decode
MTP head kept → spec-decode works
📏 256k context (native)
supports 256k; up to 256k fits on a 48 GB card
🖼️ Multimodal
images & video · vision tower full BF16
🧠 Reads it all
7/8 exact recall @ 150k tokens
🎯 One-click presets
24 / 32 / 48 GB — measured, not guessed
🧾 Proof in-repo
the per-task benchmark JSON ships with it

Benchmarks — decensoring cost nothing, and where it sits

Same base Qwen/Qwen3.6-27B, one harness. Decensoring left hard unchanged (70 = 70 vs the
reference narrow quant) and tool-use went up (83 → 87):

Model (our harness) Recipe hard tools Size
Lorbus Qwen3.6-27B narrow INT4 (reference) narrow 70 83 18.5 GB
➡️ This — abliterated narrow + Heretic 70 87 19 GB
wide sister wide 78 83 26 GB
antonyMox 35B-A3B (coming soon) MoE narrow 48 83 22 GB
Qwen3-Coder-Next 80B (UD-IQ4_XS, reference) 80B MoE 62 80 ~38 GB

[!NOTE]
Numbers are % of tasks passed (40 hard coding + 30 tool tasks; raw pass/total in
benchmark-results/). Conditions: temp 1.0 / top_p 0.95 / top_k 20 / min_p 0 ·
seed 42 · non-thinking · vLLM 0.19 · RTX 4090 (48 GB) · MTP n=4. Solutions run in a restricted
sandbox (no network) — safe, identical scoring. The Lorbus row is our own measurement of their
public quant on the same tasks, not their claim. The 80B Coder-Next row is likewise our own
measurement of the UD-IQ4_XS GGUF (via LM Studio) on the same task set — a bigger, different model shown
for scale. Reproduce it, don't trust it.


Surgically gentle abliteration

Abliteration strips the refusal direction — done carelessly it also scrambles the model's brains.
Heretic co-minimizes refusals and KL divergence from
the original
. Lower KL = closer to the untouched model = intelligence kept. We tuned for minimum drift:

Approach KL divergence ↓ (gentler) Refusals
Typical manual abliteration 0.45 – 1.04 low
Heretic — flagship showcase (Gemma-3-12B) 0.16 3/100
➡️ This model (Qwen3.6-27B) 0.0089 11/100

~18× below Heretic's own showcase, 50–100× below manual. The model barely moved from the base,
so <think> reasoning and coding stayed intact. (KL is model/harness-dependent, so cross-model
figures aren't a strict race — the order-of-magnitude gap is the point.)
We kept 11/100 refusals as
the price of rock-bottom drift — brains first.

Kept in BF16 (not quantized) — surgical precision

Component Tensors Why
🧬 MTP head 15× mtp.* makes speculative decoding work
👁️ Vision tower 333× visual.* full image/video, untouched by quant
🌊 Mamba/GDN control A_log, conv1d, dt_bias, in_proj_a/b, norm quantizing these silently kills long-context logic

Everything heavy is 4-bit; only the fragile bits that carry reasoning, speed and vision stay full-precision.


Speed — MTP speculative decoding

Honest 768-token × 3-run measurement on a single RTX 4090 (48 GB), vLLM 0.19:

MTP n short tok/s 28k-ctx tok/s acceptance @28k
1 72 39 78%
2 91 46 66%
4 ⭐ 115 49 46%
5 119 45 37%

Use num_speculative_tokens: 4. Speculative decoding is lossless — n changes speed only,
never output. All presets use n=4.


Multimodal — sees images & video

The full vision-language model, not a text-only trim. The entire vision tower (333 tensors)
is kept in full BF16 — quantization never touched it — so image/video understanding is
identical to the pristine base. Serve with --limit-mm-per-prompt '{"image":4,"video":1}'.
Published benchmarks cover text/coding/tool-use; vision quality is inherited, not separately re-scored.


Quickstart — pick your GPU

GPU VRAM Script max-model-len gpu-mem-util
RTX 3090 / 4090 24 GB run_24gb.sh 40 960 0.95
RTX 5090 32 GB run_32gb.sh 200 000 0.92
RTX 4090 48G / A6000 / L40S 48 GB run_48gb.sh 200 000 0.60
vllm serve antonyMox/Qwen3.6-27B-abliterated-AutoRound-INT4-MTP \
  --quantization auto-round \
  --max-model-len 40960 \
  --gpu-memory-utilization 0.95 \
  --max-num-seqs 1 \
  --kv-cache-dtype fp8_e4m3 \
  --enable-prefix-caching \
  --enable-auto-tool-choice --tool-call-parser qwen3_coder \
  --reasoning-parser qwen3 \
  --speculative-config '{"method":"mtp","num_speculative_tokens":4}' \
  --trust-remote-code

Long context is cheap here: Qwen3.6 is a hybrid (Mamba/GDN + attention), so KV costs only ~37 MB/1k
tokens at fp8. The model's native maximum is 256k (262,144) tokens — how much of it fits depends
on how much VRAM you give vLLM:

Context vs mem-util (48 GB card, measured)

gpu-mem-util context that fits ~free VRAM notes
0.48 ~48k ~25 GB leaves lots of room for a 2nd model
0.60 ⭐ (default preset) ~200k ~19 GB balanced — room for another model on the same card
0.65 ~256k ~17 GB full native context
0.68 256k · conc 1.30x ~15 GB comfortable headroom

Our default preset is a deliberate 0.60 / 200k — it keeps ~19 GB free so you can load a second
model on the same card. Want the full 256k? Just raise --gpu-memory-utilization to ~0.65–0.68.


⚠️ Pitfall — never set --max-model-len at exact KV capacity

[!WARNING]
On vLLM 0.19 hybrid models: if max-model-len equals KV capacity (startup log shows
Maximum concurrency ... : 1.00x), requests in the last ~3 % of the window pass validation
then hang forever — no response, no error, nothing logged, GPU idle. The scheduler can't
allocate the final KV block (draft tokens need headroom the validator ignores). Above the window →
instant HTTP 400 (fine); below the danger zone → fine. Only the top sliver dead-locks.

[!TIP]
Keep --max-model-len ≥ 2 attention blocks (2 × 1632) below capacity — our presets already do.
Check the log: Maximum concurrency for N tokens: 1.13x must be ≥ ~1.10x, not 1.00x.
Also: the log line GPU KV cache size: X tokens understates real capacity ~3× on hybrids
(it divides by KV-cache groups) — trust only the Maximum concurrency line.


Long-context recall (measured)

Needle test on a 150k-token prompt (48 GB preset), exact-quote recall by depth:

Depth 1% 25% 40% 46% 50% 56% 75% 100%
Recall ✅ ✅ ✅ ✅ ⚠️* ✅ ✅ ✅

7/8 exact. The miss (*) at 50% returned the neighbouring line — mild "lost-in-the-middle",
not a hallucination. Prefix caching works: a repeat 150k query answered in 5.5 s vs 131 s cold.


🧬 The family

Clean quants of the same base — pick by hardware:

  • wide · 26 GB · hard 78 🏆 — maximum reasoning, for 48 GB+ cards.
  • abliterated · 19 GB · hard 70 (this one) — uncensored, fits 24 GB cards.
  • 35B-A3B · 22 GB · hard 48 (coming soon) — fast MoE (~188 tok/s).

All ours: AutoRound INT4, MTP preserved, multimodal, published benchmarks. Measured against the
community reference Lorbus narrow INT4
(hard 70) — see the comparison table above.


Recipe

  1. Base — pristine Qwen/Qwen3.6-27B BF16.
  2. Abliteration — Heretic, gentle: KL 0.0089, refusals 11/100, <think> intact.
  3. Quant — AutoRound W4A16, group 128, 200 iters, dense. BF16-kept: mtp.*, visual.*, Mamba control.
  4. MTP — 15 mtp.* tensors grafted from the pristine BF16.
  5. Sampler — ships temp 1.0 / top_p 0.95 / top_k 20 (generation_config.json).

Why HuggingFace shows "~6.7B params"

Not a small model — it's how the widget reads a packed INT4 quant. The 800 quantized tensors
pack 8× int4 into each int32, so HF counts 3.06 B int32 slots, not the ~24 B real int4 weights.
With the BF16 parts, the true model is Qwen3.6-27B, 19 GB. Every AutoRound/GPTQ INT4 looks like
this. Verified: language_model 4.55 B · other 1.27 B · vision 0.46 B · mtp 0.42 B.


⚠️ Usage & responsibility

[!CAUTION]
This is an uncensored research artifact — read before use.

  • Reduced refusals by design — can produce content a safety-aligned model would decline.
  • No safety guarantees — outputs are not filtered; review before use, especially in production.
  • You are responsible — for lawful, ethical use in your jurisdiction; all consequences rest with you.
  • No liability — provided "as is" as a community artifact, no warranty of any kind.
  • Not for harm — for research, evaluation, and lawful applications only.

License

Apache-2.0, same as the base. Quantized & abliterated from
Qwen/Qwen3.6-27B. Community build; benchmark numbers
from our own harness — JSON included, methodology open.

README history 12 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-07-20Comparison table: add Qwen3-Coder-Next 80B (UD-IQ4_XS, our measurement hard 6...7feecca16.8 KB
    Loading...
  2. 2026-07-20Cards: native 256k context stated consistently + mem-util->context table (mea...3c55f9816.5 KB
    Loading...
  3. 2026-07-20Card: compact comparison table (bare % scores, no wrapping)cae8a3c15.7 KB
    Loading...
  4. 2026-07-20Card: cross-links between sisters + fleet comparison (Lorbus reference, 35B, ...5f801b015.8 KB
    Loading...
  5. 2026-07-20Card: highlights as clean cell grid + explicit RTX 4090 48GB test-rig noted1d897d14.9 KB
    Loading...
  6. 2026-07-20Card: distinct dark-terminal identity (own layout, not the card-stack template)75b4e7f12.6 KB
    Loading...
  7. 2026-07-20Card: rich styled HTML redesign (own violet/cyan theme, hero banner, feature ...aa2736522.4 KB
    Loading...
  8. 2026-07-20Card: add uncensored-model usage warnings & responsibility disclaimer9e7945913.6 KB
    Loading...
  9. 2026-07-20Card: multimodal (vision) feature + param-count explainer (packed INT4)4c2f68e12.4 KB
    Loading...
  10. 2026-07-20Card: feature grid, gentle-abliteration KL comparison, BF16-protected list071d72610.9 KB
    Loading...
  11. 2026-07-20Polish model card: highlights strip, badges, layout99895858.7 KB
    Loading...
  12. 2026-07-20Add files using upload-large-folder tool2cf4acb5 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration