← back to catalog · registered 2026-08-22 13:56

lambsea/Qwen3.6-27B-AEON-Ultimate-Uncensored-UD-GGUF

lambsea Qwen 27B GGUF multimodal second-order 262K ctx
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/lambsea%2FQwen3.6-27B-AEON-Ultimate-Uncensored-UD-GGUF"
Response includes
  • classification m-uncensored
  • files 8
  • hub_downloads_all_time 4,771
  • author_summary 4 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M-U
Primary method

Uncensored (method unknown)

No other method signals detected in this model.
Confidence
LOW
Why this label 3 signals
Weak or ambiguous signals. Best guess based on catalog patterns; treat as tentative and check the evidence below.
  • 'uncensored' in name/tags but no 'abliterated' marker
  • method not identifiable from author declaration alone
  • may be DPO fine-tune, prompt engineering, or unknown technique
Refusal direction extraction

No specific extraction method could be identified for this model. The producer either did not document it or used a proprietary pipeline.

What is a refusal direction? →
Downloads · lifetime
5K
213 last 30d - cooling
Likes
1
Model age
4mo ago
created 2026-05-28
Downloads over time
Now4.8K→from1K↑374%
8282.3K3.8K5.2K1K on Jun 44.8K on Oct 11JunJulAugSepOct
Jun 4 → Oct 11 · 59 snapshots · spans 129 days

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
apache-2.0
Languages
en zh
Quantizations
IQ4 Q5_K Q6_K Q8_0
Tags
gguf quantized qwen3 qwen3.6 hybrid ssm gated-delta-net unsloth-dynamic imatrix mtp vision text-generation

Related

Total size
120 GB
Files
8
Quantizations
6
Registered
2026-08-22 13:56
Last updated on HF
2026-08-17 06:02

Files by quantization

Q8_0 1 file 34.8 GB
Qwen3.6-27B-AEON-UD-Q8_0.gguf 34.8 GB 8fc01dab download
Q6_K 1 file 30.6 GB
Qwen3.6-27B-AEON-UD-Q6_K.gguf 30.6 GB ea75b9c6 download
Q5_K 1 file 28.7 GB
Qwen3.6-27B-AEON-UD-Q5_K_M.gguf 28.7 GB 27b4f4f8 download
IQ4 1 file 25.9 GB
Qwen3.6-27B-AEON-UD-IQ4_XS.gguf 25.9 GB f0f0e48a download
F16 1 file 885 MB
Qwen3.6-27B-AEON-mmproj-F16.gguf 885 MB 4b9a710a download
Auxiliary files 3 files 13.0 MB
imatrix_merged.dat 13.0 MB 6b5d7a0f download
README.md 8.95 KB ddc32fe6 download
.gitattributes 1.87 KB fbacfa3c download

README current version from Hugging Face


license: apache-2.0
language:

  • en
  • zh
    base_model: AEON-7/Qwen3.6-27B-AEON-Ultimate-Uncensored-BF16
    tags:
  • gguf
  • quantized
  • qwen3
  • qwen3.6
  • hybrid
  • ssm
  • gated-delta-net
  • unsloth-dynamic
  • imatrix
  • mtp
  • vision
    model_name: Qwen3.6-27B-AEON-Ultimate-Uncensored-UD-GGUF
    pipeline_tag: text-generation

Qwen3.6-27B-AEON-Ultimate-Uncensored — GGUF (UD Quants)

Unsloth Dynamic-style (UD) GGUF quantizations of AEON-7/Qwen3.6-27B-AEON-Ultimate-Uncensored-BF16

Every quant uses per-tensor overrides (sensitivity-driven) + importance matrix (multi-domain calibration). All SSM recurrence tensors are preserved at source precision. MTP speculative decoding and vision (mmproj) are preserved.


Quant Comparison

File Quant Size tg t/s PPL KL mean KL max KL p99.9
F16 F16 50.9 GB 30.9 5.7215 — — —
UD-Q8_0 Q8_0 34.7 GB 44.7 5.7055 0.0029 5.23 0.22
UD-Q6_K Q6_K 30.6 GB 49.1 5.7046 0.0045 7.64 0.33
UD-Q5_K_M Q5_K_M 28.7 GB 46.9 5.7589 0.0111 5.09 2.04
UD-IQ4_XS IQ4_XS 25.9 GB 56.5 5.7630 0.0236 4.53 1.97

Benchmarked on NVIDIA RTX PRO 6000 Blackwell (96 GB VRAM), llama.cpp fork (a4501150/llama.cpp), pp=512, tg=128.


What Makes These Different

SSM Recurrence Preservation

Qwen3.6 is a hybrid GatedDeltaNet + attention model. 48 of 64 layers use a recurrent SSM where quantization error compounds across token positions. All SSM recurrence tensors are preserved at source precision (F16) — never quantized.

Tensor Count Precision Rationale
ssm_alpha, ssm_beta 96 F16 State update projections — error accumulates in recurrence
ssm_out 48 F16 Output projection feeds directly into residual stream
ssm_a, ssm_conv1d, ssm_dt, ssm_norm 192 F32 Small state tensors (llama-quantize keeps 1D/small tensors at F32)
attn_qkv (SSM input projection) 48 F16 Highest measured KL sensitivity
attn_gate (SSM gate projection) 48 F16 Second-highest measured KL sensitivity

Per-Tensor Sensitivity Analysis

Each tensor group was probed by quantizing only that group to Q4_0 while keeping the rest at F16, then measuring KL divergence. The override generator assigns precision based on measured sensitivity:

Precision Tensor Groups Override Count
F16 SSM recurrence, norms, biases, MTP layer 512
F16 All attention tensors (attn_qkv, attn_gate, attn_v, attn_q, attn_k, attn_output), ffn_down edge 173
Base quant FFN middle layers, FFN edge gate/up, embeddings ~181

685 total overrides — every non-FFN tensor has an explicit precision assignment. No dependence on llama-quantize's internal promotion rules.

Multi-Domain Calibration + GPU Imatrix

Calibrated on a balanced mix across 4 domains from 13 HF datasets:

Domain Token Budget Sources
General 1M ultrachat, OpenHermes, COIG-CQIA, LongAlpaca, pg19, froggeric/imatrix
Code 750K Magicoder-Evol-Instruct-110K
Reasoning 750K OpenMathInstruct-2, OpenR1-Math-220k
Agentic 500K glaive-function-calling-v2, xlam-function-calling-60k, hermes-function-calling-v1

Special tokens from source datasets are stripped automatically. Samples are kept whole — never truncated mid-conversation.

The importance matrix is generated with a PyTorch GPU-native generator (src/generate_imatrix.py) at 65,536 context — uses forward hooks to accumulate squared activations on GPU with zero PCIe D2H copies. Supports multi-GPU via device_map="auto".

Per-domain imatrices are merged with equal weights (DI-MATRIX approach).

MTP + Vision Preserved

  • MTP (Multi-Token Prediction): Draft head (blk.64) pinned at F16. Use --spec-type draft-mtp --spec-draft-n-max 3 for ~1.5-2x faster generation.
  • Vision: mmproj file contains the full vision encoder. Use --mmproj flag with llama-server for image/video understanding.

Files

File Description Size
Qwen3.6-27B-AEON-UD-Q8_0.gguf Highest quality quantization 34.7 GB
Qwen3.6-27B-AEON-UD-Q6_K.gguf Recommended — best quality/size 30.6 GB
Qwen3.6-27B-AEON-UD-Q5_K_M.gguf Balanced 28.7 GB
Qwen3.6-27B-AEON-UD-IQ4_XS.gguf Smallest, for constrained VRAM 25.9 GB
Qwen3.6-27B-AEON-mmproj-F16.gguf Vision encoder (use with --mmproj) 885 MB
imatrix_merged.dat Importance matrix for requantization 13 MB

Usage

llama-server (recommended)

# Q6_K with YaRN 512k context, 5 concurrent slots, MTP + vision
llama-server \
    -m Qwen3.6-27B-AEON-UD-Q6_K.gguf \
    --mmproj Qwen3.6-27B-AEON-mmproj-F16.gguf \
    -ngl 99 \
    --flash-attn \
    -c 524288 \
    --parallel 5 \
    --cache-type-k q8_0 \
    --cache-type-v q8_0 \
    -kvu \
    --cache-ram -1 \
    --rope-scaling yarn \
    --rope-scale 2.0 \
    --yarn-orig-ctx 262144 \
    --override-kv "qwen35.context_length=int:524288" \
    --spec-type draft-mtp \
    --spec-draft-n-max 3 \
    --jinja \
    --chat-template-kwargs '{"enable_thinking":true,"preserve_thinking":true}' \
    --host 0.0.0.0 --port 8080

Note: --spec-type draft-mtp requires llama.cpp b9375+. A custom fork adds DFlash speculative decoding and Blackwell-tuned flash attention.

llama-cli

llama-cli \
    -m Qwen3.6-27B-AEON-UD-Q6_K.gguf \
    -ngl 99 \
    --flash-attn \
    -c 524288 \
    --rope-scaling yarn \
    --rope-scale 2.0 \
    --yarn-orig-ctx 262144 \
    --jinja \
    --chat-template-kwargs '{"enable_thinking":true,"preserve_thinking":true}'

Chat Template Notes

  • enable_thinking activates reasoning mode (chain-of-thought in <think> blocks)
  • preserve_thinking retains reasoning blocks in conversation history
  • No spaces after colons in the JSON — Qwen3.6's template parser is whitespace-sensitive

Architecture

Qwen3.6-27B is a hybrid SSM-attention model:

  • 64 transformer layers + 1 MTP layer (blk.0-64)
  • 48 SSM layers (GatedDeltaNet, no KV cache) + 16 full attention layers (every 4th layer)
  • 27B parameters, 24 attention heads, 4 KV heads, head dim 256
  • Vocab: 248,320 tokens, native context: 262,144 tokens

Quantization Pipeline

Built with super-quant:

  1. Convert HF to F16 GGUF (with MTP tensors) + mmproj GGUF (vision)
  2. Multi-domain calibration data from 13 HF datasets, special tokens stripped
  3. GPU-native importance matrix generation (PyTorch, 65k context) + weighted merge
  4. Per-tensor sensitivity analysis (KL divergence probing against F16 logits)
  5. Hybrid override generation — SSM at source precision, sensitivity-driven for the rest
  6. Quantize with per-tensor overrides + imatrix
  7. Benchmark: throughput + perplexity + KL divergence vs F16

Key Differences from Previous Release

Aspect Previous Current
Calibration Had <|endoftext|> token leak (2,628 occurrences) Special tokens stripped automatically
Imatrix context 32,768 65,536
Imatrix generator llama-imatrix (slow — PCIe D2H per tensor per chunk) PyTorch GPU-native (zero D2H during generation)
Sensitivity analysis llama.cpp subprocesses, ~600 GB disk I/O per run PyTorch in-place weight perturbation, zero disk I/O
SSM alpha/beta F32 (beyond source precision) F16 (matches source BF16)
SSM out Q8_0 (quantized) F16 (source precision preserved)
Attention tensors Mixed (f16/q8_0/q6_k) All F16 (sensitivity-confirmed)
Norms/biases Implicit (llama-quantize internal rules) Explicit F16 overrides
Total overrides 339 685

Links

Credits


License: Apache-2.0

README history 7 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-08-17Upload README.md with huggingface_hub7bed3a28.9 KB
    Loading...
  2. 2026-08-17Add YAML metadata to model card7b0d1d99.1 KB
    Loading...
  3. 2026-08-17Update model card: GPU-native pipeline, SSM preservation, clean calibration206631a8.7 KB
    Loading...
  4. 2026-05-28Update usage: YaRN 512k, q8_0 KV, unified KV, 5 slots, draft-mtp4e601676.7 KB
    Loading...
  5. 2026-05-28Update usage: YaRN 512k, turbo4 KV, unified KV, 5 parallel slots, draft-mtp411b1546.8 KB
    Loading...
  6. 2026-05-28Update README: add preserve_thinking kwarg, HF/GitHub URLs, chat template notesb179e8e6.2 KB
    Loading...
  7. 2026-05-28Upload README.md with huggingface_hub7f5234f5.4 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration