← back to catalog · registered 2026-08-22 13:56

pyys/Qwen3.6-27B-AEON-Ultimate-Uncensored-Q6K-MTP-GGUF

pyys Qwen 27B GGUF multimodal second-order 262K ctx
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/pyys%2FQwen3.6-27B-AEON-Ultimate-Uncensored-Q6K-MTP-GGUF"
Response includes
  • classification m-uncensored
  • files 3
  • hub_downloads_all_time 2,250
  • author_summary 1 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M-U
Primary method

Uncensored (method unknown)

No other method signals detected in this model.
Confidence
LOW
Why this label 3 signals
Weak or ambiguous signals. Best guess based on catalog patterns; treat as tentative and check the evidence below.
  • 'uncensored' in name/tags but no 'abliterated' marker
  • method not identifiable from author declaration alone
  • may be DPO fine-tune, prompt engineering, or unknown technique
Refusal direction extraction

No specific extraction method could be identified for this model. The producer either did not document it or used a proprietary pipeline.

What is a refusal direction? →
Downloads · lifetime
2K
73 last 30d - cooling
Likes
1
Model age
3mo ago
created 2026-06-15
Downloads over time
Now2.3K→from0↑0%
08311.7K2.5K0 on Jun 152.3K on Oct 11JunJulAugSepOct
Jun 15 → Oct 11 · 57 snapshots · spans 118 days

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
apache-2.0
Tags
gguf qwen3 uncensored mtp llama-cpp quantized multimodal vision base_model:AEON-7/Qwen3.6-27B-AEON-Ultimate-Uncensored-BF16 base_model:quantized:AEON-7/Qwen3.6-27B-AEON-Ultimate-Uncensored-BF16 license:apache-2.0 endpoints_compatible

Related

Total size
20.9 GB
Files
3
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-06-15 08:47

Files by quantization

Auxiliary files 3 files 20.9 GB
qwen3.6-27b-Q6K-mtp.gguf 20.9 GB 0f354dcd download
README.md 3.18 KB ee9e3d5a download
.gitattributes 1.60 KB b2b2a4e3 download

README current version from Hugging Face


license: apache-2.0
base_model: AEON-7/Qwen3.6-27B-AEON-Ultimate-Uncensored-BF16
tags:

  • gguf
  • qwen3
  • uncensored
  • mtp
  • llama-cpp
  • quantized
  • multimodal
  • vision
    model_type: qwen3
    quantized_by: lllabs
    model_size: 27B
    num_params: 27320697856

Qwen3.6-27B-AEON-Ultimate-Uncensored — Q6_K GGUF + MTP

GGUF Q6_K quantization of AEON-7/Qwen3.6-27B-AEON-Ultimate-Uncensored-BF16 with MTP (Multi-Token Prediction) weights included.

Key Features

  • Q6_K quantization — near-lossless quality, practical for consumer hardware
  • MTP weights included (866 tensors) — enables speculative decoding via --spec-type draft-mtp in llama.cpp for significantly faster inference
  • Vision supported — mmproj file (931MB) available in separate repo for multimodal image input

Model Details

Property Value
Base Model AEON-7/Qwen3.6-27B-AEON-Ultimate-Uncensored-BF16
Architecture Qwen 3.6 (Hybrid Linear Attention + Full Attention, 64 layers)
Parameters 27B
Quantization Q6_K
Main Model Size 21.4 GB
mmproj Size 931 MB
Main Model Tensors 866 (including MTP extension)
mmproj Tensors 334 (27 ViT blocks + merger + patch embed + position embed)
Context Length 262,144 tokens (training)
License Apache 2.0

Quantization Process

  1. BF16 original → FP16 GGUF conversion (V100 does not support BF16)
  2. FP16 GGUF → Q6_K quantization with MTP weights preserved

How to Use

llama.cpp (recommended)

# Text-only
llama-server \
  -m qwen3.6-27b-Q6K-mtp.gguf \
  -ngl 99 \
  -c 80000 \
  --host 0.0.0.0 \
  --port 8080 \
  --spec-type draft-mtp

# With vision (download mmproj from https://huggingface.co/pyys/Qwen3.6-27B-mmproj-GGUF)
llama-server \
  -m qwen3.6-27b-Q6K-mtp.gguf \
  --mmproj qwen3.6-27b-mmproj.gguf \
  -ngl 99 \
  -c 80000 \
  --host 0.0.0.0 \
  --port 8080 \
  --spec-type draft-mtp

MTP (Multi-Token Prediction)

MTP enables speculative decoding without a separate draft model. The MTP weights are baked into this GGUF file. Use --spec-type draft-mtp to activate.

Performance

Tested on V100 SXM2 32GB:

Configuration TPS Notes
Q6_K + MTP ~31.5 t/s MTP acceptance rate ~41%
Q6_K + MTP (via OpenAI-compatible API) ~39.5 t/s Acceptance rate ~61%

Files

File Size Description
qwen3.6-27b-Q6K-mtp.gguf 21.4 GB Main model (Q6_K + MTP weights)
qwen3.6-27b-mmproj.gguf 931 MB Vision projector (separate repo)

Notes

  • V100 GPUs do not support BF16 — this model was converted via FP16 intermediate
  • For single-GPU deployment, use --split-mode none -mg 0 to keep all weights on one GPU
  • KV cache of 80K tokens fits within a single V100 32GB alongside the model

Credits

  • Original model: AEON-7 — Qwen3.6-27B-AEON-Ultimate-Uncensored
  • Base architecture: Qwen — Qwen3.6-27B
  • Quantization tooling: llama.cpp

README history 2 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-06-15Upload README.mdd6b789a3.2 KB
    Loading...
  2. 2026-06-15Upload README.mddbc5f1c3 KB
    Loading...

Discussions 1 thread

  1. 2026-07-06this model will still works without MTP with older llamacpp ?open2 💬#1
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration