← back to catalog · registered 2026-08-22 13:56

akqmffl/qwen3.6-27b-Q8_0-MTP-Claude-Opus-Reasoning-uncensored-GGUF

akqmffl Qwen 27B GGUF 262K ctx
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/akqmffl%2Fqwen3.6-27b-Q8_0-MTP-Claude-Opus-Reasoning-uncensored-GGUF"
Response includes
  • classification m-uncensored
  • files 3
  • hub_downloads_all_time 2,002
  • author_summary 1 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M-U
Primary method

Uncensored (method unknown)

No other method signals detected in this model.
Confidence
LOW
Why this label 3 signals
Weak or ambiguous signals. Best guess based on catalog patterns; treat as tentative and check the evidence below.
  • 'uncensored' in name/tags but no 'abliterated' marker
  • method not identifiable from author declaration alone
  • may be DPO fine-tune, prompt engineering, or unknown technique
Refusal direction extraction

No specific extraction method could be identified for this model. The producer either did not document it or used a proprietary pipeline.

What is a refusal direction? →
Downloads · lifetime
2K
63 last 30d - cooling
Likes
9
Model age
3mo ago
created 2026-06-17
Downloads over time
Now2K→from242↑740%
07451.5K2.2K242 on Jun 172K on Oct 112K on Oct 10JunJulAugSepOct
Jun 17 → Oct 11 · 57 snapshots · spans 116 days

Metadata

License
apache-2.0
Tags
gguf qwen qwen3 qwen3.6 mtp native-mtp target-model-aligned-mtp-distillation q8_0 llama.cpp lemonade rocm gfx1151
Total size
27.1 GB
Files
3
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-06-17 14:42

Files by quantization

Auxiliary files 3 files 27.1 GB
qwen3.6-27b-Q8_0-MTP-Claude-Opus-Reasoning-uncensored.gguf 27.1 GB bb0a4a05 download
README.md 7.23 KB 4ca46550 download
.gitattributes 43.0 B ae756a3f download

README current version from Hugging Face


license: apache-2.0
library_name: gguf
pipeline_tag: text-generation
tags:

  • gguf
  • qwen
  • qwen3
  • qwen3.6
  • mtp
  • native-mtp
  • target-model-aligned-mtp-distillation
  • q8_0
  • llama.cpp
  • lemonade
  • rocm
  • gfx1151
  • rdna3.5
  • ryzen-ai-max
  • uncensored
  • conversational

Qwen3.6 27B Q8_0 Native MTP Claude-Opus-Reasoning Uncensored GGUF

Important runtime requirement: to reproduce the benchmark performance for
this GGUF, download and configure
akqmffl/llama.cpp-lemonade-Runtime-for-qwen-3.6-MTP.
Benchmark settings, release downloads, launch scripts, Lemonade integration,
runtime environment variables, and supported hardware notes are maintained in
that GitHub repository. Visit the GitHub README and release notes for the
current setup procedure and benchmark configuration.

This repository contains a Q8_0 GGUF build of
qwen3.6-27b-Q8_0-MTP-Claude-Opus-Reasoning-uncensored.gguf. It is intended
for the custom llama.cpp / Lemonade ROCm runtime above, especially the Windows
AMD Ryzen AI Max Series APU / RDNA 3.5 / gfx1151 path documented there.

Preserved MTPs

This GGUF preserves the Native MTP tensor set. The model file was checked at the
GGUF tensor table level and contains 15 MTP-related tensors:

  1. blk.64.attn_k.weight
  2. blk.64.attn_k_norm.weight
  3. blk.64.attn_norm.weight
  4. blk.64.attn_output.weight
  5. blk.64.attn_q.weight
  6. blk.64.attn_q_norm.weight
  7. blk.64.attn_v.weight
  8. blk.64.ffn_down.weight
  9. blk.64.ffn_gate.weight
  10. blk.64.ffn_up.weight
  11. blk.64.nextn.eh_proj.weight
  12. blk.64.nextn.enorm.weight
  13. blk.64.nextn.hnorm.weight
  14. blk.64.nextn.shared_head_norm.weight
  15. blk.64.post_attention_norm.weight

The MTP component was trained and applied using target-model-aligned MTP
distillation
against the qwen3.6-27b-Claude-Opus-Reasoning-uncensored
target model. In other words, the MTP path is intended to follow the target
model's next-token distribution for speculative decoding rather than behaving
as an unrelated draft model.

Model Overview

Item Value
Format GGUF v3
Quantization Q8_0
File qwen3.6-27b-Q8_0-MTP-Claude-Opus-Reasoning-uncensored.gguf
File size 27.05 GiB
SHA-256 BB0A4A050D598FD88D570B326B8331B9B11767ACB10E9CC8B0B899F4C3350539
Tensor count 866
Metadata KV count 40
Native MTP tensor count 15
Target alignment Target-model-aligned MTP distillation
Intended runtime Custom llama.cpp / Lemonade ROCm runtime for gfx1151

Abliteration Parameters

This GGUF is packaged from the
qwen3.6-27b-Claude-Opus-Reasoning-uncensored target-model line. This model
card does not claim a new ablation pass or publish separate ablation parameter
values for this GGUF. The important alignment note for this upload is that the
Native MTP component was trained against that uncensored target model by
target-model-aligned MTP distillation before being applied to the Q8_0 GGUF.

Performance

The performance-sensitive path depends on both the GGUF and the custom runtime.
Use the GitHub package for exact setup and benchmark reproduction:

Recorded runtime-package anchors:

Lane Shape Recorded result
n=2 production 128K context, 4096 generated tokens, FA off 15.131 tok/s, 92.286% acceptance
n=2 reasoning off 128K context, 4096 generated tokens, FA off 15.413 tok/s, 93.182% acceptance
n=3 chat stream 8192 context, 100 generated tokens, FA off 18.050 tok/s, 97.333% acceptance, exact at 100 tokens

AMD Ryzen AI Max / RDNA 3.5 / gfx1151 Q8 Baselines

The comparison set below is restricted to Qwen3.6-27B-MTP-Q8 and
Qwen3.6-27B-MTP-UD-Q8_K_XL public rows on AMD Ryzen AI Max / Strix Halo /
Radeon 8060S / RDNA 3.5 / gfx1151 class hardware that report both generation
speed and MTP draft acceptance. Non-Q8 rows are intentionally excluded from this
table.

Source Hardware / backend Model and runtime Workload tok/s MTP acceptance Comparison note
This upload, short stream AMD Ryzen AI Max Series APU / RDNA 3.5 / gfx1151, custom ROCm runtime Qwen3.6 27B Q8_0 Native MTP, n=3, FA off 8192 context, 100 generated tokens 18.050 97.333%, exact at 100 tokens Release anchor for short interactive generation.
This upload, long output AMD Ryzen AI Max Series APU / RDNA 3.5 / gfx1151, custom ROCm runtime Qwen3.6 27B Q8_0 Native MTP, n=2, FA off 128K context, 4096 generated tokens 15.413 93.182% Release anchor for long 128K-context output.
Reddit AMD ROCm Windows PR 22673 report Ryzen AI Max+ 395, Radeon 8060S iGPU, RDNA 3.5 gfx1151, Windows 11 ROCm Qwen3.6-27B-MTP-UD-Q8_K_XL.gguf, PR 22673 llama.cpp build 128K context, q8_0 KV, thinking on, draft-MTP 12.13 tokens/sec 64-69% Public same-hardware UD-Q8_K_XL row with both speed and acceptance.
Qiita AMD Ryzen AI MAX+ 395 Qwen3.6 MTP Q8_0 run AMD Ryzen AI MAX+ 395 / gfx1151, llama.cpp ROCm Qwen3.6-27B-MTP-Q8_0.gguf, draft-MTP 8192 context, 142 output tokens, temperature 0.7, seed 42 10.15 tokens/sec 60/243 = 24.7% Public Q8_0 same-hardware row; acceptance is much lower than this release anchor.

A static HTML comparison view is included in
assets/amd-gfx1151-benchmark-comparison.html.

These anchors are hardware- and runtime-specific. A generic llama.cpp build may
load the GGUF, but it should not be expected to match the recorded MTP
throughput without the runtime package and settings documented in the GitHub
repository.

Files

File Description
qwen3.6-27b-Q8_0-MTP-Claude-Opus-Reasoning-uncensored.gguf Q8_0 GGUF with Native MTP preserved and target-model-aligned MTP distillation applied

Usage

  1. Download the GGUF from this repository.
  2. Download the matching runtime package from
    akqmffl/llama.cpp-lemonade-Runtime-for-qwen-3.6-MTP.
  3. Follow the GitHub README for standalone llama.cpp or Lemonade setup.
  4. Use this GGUF path as the model path in the GitHub package's launch scripts
    or Lemonade configuration.

Reference

This model-card structure was prepared after reviewing the Hugging Face GGUF
reference repository
llmfan46/Qwen3.6-27B-uncensored-heretic-v2-Native-MTP-Preserved-NVFP4-GGUF.
Only the structure was used; benchmark/quality metrics from that repository are
not copied into this card.

License

Apache-2.0 metadata is declared to match the referenced Qwen/GGUF lineage style.
Users remain responsible for verifying upstream model licensing, local
deployment requirements, and acceptable-use constraints for their own
environment.

README history 5 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-06-17Restore Reddit UD-Q8_K_XL comparison row6dc6f247.2 KB
    Loading...
  2. 2026-06-17Keep Qwen3.6 MTP Q8 comparison rows onlydfbfeb26.8 KB
    Loading...
  3. 2026-06-17Use direct AMD gfx1151 tok/s acceptance baselinese061f539.4 KB
    Loading...
  4. 2026-06-17Add AMD gfx1151 benchmark baselinesdda71e58.4 KB
    Loading...
  5. 2026-06-17Add files using upload-large-folder tool7d472225.3 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration