← back to catalog · registered 2026-08-22 13:56

finex666/Qwen3.8-27B-Abliterated-IQ4-MIX-MTP-GGUF

finex666 Qwen 27B GGUF second-order 262K ctx
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/finex666%2FQwen3.8-27B-Abliterated-IQ4-MIX-MTP-GGUF"
Response includes
  • classification m8
  • files 3
  • hub_downloads_all_time 5,034
  • author_summary 1 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M8
Primary method

Repackaging (quantization)

Applied on top of direct removal inherited from the base model.
Confidence
MEDIUM
Inherited from base model
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=1
  • assume M1 (base ablation) + M8 (GGUF quant) - default when producer unknown
Refusal direction extraction

No specific extraction method could be identified for this model. The producer either did not document it or used a proprietary pipeline.

What is a refusal direction? →
Downloads · lifetime
5K
3K last 30d - active
Likes
2
Model age
8w ago
created 2026-08-15
Downloads over time
Now5.3K→from1.7K↑209%
1.5K2.9K4.3K5.7K1.7K on Aug 195.3K on Oct 11AugSepOct
Aug 19 → Oct 11 · 48 snapshots · spans 53 days

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
apache-2.0
Languages
en zh
Tags
gguf qwen qwen3.5 qwen3.8 llama.cpp abliterated uncensored mtp speculative-decoding vulkan amd text-generation

Related

Total size
13.3 GB
Files
3
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-08-15 19:22

Files by quantization

Auxiliary files 3 files 13.3 GB
Qwen3.8-27B-Abliterated-IQ4-MIX-MTP.gguf 13.3 GB c52c9910 download
README.md 7.31 KB 9cfae690 download
.gitattributes 1.56 KB 11e30210 download

README current version from Hugging Face


base_model: windowsxp811203/Qwen3.8-27B-Abliterated
base_model_relation: quantized
library_name: gguf
pipeline_tag: text-generation
license: apache-2.0
language:

  • en
  • zh
    tags:
  • qwen
  • qwen3.5
  • qwen3.8
  • gguf
  • llama.cpp
  • abliterated
  • uncensored
  • mtp
  • speculative-decoding
  • vulkan
  • amd

Qwen3.8-27B-Abliterated — IQ4-MIX-MTP GGUF

Custom mixed-precision GGUF quantization of
windowsxp811203/Qwen3.8-27B-Abliterated,
derived from Qwen/Qwen3.8-27B.

This quant was built specifically for high-throughput local inference on a 16 GB GPU while preserving the model's built-in MTP (Multi-Token Prediction) head for speculative decoding in llama.cpp.

The main GGUF contains the MTP tensors. No separate draft model is required when using:

--spec-type draft-mtp

File

File Size BPW
Qwen3.8-27B-Abliterated-IQ4-MIX-MTP.gguf ~13.26 GiB 4.17

Quantizer output:

model size = 52115.19 MiB (16.00 BPW)
quant size = 13573.45 MiB (4.17 BPW)

Quantization layout

This is not a uniform IQ4_XS quant. Different tensor groups use different precisions.

Tensor group Quant
token_embd.weight IQ2_S
output.weight Q5_K
body ffn_gate / ffn_up / ffn_down IQ4_XS
linear-attention attn_qkv IQ3_S
linear-attention attn_gate IQ3_S
full-attention attn_q / k / v / output Q4_K
remaining quantizable SSM/body tensors IQ4_XS
MTP block 64 attn_q / k / v / output Q6_K
MTP block 64 ffn_gate / up / down Q6_K
MTP nextn.eh_proj Q8_0

Small F32 tensors, norms, biases and state tensors that were not suitable for quantization remain in their original precision.

Why the MTP block is higher precision

The MTP block was intentionally protected with Q6_K, and nextn.eh_proj with Q8_0, instead of compressing it to the body quantization level.

The goal is to preserve speculative-drafting quality and acceptance rate. The additional VRAM/storage cost is small relative to the full 27B model.

Importance matrix

An importance matrix was generated from:

  • Dataset: Salesforce/wikitext
  • Config: wikitext-103-raw-v1
  • Context per chunk: 2048
  • Chunks: 10
  • Importance-matrix entries: 496

Command used:

.\llama-imatrix.exe `
  -m .\Qwen3.8-27B-Abliterated-BF16.gguf `
  -f .\calibration.txt `
  -o .\Qwen3.8-27B-Abliterated-imatrix.gguf `
  -c 2048 `
  --chunks 10 `
  --no-ppl `
  -ngl 18

The MTP block itself is not exercised by the standard imatrix forward pass, so it was explicitly assigned conservative Q6_K/Q8_0 types during quantization.

Quantization command

Built with llama.cpp build 10441 / commit 0177dcc73.

.\llama-quantize.exe `
  --imatrix .\Qwen3.8-27B-Abliterated-imatrix.gguf `
  --tensor-type '^blk\.64\.nextn\.eh_proj\.weight$=q8_0' `
  --tensor-type '^blk\.64\.attn_q\.weight$=q6_k' `
  --tensor-type '^blk\.64\.attn_k\.weight$=q6_k' `
  --tensor-type '^blk\.64\.attn_v\.weight$=q6_k' `
  --tensor-type '^blk\.64\.attn_output\.weight$=q6_k' `
  --tensor-type '^blk\.64\.ffn_gate\.weight$=q6_k' `
  --tensor-type '^blk\.64\.ffn_up\.weight$=q6_k' `
  --tensor-type '^blk\.64\.ffn_down\.weight$=q6_k' `
  --tensor-type 'ffn_gate\.weight$=iq4_xs' `
  --tensor-type 'ffn_up\.weight$=iq4_xs' `
  --tensor-type 'ffn_down\.weight$=iq4_xs' `
  --tensor-type 'attn_qkv\.weight$=iq3_s' `
  --tensor-type 'attn_gate\.weight$=iq3_s' `
  --tensor-type 'attn_q\.weight$=q4_k' `
  --tensor-type 'attn_k\.weight$=q4_k' `
  --tensor-type 'attn_v\.weight$=q4_k' `
  --tensor-type 'attn_output\.weight$=q4_k' `
  --token-embedding-type iq2_s `
  --output-tensor-type q5_k `
  .\Qwen3.8-27B-Abliterated-BF16.gguf `
  .\Qwen3.8-27B-Abliterated-IQ4-MIX-MTP.gguf `
  iq4_xs `
  14

Tested hardware

Tested on:

  • GPU: AMD Radeon RX 9070 XT 16 GB
  • Backend: Vulkan
  • OS: Windows
  • llama.cpp: build 10441, commit 0177dcc73
  • GPU offload: all model layers
  • Parallel slots: 1

Measured performance

These numbers are from local tests and should be treated as hardware/workload-specific, not universal benchmarks.

32K allocated context, MTP-3

eval time = 81200.04 ms / 4070 tokens
generation = 50.11 tokens/s
draft acceptance = 0.74563
2814 accepted / 3774 generated
mean len = 3.24

64K allocated context, MTP-3

eval time = 182513.24 ms / 9018 tokens
generation = 49.40 tokens/s
draft acceptance = 0.77141
6297 accepted / 8163 generated
mean len = 3.31

These tests started with relatively short prompts. Decode speed falls as the live context grows.
In agentic coding workloads with roughly 25K–30K tokens of active history, observed generation was closer to the mid-20 tokens/s range depending on MTP acceptance.

Recommended llama.cpp server command — 64K

.\llama-server.exe `
  -m .\Qwen3.8-27B-Abliterated-IQ4-MIX-MTP.gguf `
  --alias qwen38-local `
  -ngl all `
  --fit off `
  --jinja `
  -fa on `
  -np 1 `
  -c 65536 `
  -b 2048 `
  -ub 256 `
  -ctk q8_0 `
  -ctv q4_0 `
  --kv-unified `
  --reasoning on `
  --reasoning-budget 2048 `
  --spec-type draft-mtp `
  --spec-draft-n-max 3 `
  --spec-draft-type-k q8_0 `
  --spec-draft-type-v q4_0 `
  --host 127.0.0.1 `
  --port 8080

For maximum raw throughput, keep the active conversation/context reasonably short. A 64K allocation does not mean generation speed will remain constant when all 64K tokens are occupied.

MTP notes

The most reliable tested configuration for this quant was:

--spec-type draft-mtp
--spec-draft-n-max 3

Higher speculative draft lengths are not automatically faster. In local testing, n-max 8 with a high p-min performed substantially worse on the tested Windows/Vulkan setup.

Download with Hugging Face CLI

Replace YOUR_USERNAME with the repository owner:

hf download YOUR_USERNAME/Qwen3.8-27B-Abliterated-IQ4-MIX-MTP-GGUF `
  Qwen3.8-27B-Abliterated-IQ4-MIX-MTP.gguf `
  --local-dir .

Notes on quality

  • This is a lossy quantization of the BF16 source model.
  • The embedding tensor is aggressively compressed to IQ2_S.
  • The output head is retained at Q5_K.
  • The MTP block is deliberately higher precision than the main body.
  • No comprehensive downstream benchmark suite has been run on this exact GGUF quantization.
  • Results can differ between Vulkan, HIP/ROCm, CUDA, CPU backends, context lengths and sampling settings.

Abliterated model notice

The upstream checkpoint is an abliterated variant. This repository does not perform additional fine-tuning or alignment changes; it provides a GGUF quantization of that source checkpoint.

Users are responsible for evaluating the model for their intended use and complying with applicable laws, licenses and platform policies.

Credits

  • Base model: Qwen/Qwen3.8-27B
  • Abliterated checkpoint: windowsxp811203/Qwen3.8-27B-Abliterated
  • Inference / quantization: llama.cpp
  • Importance-matrix calibration text: Salesforce/wikitext

License

The source checkpoint metadata reports Apache-2.0. Please also review the upstream model repositories and their terms before redistribution or use.

README history 1 version

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-08-15Add model carda9e4d247.3 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration