← back to catalog · registered 2026-08-22 13:56

kernelogic/Qwen3.8-27B-Uncensored-GPTQ-Int4-sym-G128-MTP-BF16

kernelogic Qwen 24B second-order
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/kernelogic%2FQwen3.8-27B-Uncensored-GPTQ-Int4-sym-G128-MTP-BF16"
Response includes
  • classification m1
  • files 16
  • hub_downloads_all_time 4,688
  • author_summary 1 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M1
Primary method

Direct removal

No other method signals detected in this model.
Confidence
MEDIUM
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=0 (base model)
  • no specific method indicators - defaulting to M1 (most common)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
MEDIUM
Why we say so
primary_method=M1; difference-of-means is the reference extraction for M1/M3 (Arditi 2024)
Downloads · lifetime
5K
1K last 30d - stable
Likes
1
Model age
7w ago
created 2026-08-20
Downloads over time
Now5K→from117↑4,184%
01.8K3.7K5.5K117 on Aug 195K on Oct 11AugSepOct
Aug 19 → Oct 11 · 48 snapshots · spans 53 days

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
apache-2.0
Languages
en zh
Tags
safetensors qwen3_5 gptq int4 uncensored abliterated mtp speculative-decoding intel arc intel-arc xpu

Related

Total size
18.2 GB
Files
16
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-08-20 05:46

Files by quantization

Auxiliary files 16 files 18.2 GB
model-00004-of-00005.safetensors 3.99 GB 0100cb42 download
model-00002-of-00005.safetensors 3.98 GB d8b507c3 download
model-00003-of-00005.safetensors 3.97 GB 5cf71533 download
model-00001-of-00005.safetensors 3.28 GB 30c8e2b1 download
model-00005-of-00005.safetensors 3.00 GB bfda560b download
tokenizer.json 19.1 MB f399b3cd download
model.safetensors.index.json 224 KB 428f0b2a download
quant_log.csv 18.9 KB 5278f341 download
chat_template.jinja 8.74 KB c0c686f9 download
README.md 5.56 KB 025724be download
config.json 4.97 KB 0c0e708d download
.gitattributes 1.53 KB 52373fe2 download
processor_config.json 1.16 KB 33818c7f download
quantize_config.json 1.15 KB 5597eba1 download
tokenizer_config.json 1.14 KB 1d134cd2 download
generation_config.json 214 B 8b9f95da download

README current version from Hugging Face


license: apache-2.0
base_model: JonathanColetti/Qwen3.8-27B-Uncensored
base_model_relation: quantized
language:

  • en
  • zh
    pipeline_tag: text-generation
    tags:
  • gptq
  • int4
  • uncensored
  • abliterated
  • mtp
  • speculative-decoding
  • intel
  • arc
  • intel-arc
  • xpu
  • vllm
  • qwen3.8

Qwen3.8-27B-Uncensored-GPTQ-Int4-sym-G128-MTP-BF16

GPTQ-Int4 quantization of JonathanColetti/Qwen3.8-27B-Uncensored,
built for Intel Arc / vLLM XPU with the MTP head preserved in BF16 so
native speculative decoding still works.

At the time of upload this is the only GPTQ build of an uncensored Qwen3.8-27B —
existing variants are GGUF (llama.cpp), AWQ, or unquantized safetensors. GPTQ-Int4
is the format Intel's integer-first XMX hardware actually wants.

Why this exists

Quantizing this model naively breaks speculative decoding. The 15 mtp.*
tensors are the draft head; if they get quantized along with everything else,
draft acceptance collapses and you lose roughly half your decode speed. The fix is
one line of quantize config:

dynamic={"-:.*mtp.*": {}}   # exclude mtp.* from quantization -> stays BF16

The result has 400 quantized weight tensors (I32) + 15 preserved BF16 MTP tensors.

Provenance

Qwen/Qwen3.8-27B                        (Apache-2.0, base)
  └─ JonathanColetti/Qwen3.8-27B-Uncensored   (abliterated, BF16, 55.6 GB)
       └─ this repo                            (GPTQ-Int4, 19 GB)

Quantization

Tool gptqmodel 7.3.2
Hardware Intel Arc Pro B70 (32 GB, Xe2) — quantized on the target GPU
Config bits=4, group_size=128, sym=True, desc_act=False, lm_head=False, pack_dtype=int32
MTP dynamic={"-:.*mtp.*": {}} → 15 tensors kept BF16
Calibration 256 samples @ 2048 tokens from allenai/c4 (general web text)
Time 82.3 min (63 layers)
Output 5 shards, 19 GB

Measured performance

Intel Arc Pro B70, 230 W, vLLM 0.27.2rc1.dev77+gac7509e2b.xpu,
vllm-xpu-kernels 0.1.12.3, MTP4, fp8 KV, 131,072 context, prefix caching off,
client post-first tok/s, median of n=3–5.

Workload tok/s MTP acceptance
Code generation 68.5 73.3%
Prose 52.4 49.9%

Cold input (prompt tokens / TTFT): 1808 @ p2k · 1810 @ p4k · 1777 @ p6k · 1738 @ p8k.
Long context verified to 120,532 tokens.

Throughput tracks MTP acceptance, which depends on how predictable the output is —
code drafts well, freeform prose less so. Treat 68.5 as the good case, not a floor.

Serving (vLLM XPU)

Requires the two MTP patches from the
Intel Arc Pro B70 cookbook
(patch_mtp_nightly.py, then patch_mtp_boundary.py), applied to
vllm/vllm-openai-xpu@sha256:f01e24f6c7ff01f1e0662234255a1372297d1dbd89d003cf13c8fad3eab1ba4f.

vllm serve /model \
  --quantization gptq --dtype float16 --max-model-len 131072 \
  --gpu-memory-utilization 0.88 --kv-cache-dtype fp8 \
  --max-num-seqs 1 --max-num-batched-tokens 8192 \
  --no-enable-prefix-caching --served-model-name qwen38 --language-model-only \
  --speculative-config '{"method":"mtp","num_speculative_tokens":4}' \
  --enable-auto-tool-choice --tool-call-parser qwen3_xml

-e B70_MTP_BF16_DRAFT=1 is required for the BF16 draft gate.

Three things that will bite you

  1. --max-num-seqs 1 is mandatory with any MTP mode. With more, a batch can mix
    a prefill with a spec-decode step and the GDN kernel aborts the entire engine:
    causal_conv1d does not support spec-decode and non-spec (prefill + decode) tokens in the same invocation. Reproducible with 2 concurrent requests. For real
    parallel batching you must drop speculation entirely.
  2. Prefix caching does nothing. Qwen3.8 is a hybrid GDN/Mamba architecture and
    vLLM's supports_mamba_prefix_caching flag is not declared for it, so enabling
    it yields 0 cache hits while costing KV space and ~4% decode. Measured: 3,036
    queries → 0 hits. Not hardware-specific; the same applies on CUDA.
  3. --kv-cache-dtype fp8 is required at 128K — fp16 KV does not fit in 32 GB.

Limitations — please read

  • No quality evaluations were run. No perplexity comparison against the BF16
    source, no coding or reasoning benchmarks, no quantitative refusal-rate testing.
    What was verified: tensor structure, quantization config, inference speed, MTP
    acceptance, tool calling, 120K context, and light generation smoke tests. If you
    need quality guarantees, measure before relying on it.
  • Calibration used general web text with no code. If coding quality matters to
    you, a code-inclusive calibration set would likely be better.
  • Abliteration behaviour is inherited, not verified. In informal testing it
    complies without refusal or moralizing preamble, but stays fairly mild in
    register — it removes refusals more than it changes tone. All decensoring
    properties come from the upstream abliteration; this repo only changes numeric
    precision.
  • This is an uncensored model. It will attempt requests an aligned model
    declines. You are responsible for how you use it.

Credits

README history 2 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-08-20card: add intel / arc tagsbd3f8d55.6 KB
    Loading...
  2. 2026-08-20GPTQ-Int4 sym G128, MTP head preserved in BF16 (quantized on Intel Arc Pro B70)891acc15.5 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration