license: apache-2.0
base_model: JonathanColetti/Qwen3.8-27B-Uncensored
base_model_relation: quantized
language:
- en
- zh
pipeline_tag: text-generation
tags: - gptq
- int4
- uncensored
- abliterated
- mtp
- speculative-decoding
- intel
- arc
- intel-arc
- xpu
- vllm
- qwen3.8
Qwen3.8-27B-Uncensored-GPTQ-Int4-sym-G128-MTP-BF16
GPTQ-Int4 quantization of JonathanColetti/Qwen3.8-27B-Uncensored,
built for Intel Arc / vLLM XPU with the MTP head preserved in BF16 so
native speculative decoding still works.
At the time of upload this is the only GPTQ build of an uncensored Qwen3.8-27B —
existing variants are GGUF (llama.cpp), AWQ, or unquantized safetensors. GPTQ-Int4
is the format Intel's integer-first XMX hardware actually wants.
Why this exists
Quantizing this model naively breaks speculative decoding. The 15 mtp.*
tensors are the draft head; if they get quantized along with everything else,
draft acceptance collapses and you lose roughly half your decode speed. The fix is
one line of quantize config:
dynamic={"-:.*mtp.*": {}} # exclude mtp.* from quantization -> stays BF16
The result has 400 quantized weight tensors (I32) + 15 preserved BF16 MTP tensors.
Provenance
Qwen/Qwen3.8-27B (Apache-2.0, base)
└─ JonathanColetti/Qwen3.8-27B-Uncensored (abliterated, BF16, 55.6 GB)
└─ this repo (GPTQ-Int4, 19 GB)
Quantization
| Tool | gptqmodel 7.3.2 |
| Hardware | Intel Arc Pro B70 (32 GB, Xe2) — quantized on the target GPU |
| Config | bits=4, group_size=128, sym=True, desc_act=False, lm_head=False, pack_dtype=int32 |
| MTP | dynamic={"-:.*mtp.*": {}} → 15 tensors kept BF16 |
| Calibration | 256 samples @ 2048 tokens from allenai/c4 (general web text) |
| Time | 82.3 min (63 layers) |
| Output | 5 shards, 19 GB |
Measured performance
Intel Arc Pro B70, 230 W, vLLM 0.27.2rc1.dev77+gac7509e2b.xpu,vllm-xpu-kernels 0.1.12.3, MTP4, fp8 KV, 131,072 context, prefix caching off,
client post-first tok/s, median of n=3–5.
| Workload | tok/s | MTP acceptance |
|---|---|---|
| Code generation | 68.5 | 73.3% |
| Prose | 52.4 | 49.9% |
Cold input (prompt tokens / TTFT): 1808 @ p2k · 1810 @ p4k · 1777 @ p6k · 1738 @ p8k.
Long context verified to 120,532 tokens.
Throughput tracks MTP acceptance, which depends on how predictable the output is —
code drafts well, freeform prose less so. Treat 68.5 as the good case, not a floor.
Serving (vLLM XPU)
Requires the two MTP patches from the
Intel Arc Pro B70 cookbook
(patch_mtp_nightly.py, then patch_mtp_boundary.py), applied tovllm/vllm-openai-xpu@sha256:f01e24f6c7ff01f1e0662234255a1372297d1dbd89d003cf13c8fad3eab1ba4f.
vllm serve /model \
--quantization gptq --dtype float16 --max-model-len 131072 \
--gpu-memory-utilization 0.88 --kv-cache-dtype fp8 \
--max-num-seqs 1 --max-num-batched-tokens 8192 \
--no-enable-prefix-caching --served-model-name qwen38 --language-model-only \
--speculative-config '{"method":"mtp","num_speculative_tokens":4}' \
--enable-auto-tool-choice --tool-call-parser qwen3_xml
-e B70_MTP_BF16_DRAFT=1 is required for the BF16 draft gate.
Three things that will bite you
--max-num-seqs 1is mandatory with any MTP mode. With more, a batch can mix
a prefill with a spec-decode step and the GDN kernel aborts the entire engine:causal_conv1d does not support spec-decode and non-spec (prefill + decode) tokens in the same invocation. Reproducible with 2 concurrent requests. For real
parallel batching you must drop speculation entirely.- Prefix caching does nothing. Qwen3.8 is a hybrid GDN/Mamba architecture and
vLLM'ssupports_mamba_prefix_cachingflag is not declared for it, so enabling
it yields 0 cache hits while costing KV space and ~4% decode. Measured: 3,036
queries → 0 hits. Not hardware-specific; the same applies on CUDA. --kv-cache-dtype fp8is required at 128K — fp16 KV does not fit in 32 GB.
Limitations — please read
- No quality evaluations were run. No perplexity comparison against the BF16
source, no coding or reasoning benchmarks, no quantitative refusal-rate testing.
What was verified: tensor structure, quantization config, inference speed, MTP
acceptance, tool calling, 120K context, and light generation smoke tests. If you
need quality guarantees, measure before relying on it. - Calibration used general web text with no code. If coding quality matters to
you, a code-inclusive calibration set would likely be better. - Abliteration behaviour is inherited, not verified. In informal testing it
complies without refusal or moralizing preamble, but stays fairly mild in
register — it removes refusals more than it changes tone. All decensoring
properties come from the upstream abliteration; this repo only changes numeric
precision. - This is an uncensored model. It will attempt requests an aligned model
declines. You are responsible for how you use it.
Credits
- Qwen — base model (Apache-2.0)
- JonathanColetti — abliteration
- Intel Arc Pro B70 inference cookbook — MTP patches and serving recipe
- gptqmodel — quantization