library_name: transformers
license: apache-2.0
pipeline_tag: image-text-to-text
language:
- en
- zh
base_model: huihui-ai/Huihui-Qwen3.8-27B-abliterated
base_model_relation: quantized
quantized_by: ericmey
tags: - gptq
- int4
- w4a16
- vllm
- xpu
- intel-arc
- mtp
- vision
- abliterated
- uncensored
- qwen3
Qwen3.8-27B-abliterated — GPTQ INT4 (sym, G128) + BF16 MTP
GPTQ-Int4 quantization of
huihui-ai/Huihui-Qwen3.8-27B-abliterated,
packaged for vLLM serving on Intel Arc (Xe2 / B-series) with native MTP speculative
decoding and the vision tower preserved. The 15 mtp.* draft tensors are deliberately
excluded from quantization and kept in BF16, so speculative decoding loads and runs out of
the box (the Qwen3_5ForConditionalGeneration class drops mtp.* on load, so many quantizers
silently lose it — this build keeps it).
Qwen/Qwen3.8-27B (base, Qwen team)
└─ huihui-ai/Huihui-Qwen3.8-27B-abliterated (refusal-direction ablation, BF16)
└─ this repo (GPTQ-Int4 + BF16 MTP + F16 vision, for vLLM XPU)
Scope of this repo. I did the quantization and Intel-Arc packaging only. The weights are
huihui-ai's abliterated model, re-quantized — not a new fine-tune or a different abliteration.
Credit to the Qwen team (base) and huihui-ai (abliteration).
TL;DR
| Base | huihui-ai/Huihui-Qwen3.8-27B-abliterated (abliterated Qwen3.8-27B) |
| Format | GPTQ INT4, symmetric, group size 128, desc_act=false, int32 pack |
| Preserved | 15 mtp.* draft tensors (BF16) · vision tower (F16) |
| Modality | Text + vision (image), tool-calling, thinking toggle |
| Context | 262,144 native (served/tested at 131,072) |
| Size on disk | ~19 GB (5 shards) |
| Target engine | vLLM (Intel Arc / XPU W4A16; also loads on CUDA) |
| Censorship | Abliterated / uncensored (see Usage Warnings) |
Files
| File | Size |
|---|---|
model-00001-of-00005.safetensors |
3.3 GB |
model-0000{2,3,4}-of-00005.safetensors |
4.0 GB each |
model-00005-of-00005.safetensors |
3.0 GB |
config.json, quantize_config.json, generation_config.json |
configs |
tokenizer*.json, chat_template.jinja |
tokenizer + chat template |
preprocessor_config.json, video_preprocessor_config.json, processor_config.json |
vision/video processor |
Quantization recipe
Produced with gptqmodel 7.3.2:
from gptqmodel import GPTQModel, QuantizeConfig
qcfg = QuantizeConfig(
bits=4, group_size=128, sym=True, desc_act=False,
lm_head=False, true_sequential=True,
dynamic={"-:.*mtp.*": {}}, # keep the 15 mtp.* draft tensors in BF16
)
model = GPTQModel.load("huihui-ai/Huihui-Qwen3.8-27B-abliterated", qcfg)
model.quantize(calibration, batch_size=1) # ~256 general-text samples, seqlen 2048
model.save("Qwen3.8-27B-abliterated-GPTQ-Int4-MTP")
Result: 400 quantized weight tensors (int32) + 15 BF16 mtp.* tensors + the vision tower (F16).
The vision tower is left unquantized automatically (gptqmodel only quantizes the decoder Linears).
Usage (vLLM, Intel Arc / XPU)
vllm serve ericmey/Qwen3.8-27B-abliterated-GPTQ-Int4-MTP \
--quantization gptq --dtype float16 --max-model-len 131072 \
--tensor-parallel-size 2 --gpu-memory-utilization 0.92 \
--kv-cache-dtype auto --enable-prefix-caching \
--mamba-cache-mode align --mamba-ssm-cache-dtype float16 \
--limit-mm-per-prompt '{"image":4,"video":0}' \
--mm-processor-kwargs '{"max_pixels":1048576}' \
--speculative-config '{"method":"mtp","num_speculative_tokens":3}' \
--enable-auto-tool-choice --tool-call-parser qwen3_coder --reasoning-parser qwen3
Minimal (any backend — GPTQ kernels are hardware-independent, so this also loads on CUDA):
from vllm import LLM
llm = LLM("ericmey/Qwen3.8-27B-abliterated-GPTQ-Int4-MTP", tensor_parallel_size=2)
Speculative decoding (MTP)
The 15 mtp.* draft tensors are shipped in BF16 and excluded from quantization viadynamic={"-:.*mtp.*": {}}. Under vLLM the model resolves as Qwen3_5MTP and accepts--speculative-config '{"method":"mtp","num_speculative_tokens":3}'. Load- and generation-verified;
it is what produces the decode speeds below.
Prompt template & sampling
Uses the Qwen3 chat template (shipped chat_template.jinja) with thinking control and tool-calling.
Recommended sampling from the Qwen3.8-27B card:
- Non-thinking:
temperature 0.7, top_p 0.80, top_k 20 - Thinking:
temperature 1.0, top_p 0.95, top_k 20
Benchmarks
Measured on 2× Intel Arc Pro B60 (24 GB), TP=2, vLLM XPU, MTP3, client-side timing;
compared against the non-abliterated GPTQ build of the same base, same recipe/args:
| Check | This model | Non-abliterated baseline |
|---|---|---|
| Decode — code / explain / agent (tok/s) | 47.3 / 48.4 / 50.7 | 47.2 / 42.7 / 49.2 |
| Decode mean (tok/s) | 48.8 | 46.4 |
| Exact-answer sanity (20 items, greedy) | 20 / 20 | 20 / 20 |
| Tool-call well-formedness (60 ×2 temps) | 60/60 · 60/60 | 60/60 · 60/60 |
| Vision smoke test (read text + shape) | PASS | PASS |
Capabilities and speed match the non-abliterated build — abliteration changes no tensor shapes
and does not touch the MTP head, so there is no quantization or speed penalty.
Limitations
- Abliterated base. Refusal directions were removed upstream by huihui-ai; that behavior is
inherited here and is not something quantization changes. - Benchmarks above are a smoke-test suite (speed, exact-answer sanity, tool well-formedness, one
vision probe), not broad downstream task benchmarks (MMLU/GSM8K/etc. not run here). - Tuned and measured for the Intel Arc W4A16 (vLLM XPU) path; it loads on CUDA but the
recipe/args are Arc-oriented. - MTP acceptance rate not separately reported (speculative decoding is load- and speed-verified).
Usage Warnings
This is an uncensored / abliterated model: safety refusals have been removed, and outputs are
unfiltered and may be offensive, inaccurate, or otherwise objectionable. It is intended for
adults, for research, creative writing, and local/private use.
- You are solely responsible for how you use it and for compliance with all applicable laws.
- Do not use it to generate illegal content — including, unconditionally, any sexual content
involving minors — or to facilitate harm to others. - No guardrails and no warranty. Validate outputs before relying on them.
Credits & license
- Base model: Qwen/Qwen3.8-27B (Qwen team)
- Abliteration: huihui-ai/Huihui-Qwen3.8-27B-abliterated
- MTP / Intel-Arc quant recipe inspired by SergiioB's
Intel Arc Pro B70 Inference Cookbook - Quantizer: gptqmodel 7.3.2
Licensed Apache-2.0, following the base model.
Citation
@misc{qwen38_27b_abliterated_gptq_int4_mtp,
title = {Qwen3.8-27B-abliterated-GPTQ-Int4-MTP},
author = {ericmey},
year = {2026},
note = {GPTQ-Int4 (sym, g128) + BF16 MTP quantization of huihui-ai/Huihui-Qwen3.8-27B-abliterated for vLLM XPU},
howpublished = {\url{https://huggingface.co/ericmey/Qwen3.8-27B-abliterated-GPTQ-Int4-MTP}}
}