license: apache-2.0
library_name: transformers
pipeline_tag: image-text-to-text
base_model:
- huihui-ai/Huihui-Qwen3.8-27B-abliterated
tags: - qwen3.8
- abliterated
- uncensored
- autoround
- gptq
- int4
- int8-embedding
- mtp
- intel-arc
- arc-pro-b70
- xpu
- vllm
Huihui-Qwen3.8-27B-abliterated — AutoRound INT4, baked head/MTP, INT8 embedding
A 4-bit build of huihui-ai/Huihui-Qwen3.8-27B-abliterated (an abliterated, i.e. refusal-removed, version of Qwen/Qwen3.8-27B) made for a single Intel Arc Pro B70 (32 GB) under vLLM XPU. It follows the recipe of ukisai/Swift-1.5-Qwen3.8-27b-W4A16-AutoRound for the body and of Wrapzii's Swift 1.5 GPTQ-Int4-baked-v1-embed-int8 for the head, MTP and embedding — applied to the uncensored Huihui weights.
Uncensored model. Abliteration removes most refusal behaviour. It will follow requests the original model refuses. You are responsible for how you use it.
What is in the files
| part | form | how |
|---|---|---|
| 400 text body linears | INT4, symmetric, group 128 (GPTQ-compatible packing) | AutoRound / SignRound from the bf16 Huihui weights |
lm_head |
INT4 GPTQ, group 128 | Hessian GPTQ on this model's own activations |
MTP block (fc, q/k/v/o, gate/up/down) |
INT4 GPTQ, group 128 | same |
| token embedding | per-row symmetric INT8 + FP16 scales (model-embed-int8.safetensors) |
absmax/127 |
vision tower, MTP norms, GDN in_proj_a / in_proj_b, norms |
BF16 (unchanged) | — |
Weights resident on the GPU: 14.75 GiB (vs 18.28 GiB for a GPTQ build that keeps lm_head / MTP dense), which leaves room for the full 262,144-token context on one 32 GB card (vLLM reports a 319K-token KV pool with fp8 KV).
The dense embedding is still in the shards, so stock loaders work; the INT8 embedding only takes effect in a runtime that reads embed_tokens_quant (the patched image below).
Quality
WikiText-2 test perplexity, 40 chunks x 2,048 tokens (81,880 scored tokens), the same token chunks for both, computed from vLLM prompt_logprobs on the B70:
| build | perplexity |
|---|---|
| this repo (AutoRound body from bf16 + baked head/MTP + INT8 embedding) | 6.266 |
| zrlu's GPTQ-Int4 Huihui export + the same head/MTP/embedding bake | 6.346 |
Head / MTP bake on held-out activations (relative output L2 error; GPTQ vs round-to-nearest):
| linear | GPTQ | round-to-nearest |
|---|---|---|
| lm_head | 3.39% | 7.61% |
| mtp.fc | 5.67% | 11.86% |
| MTP q / k / v / o | 2.11% / 4.99% / 4.06% / 5.04% | 4.42% / 10.46% / 8.33% / 12.12% |
| MTP gate / up / down | 2.93% / 4.59% / 5.74% | 6.46% / 9.75% / 11.30% |
| embedding (INT8, weight L2) | 0.90% | — |
These are matrix and perplexity checks, not a full benchmark suite; upstream evaluation scores have not been re-established for this derivative. A quick functional check (arithmetic, thinking on/off, tool-call template, an uncensored prompt) passed.
Speed on one Arc Pro B70
Image vllm-openai-xpu:v0.30.0-k8v4-tp1 (vLLM 0.30 XPU + Wrapzii/k8v4-xpu patches incl. the INT8-embedding loader), TP1, fp8 KV, MTP with 3 speculative tokens, decode cudagraph mode FULL_DECODE_ONLY, max_num_seqs=4.
BetterBench (temperature 0.7, top-p 0.95, thinking on, 20 runs per category):
| category | decode t/s (median) | tokens per update |
|---|---|---|
| code | 73.6 | 2.53 |
| reasoning | 62.8 | 2.29 |
| prose | 64.3 | 2.27 |
| json | 92.6 | 3.36 |
| file_edit | 86.5 | 3.06 |
| summarization | 88.6 | 3.15 |
| math | 88.9 | 3.17 |
| chat | 68.4 | 2.37 |
| weighted combined | 75.7 |
TTFT p50 ~160 ms; update gap p99 37.4 ms.
| concurrent requests | aggregate t/s | per-request decode t/s |
|---|---|---|
| 1 | 72.0 | 81.7 |
| 2 | 129.7 | 74.4 |
| 4 | 196.0 | 61.6 |
| 8 / 16 | ~200 (capped by max_num_seqs=4) |
~62 |
| prompt length | prefill t/s (median) |
|---|---|
| 1.6K | 1,848 |
| 6.0K | 1,888 |
| 11.8K | 1,772 |
| 23.6K | 1,566 |
| 47.1K | 1,285 |
Greedy, thinking off (single runs): prose ~72 t/s, JSON ~104 t/s; MTP draft acceptance ~60%.
For reference, the Swift 1.5 bake (a different fine-tune, run with 6 speculative tokens) measured 12–34% faster per BetterBench category on the same card, except prose, where this build was 5% faster. This build has not been tuned for speculative depth yet.
Serving (vLLM XPU, one Arc Pro B70)
docker run -d --name qwen27-b70 \
--device /dev/dri/card1 --device /dev/dri/renderD128 --group-add render \
--ipc=host --shm-size=8g -p 8200:8000 \
-v /path/to/this/repo:/model:ro \
-e VLLM_TARGET_DEVICE=xpu -e VLLM_XPU_ENABLE_XPU_GRAPH=1 \
vllm-openai-xpu:v0.30.0-k8v4-tp1 /model \
--host 0.0.0.0 --port 8000 --served-model-name qwen27-huihui-unc \
--gpu-memory-utilization 0.92 --dtype bfloat16 \
--max-model-len 262144 --kv-cache-dtype fp8 --tensor-parallel-size 1 \
--max-num-seqs 4 --max-num-batched-tokens 4224 \
--enable-auto-tool-choice --tool-call-parser qwen3_xml --reasoning-parser qwen3 \
--enable-prefix-caching --trust-remote-code \
--limit-mm-per-prompt '{"image":4,"video":0}' --mm-processor-kwargs '{"max_pixels":1048576}' \
--speculative-config '{"method":"mtp","num_speculative_tokens":3}' \
--compilation-config '{"cudagraph_mode":"FULL_DECODE_ONLY","cudagraph_capture_sizes":[4,8,12,16]}'
Build the image from Wrapzii/k8v4-xpu (k8v4_v030/, plus Dockerfile.swift-bake for the INT8-embedding loader). A stock vLLM XPU image loads the checkpoint too, with the embedding dense.
Thinking: chat_template_kwargs: {"enable_thinking": false} turns it off; reasoning_effort takes low / medium / xhigh (the bundled chat_template_low.jinja defaults to low). Tool calls use the qwen3 XML format.
How it was made
- Body — AutoRound 0.16.0 on the bf16 checkpoint, on the B70 (2.5 h, 27.5 GB peak VRAM):
AutoRound writes the 15 MTP tensors (dropped by the Transformers loader) back from the source intoauto-round --model huihui-ai/Huihui-Qwen3.8-27B-abliterated --scheme W4A16 --bits 4 --group_size 128 \ --iters 200 --batch_size 8 --nsamples 256 --seqlen 2048 --seed 42 \ --dataset HuggingFaceH4/ultrachat_200k \ --ignore_layers "lm_head,.*visual.*,.*mtp.*,.*in_proj_a,.*in_proj_b" \ --format auto_round --low_gpu_mem_usagemodel_extra_tensors.safetensors. - GPTQ compatibility —
tools/prepare_swift_autoround.pyfrom k8v4-xpu: verifies the 400 packed linears and zero points, adds deterministicg_idx, switches the config togptq. No weight values change. - Calibration —
tools/calibrate_swift_bake.py(run at TP1 with fp8 KV on the single card): 128 WikiText-2 validation prompts of 384 tokens, 64 generated tokens each, MTP on; Hessians forlm_headand the MTP inputs, every 64th row held out. - Bake —
tools/bake_gptq_resume.py: symmetric GPTQ INT4, group 128, 1% damping, no activation ordering. - Assemble —
tools/assemble_swift_bake.py: rewritten shards without the dense replaced tensors, the 36 GPTQ arrays inmodel-bake-int4.safetensors, the INT8 embedding side file. Settings and per-linear errors are inBAKE.json.
Credits and licenses
- Base model: Qwen/Qwen3.8-27B (Apache-2.0); abliteration: huihui-ai (Apache-2.0,
LICENSE). - Body quantization recipe: ukisai/Swift-1.5-Qwen3.8-27b-W4A16-AutoRound; tool: Intel AutoRound.
- Head/MTP bake, INT8 embedding and XPU serving stack: Wrapzii/k8v4-xpu; its GPTQ bake tooling is adapted from launch80/B65 (MIT,
B65-LICENSE). - This derivative: weights requantized only — no fine-tuning, no merging.