license: apache-2.0
tags:
- qwen3_5
- abliterated
- uncensored
- compressed-tensors
- int4
- w4a16
- vllm
- quantized
base_model: huihui-ai/Huihui-Qwen3.8-27B-abliterated
Huihui-Qwen3.8-27B-abliterated-INT4-W4A16
INT4 W4A16 quantization of huihui-ai/Huihui-Qwen3.8-27B-abliterated, produced with Intel AutoRound and exported in compressed-tensors (pack-quantized) format.
Key facts
| Property | Value |
|---|---|
| Base model | huihui-ai/Huihui-Qwen3.8-27B-abliterated |
| Quantization | INT4 weight-only (W4A16), group_size=128, symmetric |
| Algorithm | AutoRound (Intel) |
| Format | compressed-tensors / pack-quantized (auto-detected by vLLM) |
| lm_head | Quantized to INT4 |
| Weight size | ~15.6 GiB (4 shards) |
| Architecture | Qwen3_5ForConditionalGeneration (multimodal, vision tower kept BF16) |
| Context | 262,144 native |
| License | Apache-2.0 |
Why this quant
This checkpoint is specifically sized to run fast on a single 24 GB GPU (e.g. RTX 3090):
- Quantizing
lm_headto INT4 frees ~1.9 GB of VRAM compared to BF16-lm_head quants. - That headroom is what enables speculative decoding (MTP / DFlash2) and CUDA graphs on 24 GB cards.
- Weight-only INT4 keeps activations in BF16 → minimal quality loss.
The language-model transformer layers are quantized; the vision tower, linear_attn.in_proj_a/b and the MTP fusion layer (mtp.fc) remain in BF16.
Usage with vLLM
vllm serve Yirasumi/Huihui-Qwen3.8-27B-abliterated-INT4-W4A16 \
--served-model-name qwen3.8-27b \
--max-model-len 40000 \
--gpu-memory-utilization 0.93 \
--kv-cache-dtype fp8_e5m2 \
--reasoning-parser qwen3 \
--enable-auto-tool-choice --tool-call-parser qwen3_xml \
--enable-prefix-caching --enable-chunked-prefill
MTP speculative decoding (the checkpoint ships the Qwen3.5 MTP head in model_extra_tensors.safetensors):
vllm serve Yirasumi/Huihui-Qwen3.8-27B-abliterated-INT4-W4A16 \
--speculative-config '{"method":"mtp","num_speculative_tokens":4}' \
# ... same flags as above
On a 24 GB RTX 3090 the tuned serving stack from syv-ai/qwen38-27b-rtx3090 applies directly — see that repo's prepare/ scripts for the optional lm_head/embed requantization and draft-vocab steps.
Quantization details
- Toolchain: AutoRound 0.14.2, transformers 5.15.x, torch 2.13 (cu13)
- Calibration: NeelNanda/pile-10k, 128 samples × 2048 tokens, 200 iters
- Scheme: W4A16, group_size=128, symmetric (
auto_round:llm_compressorexport) - Hardware: NVIDIA H100 80 GB, ~55 minutes total
Verification
model.safetensors.index.json+ 4 shards +model_extra_tensors.safetensors(MTP head) all present- Loads cleanly with vLLM 0.27.1 via the
compressed-tensorspath (Marlin INT4 kernels) quant_method: compressed-tensors,format: pack-quantized
Credits
- Base model & abliteration: huihui-ai
- Original model: Qwen/Qwen3.8-27B (Apache-2.0)
- Quantization tool: Intel AutoRound
Usage warning: This is an abliterated model. Safety filtering has been significantly reduced. Use responsibly and in compliance with applicable laws.
See also
- huihui-ai/Huihui-Qwen3.8-27B-abliterated (BF16 source)
- huihui-ai/Huihui-Qwen3.8-27B-abliterated-GGUF (GGUF quants for llama.cpp)
- syv-ai/qwen38-27b-rtx3090 (single-3090 serving stack)