license: apache-2.0
base_model: huihui-ai/Huihui-Qwen3.8-27B-abliterated
base_model_relation: quantized
library_name: vllm
tags:
- nvfp4
- compressed-tensors
- mtp
- speculative-decoding
- blackwell
- qwen3_5
- abliterated
- uncensored
pipeline_tag: text-generation
Huihui-Qwen3.8-27B-abliterated-NVFP4
NVFP4 (W4A4, group 16) quantization of huihui-ai/Huihui-Qwen3.8-27B-abliterated — all credit for the model goes to @huihui-ai. This repo carries only the quantized weights.
55.6 GB → 20.0 GB. Fits on two 16 GB cards with real KV headroom.
- MTP draft head preserved in bf16 and wired into the index — speculative decoding works out of the box.
- bf16 kept for:
lm_head, the vision tower, the DeltaNetconv1d, and the MTP head. Everything else is NVFP4 W4A4. - Built for SM120 / Blackwell with vanilla vLLM v0.22.0 — compressed-tensors is auto-detected, no
--quantizationflag. - Quantized GPU-resident (full bf16 dispatched across 7× RTX PRO 2000 Blackwell via
device_map=auto): 119 seconds end to end.
Measured (TP=4, 32k ctx, KV fp8, RTX PRO 2000 Blackwell ×4, MTP n=3)
| concurrency | aggregate t/s |
|---|---|
| 1 | 77.3 |
| 4 | 204.8 |
| 8 | 380.9 |
Single-stream prefill: 3,590 tok/s on an 8k prompt (prefix cache disabled, mean of 3). GPU KV cache at this config: 613,655 tokens.
Capability after abliteration + 4-bit: an 8-probe set (Japanese fluency, English code, arithmetic, instruction-following, logic puzzle, domain knowledge, defensive-security explanation, and a trading backtest that punishes look-ahead bias) scores 8/8 — identical to the unmodified Qwen3.8-27B quantized with the same recipe, at 87.2 t/s vs 85.7 t/s. No measurable degradation.
Serve
vllm serve sakamakismile/Huihui-Qwen3.8-27B-abliterated-NVFP4 \
--trust-remote-code --tensor-parallel-size 4 \
--max-model-len 32768 --kv-cache-dtype fp8 --reasoning-parser qwen3 \
--speculative-config '{"method":"qwen3_5_mtp","num_speculative_tokens":3}'
On boards without P2P add NCCL_P2P_DISABLE=1 and --disable-custom-all-reduce.
⚠️ Gotchas
Send
reasoning_effort: "medium"for long-form work. This is the one that will bite you.reasoning_effortdefaults toxhigh, and at that setting the<think>phase sometimes never terminates: it grows past 19,000 characters, degenerates into repeating a single line, and the token budget is gone before any answer is emitted. On a 9-case gate (French/English long-form, temperature 0 and 0.7) this model fails 1 of 9 atxhighand passes 9 of 9 atmedium, where thinking stays near 1k characters. The unmodified Qwen3.8-27B passes the same gate at both settings, so the abliterated fine-tune is more prone to it — but medium is the right default either way."chat_template_kwargs": {"reasoning_effort": "medium"}Worth knowing what these modes actually are: they are not three levels of capability, they are three system prompts. Reading the chat template —
xhighinjects "think carefully through the task, validate key assumptions, consider plausible alternatives…",lowinjects "keep your thinking brief and focused, moving directly to the conclusion", andmediuminjects nothing at all — it is simply the model with no deliberation instruction. That matches what I measured:<think>runs a few hundred characters atlow, ~1k atmedium, 5–19k atxhigh, while answer quality was indistinguishable across all three on short verifiable tasks (96 runs, no errors at any setting). The risk atxhighis not worse reasoning — it is deliberation outliving the token budget.The 15
mtp.*modules are listed inquantization_config.ignore— do not remove them. If vLLM treats the bf16 MTP head as NVFP4 the draft breaks silently: 0% acceptance and slower than no MTP.W4A16 (NVFP4A16) does not serve on this architecture in vLLM 0.22 (
gptq_marlin_repack: size_n=24 not divisible by tile_n_size=64). W4A4 only.Long single-file code generation drops a closing paren roughly 1–2 times in 14 regardless of sampling temperature (measured with a JS parser on the base model). Put a syntax check in the loop rather than tuning temperature.
Text-only serving shown above; the vision tower ships in bf16 but multimodal serving was not benchmarked here.
Recipe
llm-compressor NVFP4 (W4A4, group 16), targets: [Linear], ignore: [lm_head, re:.*visual.*, re:.*conv1d.*, re:.*mtp.*], 32 calibration samples × 8192 tok from neuralmagic/calibration. MTP tensors grafted back in bf16 after saving and appended to quantization_config.ignore.
🙏 @huihui-ai for the model, Qwen team for the base, vLLM & llm-compressor teams for the tooling.