license: gemma
base_model:
- google/gemma-4-31B-it-qat-q4_0-unquantized
- huihui-ai/Huihui-gemma-4-31B-it-qat-q4_0-unquantized-abliterated
base_model_relation: quantized
language: - ja
- en
tags: - gemma4
- nvfp4
- w4a4
- qat
- quantized
- abliterated
- vllm
- compressed-tensors
- blackwell
- sm120
library_name: transformers
pipeline_tag: image-text-to-text
model_type: gemma4
quantized_by: Lna-Lab
Huihui-gemma-4-31B-it-qat-abliterated-NVFP4
推奨 / Recommended: the MTP bundle → Huihui-gemma-4-31B-it-qat-abliterated-MTP-NVFP4 — same body + the
gemma4_mtpassistant included inassistant/, one download for spec-decode (JA 2.1–2.4× / EN 2.5–2.9×: 86 / 106 tok/s on TP=4).
NVFP4 (full W4A4) quantization of huihui-ai/Huihui-gemma-4-31B-it-qat-q4_0-unquantized-abliterated — the abliterated, QAT-q4_0-origin Gemma 4 31B instruct model (text + vision).
Lineage: google/gemma-4-31B-it-qat-q4_0-unquantized (QAT q4_0 → bf16) → huihui-ai abliteration → this NVFP4 (W4A4).
62.6 GB → 20.4 GB. Serves on 2× 16 GB Blackwell GPUs (TP=2); 4× (TP=4) is markedly faster and roomier — 37.3 vs 22.7 tok/s single-stream, 5× the KV cache (see Measured below).
| Base | huihui-ai/Huihui-gemma-4-31B-it-qat-q4_0-unquantized-abliterated (QAT q4_0 → bf16, abliterated google/gemma-4-31B-it) |
| Architecture | Gemma4ForConditionalGeneration — 31B dense, 60 text layers (hidden 5376) + 27-layer vision tower |
| Quantization | NVFP4 (W4A4) — weights FP4 and activations FP4 (group 16, FP8 scales) |
| Format | compressed-tensors / nvfp4-pack-quantized (native vLLM auto-detect) |
| Tool | llm-compressor 0.11.0 |
| Size | 20.4 GB · Requires NVIDIA Blackwell (SM120) |
The finding: QAT checkpoints survive full W4A4
Non-QAT gemma-4-12B collapsed at full W4A4 on this exact recipe — it needed weight-only W4A16 to stay coherent. This 31B's weights were trained quantization-aware (q4_0), and the result holds at full W4A4: fluent Japanese, correct multi-step logic, valid haiku, zero repetition/mojibake artifacts — with nothing fancier than a plain 256×2048 ultrachat_200k calibration. The q4_0-shaped weight distribution appears to be exactly the prior NVFP4 wants.
Reproducible takeaway: if you want gemma-4 in NVFP4 (W4A4), go through a QAT checkpoint. The non-QAT instruct weights will not take it.
Quality evidence (Japanese, greedy/low-temp — verbatim outputs)
- 「こんにちは。一文で自己紹介して。」→ 「私は、あなたの質問に答え、思考をサポートするAIアシスタントです。」
- 太郎>花子>次郎 height reasoning → 「一番背が低いのは次郎です。… 太郎 > 花子 > 次郎 という順番になるため、最後にある次郎が一番低い…」 (correct answer, clean chain)
- 春の俳句 → 「ひだまりに 眠る子猫の あくびかな」 (valid 5-7-5, spring kigo)
No repetition loops, no garbled output, no empty/pad responses.
Serving with vLLM
Requires a Blackwell GPU (SM120: RTX 50-series / RTX PRO Blackwell / GB10 / B100/B200) and vLLM ≥ 0.21 (Gemma4ForConditionalGeneration + compressed-tensors NVFP4 auto-detect — no --quantization flag needed).
Recommended: TP=4 (4× 16 GB)
vllm serve sakamakismile/Huihui-gemma-4-31B-it-qat-abliterated-NVFP4 \
--served-model-name gemma4-31b \
--tensor-parallel-size 4 \
--max-model-len 16384 \
--gpu-memory-utilization 0.90 \
--max-num-batched-tokens 8192 \
--limit-mm-per-prompt '{"image":0}'
Attention heads divide cleanly for TP=4 (32 attn / 16 kv). This config measured 43,123 KV-cache tokens — comfortable for batch serving.
Minimum footprint: TP=2 (2× 16 GB)
20.4 GB of weights / TP2 ≈ 10.2 GB per GPU, so KV is tight — squeeze:
vllm serve sakamakismile/Huihui-gemma-4-31B-it-qat-abliterated-NVFP4 \
--served-model-name gemma4-31b \
--tensor-parallel-size 2 \
--max-model-len 8192 \
--gpu-memory-utilization 0.95 \
--max-num-batched-tokens 2560 \
--limit-mm-per-prompt '{"image":0}'
This yields an 8,577-token KV cache — enough for the 8192 context, no more.
Gotchas
- vLLM 0.21 multimodal budget trap: even with
--limit-mm-per-prompt '{"image":0}', vLLM validatesmax_tokens_per_mm_item(2496 for this model) against--max-num-batched-tokens. Keep MBT ≥ 2496 (hence the 2560 above) or startup fails. - Vision inputs: the examples above run text-only. To accept images, set
'{"image":1}'and drop--max-model-lento ~4096 on 16 GB cards. - Multi-GPU boxes without NVLink/P2P only (e.g. consumer/entry Blackwell on plain PCIe): vLLM tensor-parallel hangs unless you add both
NCCL_P2P_DISABLE=1(env) and--disable-custom-all-reduce. If your GPUs have NVLink or working P2P, skip both — they only cost you speed there.
# no-P2P variant (prepend/append to either command above)
NCCL_P2P_DISABLE=1 vllm serve ... --disable-custom-all-reduce
Query it (OpenAI-compatible)
curl -s http://localhost:8000/v1/chat/completions \
-H 'Content-Type: application/json' \
-d '{
"model": "gemma4-31b",
"messages": [{"role": "user", "content": "こんにちは。一文で自己紹介して。"}],
"max_tokens": 128
}'
Measured (RTX PRO 2000 Blackwell 16 GB ×N, PCIe no-NVLink, CUDA graphs default ON)
| metric (tok/s) | TP=2 (2 GPU) | TP=4 (4 GPU) |
|---|---|---|
| single-stream, 128 tok ×3 | 22.7 | 37.3 |
| single-stream, 512 tok ×3 | — | 36.8 |
| 4 concurrent ×256 tok, aggregate | — | 139.5 (34.9/stream) |
| 8 concurrent ×256 tok, aggregate | — | 252.9 (31.6/stream) |
| KV cache | 8,577 tok | 43,123 tok |
| max-model-len | 8192 | 16384 |
Even on this no-P2P box (host-memory all-reduce), the dense 31B scales: TP=4 beats TP=2 single-stream by +64%, and continuous batching is near-linear out to 8 streams (per-stream 37.3 → 31.6).
Speculative Decoding (measured 2026-06-12)
Measured single-stream (T=0, chat completions, ×3 each) on TP=4 GPU2,3,5,6 (maxlen 8192 / GMU 0.90 / MBT 8192, bf16 KV), vLLM 0.21.0. TP=2 + draft does not fit: the AEON-7 NVFP4 draft (3.3 GB safetensors) leaves only 0.23 GiB KV on 2×16 GB even with fp8 KV — spec-decode on this model is a TP=4 game on this box.
| config | JA 128 | JA 512 | EN 128 | EN 512 | acceptance JA / EN |
|---|---|---|---|---|---|
| baseline (no spec) | 36.4 | 36.3 | 36.7 | — | — |
EAGLE-3 AEON-7/gemma-4-31B-it-speculator.eagle3-NVFP4 N=3 |
33.5 | 34.0 | 46.9 | — | 1–3% / 16% |
| native MTP (gemma4_mtp) N=4 | 85.7 | 75.6 | 106.1 | 91.0 | 41–51% / 55–71% |
Winner: native MTP, and it is dramatic — JA 2.1–2.4×, EN 2.5–2.9×. Draft = google/gemma-4-31B-it-qat-q4_0-unquantized-assistant (bf16, 351 MB), method auto-normalized to gemma4_mtp:
--speculative-config '{"method":"gemma4_mtp","model":"<drafts>/google-31b-mtp-assistant","num_speculative_tokens":4}'
# TP=4, maxlen 8192, GMU 0.90, MBT 8192 → KV 17,385 tok
- Concurrent (MTP, aggregate): 4 / 8 streams × 256 tok (×3 avg, diverse prompts): MTP 211.6 / 320.1 tok/s JA (251.9 / 387.2 EN) vs baseline 139.5 / 252.9 → +52% / +27% JA (+81% / +53% EN) — MTP keeps winning at every concurrency this box can reach; acceptance holds ~40% JA / ~53% EN under batch. No regime to turn it off on this dense 31B.
- Japanese caveat: EAGLE-3 (vanilla-31B-trained, English data) is dead on arrival against this abliterated QAT body — JA acceptance 1–3% lands it below baseline; even EN only reaches 16% (distribution shift: vanilla-trained drafter vs abliterated verifier, same failure mode coolthor documented for 26B). The google MTP assistant shrugs both problems off: 41–51% JA acceptance despite vanilla training, because the 4-layer MTP head re-uses the target's own hidden states.
- The dense 31B at 37 t/s leaves the GPUs verification-hungry — that is why MTP nearly triples it while the MoE 26B (already 108 t/s) only gains ~1.2–1.5×.
- vLLM 0.21 quantization-inheritance trap does not fire with an explicit draft
modelpath (only themodel:nullMTP-from-target path inherits target quantization).
Bake recipe (key points)
QuantizationModifier(targets=Linear, scheme=NVFP4, ignore=[lm_head, re:.*embed.*, re:.*vision_tower.*])— vision tower, embeddings, lm_head kept BF16- Calibration:
HuggingFaceH4/ultrachat_200ktrain_sft, 256 samples × 2048 tok, driven through the multimodalAutoProcessor— calibrating through the bare tokenizer leavesinput_global_scaleuncalibrated and the model degenerates to<pad>spam pipeline="basic"— gemma4 is fx-untraceable (shared-KVUserDictpassed between layers breaks llm-compressor's sequential pipeline)- Pure-CPU calibration (
device_map="cpu"): ~2.2 h on a 48-core CPU, ~80 GB RAM. Also a correctness measure: on this no-P2P box, multi-GPU accelerate dispatch silently corrupts gemma4 activations — never calibrate (or judge a base) through it
Notes
- Abliterated (uncensored). Refusal behavior has been removed upstream — you are responsible for your deployment. Use responsibly and lawfully.
- NVFP4 is Blackwell-specific; it will not run on Ampere/Hopper.
- Gemma is provided under and subject to the Gemma Terms of Use.
Credits
- Base model, QAT-unquantize & abliteration: huihui-ai
- Original model: Google DeepMind (Gemma 4, QAT q4_0)
- Quantization & serving recipe: Lna-Lab · Tooling: llm-compressor / vLLM
Support the Base Model Author (huihui-ai)
If you find the abliterated base useful, please support huihui-ai:
- Ko-fi: https://ko-fi.com/huihuiai
- Bitcoin:
bc1qqnkhuchxw0zqjh2ku3lu4hq45hc6gy84uk70ge