license: gemma
base_model:
- google/gemma-4-26B-A4B-it-qat-q4_0-unquantized
- huihui-ai/Huihui-gemma-4-26B-A4B-it-qat-q4_0-unquantized-abliterated
- google/gemma-4-26B-A4B-it-qat-q4_0-unquantized-assistant
base_model_relation: quantized
language: - ja
- en
tags: - gemma4
- moe
- nvfp4
- w4a4
- qat
- quantized
- abliterated
- vllm
- compressed-tensors
- blackwell
- sm120
- speculative-decoding
- mtp
- gemma4-assistant
library_name: transformers
pipeline_tag: image-text-to-text
model_type: gemma4
quantized_by: Lna-Lab
Huihui-gemma-4-26B-A4B-it-qat-abliterated-MTP-NVFP4
One repo, speculative decoding included. This is the MTP bundle of Huihui-gemma-4-26B-A4B-it-qat-abliterated-pad768-NVFP4: the NVFP4 (full W4A4) body plus the matching gemma4_assistant MTP draft checkpoint in assistant/, so a single hf download gives you everything vllm serve --speculative-config needs.
Measured: Japanese 134 tok/s · English 163 tok/s single-stream on 2× RTX PRO 2000 Blackwell 16 GB (vs 108 baseline) — a QAT-origin, abliterated, Japanese-safe Gemma 4 26B-A4B MoE in 17.6 GB + a 0.84 GB draft, in the practical zone on 2–4 entry-level Blackwell cards.
Lineage: google/gemma-4-26B-A4B-it-qat-q4_0-unquantized (QAT q4_0 → bf16) → huihui-ai abliteration → Lna-Lab NVFP4 W4A4 + loss-less 704→768 MoE pad → this bundle, adding google/gemma-4-26B-A4B-it-qat-q4_0-unquantized-assistant (bf16, unmodified) as the speculative draft.
| Body | gemma4 MoE 26B-A4B: 128 experts / top-8, 30 text layers, moe_intermediate padded 704→768 for stock-vLLM CUTLASS alignment · NVFP4 W4A4 (compressed-tensors / nvfp4-pack-quantized) · 17.6 GB |
Draft (assistant/) |
Gemma4AssistantForCausalLM (model_type: gemma4_assistant) — 4-layer MTP head riding the target's hidden states · bf16 · 0.84 GB (~0.4 GB VRAM per GPU at TP=2) |
| Spec method | vLLM gemma4_mtp, num_speculative_tokens: 4 |
| Hardware | NVIDIA Blackwell (SM120) required · 2× 16 GB (TP=2) is the sweet spot |
| vLLM | ≥ 0.21 (compressed-tensors NVFP4 auto-detect + gemma4_mtp; measured on 0.21.0) |
Why MTP, and why a bundle
Gemma 4's multi-token-prediction is not a head baked into the main checkpoint (Qwen3.6-style). Google ships it as a separate assistant checkpoint that vLLM's --speculative-config loads alongside the target. That means spec-decode normally costs you a second hf download and a path dance. This repo ends that: the assistant lives in assistant/ and you point the config at the local subfolder.
Huihui-gemma-4-26B-A4B-it-qat-abliterated-MTP-NVFP4/
├── model.safetensors # NVFP4 W4A4 body (17.6 GB)
├── config.json / generation_config.json / processor_config.json
├── tokenizer.json / tokenizer_config.json / chat_template.jinja
├── recipe.yaml # llm-compressor recipe
└── assistant/ # gemma4_mtp draft (bf16, 0.84 GB)
├── model.safetensors
├── config.json # model_type: gemma4_assistant
└── tokenizer / chat_template
Quickstart
hf download sakamakismile/Huihui-gemma-4-26B-A4B-it-qat-abliterated-MTP-NVFP4 \
--local-dir gemma4-26b-mtp
DIR=$(realpath gemma4-26b-mtp)
NCCL_P2P_DISABLE=1 vllm serve "$DIR" \
--served-model-name gemma4-26b-qat-mtp \
--tensor-parallel-size 2 \
--disable-custom-all-reduce \
--kv-cache-dtype fp8 \
--max-model-len 16384 \
--gpu-memory-utilization 0.92 \
--max-num-batched-tokens 8192 \
--limit-mm-per-prompt '{"image":0}' \
--speculative-config "{\"method\":\"gemma4_mtp\",\"model\":\"$DIR/assistant\",\"num_speculative_tokens\":4}"
- The point:
modelin--speculative-configis the bundled local path — no second download, no HF resolution at serve time. - TP=2 (2× 16 GB) recommended. Weights ~16.4 GiB don't fit one 16 GB card; TP=4 buys only +6% single-stream on this MoE — spend extra GPUs on a second replica.
NCCL_P2P_DISABLE=1+--disable-custom-all-reduceare required on PCIe no-NVLink boxes (TP hangs without them); drop both if you have NVLink/P2P.--kv-cache-dtype fp8doubles KV; with the assistant on board this config still holds maxlen 16384 / KV 23,560 tok at GMU 0.92. Keep CUDA graphs ON (no--enforce-eager).- vLLM 0.21's quantization-inheritance trap does not fire here: with an explicit draft
modelpath the draft's own config decides (bf16). Only themodel:nullMTP-from-target path inherits target quantization.
Measured (RTX PRO 2000 Blackwell 16 GB ×2, TP=2, PCIe no-NVLink, fp8 KV, vLLM 0.21.0, 2026-06-12)
Single-stream, T=0 chat completions, ×3 each. Acceptance = accepted/drafted from /metrics.
| config | JA 128 | JA 512 | EN 128 | EN 512 | acceptance JA / EN |
|---|---|---|---|---|---|
| baseline (no spec) | 108.5 | 108.9 | 108.2 | — | — |
| native MTP (this bundle, N=4) | 133.6 | 121.0 | 163.2 | 142.4 | 35–44% / 50–64% |
| EAGLE-3 (English-trained draft) N=3 | 73.2 | 75.7 | 158.0 | 129.9 | 2.7–3.9% / 33–49% |
| ngram N=4 (lookup 2–4) | 67.5 | 74.4 | 70.4 | 69.4 | 10–25% / 13–21% |
JA +12–23%, EN +32–51% over baseline. Two lessons paid for in benchmarks:
- EAGLE-3 collapses on Japanese — the English-Magpie-trained draft gets 3–4% JA acceptance and lands below baseline (0.69×). The google MTP assistant holds 35–44% JA acceptance because its 4-layer head re-uses the target's own hidden states instead of imitating its distribution from scratch. If your traffic is non-English, MTP is the only one of the three that pays.
- ngram never pays for itself on free-form chat in either language.
The MoE body already runs 108 tok/s, so MTP's gain here is "fast → faster" (1.2–1.5×). On the dense 31B sibling the same assistant nearly triples throughput — see Huihui-gemma-4-31B-it-qat-abliterated-MTP-NVFP4.
Concurrent (aggregate throughput)
4 / 8 concurrent streams × 256 tok each (T=0, diverse prompts, prefix-cache busted, ×3 averaged). Baseline = same body, no spec-decode (measured with JA prompts; baseline JA≈EN single-stream).
| streams | baseline tok/s | MTP JA tok/s | MTP EN tok/s | acceptance JA / EN |
|---|---|---|---|---|
| 1 | 108.5 | 133.6 (+23%) | 163.2 (+51%) | 35–44% / 50–64% |
| 4 | 326.3 | 341.5 (+4.7%) | 363.4 | ~36% / ~47% |
| 8 | 571.1 | 557.6 (−2.4%) | 566.4 (−0.8%) | ~39% / ~45% |
On this MoE, MTP pays up to ~4 concurrent streams; at 8 it is a wash (−2%, run-noise territory). Acceptance stays flat under batch — the shrinking gain isn't the drafter failing, it's the GPUs reaching compute saturation (571 tok/s aggregate), where verifying rejected draft tokens competes with real batch work instead of filling decode bubbles. Worst case is break-even, so keeping MTP resident costs ~nothing at peak load and buys 1.2–1.5× whenever concurrency drops. (The dense 31B sibling, still verification-hungry at 8 streams, keeps +27% JA / +53% EN there.)
The body: QAT × NVFP4 (the finding, in short)
Full W4A4 NVFP4 breaks non-QAT gemma-4 on Japanese long-form — the non-QAT 26B sibling intermittently collapses into repetition loops past ~500 tokens. This QAT-origin body does not: adversarial 1300+-token Japanese essays run to natural EOS with zero loop-detector hits, while English (HumanEval-class) is unaffected. Pattern holds across dense (31B/12B) and MoE: if you want gemma-4 in NVFP4 W4A4, go through a QAT checkpoint — the q4_0-shaped weight distribution is the prior FP4 wants. Full evidence, bake recipe (pure-CPU calibration through the multimodal processor, MoECalibrationModule, pad768 surgery) in the non-MTP card.
Notes
- Abliterated (uncensored). Refusal behavior removed upstream — you are responsible for your deployment. Use responsibly and lawfully.
- NVFP4 is Blackwell-specific; the body will not run on Ampere/Hopper. The bf16 assistant inherits the body's GPU anyway.
- The
assistant/checkpoint is google's, redistributed unmodified under the same Gemma terms; its original model card is included asassistant/README.md. - Gemma is provided under and subject to the Gemma Terms of Use.
Credits
- Original model & MTP assistant: Google DeepMind (Gemma 4, QAT q4_0)
- QAT-unquantize & abliteration: huihui-ai
- NVFP4 quantization, pad768 surgery, spec-decode measurement & bundle: Lna-Lab · Tooling: llm-compressor / vLLM
Support the Base Model Author (huihui-ai)
If you find the abliterated base useful, please support huihui-ai:
- Ko-fi: https://ko-fi.com/huihuiai
- Bitcoin:
bc1qqnkhuchxw0zqjh2ku3lu4hq45hc6gy84uk70ge