← back to catalog · registered 2026-08-22 13:56

sakamakismile/Huihui-gemma-4-31B-it-qat-abliterated-NVFP4

sakamakismile Gemma 15B multimodal
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/sakamakismile%2FHuihui-gemma-4-31B-it-qat-abliterated-NVFP4"
Response includes
  • classification m1
  • files 10
  • hub_downloads_all_time 1,433
  • author_summary 34 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M1
Primary method

Direct removal

No other method signals detected in this model.
Confidence
MEDIUM
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=0 (base model)
  • no specific method indicators - defaulting to M1 (most common)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
MEDIUM
Why we say so
primary_method=M1; difference-of-means is the reference extraction for M1/M3 (Arditi 2024)
Downloads · lifetime
1K
22 last 30d - cooling
Likes
0
Model age
4mo ago
created 2026-06-11
Downloads over time
Now1.4K→from143↑907%
785751.1K1.6K143 on Jun 101.4K on Oct 111.4K on Oct 7JunJulAugSepOct
Jun 10 → Oct 11 · 57 snapshots · spans 123 days

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
gemma
Languages
ja en
Tags
transformers safetensors gemma4 image-text-to-text nvfp4 w4a4 qat quantized abliterated vllm compressed-tensors blackwell

Related

Total size
19.0 GB
Files
10
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-06-12 01:28

Files by quantization

Auxiliary files 10 files 19.1 GB
model.safetensors 19.0 GB 8d2916f8 download
tokenizer.json 30.7 MB a43152a9 download
config.json 18.5 KB f2385479 download
chat_template.jinja 16.5 KB f62ca843 download
README.md 10.4 KB b8f4f308 download
tokenizer_config.json 2.68 KB af7f2586 download
processor_config.json 1.65 KB 5465974d download
.gitattributes 1.53 KB 52373fe2 download
recipe.yaml 224 B f790e78b download
generation_config.json 203 B a5245e87 download

README current version from Hugging Face


license: gemma
base_model:

  • google/gemma-4-31B-it-qat-q4_0-unquantized
  • huihui-ai/Huihui-gemma-4-31B-it-qat-q4_0-unquantized-abliterated
    base_model_relation: quantized
    language:
  • ja
  • en
    tags:
  • gemma4
  • nvfp4
  • w4a4
  • qat
  • quantized
  • abliterated
  • vllm
  • compressed-tensors
  • blackwell
  • sm120
    library_name: transformers
    pipeline_tag: image-text-to-text
    model_type: gemma4
    quantized_by: Lna-Lab

Huihui-gemma-4-31B-it-qat-abliterated-NVFP4

推奨 / Recommended: the MTP bundle → Huihui-gemma-4-31B-it-qat-abliterated-MTP-NVFP4 — same body + the gemma4_mtp assistant included in assistant/, one download for spec-decode (JA 2.1–2.4× / EN 2.5–2.9×: 86 / 106 tok/s on TP=4).

NVFP4 (full W4A4) quantization of huihui-ai/Huihui-gemma-4-31B-it-qat-q4_0-unquantized-abliterated — the abliterated, QAT-q4_0-origin Gemma 4 31B instruct model (text + vision).

Lineage: google/gemma-4-31B-it-qat-q4_0-unquantized (QAT q4_0 → bf16) → huihui-ai abliteration → this NVFP4 (W4A4).

62.6 GB → 20.4 GB. Serves on 2× 16 GB Blackwell GPUs (TP=2); 4× (TP=4) is markedly faster and roomier — 37.3 vs 22.7 tok/s single-stream, 5× the KV cache (see Measured below).

Base huihui-ai/Huihui-gemma-4-31B-it-qat-q4_0-unquantized-abliterated (QAT q4_0 → bf16, abliterated google/gemma-4-31B-it)
Architecture Gemma4ForConditionalGeneration — 31B dense, 60 text layers (hidden 5376) + 27-layer vision tower
Quantization NVFP4 (W4A4) — weights FP4 and activations FP4 (group 16, FP8 scales)
Format compressed-tensors / nvfp4-pack-quantized (native vLLM auto-detect)
Tool llm-compressor 0.11.0
Size 20.4 GB · Requires NVIDIA Blackwell (SM120)

The finding: QAT checkpoints survive full W4A4

Non-QAT gemma-4-12B collapsed at full W4A4 on this exact recipe — it needed weight-only W4A16 to stay coherent. This 31B's weights were trained quantization-aware (q4_0), and the result holds at full W4A4: fluent Japanese, correct multi-step logic, valid haiku, zero repetition/mojibake artifacts — with nothing fancier than a plain 256×2048 ultrachat_200k calibration. The q4_0-shaped weight distribution appears to be exactly the prior NVFP4 wants.

Reproducible takeaway: if you want gemma-4 in NVFP4 (W4A4), go through a QAT checkpoint. The non-QAT instruct weights will not take it.

Quality evidence (Japanese, greedy/low-temp — verbatim outputs)

  • 「こんにちは。一文で自己紹介して。」→ 「私は、あなたの質問に答え、思考をサポートするAIアシスタントです。」
  • 太郎>花子>次郎 height reasoning → 「一番背が低いのは次郎です。… 太郎 > 花子 > 次郎 という順番になるため、最後にある次郎が一番低い…」 (correct answer, clean chain)
  • 春の俳句 → 「ひだまりに 眠る子猫の あくびかな」 (valid 5-7-5, spring kigo)

No repetition loops, no garbled output, no empty/pad responses.

Serving with vLLM

Requires a Blackwell GPU (SM120: RTX 50-series / RTX PRO Blackwell / GB10 / B100/B200) and vLLM ≥ 0.21 (Gemma4ForConditionalGeneration + compressed-tensors NVFP4 auto-detect — no --quantization flag needed).

Recommended: TP=4 (4× 16 GB)

vllm serve sakamakismile/Huihui-gemma-4-31B-it-qat-abliterated-NVFP4 \
  --served-model-name gemma4-31b \
  --tensor-parallel-size 4 \
  --max-model-len 16384 \
  --gpu-memory-utilization 0.90 \
  --max-num-batched-tokens 8192 \
  --limit-mm-per-prompt '{"image":0}'

Attention heads divide cleanly for TP=4 (32 attn / 16 kv). This config measured 43,123 KV-cache tokens — comfortable for batch serving.

Minimum footprint: TP=2 (2× 16 GB)

20.4 GB of weights / TP2 ≈ 10.2 GB per GPU, so KV is tight — squeeze:

vllm serve sakamakismile/Huihui-gemma-4-31B-it-qat-abliterated-NVFP4 \
  --served-model-name gemma4-31b \
  --tensor-parallel-size 2 \
  --max-model-len 8192 \
  --gpu-memory-utilization 0.95 \
  --max-num-batched-tokens 2560 \
  --limit-mm-per-prompt '{"image":0}'

This yields an 8,577-token KV cache — enough for the 8192 context, no more.

Gotchas

  • vLLM 0.21 multimodal budget trap: even with --limit-mm-per-prompt '{"image":0}', vLLM validates max_tokens_per_mm_item (2496 for this model) against --max-num-batched-tokens. Keep MBT ≥ 2496 (hence the 2560 above) or startup fails.
  • Vision inputs: the examples above run text-only. To accept images, set '{"image":1}' and drop --max-model-len to ~4096 on 16 GB cards.
  • Multi-GPU boxes without NVLink/P2P only (e.g. consumer/entry Blackwell on plain PCIe): vLLM tensor-parallel hangs unless you add both NCCL_P2P_DISABLE=1 (env) and --disable-custom-all-reduce. If your GPUs have NVLink or working P2P, skip both — they only cost you speed there.
# no-P2P variant (prepend/append to either command above)
NCCL_P2P_DISABLE=1 vllm serve ... --disable-custom-all-reduce

Query it (OpenAI-compatible)

curl -s http://localhost:8000/v1/chat/completions \
  -H 'Content-Type: application/json' \
  -d '{
    "model": "gemma4-31b",
    "messages": [{"role": "user", "content": "こんにちは。一文で自己紹介して。"}],
    "max_tokens": 128
  }'

Measured (RTX PRO 2000 Blackwell 16 GB ×N, PCIe no-NVLink, CUDA graphs default ON)

metric (tok/s) TP=2 (2 GPU) TP=4 (4 GPU)
single-stream, 128 tok ×3 22.7 37.3
single-stream, 512 tok ×3 — 36.8
4 concurrent ×256 tok, aggregate — 139.5 (34.9/stream)
8 concurrent ×256 tok, aggregate — 252.9 (31.6/stream)
KV cache 8,577 tok 43,123 tok
max-model-len 8192 16384

Even on this no-P2P box (host-memory all-reduce), the dense 31B scales: TP=4 beats TP=2 single-stream by +64%, and continuous batching is near-linear out to 8 streams (per-stream 37.3 → 31.6).

Speculative Decoding (measured 2026-06-12)

Measured single-stream (T=0, chat completions, ×3 each) on TP=4 GPU2,3,5,6 (maxlen 8192 / GMU 0.90 / MBT 8192, bf16 KV), vLLM 0.21.0. TP=2 + draft does not fit: the AEON-7 NVFP4 draft (3.3 GB safetensors) leaves only 0.23 GiB KV on 2×16 GB even with fp8 KV — spec-decode on this model is a TP=4 game on this box.

config JA 128 JA 512 EN 128 EN 512 acceptance JA / EN
baseline (no spec) 36.4 36.3 36.7 — —
EAGLE-3 AEON-7/gemma-4-31B-it-speculator.eagle3-NVFP4 N=3 33.5 34.0 46.9 — 1–3% / 16%
native MTP (gemma4_mtp) N=4 85.7 75.6 106.1 91.0 41–51% / 55–71%

Winner: native MTP, and it is dramatic — JA 2.1–2.4×, EN 2.5–2.9×. Draft = google/gemma-4-31B-it-qat-q4_0-unquantized-assistant (bf16, 351 MB), method auto-normalized to gemma4_mtp:

--speculative-config '{"method":"gemma4_mtp","model":"<drafts>/google-31b-mtp-assistant","num_speculative_tokens":4}'
# TP=4, maxlen 8192, GMU 0.90, MBT 8192 → KV 17,385 tok
  • Concurrent (MTP, aggregate): 4 / 8 streams × 256 tok (×3 avg, diverse prompts): MTP 211.6 / 320.1 tok/s JA (251.9 / 387.2 EN) vs baseline 139.5 / 252.9 → +52% / +27% JA (+81% / +53% EN) — MTP keeps winning at every concurrency this box can reach; acceptance holds ~40% JA / ~53% EN under batch. No regime to turn it off on this dense 31B.
  • Japanese caveat: EAGLE-3 (vanilla-31B-trained, English data) is dead on arrival against this abliterated QAT body — JA acceptance 1–3% lands it below baseline; even EN only reaches 16% (distribution shift: vanilla-trained drafter vs abliterated verifier, same failure mode coolthor documented for 26B). The google MTP assistant shrugs both problems off: 41–51% JA acceptance despite vanilla training, because the 4-layer MTP head re-uses the target's own hidden states.
  • The dense 31B at 37 t/s leaves the GPUs verification-hungry — that is why MTP nearly triples it while the MoE 26B (already 108 t/s) only gains ~1.2–1.5×.
  • vLLM 0.21 quantization-inheritance trap does not fire with an explicit draft model path (only the model:null MTP-from-target path inherits target quantization).

Bake recipe (key points)

  • QuantizationModifier(targets=Linear, scheme=NVFP4, ignore=[lm_head, re:.*embed.*, re:.*vision_tower.*]) — vision tower, embeddings, lm_head kept BF16
  • Calibration: HuggingFaceH4/ultrachat_200k train_sft, 256 samples × 2048 tok, driven through the multimodal AutoProcessor — calibrating through the bare tokenizer leaves input_global_scale uncalibrated and the model degenerates to <pad> spam
  • pipeline="basic" — gemma4 is fx-untraceable (shared-KV UserDict passed between layers breaks llm-compressor's sequential pipeline)
  • Pure-CPU calibration (device_map="cpu"): ~2.2 h on a 48-core CPU, ~80 GB RAM. Also a correctness measure: on this no-P2P box, multi-GPU accelerate dispatch silently corrupts gemma4 activations — never calibrate (or judge a base) through it

Notes

  • Abliterated (uncensored). Refusal behavior has been removed upstream — you are responsible for your deployment. Use responsibly and lawfully.
  • NVFP4 is Blackwell-specific; it will not run on Ampere/Hopper.
  • Gemma is provided under and subject to the Gemma Terms of Use.

Credits

Support the Base Model Author (huihui-ai)

If you find the abliterated base useful, please support huihui-ai:

README history 4 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-06-12Spec-decode section: add concurrent MTP results (211.6/320.1 JA agg; MTP wins...257979b10.4 KB
    Loading...
  2. 2026-06-12README: point to the MTP bundle repoabfc1f010.1 KB
    Loading...
  3. 2026-06-12Add Speculative Decoding section: native gemma4_mtp (QAT assistant) = 2.4-2.9...56ead8f9.8 KB
    Loading...
  4. 2026-06-11Initial release: NVFP4 (W4A4) of QAT-q4_0-origin gemma-4-31B — QAT checkpoint...69dd1cf7.9 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration