← back to catalog · registered 2026-08-22 13:56

sakamakismile/Huihui-gemma-4-26B-A4B-it-qat-abliterated-MTP-NVFP4

sakamakismile Gemma 13B MoE multimodal
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/sakamakismile%2FHuihui-gemma-4-26B-A4B-it-qat-abliterated-MTP-NVFP4"
Response includes
  • classification m1
  • files 10
  • hub_downloads_all_time 2,790
  • author_summary 34 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M1
Primary method

Direct removal

No other method signals detected in this model.
Confidence
MEDIUM
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=0 (base model)
  • no specific method indicators - defaulting to M1 (most common)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
MEDIUM
Why we say so
primary_method=M1; difference-of-means is the reference extraction for M1/M3 (Arditi 2024)
Downloads · lifetime
3K
565 last 30d - stable
Likes
5
Model age
4mo ago
created 2026-06-12
Downloads over time
Now3.1K→from23↑13,178%
01.1K2.2K3.4K23 on Jun 103.1K on Oct 11JunJulAugSepOct
Jun 10 → Oct 11 · 57 snapshots · spans 123 days

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
gemma
Languages
ja en
Tags
transformers safetensors gemma4 image-text-to-text moe nvfp4 w4a4 qat quantized abliterated vllm compressed-tensors

Related

Total size
16.4 GB
Files
10
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-06-12 01:28

Files by quantization

Auxiliary files 10 files 16.4 GB
model.safetensors 16.4 GB 6bb1f3d0 download
tokenizer.json 30.7 MB a43152a9 download
config.json 19.2 KB ae84f8be download
chat_template.jinja 16.5 KB f62ca843 download
README.md 9.38 KB 721db99f download
tokenizer_config.json 2.68 KB af7f2586 download
processor_config.json 1.65 KB 5465974d download
.gitattributes 1.59 KB b50a4d2f download
recipe.yaml 224 B f790e78b download
generation_config.json 203 B a5245e87 download

README current version from Hugging Face


license: gemma
base_model:

  • google/gemma-4-26B-A4B-it-qat-q4_0-unquantized
  • huihui-ai/Huihui-gemma-4-26B-A4B-it-qat-q4_0-unquantized-abliterated
  • google/gemma-4-26B-A4B-it-qat-q4_0-unquantized-assistant
    base_model_relation: quantized
    language:
  • ja
  • en
    tags:
  • gemma4
  • moe
  • nvfp4
  • w4a4
  • qat
  • quantized
  • abliterated
  • vllm
  • compressed-tensors
  • blackwell
  • sm120
  • speculative-decoding
  • mtp
  • gemma4-assistant
    library_name: transformers
    pipeline_tag: image-text-to-text
    model_type: gemma4
    quantized_by: Lna-Lab

Huihui-gemma-4-26B-A4B-it-qat-abliterated-MTP-NVFP4

One repo, speculative decoding included. This is the MTP bundle of Huihui-gemma-4-26B-A4B-it-qat-abliterated-pad768-NVFP4: the NVFP4 (full W4A4) body plus the matching gemma4_assistant MTP draft checkpoint in assistant/, so a single hf download gives you everything vllm serve --speculative-config needs.

Measured: Japanese 134 tok/s · English 163 tok/s single-stream on 2× RTX PRO 2000 Blackwell 16 GB (vs 108 baseline) — a QAT-origin, abliterated, Japanese-safe Gemma 4 26B-A4B MoE in 17.6 GB + a 0.84 GB draft, in the practical zone on 2–4 entry-level Blackwell cards.

Lineage: google/gemma-4-26B-A4B-it-qat-q4_0-unquantized (QAT q4_0 → bf16) → huihui-ai abliteration → Lna-Lab NVFP4 W4A4 + loss-less 704→768 MoE pad → this bundle, adding google/gemma-4-26B-A4B-it-qat-q4_0-unquantized-assistant (bf16, unmodified) as the speculative draft.

Body gemma4 MoE 26B-A4B: 128 experts / top-8, 30 text layers, moe_intermediate padded 704→768 for stock-vLLM CUTLASS alignment · NVFP4 W4A4 (compressed-tensors / nvfp4-pack-quantized) · 17.6 GB
Draft (assistant/) Gemma4AssistantForCausalLM (model_type: gemma4_assistant) — 4-layer MTP head riding the target's hidden states · bf16 · 0.84 GB (~0.4 GB VRAM per GPU at TP=2)
Spec method vLLM gemma4_mtp, num_speculative_tokens: 4
Hardware NVIDIA Blackwell (SM120) required · 2× 16 GB (TP=2) is the sweet spot
vLLM ≥ 0.21 (compressed-tensors NVFP4 auto-detect + gemma4_mtp; measured on 0.21.0)

Why MTP, and why a bundle

Gemma 4's multi-token-prediction is not a head baked into the main checkpoint (Qwen3.6-style). Google ships it as a separate assistant checkpoint that vLLM's --speculative-config loads alongside the target. That means spec-decode normally costs you a second hf download and a path dance. This repo ends that: the assistant lives in assistant/ and you point the config at the local subfolder.

Huihui-gemma-4-26B-A4B-it-qat-abliterated-MTP-NVFP4/
├── model.safetensors          # NVFP4 W4A4 body (17.6 GB)
├── config.json / generation_config.json / processor_config.json
├── tokenizer.json / tokenizer_config.json / chat_template.jinja
├── recipe.yaml                # llm-compressor recipe
└── assistant/                 # gemma4_mtp draft (bf16, 0.84 GB)
    ├── model.safetensors
    ├── config.json            # model_type: gemma4_assistant
    └── tokenizer / chat_template

Quickstart

hf download sakamakismile/Huihui-gemma-4-26B-A4B-it-qat-abliterated-MTP-NVFP4 \
  --local-dir gemma4-26b-mtp
DIR=$(realpath gemma4-26b-mtp)

NCCL_P2P_DISABLE=1 vllm serve "$DIR" \
  --served-model-name gemma4-26b-qat-mtp \
  --tensor-parallel-size 2 \
  --disable-custom-all-reduce \
  --kv-cache-dtype fp8 \
  --max-model-len 16384 \
  --gpu-memory-utilization 0.92 \
  --max-num-batched-tokens 8192 \
  --limit-mm-per-prompt '{"image":0}' \
  --speculative-config "{\"method\":\"gemma4_mtp\",\"model\":\"$DIR/assistant\",\"num_speculative_tokens\":4}"
  • The point: model in --speculative-config is the bundled local path — no second download, no HF resolution at serve time.
  • TP=2 (2× 16 GB) recommended. Weights ~16.4 GiB don't fit one 16 GB card; TP=4 buys only +6% single-stream on this MoE — spend extra GPUs on a second replica.
  • NCCL_P2P_DISABLE=1 + --disable-custom-all-reduce are required on PCIe no-NVLink boxes (TP hangs without them); drop both if you have NVLink/P2P.
  • --kv-cache-dtype fp8 doubles KV; with the assistant on board this config still holds maxlen 16384 / KV 23,560 tok at GMU 0.92. Keep CUDA graphs ON (no --enforce-eager).
  • vLLM 0.21's quantization-inheritance trap does not fire here: with an explicit draft model path the draft's own config decides (bf16). Only the model:null MTP-from-target path inherits target quantization.

Measured (RTX PRO 2000 Blackwell 16 GB ×2, TP=2, PCIe no-NVLink, fp8 KV, vLLM 0.21.0, 2026-06-12)

Single-stream, T=0 chat completions, ×3 each. Acceptance = accepted/drafted from /metrics.

config JA 128 JA 512 EN 128 EN 512 acceptance JA / EN
baseline (no spec) 108.5 108.9 108.2 — —
native MTP (this bundle, N=4) 133.6 121.0 163.2 142.4 35–44% / 50–64%
EAGLE-3 (English-trained draft) N=3 73.2 75.7 158.0 129.9 2.7–3.9% / 33–49%
ngram N=4 (lookup 2–4) 67.5 74.4 70.4 69.4 10–25% / 13–21%

JA +12–23%, EN +32–51% over baseline. Two lessons paid for in benchmarks:

  • EAGLE-3 collapses on Japanese — the English-Magpie-trained draft gets 3–4% JA acceptance and lands below baseline (0.69×). The google MTP assistant holds 35–44% JA acceptance because its 4-layer head re-uses the target's own hidden states instead of imitating its distribution from scratch. If your traffic is non-English, MTP is the only one of the three that pays.
  • ngram never pays for itself on free-form chat in either language.

The MoE body already runs 108 tok/s, so MTP's gain here is "fast → faster" (1.2–1.5×). On the dense 31B sibling the same assistant nearly triples throughput — see Huihui-gemma-4-31B-it-qat-abliterated-MTP-NVFP4.

Concurrent (aggregate throughput)

4 / 8 concurrent streams × 256 tok each (T=0, diverse prompts, prefix-cache busted, ×3 averaged). Baseline = same body, no spec-decode (measured with JA prompts; baseline JA≈EN single-stream).

streams baseline tok/s MTP JA tok/s MTP EN tok/s acceptance JA / EN
1 108.5 133.6 (+23%) 163.2 (+51%) 35–44% / 50–64%
4 326.3 341.5 (+4.7%) 363.4 ~36% / ~47%
8 571.1 557.6 (−2.4%) 566.4 (−0.8%) ~39% / ~45%

On this MoE, MTP pays up to ~4 concurrent streams; at 8 it is a wash (−2%, run-noise territory). Acceptance stays flat under batch — the shrinking gain isn't the drafter failing, it's the GPUs reaching compute saturation (571 tok/s aggregate), where verifying rejected draft tokens competes with real batch work instead of filling decode bubbles. Worst case is break-even, so keeping MTP resident costs ~nothing at peak load and buys 1.2–1.5× whenever concurrency drops. (The dense 31B sibling, still verification-hungry at 8 streams, keeps +27% JA / +53% EN there.)

The body: QAT × NVFP4 (the finding, in short)

Full W4A4 NVFP4 breaks non-QAT gemma-4 on Japanese long-form — the non-QAT 26B sibling intermittently collapses into repetition loops past ~500 tokens. This QAT-origin body does not: adversarial 1300+-token Japanese essays run to natural EOS with zero loop-detector hits, while English (HumanEval-class) is unaffected. Pattern holds across dense (31B/12B) and MoE: if you want gemma-4 in NVFP4 W4A4, go through a QAT checkpoint — the q4_0-shaped weight distribution is the prior FP4 wants. Full evidence, bake recipe (pure-CPU calibration through the multimodal processor, MoECalibrationModule, pad768 surgery) in the non-MTP card.

Notes

  • Abliterated (uncensored). Refusal behavior removed upstream — you are responsible for your deployment. Use responsibly and lawfully.
  • NVFP4 is Blackwell-specific; the body will not run on Ampere/Hopper. The bf16 assistant inherits the body's GPU anyway.
  • The assistant/ checkpoint is google's, redistributed unmodified under the same Gemma terms; its original model card is included as assistant/README.md.
  • Gemma is provided under and subject to the Gemma Terms of Use.

Credits

  • Original model & MTP assistant: Google DeepMind (Gemma 4, QAT q4_0)
  • QAT-unquantize & abliteration: huihui-ai
  • NVFP4 quantization, pad768 surgery, spec-decode measurement & bundle: Lna-Lab · Tooling: llm-compressor / vLLM

Support the Base Model Author (huihui-ai)

If you find the abliterated base useful, please support huihui-ai:

README history 2 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-06-12Add concurrent aggregate TPS: MTP pays to ~4 streams, break-even at 8 (341.5/...2fc72e99.4 KB
    Loading...
  2. 2026-06-12Upload folder using huggingface_hubb7a9e1e8.3 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration