license: gemma
base_model:
- sakamakismile/SuperGemma4-26B-Abliterated-Multimodal-NVFP4
- Jiunsong/supergemma4-26b-abliterated-multimodal
tags: - gemma4
- nvfp4
- vllm
- compressed-tensors
- moe
- blackwell
- abliterated
- text-generation
pipeline_tag: text-generation
SuperGemma4-26B-A4B · Abliterated · Multimodal · NVFP4 (vLLM-ready, 768)
A drop-in, serve-it-now build of SuperGemma4-26B in NVFP4 (W4A4) that loads
on a stock vLLM with no source patches.
The excellent existing NVFP4 quant of this model has moe_intermediate_size = 704.
704 is not aligned for vLLM's fast FP4 MoE kernels (Marlin / CUTLASS / FlashInfer),
so on a stock vLLM it fails at load with:
NotImplementedError: ('Intermediate size padding for w1 and w3, for %s NvFp4 backend,
but this is not currently supported', 'VLLM_CUTLASS')
This repo fixes that by baking the alignment in: the MoE intermediate is
zero-padded 704 → 768 offline, on the packed FP4 weights — a mathematically
loss-less transform that inherits the original (already-correct) activation
scales and needs no re-quantization, no GPU, no calibration. 768 is a
multiple of 128, so vLLM's CUTLASS NVFP4 MoE kernel accepts it directly (and the
runtime swizzle-pad becomes a no-op, which also avoids the TP>1 / PP>1 gated-MoENotImplementedError).
If you tried to serve the 704 build and hit the CUTLASS error above — use this
one instead. Same weights, same scales, just kernel-aligned.
TL;DR specs
Format: NVFP4 (FP4 weights and activations,
compressed-tensors).Arch: gemma4 MoE — 128 experts, top-8, hidden 2816,
moe_intermediate 768, 30 layers, sliding-window attention. Context up to 256K (max_position_embeddings).VRAM: weights are ~15.3 GiB — they do not fit a single 16 GB card (the weights fill it, leaving no room for the KV cache → OOM). Run on 2× 16 GB (
--tensor-parallel-size 2) or a single ≥ 20 GB GPU.Single-user speed: ≈ 99 tok/s single-stream (1 request at a time), measured on 2× RTX PRO 2000 Blackwell 16 GB · TP=2 · CUDA graph · fp8 KV (very stable: 99.1–99.4).
Server throughput (many concurrent users): measured on 4× 16 GB (TP=4) — these are 4-GPU aggregate figures over concurrent requests, not single-stream:
concurrent requests 1 2 4 8 16 32 64 aggregate tok/s 127 212 368 664 1056 1527 1990 tok/joule 0.54 0.87 1.39 2.42 3.77 5.47 7.19 Power stays ~flat (~235–280 W for the 4 cards) while aggregate throughput scales ~16× — batching is the efficiency win. ~11 full-128K-context requests fit on 4×16 GB with fp8 KV.
Quick start — from zero to inference
You need NVIDIA Blackwell GPU(s) and Docker with the NVIDIA Container Toolkit.
NVFP4 is auto-detected (compressed-tensors); no --quantization flag.
Single 24 GB+ Blackwell GPU:
docker run --gpus all -p 8000:8000 -v $(pwd):/model:ro \
vllm/vllm-openai:cu130-nightly \
/model --served-model-name supergemma4 \
--max-model-len 8192 \
--limit-mm-per-prompt '{"image":0,"video":0,"audio":0}'
Multiple 16 GB Blackwell GPUs (e.g. 2× via tensor parallel, no NVLink):
docker run --gpus all -p 8000:8000 -v $(pwd):/model:ro \
-e NCCL_P2P_DISABLE=1 -e PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True \
vllm/vllm-openai:cu130-nightly \
/model --served-model-name supergemma4 \
--tensor-parallel-size 2 --disable-custom-all-reduce \
--gpu-memory-utilization 0.85 --max-num-seqs 16 --max-num-batched-tokens 2048 \
--max-model-len 8192 \
--limit-mm-per-prompt '{"image":0,"video":0,"audio":0}'
Notes for 16 GB cards: keep CUDA graph on (do not pass --enforce-eager) —
the modest --gpu-memory-utilization 0.85 + --max-num-seqs 16 leave room for the
~0.2 GiB graph capture, and you get ~10× faster decode than eager. NCCL_P2P_DISABLE=1
--disable-custom-all-reduceare required for tensor parallel on PCIe/NODE topology
(no NVLink). Add--kv-cache-dtype fp8for long-context capacity.
Talk to it (this is an instruct model — use the chat endpoint):
curl -s localhost:8000/v1/chat/completions -H 'Content-Type: application/json' -d '{
"model":"supergemma4",
"messages":[{"role":"user","content":"侘び寂びを一段落で説明して。"}],
"max_tokens":300, "temperature":0.7
}'
Use /v1/chat/completions (the Gemma chat template is applied). Raw /v1/completions
on an instruct model produces degenerate output.
How it was made
Offline FP4 surgery on the 704 NVFP4 checkpoint: for every expert,{gate,up}_proj weight+scale are padded 704→768 along the output dim anddown_proj along the input dim, with FP4/FP8 0x00 (= +0.0); every *_global_scale
and all non-expert tensors are copied verbatim. Loss-less because the MoE
intermediate has no norm and gelu(0)·0 = 0; the padded down_proj columns
multiply zero weights. Verified: input_global_scale median ≈ 330, zero 1.0/0.0
sentinels, all expert shapes 768-aligned.
License & credits
Built on Google Gemma — your use is subject to the
Gemma Terms of Use. SuperGemma4 enhancement by
Jiunsong (Jiunsong/supergemma4-26b-abliterated-multimodal); abliterated +
multimodal NVFP4 by sakamakismile. This repo adds only the loss-less 704→768
kernel-alignment padding. Abliterated (reduced refusals) — use responsibly.