← back to catalog · registered 2026-08-22 13:56

sakamakismile/SuperGemma4-26B-Abliterated-Multimodal-NVFP4-vLLM

sakamakismile Gemma 13B MoE multimodal second-order
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/sakamakismile%2FSuperGemma4-26B-Abliterated-Multimodal-NVFP4-vLLM"
Response includes
  • classification m1
  • files 10
  • hub_downloads_all_time 3,030
  • author_summary 34 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M1
Primary method

Direct removal

No other method signals detected in this model.
Confidence
MEDIUM
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=0 (base model)
  • no specific method indicators - defaulting to M1 (most common)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
MEDIUM
Why we say so
primary_method=M1; difference-of-means is the reference extraction for M1/M3 (Arditi 2024)
Downloads · lifetime
3K
299 last 30d - cooling
Likes
0
Model age
4mo ago
created 2026-06-04
Downloads over time
Now3.2K→from50↑6,208%
01.2K2.3K3.5K50 on Jun 103.2K on Oct 11JunJulAugSepOct
Jun 10 → Oct 11 · 57 snapshots · spans 123 days

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
gemma
Tags
safetensors gemma4 nvfp4 vllm compressed-tensors moe blackwell abliterated text-generation conversational base_model:Jiunsong/supergemma4-26b-abliterated-multimodal base_model:quantized:Jiunsong/supergemma4-26b-abliterated-multimodal

Related

Total size
16.4 GB
Files
10
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-06-05 00:10

Files by quantization

Auxiliary files 10 files 16.4 GB
model.safetensors 16.4 GB 79af154d download
tokenizer.json 30.7 MB d93b1947 download
config.json 19.3 KB 42e8aef8 download
chat_template.jinja 16.1 KB 98da08eb download
README.md 5.46 KB 013a26ac download
tokenizer_config.json 2.62 KB 59dd4b62 download
processor_config.json 1.65 KB 5465974d download
.gitattributes 1.53 KB 52373fe2 download
recipe.yaml 224 B f790e78b download
generation_config.json 203 B 3110a9b4 download

README current version from Hugging Face


license: gemma
base_model:

  • sakamakismile/SuperGemma4-26B-Abliterated-Multimodal-NVFP4
  • Jiunsong/supergemma4-26b-abliterated-multimodal
    tags:
  • gemma4
  • nvfp4
  • vllm
  • compressed-tensors
  • moe
  • blackwell
  • abliterated
  • text-generation
    pipeline_tag: text-generation

SuperGemma4-26B-A4B · Abliterated · Multimodal · NVFP4 (vLLM-ready, 768)

A drop-in, serve-it-now build of SuperGemma4-26B in NVFP4 (W4A4) that loads
on a stock vLLM with no source patches.

The excellent existing NVFP4 quant of this model has moe_intermediate_size = 704.
704 is not aligned for vLLM's fast FP4 MoE kernels (Marlin / CUTLASS / FlashInfer),
so on a stock vLLM it fails at load with:

NotImplementedError: ('Intermediate size padding for w1 and w3, for %s NvFp4 backend,
but this is not currently supported', 'VLLM_CUTLASS')

This repo fixes that by baking the alignment in: the MoE intermediate is
zero-padded 704 → 768 offline, on the packed FP4 weights — a mathematically
loss-less transform that inherits the original (already-correct) activation
scales
and needs no re-quantization, no GPU, no calibration. 768 is a
multiple of 128, so vLLM's CUTLASS NVFP4 MoE kernel accepts it directly (and the
runtime swizzle-pad becomes a no-op, which also avoids the TP>1 / PP>1 gated-MoE
NotImplementedError).

If you tried to serve the 704 build and hit the CUTLASS error above — use this
one instead.
Same weights, same scales, just kernel-aligned.

TL;DR specs

  • Format: NVFP4 (FP4 weights and activations, compressed-tensors).

  • Arch: gemma4 MoE — 128 experts, top-8, hidden 2816, moe_intermediate 768, 30 layers, sliding-window attention. Context up to 256K (max_position_embeddings).

  • VRAM: weights are ~15.3 GiB — they do not fit a single 16 GB card (the weights fill it, leaving no room for the KV cache → OOM). Run on 2× 16 GB (--tensor-parallel-size 2) or a single ≥ 20 GB GPU.

  • Single-user speed: ≈ 99 tok/s single-stream (1 request at a time), measured on 2× RTX PRO 2000 Blackwell 16 GB · TP=2 · CUDA graph · fp8 KV (very stable: 99.1–99.4).

  • Server throughput (many concurrent users): measured on 4× 16 GB (TP=4) — these are 4-GPU aggregate figures over concurrent requests, not single-stream:

    concurrent requests 1 2 4 8 16 32 64
    aggregate tok/s 127 212 368 664 1056 1527 1990
    tok/joule 0.54 0.87 1.39 2.42 3.77 5.47 7.19

    Power stays ~flat (~235–280 W for the 4 cards) while aggregate throughput scales ~16× — batching is the efficiency win. ~11 full-128K-context requests fit on 4×16 GB with fp8 KV.

Quick start — from zero to inference

You need NVIDIA Blackwell GPU(s) and Docker with the NVIDIA Container Toolkit.
NVFP4 is auto-detected (compressed-tensors); no --quantization flag.

Single 24 GB+ Blackwell GPU:

docker run --gpus all -p 8000:8000 -v $(pwd):/model:ro \
  vllm/vllm-openai:cu130-nightly \
  /model --served-model-name supergemma4 \
  --max-model-len 8192 \
  --limit-mm-per-prompt '{"image":0,"video":0,"audio":0}'

Multiple 16 GB Blackwell GPUs (e.g. 2× via tensor parallel, no NVLink):

docker run --gpus all -p 8000:8000 -v $(pwd):/model:ro \
  -e NCCL_P2P_DISABLE=1 -e PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True \
  vllm/vllm-openai:cu130-nightly \
  /model --served-model-name supergemma4 \
  --tensor-parallel-size 2 --disable-custom-all-reduce \
  --gpu-memory-utilization 0.85 --max-num-seqs 16 --max-num-batched-tokens 2048 \
  --max-model-len 8192 \
  --limit-mm-per-prompt '{"image":0,"video":0,"audio":0}'

Notes for 16 GB cards: keep CUDA graph on (do not pass --enforce-eager) —
the modest --gpu-memory-utilization 0.85 + --max-num-seqs 16 leave room for the
~0.2 GiB graph capture, and you get ~10× faster decode than eager. NCCL_P2P_DISABLE=1

  • --disable-custom-all-reduce are required for tensor parallel on PCIe/NODE topology
    (no NVLink). Add --kv-cache-dtype fp8 for long-context capacity.

Talk to it (this is an instruct model — use the chat endpoint):

curl -s localhost:8000/v1/chat/completions -H 'Content-Type: application/json' -d '{
  "model":"supergemma4",
  "messages":[{"role":"user","content":"侘び寂びを一段落で説明して。"}],
  "max_tokens":300, "temperature":0.7
}'

Use /v1/chat/completions (the Gemma chat template is applied). Raw /v1/completions
on an instruct model produces degenerate output.

How it was made

Offline FP4 surgery on the 704 NVFP4 checkpoint: for every expert,
{gate,up}_proj weight+scale are padded 704→768 along the output dim and
down_proj along the input dim, with FP4/FP8 0x00 (= +0.0); every *_global_scale
and all non-expert tensors are copied verbatim. Loss-less because the MoE
intermediate has no norm and gelu(0)·0 = 0; the padded down_proj columns
multiply zero weights. Verified: input_global_scale median ≈ 330, zero 1.0/0.0
sentinels, all expert shapes 768-aligned.

License & credits

Built on Google Gemma — your use is subject to the
Gemma Terms of Use. SuperGemma4 enhancement by
Jiunsong (Jiunsong/supergemma4-26b-abliterated-multimodal); abliterated +
multimodal NVFP4 by sakamakismile. This repo adds only the loss-less 704→768
kernel-alignment padding. Abliterated (reduced refusals) — use responsibly.

README history 2 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-06-05Upload README.md with huggingface_hub1737ff35.5 KB
    Loading...
  2. 2026-06-04Upload folder using huggingface_hube15ee955 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration