← back to catalog · registered 2026-08-22 13:56

sakamakismile/Huihui-gemma-4-12B-it-abliterated-NVFP4A16

sakamakismile Gemma 5.5B multimodal second-order
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/sakamakismile%2FHuihui-gemma-4-12B-it-abliterated-NVFP4A16"
Response includes
  • classification m1
  • files 10
  • hub_downloads_all_time 3,806
  • author_summary 34 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M1
Primary method

Direct removal

No other method signals detected in this model.
Confidence
MEDIUM
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=0 (base model)
  • no specific method indicators - defaulting to M1 (most common)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
MEDIUM
Why we say so
primary_method=M1; difference-of-means is the reference extraction for M1/M3 (Arditi 2024)
Downloads · lifetime
4K
395 last 30d - stable
Likes
4
Model age
4mo ago
created 2026-06-07
Downloads over time
Now4K→from25↑15,936%
01.5K2.9K4.4K25 on Jun 84K on Oct 11JunJulAugSepOct
Jun 8 → Oct 11 · 58 snapshots · spans 125 days

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
gemma
Tags
transformers safetensors gemma4_unified image-text-to-text gemma4 nvfp4 w4a16 quantized abliterated vllm compressed-tensors blackwell

Related

Total size
7.65 GB
Files
10
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-06-07 23:53

Files by quantization

Auxiliary files 10 files 7.68 GB
model.safetensors 7.65 GB 10a8f6d1 download
tokenizer.json 30.7 MB 88e71407 download
chat_template.jinja 17.1 KB e61bbfe9 download
README.md 7.98 KB 07438636 download
config.json 5.31 KB 9ab7528a download
tokenizer_config.json 2.68 KB 18faad3a download
.gitattributes 1.53 KB 52373fe2 download
processor_config.json 1.35 KB b889adcd download
generation_config.json 255 B 683ff358 download
recipe.yaml 204 B 9a617c03 download

README current version from Hugging Face


license: gemma
base_model: huihui-ai/Huihui-gemma-4-12B-it-abliterated
base_model_relation: quantized
tags:

  • gemma4
  • gemma4_unified
  • nvfp4
  • w4a16
  • quantized
  • abliterated
  • vllm
  • compressed-tensors
  • blackwell
  • sm120
  • multimodal
  • speculative-decoding
    library_name: transformers
    pipeline_tag: image-text-to-text
    model_type: gemma4_unified
    quantized_by: Lna-Lab

Huihui-gemma-4-12B-it-abliterated-NVFP4A16

NVFP4 (W4A16) quantization of huihui-ai/Huihui-gemma-4-12B-it-abliterated — the abliterated (uncensored) Gemma 4 12B unified model (text + vision + audio).

24 GB → 7.7 GB. Runs on a single 16 GB Blackwell GPU, or shards across several for higher throughput. Up to 118 tok/s single-stream (TP=4 + MTP speculative decode) and ~1117 tok/s aggregate.

Base huihui-ai/Huihui-gemma-4-12B-it-abliterated (abliterated google/gemma-4-12B-it)
Architecture Gemma4UnifiedForConditionalGeneration — 12B dense, 48 layers, 131K ctx
Quantization NVFP4A16 — weights FP4 (group 16, FP8 scales), activations BF16
Format compressed-tensors / nvfp4-pack-quantized (native vLLM)
Tool llm-compressor
Size 7.7 GB · Requires NVIDIA Blackwell (SM120)

Weight-only FP4 (W4A16) keeps activations at BF16, so it is robust where full W4A4 NVFP4 collapses on this architecture.


Quickstart

Requires a Blackwell GPU (SM120 / RTX 50-series / GB10 / B100/B200), Docker with the NVIDIA runtime, and the hf CLI. Gemma 4 unified is brand new — you need vLLM nightly (released ≤ 0.22.1 lack the Gemma4Unified class).

# 1) Download this model (7.7 GB). For spec-decode, also grab the 0.4B MTP draft.
hf download sakamakismile/Huihui-gemma-4-12B-it-abliterated-NVFP4A16 --local-dir ./model
hf download google/gemma-4-12B-it-assistant --local-dir ./draft   # optional, for spec-decode

# 2a) Simplest — single GPU, no speculative decode
docker run --rm --gpus '"device=0"' --ipc=host --shm-size 16gb -p 8000:8000 \
  -v $PWD/model:/model:ro \
  vllm/vllm-openai:nightly \
  --model /model --served-model-name gemma4-12b --max-model-len 65536 \
  --gpu-memory-utilization 0.92 --trust-remote-code

Multi-GPU — read this if your box has no NVLink

On consumer/entry Blackwell (e.g. RTX PRO 2000) over plain PCIe there is no working GPU P2P, and vLLM tensor-parallel hangs unless you disable both NCCL P2P and vLLM's custom all-reduce:

docker run --rm --gpus '"device=0,1,2,3"' --ipc=host --shm-size 16gb -p 8000:8000 \
  -e NCCL_P2P_DISABLE=1 \                          # <-- without this, hangs at NCCL init
  -v $PWD/model:/model:ro \
  vllm/vllm-openai:nightly \
  --model /model --served-model-name gemma4-12b \
  --tensor-parallel-size 4 \
  --disable-custom-all-reduce \                     # <-- without this, the forward deadlocks
  --max-model-len 65536 --gpu-memory-utilization 0.85 --trust-remote-code

Maximum interactive speed — TP=4 + MTP speculative decode

Google ships a 0.4B MTP draft (google/gemma-4-12B-it-assistant). It nearly doubles single-stream throughput (lossless — the target verifies every token). Use num_speculative_tokens: 3 (the stable optimum; k≥5 collapses acceptance) and --kv-cache-dtype fp8 (NVFP4 KV would break the draft):

docker run --rm --gpus '"device=0,1,2,3"' --ipc=host --shm-size 16gb -p 8000:8000 \
  -e NCCL_P2P_DISABLE=1 \
  -v $PWD/model:/model:ro -v $PWD/draft:/draft:ro \
  vllm/vllm-openai:nightly \
  --model /model --served-model-name gemma4-12b \
  --tensor-parallel-size 4 --disable-custom-all-reduce \
  --kv-cache-dtype fp8 \
  --speculative-config '{"method":"mtp","model":"/draft","num_speculative_tokens":3}' \
  --max-model-len 65536 --gpu-memory-utilization 0.85 --trust-remote-code

Test it:

curl -s localhost:8000/v1/chat/completions -H 'Content-Type: application/json' -d \
 '{"model":"gemma4-12b","messages":[{"role":"user","content":"Explain the CAP theorem in one sentence."}]}'

Flag cheat-sheet

Flag / env When Why
vllm/vllm-openai:nightly always only nightly registers Gemma4UnifiedForConditionalGeneration
--trust-remote-code always new arch
NCCL_P2P_DISABLE=1 (env) TP > 1 on no-NVLink else hangs at NCCL init
--disable-custom-all-reduce TP > 1 on no-NVLink else the forward deadlocks
--ipc=host --shm-size 16gb TP > 1 (docker) host-path NCCL needs shared memory
--speculative-config '{"method":"mtp",…,"num_speculative_tokens":3}' interactive ~1.6–1.7× single-stream
--kv-cache-dtype fp8 with spec-decode nvfp4 KV collapses draft acceptance
--max-num-seqs 4 (+ --gpu-memory-utilization 0.95) single GPU, long ctx frees KV room for up to -c 32768 on 16 GB

Benchmarks

Measured on 4× RTX PRO 2000 Blackwell (16 GB, SM120, 288 GB/s, PCIe — no NVLink), TP=4, -c 65536.

Single-stream decode (interactive) — TP sweep, 1 request × 512 tok:

TP GPUs no spec + MTP (k=3) MTP gain
1 1 30.5 55.0 1.80×
2 2 53.2 94.8 1.78×
4 4 73.3 118.5 1.62×

(TP=4 + MTP peaks at 121.0 with k=4, but k=3 is the stable optimum.) MTP gives a steady ~1.6–1.8× at every TP. TP scaling is sub-linear on this no-NVLink box (host-memory all-reduce). Pick by what you have:

goal config single-stream GPUs freed
low-power, 1-GPU resident TP=1 + MTP 55 5
balanced TP=2 + MTP 95 4
fastest interactive TP=4 + MTP 118 2

Aggregate throughput (concurrency sweep, no spec-decode):

concurrency 1 2 4 8 16 32
tok/s (-c65536) 73 145 274 487 796 1117
tok/s (-c131072) 74 145 275 498 792 1100

64K and 128K context decode identically (sliding-window KV). Rule: MTP spec-decode for low concurrency (≤8); turn it off for high-concurrency batch serving (it costs throughput once the batch saturates).

Quality — measured vs BF16 base and an FP8 build (same huihui base)

Greedy side-by-side on EN / 繁體中文 / 日本語 / code / facts / reasoning traps:

  • Standard tasks: identical. Facts (Chernobyl: April 1986, reactor 4), Traditional-Chinese & Japanese explanations, 17×23−100 = 291, 60 km / 45 min = 80 km/h, code — NVFP4 = FP8 = BF16 base, no collapse, no drift.
  • Hard reasoning traps (7 tested): a small, real W4A16 tax. FP8 matched the BF16 base on every trap the base got right; NVFP4 slipped on ~1 of 7 (it answered a Barbara-type syllogism "Yes" where No is correct, plus one minor secondary-detail slip). One age-word-problem even the BF16 base fails — a model limit, not a quant artifact.

Verdict: half the size and faster than FP8, at standard-task parity. Choose FP8 for maximum reasoning fidelity; choose this NVFP4A16 for the best size/speed at ~85–90% reasoning parity — the right default for most local-agent and chat workloads.

Notes

  • Abliterated (uncensored). Use responsibly.
  • NVFP4 is Blackwell-specific; it will not run on Ampere/Hopper.
  • Multimodal vision/audio embedders kept in BF16.

Credits

Support the Base Model Author (huihui-ai)

If you find the abliterated base useful, please support huihui-ai:

README history 2 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-06-07Add single-stream TP=1/2/4 benchmark sweepc31260c8 KB
    Loading...
  2. 2026-06-07NVFP4A16 (W4A16) quant of huihui-gemma-4-12B-it-abliterated + serving guidecc067717.5 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration