← back to catalog · registered 2026-08-22 13:56

sakamakismile/Huihui-gemma-4-26B-A4B-it-abliterated-NVFP4-vLLM

sakamakismile Gemma 13B MoE second-order
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/sakamakismile%2FHuihui-gemma-4-26B-A4B-it-abliterated-NVFP4-vLLM"
Response includes
  • classification m1
  • files 10
  • hub_downloads_all_time 789
  • author_summary 34 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M1
Primary method

Direct removal

No other method signals detected in this model.
Confidence
MEDIUM
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=0 (base model)
  • no specific method indicators - defaulting to M1 (most common)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
MEDIUM
Why we say so
primary_method=M1; difference-of-means is the reference extraction for M1/M3 (Arditi 2024)
Downloads · lifetime
789
90 last 30d - stable
Likes
0
Model age
4mo ago
created 2026-06-05
Downloads over time
Now812→from58↑1,300%
2030959888758 on Jun 10812 on Oct 11812 on Oct 10JunJulAugSepOct
Jun 10 → Oct 11 · 57 snapshots · spans 123 days

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
apache-2.0
Tags
safetensors gemma4 nvfp4 vllm compressed-tensors moe blackwell abliterated uncensored text-generation conversational base_model:huihui-ai/Huihui-gemma-4-26B-A4B-it-abliterated

Related

Total size
16.4 GB
Files
10
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-06-05 02:05

Files by quantization

Auxiliary files 10 files 16.4 GB
model.safetensors 16.4 GB 90589181 download
tokenizer.json 30.7 MB d93b1947 download
config.json 19.2 KB e8c5d8f3 download
chat_template.jinja 11.8 KB 33c51c2d download
README.md 6.39 KB f4a3417d download
tokenizer_config.json 2.62 KB 59dd4b62 download
processor_config.json 1.65 KB 5465974d download
.gitattributes 1.53 KB 52373fe2 download
recipe.yaml 224 B f790e78b download
generation_config.json 203 B 3110a9b4 download

README current version from Hugging Face


license: apache-2.0
base_model:

  • huihui-ai/Huihui-gemma-4-26B-A4B-it-abliterated
  • sakamakismile/Huihui-gemma-4-26B-A4B-it-abliterated-NVFP4
    tags:
  • gemma4
  • nvfp4
  • vllm
  • compressed-tensors
  • moe
  • blackwell
  • abliterated
  • uncensored
  • text-generation
    pipeline_tag: text-generation

Huihui gemma-4-26B-A4B-it abliterated · NVFP4 (vLLM-ready, 768)

An uncensored and full-strength build of gemma-4-26B-A4B in NVFP4 (W4A4)
that loads on a stock vLLM with no source patches.

Two things make this build worth picking:

1. It serves on stock vLLM. The existing 704-intermediate NVFP4 quant fails to
load on a stock vLLM with:

NotImplementedError: ('Intermediate size padding for w1 and w3, for %s NvFp4 backend,
but this is not currently supported', 'VLLM_CUTLASS')

This repo bakes the kernel alignment in: the MoE intermediate is zero-padded
704 → 768
offline, on the packed FP4 weights — a mathematically loss-less
transform that inherits the original (already-correct) activation scales and
needs no re-quant, no GPU, no calibration. 768 is a multiple of 128, so the
CUTLASS NVFP4 MoE kernel accepts it directly (and the TP>1 / PP>1 gated-MoE
NotImplementedError is sidestepped).

2. Abliterated without the usual quality tax. A common worry with uncensored
models is that the modification hurts capability. We measured it. On HumanEval+
(EvalPlus strict tests, 163 problems) this lightly abliterated build scores the
same as the unmodified official model — far above a heavily-modified variant:

gemma-4-26B-A4B build HumanEval+ pass@1
heavily-modified "super" abliteration 77.3 %
official (no abliteration) 90.8 %
this — huihui abliteration 90.8 %

You keep the uncensored behavior and the full coding ability.

Specs

  • Format: NVFP4 (FP4 weights and activations, compressed-tensors).
  • Arch: gemma4 MoE — 128 experts, top-8, hidden 2816, moe_intermediate 768, 30 layers, sliding-window attention. Context up to 256K.
  • VRAM: weights are ~15.3 GiB — they do not fit a single 16 GB card (no room left for KV → OOM). Run on 2× 16 GB (--tensor-parallel-size 2) or a single ≥ 20 GB GPU.

Throughput (measured, RTX PRO 2000 Blackwell 16 GB, CUDA graph, fp8 KV)

TP=2 (2× 16 GB): single-user ≈ 99 tok/s; energy-efficient.

concurrent 1 2 4 8 16 32
aggregate tok/s 98.8 161.9 278.1 489.4 784.1 1186.3
tok/joule 0.72 1.18 2.00 3.52 5.67 8.54
~watts (2 GPU) 137 137 139 139 138 139

TP=4 (4× 16 GB): faster single-stream + higher peak aggregate.

concurrent 1 2 4 8 16 32
aggregate tok/s 127.4 213.6 377.7 683.9 1089.3 1571.6
tok/joule 0.54 0.89 1.41 2.46 3.98 5.63
~watts (4 GPU) 236 240 268 278 274 279

(Aggregate = summed over concurrent requests. TP=2 wins tok/joule — half the GPUs,
half the power; TP=4 wins single-stream latency and peak throughput. Power stays
~flat across concurrency: batching is the efficiency win.)

Quick start — from zero to inference

NVIDIA Blackwell GPU(s) + Docker with the NVIDIA Container Toolkit. NVFP4 is
auto-detected; no --quantization flag.

2× 16 GB (tensor parallel, no NVLink) — the typical setup:

docker run --gpus all -p 8000:8000 -v $(pwd):/model:ro \
  -e NCCL_P2P_DISABLE=1 -e PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True \
  vllm/vllm-openai:cu130-nightly \
  /model --served-model-name gemma \
  --tensor-parallel-size 2 --disable-custom-all-reduce \
  --gpu-memory-utilization 0.85 --max-num-seqs 32 --max-num-batched-tokens 4096 \
  --max-model-len 8192 --kv-cache-dtype fp8 \
  --limit-mm-per-prompt '{"image":0,"video":0,"audio":0}'

For TP=4 set --tensor-parallel-size 4 and --gpu-memory-utilization 0.90.
For a single ≥24 GB GPU drop the --tensor-parallel-size/NCCL lines.

Inference flags that matter:

  • Keep CUDA graph ON (do not pass --enforce-eager) — --gpu-memory-utilization 0.85 + --max-num-seqs 32 leave room for the ~0.3 GiB capture and give ~10× faster decode than eager.
  • NCCL_P2P_DISABLE=1 + --disable-custom-all-reduce are required for tensor-parallel on PCIe/NODE topology (no NVLink).
  • --kv-cache-dtype fp8 for long-context capacity (~11 full-128K requests fit on 4×16 GB).
  • --limit-mm-per-prompt '{"image":0,"video":0,"audio":0}' to serve text-only and skip multimodal profiling.

Talk to it (instruct model — use the chat endpoint):

curl -s localhost:8000/v1/chat/completions -H 'Content-Type: application/json' -d '{
  "model":"gemma","messages":[{"role":"user","content":"侘び寂びを一文で。"}],
  "max_tokens":256,"temperature":0.7}'

Use /v1/chat/completions (the Gemma chat template is applied). Raw /v1/completions
on an instruct model yields degenerate output.

How it was made

Offline FP4 surgery on the 704 NVFP4 checkpoint: per expert, {gate,up}_proj
weight+scale are padded 704→768 on the output dim and down_proj on the input dim
with FP4/FP8 0x00 (=+0.0); every *_global_scale and all non-expert tensors are
copied verbatim. Loss-less (no MoE-intermediate norm; gelu(0)·0=0; padded
down_proj columns multiply zero weights). Verified: input_global_scale
median ≈ 322, zero 1.0/0.0 sentinels, expert shapes 768-aligned.

This is an abliterated (uncensored) model — safety training removed. Use responsibly.

Credits

Support the Base Model Author

If you find this model useful, please consider supporting huihui-ai — the creator of the abliterated base model:

README history 1 version

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-06-05Upload folder using huggingface_hub3bb19186.4 KB
    Loading...

Discussions 1 thread

  1. 2026-08-05rapid degrade into psychosis.open1 💬#1
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration