← back to catalog · registered 2026-08-22 13:56

lyf/Huihui-gemma-4-31B-it-abliterated-v2-NVFP4

lyf Gemma 15B multimodal second-order
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/lyf%2FHuihui-gemma-4-31B-it-abliterated-v2-NVFP4"
Response includes
  • classification m1
  • files 10
  • hub_downloads_all_time 132,080
  • author_summary 10 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M1
Primary method

Direct removal

No other method signals detected in this model.
Confidence
MEDIUM
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=0 (base model)
  • no specific method indicators - defaulting to M1 (most common)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
MEDIUM
Why we say so
primary_method=M1; difference-of-means is the reference extraction for M1/M3 (Arditi 2024)
Downloads · lifetime
132K
1K last 30d - cooling
Likes
3
Model age
6mo ago
created 2026-04-11

Training datasets

Corpora the author lists in the model card. Datasets tracked in our /datasets catalog carry a category badge linking to the workflow stage. Others open on Hugging Face.

Downloads over time
Now132.3K→from329↑40,123%
048.5K97K145.5K329 on Apr 15132.3K on Oct 11AprMayJunJulAugSepOct
Apr 15 → Oct 11 · 65 snapshots · spans 179 days

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
gemma
Tags
transformers safetensors gemma4 image-text-to-text quantized nvfp4 fp4 4-bit compressed-tensors llm-compressor vllm multimodal

Related

Total size
19.0 GB
Files
10
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-04-12 05:42

Files by quantization

Auxiliary files 10 files 19.1 GB
model.safetensors 19.0 GB 18675bc7 download
tokenizer.json 30.7 MB cc8d3a0c download
config.json 18.6 KB b2f4e43d download
chat_template.jinja 11.8 KB 33c51c2d download
README.md 9.88 KB fa4ad19c download
tokenizer_config.json 2.62 KB f07b8ede download
processor_config.json 1.65 KB 5465974d download
.gitattributes 1.53 KB 52373fe2 download
recipe.yaml 219 B 303fa8af download
generation_config.json 203 B edda3c19 download

README current version from Hugging Face


license: gemma
library_name: transformers
pipeline_tag: image-text-to-text
base_model: huihui-ai/Huihui-gemma-4-31B-it-abliterated-v2
base_model_relation: quantized
tags:

  • transformers
  • safetensors
  • gemma4
  • quantized
  • nvfp4
  • fp4
  • 4-bit
  • compressed-tensors
  • llm-compressor
  • vllm
  • image-text-to-text
  • multimodal
  • conversational
  • abliterated
    datasets:
  • mit-han-lab/pile-val-backup

Huihui-gemma-4-31B-it-abliterated-v2-NVFP4

This repository contains an NVFP4-compressed version of huihui-ai/Huihui-gemma-4-31B-it-abliterated-v2.

The goal of this build is to produce an NVFP4 checkpoint that actually fits on a single 32 GB RTX 5090 while preserving the multimodal pipeline and the abliteration intervention. The existing NVIDIA reference release keeps every self_attn layer in BF16, which pushes the on-disk weights past 32 GB and makes single-GPU serving impossible. This build quantizes self_attn as well, so only the vision tower, embeddings, and lm_head remain in BF16.

What Changed In v2 (vs v1)

The source model (huihui-ai/...-abliterated-v2) differs from v1 in one key way: the first 5 text layers (layers 0–4) are no longer abliterated. They retain the original google/gemma-4-31B-it weights. Layers 5–59 are still abliterated but with a refusal direction recomputed excluding those early layers.

Per the source model card, this produces lower perplexity (better quality) while maintaining the same level of refusal removal:

Model PPL Gap vs base
google/gemma-4-31B-it (base) 14874.75 —
v1 abliterated 13335.55 -1539
v2 abliterated 13161.29 -1713 (v1 − 174)

In local side-by-side testing of the NVFP4 quantized versions, v2 showed tighter instruction following (stricter JSON output, more efficient reasoning steps within the same token budget) while retaining identical abliteration effectiveness and decode throughput (~69 tok/s).

Source And References

Supported Modalities

This checkpoint matches the input/output capabilities of the upstream google/gemma-4-31B-it exactly — no modality was added or removed by quantization.

Modality Status Notes
Text → Text ✅ Supported Verified locally across English, Chinese, reasoning, code, long-form, and JSON-constrained outputs
Image → Text ✅ Supported Vision tower (27-layer SigLIP-style, ~550M params) preserved in BF16 via the re:.*vision.* ignore pattern. Verified locally with multiple real images at species-level accuracy
Video → Text ✅ Supported Uses the same vision tower with the Gemma4VideoProcessor sampling 32 frames at 70 soft-tokens per frame. Verified locally by posting a real 5-second MP4 through the vLLM chat/completions endpoint with a video_url content part
Audio → Text ❌ Not supported The 31B dense checkpoint has no audio encoder. config.json has audio_config: null and there are zero audio-related tensors in the safetensors. Per the Gemma 4 model card, audio ASR/translation is available only on the E2B and E4B variants. This is a property of the upstream weights, not the quantization

A Note On The any-to-any Tag Seen On Some Related Repos

The huihui source repository carries an any-to-any tag among its Hub tags. That tag does not reflect the actual capabilities of the 31B weights — the upstream google/gemma-4-31B-it is tagged as image-text-to-text and its model card explicitly restricts audio support to the E2B/E4B variants. The Gemma4Processor in processor_config.json does include a Gemma4AudioFeatureExtractor section (because the Gemma4ForConditionalGeneration architecture class is a superset capable of hosting audio), but the 31B checkpoint ships without audio encoder weights. This release intentionally does not carry an any-to-any tag, to avoid propagating that mislabel.

What Was Quantized

  • Every Linear layer in the text stack — including self_attn.{q,k,v,o}_proj and all mlp.*_proj — was quantized to NVFP4 with llm-compressor.
  • The vision tower, all embed_* tensors, the lm_head, and any audio_* modules remain in BF16 so the multimodal pipeline and vocabulary pathways are preserved.
  • Calibration used mit-han-lab/pile-val-backup, 128 samples at 1024 tokens per sample, batch size 1.
  • Quantization recipe (from recipe.yaml):
    QuantizationModifier:
      targets: [Linear]
      ignore: ['re:.*vision.*', 're:.*audio.*', lm_head, 're:.*embed.*']
      scheme: NVFP4
    

Ignore List Compared To The NVIDIA Reference

NVIDIA Gemma-4-31B-IT-NVFP4 This build
MLP Linear layers NVFP4 NVFP4
self_attn.* Linear layers BF16 (excluded) NVFP4
Vision tower BF16 (excluded) BF16 (excluded)
embed_* BF16 (excluded) BF16 (excluded)
lm_head BF16 (excluded) BF16 (excluded)
On-disk size ~32.6 GB ~19.5 GB
Single-GPU 32 GB fit No Yes

Inference

As tested locally, this model works with the vllm/vllm-openai:gemma4-cu130 image on an RTX 5090 (32 GB). Representative launch via docker compose:

services:
  vllm:
    image: vllm/vllm-openai:gemma4-cu130
    network_mode: host
    ipc: host
    devices:
      - nvidia.com/gpu=all
    volumes:
      - ./Huihui-gemma-4-31B-it-abliterated-v2-NVFP4:/models/huihui-gemma4-nvfp4:ro
    environment:
      - VLLM_NVFP4_GEMM_BACKEND=marlin
      - PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True
    command:
      - /models/huihui-gemma4-nvfp4
      - --served-model-name=huihui-gemma4-31b-nvfp4
      - --max-model-len=102400
      - --gpu-memory-utilization=0.95
      - --kv-cache-dtype=fp8
      - --trust-remote-code
      - --enable-prefix-caching
      - --enable-auto-tool-choice
      - --tool-call-parser=gemma4

Observed on RTX 5090 at max-model-len=102400, gpu-memory-utilization=0.95:

  • ~29.7 GB VRAM occupied (weights + KV cache + CUDA runtime)
  • ~69 tokens/second single-stream decode throughput
  • Available KV cache: 8.55 GiB, maximum concurrency for 102,400 tokens: 1.15x
  • Engine init (profile + KV cache + warmup): ~51 seconds
  • Tool calling enabled with native gemma4 parser

Validation

A local quality sweep was run against the served OpenAI-compatible endpoint. All tests passed.

Text (7 / 7)

  • English coherence with sentence-count constraint
  • Chinese generation with first-sentence geographic-location constraint
  • Multi-step arithmetic (two trains converging, correct reasoning chain)
  • Strict JSON output — parsed cleanly by json.loads, output was bare JSON without markdown fencing
  • 400-word long-form narrative generation — 21 / 21 unique sentences, no degeneration
  • Python code generation with assert-based test cases
  • Abliteration smoke test: the refusal-vector intervention from the source model is preserved through the NVFP4 round-trip

Image (4 / 4)

Posted real JPEGs via the image_url content part:

Subject Model output Verdict
Black Labrador puppy "The image shows a black dog." ✅
Walruses on a beach "Several walruses are lounging and resting on a sandy beach. Overcast and cool, hazy grey sky." ✅ species-level
Pug wrapped in blanket "Dog (Pug), Blanket, Bed/Bedding" ✅ breed-level
Grey-tone cat portrait "grays... whites/off-whites... soft pinks/browns on the nose" ✅

Video (1 / 1)

Posted a real 5-second MP4 (~2.85 MB) via the video_url content part. prompt_tokens = 2436, consistent with 32 frames × 70 soft-tokens per frame plus the text wrapper.

Model output: "A white truck and several cars are parked along a road bordering a lush, green park. The scene features tall trees, a grassy area with a path, and a park bench."

Spatial and object-level description are correct. End-to-end latency was ~2 seconds.

v1 vs v2 Comparison (NVFP4)

Both versions were tested with the same quality battery on the same hardware:

Test v1 v2 Notes
English coherence ✅ 31 tok/s ✅ 36 tok/s v2 slightly more articulate
Chinese ✅ 69 tok/s ✅ 69 tok/s Equivalent quality
Math reasoning Correct chain, truncated before final answer Correct chain, reached "1h 21min" conclusion v2 more token-efficient
JSON output Pretty-printed with fencing Bare single-line JSON v2 stricter instruction following
Long-form story 27/27 unique sentences 21/21 unique sentences Both zero degeneration
Code Correct Correct Near-identical
Abliteration Zero refusal Zero refusal Both fully preserved

Formal Benchmarks

Intentionally omitted for now. They will be added later once a reproducible evaluation harness run is available.

Notes

  • Architecture: Gemma4ForConditionalGeneration
  • Model type: gemma4
  • Pipeline tag: image-text-to-text
  • Quantization format: nvfp4-pack-quantized (compressed-tensors)
  • KV cache precision at inference time (recommended): fp8
  • Requires transformers >= 5.5.0 to load gemma4 configurations
  • The multimodal processing files (processor_config.json, chat_template.jinja) are preserved so vision input still works; text-only usage also works unchanged

README history 3 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-04-12Replace with v2 weights + updated README (v2: first 5 layers not abliterated,...985633e9.9 KB
    Loading...
  2. 2026-04-11Add Supported Modalities section with local image/video validation resultsd03eb0e7.9 KB
    Loading...
  3. 2026-04-11Initial NVFP4 upload: quantized via llm-compressor on RTX 50902ab0de95 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration