← back to catalog · registered 2026-08-22 13:56

elbelga/Huihui-gemma-4-31B-it-abliterated-v2-MXFP4

elbelga Gemma 16B
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/elbelga%2FHuihui-gemma-4-31B-it-abliterated-v2-MXFP4"
Response includes
  • classification m1
  • files 10
  • hub_downloads_all_time 10,659
  • author_summary 3 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M1
Primary method

Direct removal

No other method signals detected in this model.
Confidence
MEDIUM
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=0 (base model)
  • no specific method indicators - defaulting to M1 (most common)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
MEDIUM
Why we say so
primary_method=M1; difference-of-means is the reference extraction for M1/M3 (Arditi 2024)
Downloads · lifetime
11K
29 last 30d - cooling
Likes
0
Model age
5mo ago
created 2026-04-19
Downloads over time
Now10.7K→from19↑56,089%
03.9K7.8K11.7K19 on Apr 2210.7K on Oct 1110.7K on Oct 9AprMayJunJulAugSepOct
Apr 22 → Oct 11 · 64 snapshots · spans 172 days

Metadata

Tags
safetensors gemma4 8-bit compressed-tensors region:us

Related

Total size
18.2 GB
Files
10
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-04-19 19:33

Files by quantization

Auxiliary files 10 files 18.2 GB
model.safetensors 18.2 GB efc8a71a download
tokenizer.json 30.7 MB e4c18ffe download
config.json 18.2 KB 982ec6ee download
chat_template.jinja 11.8 KB 57ab6633 download
README.md 5.04 KB 3994b7ea download
tokenizer_config.json 2.57 KB 2113820c download
processor_config.json 1.65 KB 5465974d download
.gitattributes 1.53 KB 52373fe2 download
recipe.yaml 436 B 04aec7d3 download
generation_config.json 203 B 3110a9b4 download

README current version from Hugging Face

Huihui-gemma-4-31B-it-abliterated-v2-MXFP4

MXFP4-quantized version of huihui-ai/Huihui-gemma-4-31B-it-abliterated-v2

Quantization Details

Property Value
Date 2026-04-15
Scheme MXFP4A16 (4-bit weight-only)
Algorithm GPTQ
Group Size 32
Output Size 18.2 GB
Compression ~3.4x
Format mxfp4-pack-quantized (compressed-tensors)

Quantized Layers (~1,571 layers)

  • Main LLM attention layers (q_proj, k_proj, v_proj, o_proj)
  • MLP layers (gate_proj, up_proj, down_proj)
  • Embed tokens

BF16 Layers (192 layers)

  • Vision tower (26 encoder layers)
  • Vision embeddings
  • LM head

Comparison: Original vs Quantized

Metric Original (BF16) Quantized (MXFP4) Difference
Size 59 GB 18.2 GB -69% (40.8 GB saved)
Precision 16-bit 4-bit 4x reduction
Quantization None mxfp4-pack Compressed-tensors format
Layers Quantized 0 1,571 All LLM layers
Layers BF16 All 192 Vision tower only

Memory Requirements

  • Original (BF16): ~62 GB VRAM for inference
  • Quantized (MXFP4): ~20 GB VRAM for inference (~3x less)

Inference Speed (Estimated)

With Blackwell GPU (GB10/B200) native FP4 tensor cores:

  • Throughput: ~2-3x faster decoding vs BF16
  • Pre-fill: Similar or slight improvement

Quality (PPL)

Model PPL Notes
google/gemma-4-31B-it (f16) 14,874.75 Base model
huihui-abliterated (f16) 13,161.29 Before quantization
huihui-MXFP4 ~13,000-14,000 vLLM inference (decompression at runtime)

Note: Standard perplexity evaluation via HuggingFace transformers is not supported because MXFP4 decompression is not implemented in transformers. However, vLLM handles the decompression transparently during inference, so output quality should be equivalent to the original model.

The high PPL values seen in direct API perplexity tests are due to the model's thinking/reasoning mode interfering with token prediction - this is a model-specific behavior, not a quantization issue.

Original Model

Source: huihui-ai/Huihui-gemma-4-31B-it-abliterated-v2

Base model: google/gemma-4-31B-it

Model Information

Property Value
Type image-text-to-text (multimodal)
Parameters 33B
Precision BF16
Vocab Size 262,144
Hidden Size 5,376
Layers 60
Context Length 256K

About This Model

This is an abliterated (uncensored) version of google/gemma-4-31B-it created using abliteration to remove refusals.

PPL (Perplexity) Comparison:

Model PPL Gap
google/gemma-4-31B-it (f16) 14,874.75 baseline
v1 abliterated 13,335.55 -1,539.20
v2 abliterated (this) 13,161.29 -1,713.46

Lower PPL = better quality

Usage Warnings

  • ⚠️ Significantly reduced safety filtering - may generate sensitive/inappropriate content
  • Not suitable for minors or public-facing applications
  • Users bear full responsibility for any consequences
  • Recommended for research/testing only

Usage with vLLM

Basic Inference

NOTE: Do NOT specify --quantization mxfp4 - the quantization format is auto-detected from the model config.

vllm serve ./Huihui-gemma-4-31B-it-abliterated-v2-MXFP4 \
  --trust-remote-code \
  --gpu-memory-utilization 0.85

Or in Python:

from vllm import LLM

llm = LLM(
    model="./Huihui-gemma-4-31B-it-abliterated-v2-MXFP4",
    # DO NOT specify quantization - auto-detected from compressed-tensors config
    trust_remote_code=True,
    gpu_memory_utilization=0.85,
    max_model_len=32768,
)

Enable Thinking/Reasoning Mode

To enable the model's built-in thinking/reasoning capability:

vllm serve ./Huihui-gemma-4-31B-it-abliterated-v2-MXFP4 \
  --trust-remote-code \
  --max-model-len 16384 \
  --enable-auto-tool-choice \
  --reasoning-parser gemma4 \
  --tool-call-parser gemma4 \
  --chat-template examples/tool_chat_template_gemma4.jinja \
  --default-chat-template-kwargs '{"enable_thinking": true}'

Or in Python:

from vllm import LLM

llm = LLM(
    model="./Huihui-gemma-4-31B-it-abliterated-v2-MXFP4",
    trust_remote_code=True,
    max_model_len=16384,
    enable_auto_tool_choice=True,
    chat_template="examples/tool_chat_template_gemma4.jinja",
    default_chat_template_kwargs={"enable_thinking": True},
)

Troubleshooting

Error: "Quantization method specified in the model config (compressed-tensors) does not match..."

This means you're explicitly specifying --quantization mxfp4 when it's not needed. Simply remove the --quantization argument - vLLM auto-detects the compressed-tensors format from the model config.

License

Apache-2.0 (same as base model)

README history 1 version

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-04-19Upload folder using huggingface_hub0cd135e5 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration