← back to catalog · registered 2026-08-22 13:56

sakamakismile/Huihui4-48B-A4B-abliterated-NVFP4

sakamakismile Gemma 24B MoE multimodal second-order
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/sakamakismile%2FHuihui4-48B-A4B-abliterated-NVFP4"
Response includes
  • classification m1
  • files 11
  • hub_downloads_all_time 701
  • author_summary 34 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M1
Primary method

Direct removal

No other method signals detected in this model.
Confidence
MEDIUM
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=0 (base model)
  • no specific method indicators - defaulting to M1 (most common)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
MEDIUM
Why we say so
primary_method=M1; difference-of-means is the reference extraction for M1/M3 (Arditi 2024)
Downloads · lifetime
701
33 last 30d - cooling
Likes
0
Model age
5mo ago
created 2026-04-22
Downloads over time
Now706→from15↑4,607%
025851777515 on Apr 22706 on Oct 11706 on Oct 9AprMayJunJulAugSepOct
Apr 22 → Oct 11 · 64 snapshots · spans 172 days

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
gemma
Languages
en ja zh ko de fr es multilingual
Tags
transformers safetensors gemma4 image-text-to-text nvfp4 quantized compressed-tensors vlm vision-language-model abliterated moe blackwell

Related

Total size
27.3 GB
Files
11
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-04-28 00:01

Files by quantization

Auxiliary files 11 files 27.3 GB
model.safetensors 27.3 GB 5d9a2ec4 download
tokenizer.json 30.7 MB d93b1947 download
config.json 19.2 KB 4d251d7d download
chat_template.jinja 11.8 KB 33c51c2d download
README.md 7.73 KB 4fce29b2 download
INFERENCE_GUIDE.md 3.74 KB 0cae82c7 download
tokenizer_config.json 2.62 KB 59dd4b62 download
processor_config.json 1.65 KB 5465974d download
.gitattributes 1.53 KB 52373fe2 download
recipe.yaml 224 B f790e78b download
generation_config.json 203 B 92b5abfd download

README current version from Hugging Face


license: gemma
library_name: transformers
base_model: huihui-ai/Huihui4-48B-A4B-abliterated
tags:

  • gemma4
  • nvfp4
  • quantized
  • compressed-tensors
  • vlm
  • vision-language-model
  • abliterated
  • moe
  • blackwell
    language:
  • en
  • ja
  • zh
  • ko
  • de
  • fr
  • es
  • multilingual
    pipeline_tag: image-text-to-text
    model-index:
  • name: Huihui4-48B-A4B-abliterated-NVFP4
    results:
    • task:
      type: text-generation
      dataset:
      name: Custom Benchmark (Lna-Lab)
      type: custom
      metrics:
      • name: Throughput (code gen)
        type: custom
        value: 150.0
        unit: tok/s
      • name: Throughput (math reasoning)
        type: custom
        value: 151.7
        unit: tok/s
      • name: Throughput (VLM)
        type: custom
        value: 145.8
        unit: tok/s

Huihui4-48B-A4B-abliterated-NVFP4

NVFP4-quantized version of huihui-ai/Huihui4-48B-A4B-abliterated — a Gemma 4 48B A4B (MoE, 256 experts / top-8) vision-language model with abliteration applied.

Quantized to NVIDIA FP4 by Lna-Lab using our custom Blackwell NVFP4 GEMM kernels (lna-lab/blackwell-geforce-nvfp4-gemm) for efficient single-GPU inference on Blackwell (RTX PRO 6000 / B200 / GB200) and Ada/Hopper GPUs with FP4 tensor core support.

Key Features

  • Single-GPU deployment — 28 GB on disk, fits within one 96 GB Blackwell GPU
  • Full VLM retained — Vision tower kept in BF16 (not quantized), image understanding works out of the box
  • Abliterated — Safety over-refusal removed for unrestricted research use
  • 125–152 tok/s on RTX PRO 6000 Blackwell depending on task

Model Architecture

Parameter Value
Architecture Gemma4ForConditionalGeneration (VLM)
Total Parameters 48B (4B active via MoE)
Experts 256 experts, top-8 routing
Layers 30 (25 sliding attention + 5 full attention)
Hidden Size 2,816
Attention Heads 16 (8 KV heads, GQA)
Head Dim 256 (global: 512)
Sliding Window 1,024 tokens
Max Position 262,144 tokens
Vocab Size 262,144
MoE Intermediate 704 per expert
Dense Intermediate 2,112
Vision Encoder 27-layer, hidden=1,152, patch=16, SigLIP-style
Quantization NVFP4 (compressed-tensors, nvfp4-pack-quantized)
Weight Format 4-bit float, group_size=16, scale=fp8_e4m3fn
Size on Disk ~27.3 GB

What's Quantized / What's Not

Component Precision
Text MoE layers (Linear) NVFP4
Router projections BF16 (excluded)
Vision tower (all layers) BF16 (excluded)
Embedding projection BF16 (excluded)
LM head BF16 (excluded)

Quantization Details

Quantized by Lna-Lab using llm-compressor with custom Blackwell NVFP4 GEMM kernels (lna-lab/blackwell-geforce-nvfp4-gemm):

default_stage:
  default_modifiers:
    QuantizationModifier:
      targets: [Linear]
      ignore: [lm_head, 're:.*embed.*', 're:.*router', 're:.*vision_tower.*']
      scheme: NVFP4
      bypass_divisibility_checks: false
  • Weights: 4-bit float, group_size=16, symmetric, static minmax observer
  • Activations: 4-bit float, group_size=16, symmetric, dynamic local quantization
  • Scales: fp8_e4m3fn

Benchmark Results

Tested on a single NVIDIA RTX PRO 6000 Blackwell (96 GB), vLLM 0.19.1+, max_model_len=8192, temperature=0.0.

Task Output Tokens Time (s) Throughput (tok/s)
Japanese essay (方丈記 analysis, 800+ chars) 757 6.08 124.5
Python code generation (LRU cache w/ TTL) 3,000 20.0 150.0
Math reasoning (calculus + AM-GM) 1,217 8.02 151.7
VLM image description 343 2.35 145.8

VRAM Usage

State GPU Memory
After model load 89,726 MiB
Peak (during inference) 89,730 MiB

Quality Assessment

  • Japanese: Coherent, accurate modern translation of classical text with cultural analysis. No repetition artifacts.
  • Code generation: Complete, well-structured Python with type hints, docstrings, and tests.
  • Math reasoning: Correct calculus derivation, verified with AM-GM inequality, includes semicircle comparison.
  • VLM: Correctly identifies geometric shapes, colors, text, and solves embedded math problems (7×8=56).

How to Use

Requirements

  • vLLM >= 0.19 with compressed-tensors support
  • Transformers >= 5.5
  • PyTorch >= 2.11 with CUDA 13.0+
  • GPU with FP4 tensor core support (Blackwell, Ada Lovelace, Hopper)

vLLM (Recommended)

# Text-only
vllm serve /path/to/Huihui4-48B-A4B-abliterated-NVFP4 \
  --max-model-len 8192 \
  --gpu-memory-utilization 0.92 \
  --dtype auto \
  --trust-remote-code

# With VLM (image input)
vllm serve /path/to/Huihui4-48B-A4B-abliterated-NVFP4 \
  --max-model-len 8192 \
  --gpu-memory-utilization 0.92 \
  --dtype auto \
  --limit-mm-per-prompt '{"image":1}' \
  --trust-remote-code

Docker (Lna-Lab image)

docker run -d --name huihui4-48b \
  --gpus '"device=0"' --shm-size=16g \
  -v /models/Huihui4-48B-A4B-abliterated-NVFP4:/models/current:ro \
  -p 8000:8000 \
  vllm/vllm-openai:cu130-nightly \
  --model /models/current \
  --trust-remote-code --quantization modelopt --language-model-only \
  --reasoning-parser qwen3 \
  --enable-auto-tool-choice --tool-call-parser qwen3_xml \
  --default-chat-template-kwargs '{"preserve_thinking":true}' \
  --enable-prefix-caching --enable-chunked-prefill \
  --max-model-len 131072 --gpu-memory-utilization 0.95 \
  --kv-cache-dtype fp8_e4m3

API Usage

from openai import OpenAI

client = OpenAI(base_url="http://localhost:8000/v1", api_key="dummy")

# Text
response = client.chat.completions.create(
    model="Huihui4-48B-A4B-abliterated-NVFP4",
    messages=[{"role": "user", "content": "Write a haiku about quantization."}],
    max_tokens=256,
)
print(response.choices[0].message.content)
# VLM (image input)
import base64
from pathlib import Path

img_b64 = base64.b64encode(Path("photo.jpg").read_bytes()).decode()

response = client.chat.completions.create(
    model="Huihui4-48B-A4B-abliterated-NVFP4",
    messages=[{
        "role": "user",
        "content": [
            {"type": "image_url", "image_url": {"url": f"data:image/jpeg;base64,{img_b64}"}},
            {"type": "text", "text": "What do you see in this image?"},
        ],
    }],
    max_tokens=1024,
)
print(response.choices[0].message.content)

Tested Environment

Component Version
vLLM 0.19.1rc1+ (nightly)
Transformers 5.5.4
PyTorch 2.11.0+cu130
CUDA 13.0
GPU NVIDIA RTX PRO 6000 Blackwell (96 GB)
OS Ubuntu 24.04, Linux 6.17

Credits & Donations

License

This model inherits the Gemma license.

README history 2 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-04-28fix: replace stale lna-lab/gemma4-inference Docker reference with upstream vl...8b4dcf87.7 KB
    Loading...
  2. 2026-04-22Add files using upload-large-folder tool75652b27.6 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration