license: gemma
library_name: transformers
base_model: huihui-ai/Huihui4-48B-A4B-abliterated
tags:
- gemma4
- nvfp4
- quantized
- compressed-tensors
- vlm
- vision-language-model
- abliterated
- moe
- blackwell
language: - en
- ja
- zh
- ko
- de
- fr
- es
- multilingual
pipeline_tag: image-text-to-text
model-index: - name: Huihui4-48B-A4B-abliterated-NVFP4
results:- task:
type: text-generation
dataset:
name: Custom Benchmark (Lna-Lab)
type: custom
metrics:- name: Throughput (code gen)
type: custom
value: 150.0
unit: tok/s - name: Throughput (math reasoning)
type: custom
value: 151.7
unit: tok/s - name: Throughput (VLM)
type: custom
value: 145.8
unit: tok/s
- name: Throughput (code gen)
- task:
Huihui4-48B-A4B-abliterated-NVFP4
NVFP4-quantized version of huihui-ai/Huihui4-48B-A4B-abliterated — a Gemma 4 48B A4B (MoE, 256 experts / top-8) vision-language model with abliteration applied.
Quantized to NVIDIA FP4 by Lna-Lab using our custom Blackwell NVFP4 GEMM kernels (lna-lab/blackwell-geforce-nvfp4-gemm) for efficient single-GPU inference on Blackwell (RTX PRO 6000 / B200 / GB200) and Ada/Hopper GPUs with FP4 tensor core support.
Key Features
- Single-GPU deployment — 28 GB on disk, fits within one 96 GB Blackwell GPU
- Full VLM retained — Vision tower kept in BF16 (not quantized), image understanding works out of the box
- Abliterated — Safety over-refusal removed for unrestricted research use
- 125–152 tok/s on RTX PRO 6000 Blackwell depending on task
Model Architecture
| Parameter | Value |
|---|---|
| Architecture | Gemma4ForConditionalGeneration (VLM) |
| Total Parameters | 48B (4B active via MoE) |
| Experts | 256 experts, top-8 routing |
| Layers | 30 (25 sliding attention + 5 full attention) |
| Hidden Size | 2,816 |
| Attention Heads | 16 (8 KV heads, GQA) |
| Head Dim | 256 (global: 512) |
| Sliding Window | 1,024 tokens |
| Max Position | 262,144 tokens |
| Vocab Size | 262,144 |
| MoE Intermediate | 704 per expert |
| Dense Intermediate | 2,112 |
| Vision Encoder | 27-layer, hidden=1,152, patch=16, SigLIP-style |
| Quantization | NVFP4 (compressed-tensors, nvfp4-pack-quantized) |
| Weight Format | 4-bit float, group_size=16, scale=fp8_e4m3fn |
| Size on Disk | ~27.3 GB |
What's Quantized / What's Not
| Component | Precision |
|---|---|
| Text MoE layers (Linear) | NVFP4 |
| Router projections | BF16 (excluded) |
| Vision tower (all layers) | BF16 (excluded) |
| Embedding projection | BF16 (excluded) |
| LM head | BF16 (excluded) |
Quantization Details
Quantized by Lna-Lab using llm-compressor with custom Blackwell NVFP4 GEMM kernels (lna-lab/blackwell-geforce-nvfp4-gemm):
default_stage:
default_modifiers:
QuantizationModifier:
targets: [Linear]
ignore: [lm_head, 're:.*embed.*', 're:.*router', 're:.*vision_tower.*']
scheme: NVFP4
bypass_divisibility_checks: false
- Weights: 4-bit float, group_size=16, symmetric, static minmax observer
- Activations: 4-bit float, group_size=16, symmetric, dynamic local quantization
- Scales:
fp8_e4m3fn
Benchmark Results
Tested on a single NVIDIA RTX PRO 6000 Blackwell (96 GB), vLLM 0.19.1+, max_model_len=8192, temperature=0.0.
| Task | Output Tokens | Time (s) | Throughput (tok/s) |
|---|---|---|---|
| Japanese essay (方丈記 analysis, 800+ chars) | 757 | 6.08 | 124.5 |
| Python code generation (LRU cache w/ TTL) | 3,000 | 20.0 | 150.0 |
| Math reasoning (calculus + AM-GM) | 1,217 | 8.02 | 151.7 |
| VLM image description | 343 | 2.35 | 145.8 |
VRAM Usage
| State | GPU Memory |
|---|---|
| After model load | 89,726 MiB |
| Peak (during inference) | 89,730 MiB |
Quality Assessment
- Japanese: Coherent, accurate modern translation of classical text with cultural analysis. No repetition artifacts.
- Code generation: Complete, well-structured Python with type hints, docstrings, and tests.
- Math reasoning: Correct calculus derivation, verified with AM-GM inequality, includes semicircle comparison.
- VLM: Correctly identifies geometric shapes, colors, text, and solves embedded math problems (7×8=56).
How to Use
Requirements
- vLLM >= 0.19 with
compressed-tensorssupport - Transformers >= 5.5
- PyTorch >= 2.11 with CUDA 13.0+
- GPU with FP4 tensor core support (Blackwell, Ada Lovelace, Hopper)
vLLM (Recommended)
# Text-only
vllm serve /path/to/Huihui4-48B-A4B-abliterated-NVFP4 \
--max-model-len 8192 \
--gpu-memory-utilization 0.92 \
--dtype auto \
--trust-remote-code
# With VLM (image input)
vllm serve /path/to/Huihui4-48B-A4B-abliterated-NVFP4 \
--max-model-len 8192 \
--gpu-memory-utilization 0.92 \
--dtype auto \
--limit-mm-per-prompt '{"image":1}' \
--trust-remote-code
Docker (Lna-Lab image)
docker run -d --name huihui4-48b \
--gpus '"device=0"' --shm-size=16g \
-v /models/Huihui4-48B-A4B-abliterated-NVFP4:/models/current:ro \
-p 8000:8000 \
vllm/vllm-openai:cu130-nightly \
--model /models/current \
--trust-remote-code --quantization modelopt --language-model-only \
--reasoning-parser qwen3 \
--enable-auto-tool-choice --tool-call-parser qwen3_xml \
--default-chat-template-kwargs '{"preserve_thinking":true}' \
--enable-prefix-caching --enable-chunked-prefill \
--max-model-len 131072 --gpu-memory-utilization 0.95 \
--kv-cache-dtype fp8_e4m3
API Usage
from openai import OpenAI
client = OpenAI(base_url="http://localhost:8000/v1", api_key="dummy")
# Text
response = client.chat.completions.create(
model="Huihui4-48B-A4B-abliterated-NVFP4",
messages=[{"role": "user", "content": "Write a haiku about quantization."}],
max_tokens=256,
)
print(response.choices[0].message.content)
# VLM (image input)
import base64
from pathlib import Path
img_b64 = base64.b64encode(Path("photo.jpg").read_bytes()).decode()
response = client.chat.completions.create(
model="Huihui4-48B-A4B-abliterated-NVFP4",
messages=[{
"role": "user",
"content": [
{"type": "image_url", "image_url": {"url": f"data:image/jpeg;base64,{img_b64}"}},
{"type": "text", "text": "What do you see in this image?"},
],
}],
max_tokens=1024,
)
print(response.choices[0].message.content)
Tested Environment
| Component | Version |
|---|---|
| vLLM | 0.19.1rc1+ (nightly) |
| Transformers | 5.5.4 |
| PyTorch | 2.11.0+cu130 |
| CUDA | 13.0 |
| GPU | NVIDIA RTX PRO 6000 Blackwell (96 GB) |
| OS | Ubuntu 24.04, Linux 6.17 |
Credits & Donations
- Base model: Google Gemma 4
- Abliteration: huihui-ai — Please consider supporting the original abliterated model author: huihui-ai/Huihui4-48B-A4B-abliterated
- NVFP4 quantization & benchmarking: Lna-Lab
- Blackwell NVFP4 GEMM kernels: lna-lab/blackwell-geforce-nvfp4-gemm
- Quantization framework: llm-compressor by vLLM Project
License
This model inherits the Gemma license.