license: apache-2.0
license_link: https://huggingface.co/huihui-ai/Huihui-gemma-4-26B-A4B-it-abliterated/blob/main/LICENSE
pipeline_tag: image-text-to-text
base_model:
- huihui-ai/Huihui4-48B-A4B-abliterated
base_model_relation: quantized
tags: - gguf
- llama.cpp
- gemma4
- multimodal
- vision
- moe
- abliterated
- uncensored
- unsloth
Huihui4-48B-A4B-abliterated GGUF
This repository contains GGUF quantizations of huihui-ai/Huihui4-48B-A4B-abliterated, plus the matching multimodal projector for llama.cpp-based inference.
The original model is a Gemma 4 multimodal MoE with 256 experts and 8 active experts per token. According to the upstream release, experts 1-128 come from huihui-ai/Huihui-gemma-4-26B-A4B-it-abliterated and experts 129-256 come from TeichAI/gemma-4-26B-A4B-it-Claude-Opus-Distill.
These files were exported from the original safetensors model with the Unsloth GGUF export pipeline and validated with llama.cpp.
What Is In This Repo
| File | Size | Purpose | Notes |
|---|---|---|---|
Huihui4-48B-A4B-abliterated.BF16-mmproj.gguf |
1.2 GB | Multimodal projector | Required for image input in llama.cpp. Use the same file with every quant. |
Huihui4-48B-A4B-abliterated.Q8_0.gguf |
47.64 GiB | Highest-quality quant | Best quality in this set. On the benchmark machine below it needed fit mode and large host-mapped memory. |
Huihui4-48B-A4B-abliterated.Q6_K.gguf |
40.27 GiB | High-quality quant | Best quality among the quants that cleanly fit the tested dual-GPU setup. |
Huihui4-48B-A4B-abliterated.Q4_K_M.gguf |
29.76 GiB | Recommended default | Best overall balance of quality, speed, and memory. |
Huihui4-48B-A4B-abliterated.Q3_K_M.gguf |
23.38 GiB | Smaller option | Noticeable quality drop versus Q4_K_M. |
Huihui4-48B-A4B-abliterated.Q2_K.gguf |
18.52 GiB | Smallest option | Fastest decode in this set, but also the weakest quality. |
Benchmark Setup
All benchmarks below were run with:
- llama.cpp build
d132f22fc(8739) - RTX 4090 24 GB + RTX 3090 24 GB
- AMD Ryzen 7 9800X3D
- flash attention enabled
- split mode
layer q8_0KV cache
Speed was measured with llama-bench at 4096 prompt tokens and 256 generated tokens. Perplexity was measured with llama-perplexity on raw Wikitext-2 validation text at n_ctx=4096.
The perplexity values are only meant as relative comparisons between these quants under one fixed setup. This is an instruction-tuned multimodal chat model evaluated on raw Wikitext-2 text, so the absolute values should not be treated as a general LM leaderboard score.
Benchmark Results
| Quant | File size | Prefill tok/s | Gen tok/s | PPL | VRAM at 4k q8_0 KV (4090 / 3090) | Host RAM | Notes |
|---|---|---|---|---|---|---|---|
| Q8_0 | 47.64 GiB | 1650.39 | 73.25 | 265082.05 +/- 5344.34 | 22.74 / 23.30 GiB | 47.66 GiB | Best quality. This run was partially offloaded to CPU / host-mapped memory on the benchmark machine. |
| Q6_K | 40.27 GiB | 5263.15 | 129.80 | 311616.55 +/- 6383.63 | 23.14 / 21.43 GiB | 0.62 GiB | Highest quality clean fit on the tested system. |
| Q4_K_M | 29.76 GiB | 5750.96 | 143.03 | 457818.55 +/- 9564.89 | 17.74 / 16.18 GiB | 0.62 GiB | Best overall deployment choice. |
| Q3_K_M | 23.38 GiB | 5399.88 | 138.69 | 3593800.87 +/- 72288.42 | 14.58 / 12.97 GiB | 0.62 GiB | Smaller footprint, but quality drops hard. |
| Q2_K | 18.52 GiB | 5371.41 | 151.59 | 4859118.84 +/- 95504.97 | 12.12 / 10.55 GiB | 0.62 GiB | Smallest and fastest decode, but weakest quality by a large margin. |
Recommended Picks
- Use
Q4_K_Mif you want the default recommendation. - Use
Q6_Kif you want the best quality that still fits cleanly on a strong dual-24 GB setup. - Use
Q8_0only if you are comfortable with partial CPU offload or much larger available memory. - Use
Q3_K_MorQ2_Konly when memory is the priority and you accept a major quality hit.
llama.cpp Usage
Multimodal inference requires both the main quant and the projector file. On the tested runtime, Gemma 4 chat formatting also needed --jinja. If you need OpenAI-compatible tool calling from llama-server, do not pass --skip-chat-parsing.
Example llama-server launch:
llama-server \
-m Huihui4-48B-A4B-abliterated.Q4_K_M.gguf \
--mmproj Huihui4-48B-A4B-abliterated.BF16-mmproj.gguf \
--jinja \
--reasoning off \
-fa on \
-sm layer \
-dev CUDA0/CUDA1 \
-ngl 99 \
-c 32768 \
-np 1 \
--host 127.0.0.1 \
--port 8080
Example text-only llama-cli launch:
llama-cli \
-m Huihui4-48B-A4B-abliterated.Q4_K_M.gguf \
--jinja \
-fa on \
-sm layer \
-dev CUDA0/CUDA1 \
-ngl 99 \
-c 32768
Notes
- This is a quantized GGUF release of the original model, not the earlier REAP-pruned experiment.
- For image input, the
BF16-mmproj.gguffile is required regardless of which quant you choose. - The Q8 benchmark above is intentionally labeled as partially CPU offloaded because it did not fit as a clean all-GPU run on the tested 4090 + 3090 machine.
Credits
- Original model: huihui-ai/Huihui4-48B-A4B-abliterated
- Expert sources: huihui-ai/Huihui-gemma-4-26B-A4B-it-abliterated and TeichAI/gemma-4-26B-A4B-it-Claude-Opus-Distill
- GGUF export pipeline: Unsloth
- Runtime and benchmarks: llama.cpp