← back to catalog · registered 2026-08-22 13:56

cesarsal1nas/Huihui4-48B-A4B-abliterated-GGUF

cesarsal1nas Gemma 48B GGUF MoE multimodal second-order 262K ctx
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/cesarsal1nas%2FHuihui4-48B-A4B-abliterated-GGUF"
Response includes
  • classification m8
  • files 8
  • hub_downloads_all_time 3,131
  • author_summary 4 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M8
Primary method

Repackaging (quantization)

Applied on top of direct removal inherited from the base model.
Confidence
MEDIUM
Inherited from base model
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=1
  • assume M1 (base ablation) + M8 (GGUF quant) - default when producer unknown
Refusal direction extraction

No specific extraction method could be identified for this model. The producer either did not document it or used a proprietary pipeline.

What is a refusal direction? →
Downloads · lifetime
3K
351 last 30d - stable
Likes
2
Model age
6mo ago
created 2026-04-13
Downloads over time
Now3.3K→from747↑337%
6211.6K2.6K3.5K747 on Apr 153.3K on Oct 11AprMayJunJulAugSepOct
Apr 15 → Oct 11 · 65 snapshots · spans 179 days

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
apache-2.0
Quantizations
Q2_K Q3_K Q4_K Q6_K Q8_0
Tags
gguf llama.cpp gemma4 multimodal vision moe abliterated uncensored unsloth image-text-to-text base_model:huihui-ai/Huihui4-48B-A4B-abliterated base_model:quantized:huihui-ai/Huihui4-48B-A4B-abliterated

Related

Total size
160 GB
Files
8
Quantizations
7
Registered
2026-08-22 13:56
Last updated on HF
2026-04-13 17:23

Files by quantization

Q8_0 1 file 47.7 GB
Huihui4-48B-A4B-abliterated.Q8_0.gguf 47.7 GB 786b5992 download
Q6_K 1 file 40.3 GB
Huihui4-48B-A4B-abliterated.Q6_K.gguf 40.3 GB 1e71b12c download
Q4_K 1 file 29.8 GB
Huihui4-48B-A4B-abliterated.Q4_K_M.gguf 29.8 GB ac7c221e download
Q3_K 1 file 23.4 GB
Huihui4-48B-A4B-abliterated.Q3_K_M.gguf 23.4 GB d8b7fe14 download
Q2_K 1 file 18.5 GB
Huihui4-48B-A4B-abliterated.Q2_K.gguf 18.5 GB 031a8f65 download
mmproj 1 file 1.11 GB
Huihui4-48B-A4B-abliterated.BF16-mmproj.gguf 1.11 GB bf8d14f9 download
Auxiliary files 2 files 7.54 KB
README.md 5.62 KB c7b576bf download
.gitattributes 1.93 KB f918e59d download

README current version from Hugging Face


license: apache-2.0
license_link: https://huggingface.co/huihui-ai/Huihui-gemma-4-26B-A4B-it-abliterated/blob/main/LICENSE
pipeline_tag: image-text-to-text
base_model:

  • huihui-ai/Huihui4-48B-A4B-abliterated
    base_model_relation: quantized
    tags:
  • gguf
  • llama.cpp
  • gemma4
  • multimodal
  • vision
  • moe
  • abliterated
  • uncensored
  • unsloth

Huihui4-48B-A4B-abliterated GGUF

This repository contains GGUF quantizations of huihui-ai/Huihui4-48B-A4B-abliterated, plus the matching multimodal projector for llama.cpp-based inference.

The original model is a Gemma 4 multimodal MoE with 256 experts and 8 active experts per token. According to the upstream release, experts 1-128 come from huihui-ai/Huihui-gemma-4-26B-A4B-it-abliterated and experts 129-256 come from TeichAI/gemma-4-26B-A4B-it-Claude-Opus-Distill.

These files were exported from the original safetensors model with the Unsloth GGUF export pipeline and validated with llama.cpp.

What Is In This Repo

File Size Purpose Notes
Huihui4-48B-A4B-abliterated.BF16-mmproj.gguf 1.2 GB Multimodal projector Required for image input in llama.cpp. Use the same file with every quant.
Huihui4-48B-A4B-abliterated.Q8_0.gguf 47.64 GiB Highest-quality quant Best quality in this set. On the benchmark machine below it needed fit mode and large host-mapped memory.
Huihui4-48B-A4B-abliterated.Q6_K.gguf 40.27 GiB High-quality quant Best quality among the quants that cleanly fit the tested dual-GPU setup.
Huihui4-48B-A4B-abliterated.Q4_K_M.gguf 29.76 GiB Recommended default Best overall balance of quality, speed, and memory.
Huihui4-48B-A4B-abliterated.Q3_K_M.gguf 23.38 GiB Smaller option Noticeable quality drop versus Q4_K_M.
Huihui4-48B-A4B-abliterated.Q2_K.gguf 18.52 GiB Smallest option Fastest decode in this set, but also the weakest quality.

Benchmark Setup

All benchmarks below were run with:

  • llama.cpp build d132f22fc (8739)
  • RTX 4090 24 GB + RTX 3090 24 GB
  • AMD Ryzen 7 9800X3D
  • flash attention enabled
  • split mode layer
  • q8_0 KV cache

Speed was measured with llama-bench at 4096 prompt tokens and 256 generated tokens. Perplexity was measured with llama-perplexity on raw Wikitext-2 validation text at n_ctx=4096.

The perplexity values are only meant as relative comparisons between these quants under one fixed setup. This is an instruction-tuned multimodal chat model evaluated on raw Wikitext-2 text, so the absolute values should not be treated as a general LM leaderboard score.

Benchmark Results

Quant File size Prefill tok/s Gen tok/s PPL VRAM at 4k q8_0 KV (4090 / 3090) Host RAM Notes
Q8_0 47.64 GiB 1650.39 73.25 265082.05 +/- 5344.34 22.74 / 23.30 GiB 47.66 GiB Best quality. This run was partially offloaded to CPU / host-mapped memory on the benchmark machine.
Q6_K 40.27 GiB 5263.15 129.80 311616.55 +/- 6383.63 23.14 / 21.43 GiB 0.62 GiB Highest quality clean fit on the tested system.
Q4_K_M 29.76 GiB 5750.96 143.03 457818.55 +/- 9564.89 17.74 / 16.18 GiB 0.62 GiB Best overall deployment choice.
Q3_K_M 23.38 GiB 5399.88 138.69 3593800.87 +/- 72288.42 14.58 / 12.97 GiB 0.62 GiB Smaller footprint, but quality drops hard.
Q2_K 18.52 GiB 5371.41 151.59 4859118.84 +/- 95504.97 12.12 / 10.55 GiB 0.62 GiB Smallest and fastest decode, but weakest quality by a large margin.

Recommended Picks

  • Use Q4_K_M if you want the default recommendation.
  • Use Q6_K if you want the best quality that still fits cleanly on a strong dual-24 GB setup.
  • Use Q8_0 only if you are comfortable with partial CPU offload or much larger available memory.
  • Use Q3_K_M or Q2_K only when memory is the priority and you accept a major quality hit.

llama.cpp Usage

Multimodal inference requires both the main quant and the projector file. On the tested runtime, Gemma 4 chat formatting also needed --jinja. If you need OpenAI-compatible tool calling from llama-server, do not pass --skip-chat-parsing.

Example llama-server launch:

llama-server \
  -m Huihui4-48B-A4B-abliterated.Q4_K_M.gguf \
  --mmproj Huihui4-48B-A4B-abliterated.BF16-mmproj.gguf \
  --jinja \
  --reasoning off \
  -fa on \
  -sm layer \
  -dev CUDA0/CUDA1 \
  -ngl 99 \
  -c 32768 \
  -np 1 \
  --host 127.0.0.1 \
  --port 8080

Example text-only llama-cli launch:

llama-cli \
  -m Huihui4-48B-A4B-abliterated.Q4_K_M.gguf \
  --jinja \
  -fa on \
  -sm layer \
  -dev CUDA0/CUDA1 \
  -ngl 99 \
  -c 32768

Notes

  • This is a quantized GGUF release of the original model, not the earlier REAP-pruned experiment.
  • For image input, the BF16-mmproj.gguf file is required regardless of which quant you choose.
  • The Q8 benchmark above is intentionally labeled as partially CPU offloaded because it did not fit as a clean all-GPU run on the tested 4090 + 3090 machine.

Credits

README history 4 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-04-13Upload README.md with huggingface_hub56c349e5.6 KB
    Loading...
  2. 2026-04-13Update README.mdbf09ec15.4 KB
    Loading...
  3. 2026-04-13Update model card682a6cb5.5 KB
    Loading...
  4. 2026-04-13Add files using upload-large-folder tool8c9a2435.7 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration