← back to catalog · registered 2026-08-22 13:56

coolthor/Huihui-gemma-4-26B-A4B-it-abliterated-FP8-Dynamic

coolthor Gemma 24B MoE multimodal second-order
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/coolthor%2FHuihui-gemma-4-26B-A4B-it-abliterated-FP8-Dynamic"
Response includes
  • classification m1
  • files 10
  • hub_downloads_all_time 989
  • author_summary 7 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M1
Primary method

Direct removal

No other method signals detected in this model.
Confidence
MEDIUM
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=0 (base model)
  • no specific method indicators - defaulting to M1 (most common)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
MEDIUM
Why we say so
primary_method=M1; difference-of-means is the reference extraction for M1/M3 (Arditi 2024)
Downloads · lifetime
989
267 last 30d - stable
Likes
0
Model age
5mo ago
created 2026-05-08
Downloads over time
Now1.2K→from35↑3,320%
04388751.3K35 on May 61.2K on Oct 11MayJunJulAugSepOct
May 6 → Oct 11 · 62 snapshots · spans 158 days

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
apache-2.0
Languages
en zh multilingual
Tags
safetensors gemma4 abliterated gemma gemma-4 fp8 fp8-dynamic compressed-tensors vllm multimodal image-text-to-text conversational

Related

Total size
26.7 GB
Files
10
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-07-01 03:41

Files by quantization

Auxiliary files 10 files 26.7 GB
model.safetensors 26.7 GB 230b1659 download
tokenizer.json 30.7 MB cc8d3a0c download
config.json 19.2 KB 6132e6d5 download
chat_template.jinja 11.8 KB 33c51c2d download
README.md 5.64 KB b0021c46 download
tokenizer_config.json 2.02 KB e5418067 download
processor_config.json 1.65 KB 5465974d download
.gitattributes 1.53 KB 52373fe2 download
recipe.yaml 664 B 064b1293 download
generation_config.json 203 B 5a376e9f download

README current version from Hugging Face


license: apache-2.0
language:

  • en
  • zh
  • multilingual
    base_model: huihui-ai/Huihui-gemma-4-26B-A4B-it-abliterated
    tags:
  • abliterated
  • gemma
  • gemma-4
  • fp8
  • fp8-dynamic
  • compressed-tensors
  • vllm
  • multimodal
    pipeline_tag: image-text-to-text

Huihui-gemma-4-26B-A4B-it-abliterated FP8-Dynamic

vLLM-compatible FP8-Dynamic quantization of huihui-ai/Huihui-gemma-4-26B-A4B-it-abliterated.

The original BF16 release (~50 GB) is not directly servable by vLLM at production speed — --quantization fp8 runtime path is ~6× slower than pre-quantized FP8. This repo fills that gap: 27 GB FP8 weights, drop-in for vLLM, multimodal vision retained, abliteration profile preserved.

Smoke test (DGX Spark / GB10)

Modality Result
Text English (haiku) ✅ coherent
Text 繁體中文 ✅ humor + structure preserved
Vision (古風美女 portrait) ✅ correctly described hanfu, bamboo, pose, expression
Audio N/A — Gemma 4 26B-A4B-it has audio_config: null (vanilla design, not quant artifact)

Speculative-decoding bench (DGX Spark, GB10, batch=1, T=0.0)

Tested pairing this model as MTP target with the vanilla google/gemma-4-26B-A4B-it-assistant draft model:

Config Acceptance Throughput
Baseline (no spec) n/a 39.3 tok/s
MTP, num_speculative_tokens=4 40% token-level (per-pos: 65 / 43 / 29 / 21) 52.4 tok/s (+33%)
MTP, num_speculative_tokens=1 69% 52.6 tok/s (+34%)

Verdict: MTP gives a real +33% gain even with vanilla draft head. num_speculative_tokens=1 is recommended — cleaner per-step acceptance (69% vs 40% token-level for n=4), same throughput. The per-position decay (65% → 21% over 4 positions) is consistent with huihui's abliteration shifting the body's prediction distribution; n=1 only uses the high-confidence first position.

For comparison, vanilla Gemma 4 + MTP achieves 108 tok/s on the same hardware (Part 27 reference). The ~50% gap is the "abliteration tax" — likely from a combination of FP8 calibration on shifted weights and MoE routing imbalance after abliteration. If you need raw speed without abliteration, use the vanilla google/gemma-4-26B-A4B-it directly.

vLLM serving

Recommended (with MTP, +34%)

vllm serve coolthor/Huihui-gemma-4-26B-A4B-it-abliterated-FP8-Dynamic \
  --speculative-config '{"method":"mtp","model":"google/gemma-4-26B-A4B-it-assistant","num_speculative_tokens":1}' \
  --kv-cache-dtype fp8 \
  --gpu-memory-utilization 0.65 \
  --max-model-len 8192 \
  --limit-mm-per-prompt '{"image":0,"audio":0,"video":0}' \
  --enable-auto-tool-choice \
  --tool-call-parser gemma4 \
  --trust-remote-code

Note: MTP requires vLLM with PR #41745's gemma4_mtp.py integration. Tested with the vllm/vllm-openai:gemma4-0505-arm64-cu130 image.

Usage (client side)

⚠️ This is an instruction-tuned model. Always use the chat completions endpoint.

curl -s -X POST http://localhost:8000/v1/chat/completions \
  -H 'Content-Type: application/json' \
  -d '{
    "model": "huihui-gemma4",
    "messages": [{"role": "user", "content": "Write a haiku about an old API."}],
    "max_tokens": 80,
    "temperature": 0.7
  }'

Do NOT use /v1/completions with a raw prompt — it bypasses the Gemma chat template (<start_of_turn>user ... <end_of_turn>\n<start_of_turn>model\n) and the model will fall off-distribution and loop garbage tokens. This is not abliteration / FP8 / MTP related; vanilla google/gemma-4-26B-A4B-it behaves the same way under /v1/completions. If you must use the completions endpoint, apply the chat template manually.

Quantization recipe

Used llmcompressor with FP8_DYNAMIC scheme. Critical ignore list — note re:.*router.* in particular: omitting this entry quantizes MoE router weights and breaks expert dispatch (verified by trial in our own testing):

ignore = [
    "re:.*router.*",        # MoE router — must NOT quantize
    "lm_head",
    "re:.*embed_tokens.*",
    "re:.*norm.*", "re:.*layernorm.*", "re:.*layer_norm.*",
    "re:.*rmsnorm.*", "re:.*rms_norm.*",
    "re:.*conv1d.*", "re:.*linear_attn.*",
    "re:visual.*", "re:model.visual.*",
    "re:.*patch_embed.*", "re:.*vision.*", "re:.*image.*",
    "re:.*video.*", "re:.*projector.*", "re:.*merger.*",
    "re:.*mlp.gate$", "re:.*shared_expert_gate.*",
    "re:.*embed_audio.*", "re:.*embed_vision.*",
    "re:.*audio_tower.*", "re:.*audio_projector.*",
]

Quantization run: 2.9 minutes on GB10 (data-free pipeline, 30 MoE module calibration, FP8 scale derivation).

Toolchain notes

llmcompressor stable releases pin transformers <= 4.57.6, but Gemma4ForConditionalGeneration requires transformers 5.0+. Workaround: patch llmcompressor/entrypoints/utils.py to remove deprecated use_auth_token=... kwargs (8 occurrences) before reinstalling editable.

Credits

License

Apache 2.0, inherited from base model and Gemma usage terms.


☕ If this saved you GPU hours, you can buy me a coffee.


📝 Quantized & benchmarked by ai-muninn — writeups on how it was built and how it actually runs.

README history 5 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-07-01Add ai-muninn linkc4bbb765.6 KB
    Loading...
  2. 2026-06-02docs: add Buy Me a Coffee link to model card9711f0d5.5 KB
    Loading...
  3. 2026-05-16docs: add client-side Usage section (chat completions, not /v1/completions)7f5f8015.4 KB
    Loading...
  4. 2026-05-08Reword router note3491e3e4.6 KB
    Loading...
  5. 2026-05-08Add README with bench results0acd4ad4.5 KB
    Loading...

Discussions 1 thread

  1. 2026-05-14Model hallucinatesopen3 💬#1
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration