← back to catalog · registered 2026-08-22 13:56

coolthor/Huihui-gemma-4-26B-A4B-it-abliterated-eagle3-draft

coolthor Gemma 928M MoE second-order
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/coolthor%2FHuihui-gemma-4-26B-A4B-it-abliterated-eagle3-draft"
Response includes
  • classification m1
  • files 5
  • hub_downloads_all_time 171
  • author_summary 7 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M1
Primary method

Direct removal

No other method signals detected in this model.
Confidence
MEDIUM
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=0 (base model)
  • no specific method indicators - defaulting to M1 (most common)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
MEDIUM
Why we say so
primary_method=M1; difference-of-means is the reference extraction for M1/M3 (Arditi 2024)
Downloads · lifetime
171
35 last 30d - stable
Likes
2
Model age
4mo ago
created 2026-05-15
Downloads over time
Now184→from42↑338%
358914419842 on May 20184 on Oct 11MayJunJulAugSepOct
May 20 → Oct 11 · 60 snapshots · spans 144 days

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
apache-2.0
Tags
speculators safetensors eagle3 speculative-decoding gemma4 abliterated draft-model vllm custom_code base_model:google/gemma-4-26B-A4B-it base_model:finetune:google/gemma-4-26B-A4B-it license:apache-2.0

Related

Total size
1.73 GB
Files
5
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-06-02 09:04

Files by quantization

Auxiliary files 5 files 1.73 GB
model.safetensors 1.73 GB fb0d2f16 download
README.md 10.7 KB 14380f2b download
config.py 3.21 KB a3201c45 download
config.json 1.55 KB 95990b3e download
.gitattributes 1.48 KB a6344aac download

README current version from Hugging Face


license: apache-2.0
base_model:

  • huihui-ai/Huihui-gemma-4-26B-A4B-it-abliterated
  • google/gemma-4-26B-A4B-it
    library_name: speculators
    tags:
  • eagle3
  • speculative-decoding
  • gemma4
  • abliterated
  • draft-model
  • vllm
  • speculators

huihui Gemma 4 26B-A4B abliterated · EAGLE-3 draft model

⚠️ 2026-05-17 endpoint correction — the ~100 tok/s and ~2× speedup numbers in this card are measured on /v1/completions (raw prompt, no chat template). A later paired bench showed that on production /v1/chat/completions workloads, this drafter delivers ~46 tok/s vs pure body ~40 tok/s (≈ +15% uplift), and vanilla MTP n=1 hits ~51 tok/s on the same workload. The acceptance-curve flattening this drafter achieves is real on raw output, but the per-position-acceptance and throughput numbers below don't translate 1:1 to chat workloads. Round 2 paired bench (with Chinese training data) will land in Part 31. Until then, for chat workloads consider using gemma4-26b-a4b-it-assistant (vanilla MTP) with num_speculative_tokens=4 instead — Round 2 paired bench shows it hits chat EN ~53 / ZH ~45 tok/s. See the Part 30 endpoint correction note for full paired numbers.

EAGLE-3 speculative-decoding draft model fine-tuned for huihui-ai/Huihui-gemma-4-26B-A4B-it-abliterated (and the FP8 quantized variant coolthor/Huihui-gemma-4-26B-A4B-it-abliterated-FP8-Dynamic).

Starts from RedHatAI/gemma-4-26B-A4B-it-speculator-eagle3 (vanilla-trained on the unmodified Gemma 4 body) and is fine-tuned for 1 epoch / 50K samples against the abliterated body's hidden-state distribution. The point is to restore deep-speculation acceptance that the vanilla draft loses once the body has been abliterated.

Why this exists

Vanilla MTP / EAGLE-3 drafters are trained against the vanilla gemma-4-26B-A4B-it body. When you swap the body for an abliterated one (refusal direction removed), the body's hidden-state distribution shifts and the draft model's predictions stop matching the body's actual outputs — especially at deeper speculation depths. Per-position acceptance collapses from ~65% (pos 0) to ~20% (pos 3) in a 4-token speculation window.

This drafter is fine-tuned to re-align with the abliterated body. On a same-stack inference benchmark (DGX Spark GB10, FP8 verifier, T=0.7, N=10 prompts, max_tokens=200), per-position acceptance recovers to a nearly-flat curve:

Position Vanilla draft (against abliterated body) This drafter Δ
pos 0 65.6% 84.4% +18.8pp
pos 1 43.3% 74.9% +31.6pp
pos 2 29.2% 74.1% +44.9pp
pos 3 20.5% 72.7% +52.2pp

Throughput at num_speculative_tokens=4 (on /v1/completions raw endpoint): ~100 tok/s vs ~50 tok/s with the vanilla draft. Same hardware, same prompts. See the endpoint correction above — these acceptance and throughput numbers come from the raw endpoint, not from chat completions.

Usage with vLLM

Requires vLLM with PR #41745 (Gemma 4 MTP / EAGLE-3 integration) merged — currently means building from main or using a preview image. PR is merged into main as of 2026-05-06; will land in the next minor release after v0.20.2.

vllm serve coolthor/Huihui-gemma-4-26B-A4B-it-abliterated-FP8-Dynamic \
  --speculative-config '{"method":"eagle3","model":"coolthor/Huihui-gemma-4-26B-A4B-it-abliterated-eagle3-draft","num_speculative_tokens":4}' \
  --kv-cache-dtype fp8 \
  --gpu-memory-utilization 0.85 \
  --max-model-len 8192 \
  --limit-mm-per-prompt '{"image":0,"audio":0,"video":0}' \
  --enable-auto-tool-choice --tool-call-parser gemma4 \
  --enable-prefix-caching \
  --trust-remote-code

num_speculative_tokens sweep (DGX Spark GB10, FP8 verifier, batch=1):

n Throughput pos-0 acceptance
1 59.04 tok/s 81.3%
2 66.96 tok/s 81.6%
3 74.90 tok/s 88.5%
4 100.36 tok/s 84.4%

n=4 is the recommended setting on this hardware. Throughput keeps increasing through n=4 because deeper positions stay highly accepted (74/74/73% at pos 1/2/3), unlike the vanilla draft where deep speculation collapses.

Training details

  • Framework: vllm-project/speculators v0.5.0.dev0
  • Hardware: NVIDIA GB10 (DGX Spark, sm_12.1), 121 GB unified memory, 273 GB/s
  • Wall time: ~11 hours (vLLM extraction + EAGLE-3 fine-tune concurrent on a single GB10)
  • Starting point: RedHatAI/gemma-4-26B-A4B-it-speculator-eagle3 (vanilla-trained pretrained)
  • Training data: Magpie-Align/Magpie-Llama-3.1-Pro-300K-Filtered — 50K instruction prompts regenerated through the FP8 abliterated body to produce on-distribution (prompt, response) pairs
  • Hyper-params: 1 epoch, total_seq_len=4096 (packed), TORCH_COMPILE_DISABLE=1, default LR schedule (cosine, warmup ~100 steps)
  • Aux hidden-state layers: [2, 15, 27] (EAGLE-3 default)
  • Trained spec depth (ttt_steps): 3 (positions 0, 1, 2). At inference, pos 3 is extrapolated from the same drafter — empirically it works (72.7% acceptance), but this is outside the training distribution and may degrade on non-Magpie-shaped workloads.

Validation metrics (Magpie held-out, teacher-forced)

Position full_acc (val) cond_acc (val)
pos 0 66.8% 66.8%
pos 1 41.4% 61.5%
pos 2 26.4% 62.6%

⚠️ Note: Validation full_acc is measured via teacher-forced argmax against Magpie ground-truth tokens — strict. Inference acceptance (table above) is measured via rejection-sampling against the body's actual sampling distribution at T=0.7 — looser. The 26.4% val pos-2 vs 74.1% inference pos-2 gap reflects this metric difference, not training failure. What matters for speculative decoding throughput is the inference number.

Limitations

  • Production chat workloads see much smaller uplift than the headline numbers suggest. Paired bench (2026-05-17) on /v1/chat/completions: this drafter ~46 tok/s, vanilla MTP n=1 ~51 tok/s, pure body (no spec) ~40 tok/s. Real uplift on chat ≈ +15%, not +100%. For production chat use cases, vanilla MTP gemma4-26b-a4b-it-assistant with num_speculative_tokens=4 outperforms this drafter (EN ~53 / ZH ~45 tok/s on chat). This drafter's headline numbers come from /v1/completions raw endpoint, where instruct-tuned bodies produce degenerate output that small drafters can trivially predict.
  • Chinese (and likely other non-English) workloads are out-of-distribution. v1 was trained only on English Magpie data — paired bench shows ZH chat pos-0 acceptance drops to ~12% (vs vanilla MTP's ~57%). Round 2 (with Chinese training data + ttt_steps=4) is in flight; results will land in Part 31.
  • Trained for 3 spec positions, not 4. Inference num_speculative_tokens=4 works well empirically but pos 3 is extrapolation. n=3 is the safest "in-distribution" setting; n=4 is what we recommend for max throughput on the tested hardware.
  • Training data is English instruction-style. Magpie-Llama-3.1 is English-leaning. Performance on Traditional Chinese, code-heavy, or domain-specific workloads may differ from the reported numbers. For Traditional Chinese workloads, Qwen 3.6 abliterated remains the better base model anyway (TMMLU+ 75% vs Gemma 4's 46%).
  • Optimizer state not included. This release ships only the inference weights (model.safetensors + config.json + config.py). To resume training, train from scratch or contact me.
  • One-epoch run. EAGLE-3 papers typically train for many more samples × epochs. Multi-epoch and/or larger training data may improve acceptance further.

License

Apache 2.0 — inherited from Gemma 4. The huihui base model is also Apache 2.0 per its model card.

Acknowledgements

Related


☕ If this saved you GPU hours, you can buy me a coffee.

README history 4 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-06-02docs: add Buy Me a Coffee link to model card8a4e24910.7 KB
    Loading...
  2. 2026-05-17README: add 2026-05-17 endpoint correction — chat workload uplift ~15% not 2x4edf79d10.6 KB
    Loading...
  3. 2026-05-16README: add bilingual deep links to Part 28 / 29 / 30 (bidirectional traffic)d3165e58.6 KB
    Loading...
  4. 2026-05-15Upload folder using huggingface_huba0e5cd18 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration