← back to catalog · registered 2026-08-22 13:56

satgeze/Gemma4-26B-A4B-Uncensored-HauhauCS-1M-GGUF

satgeze Gemma 26B GGUF MoE multimodal second-order 1.0M ctx
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/satgeze%2FGemma4-26B-A4B-Uncensored-HauhauCS-1M-GGUF"
Response includes
  • classification m-uncensored
  • files 9
  • hub_downloads_all_time 10,583
  • author_summary 7 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M-U
Primary method

Uncensored (method unknown)

No other method signals detected in this model.
Confidence
LOW
Why this label 3 signals
Weak or ambiguous signals. Best guess based on catalog patterns; treat as tentative and check the evidence below.
  • 'uncensored' in name/tags but no 'abliterated' marker
  • method not identifiable from author declaration alone
  • may be DPO fine-tune, prompt engineering, or unknown technique
Refusal direction extraction

No specific extraction method could be identified for this model. The producer either did not document it or used a proprietary pipeline.

What is a refusal direction? →
Downloads · lifetime
11K
2K last 30d - stable
Likes
8
Model age
3mo ago
created 2026-07-06
Downloads over time
Now11.8K→from514↑2,200%
04.3K8.6K13K514 on Jul 711.8K on Oct 11JulAugSepOct
Jul 7 → Oct 11 · 55 snapshots · spans 96 days

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
gemma
Quantizations
Q4_K
Tags
gguf long-context yarn gemma4 uncensored mtp speculative-decoding vision llama.cpp ollama text-generation base_model:HauhauCS/Gemma4-26B-A4B-QAT-Uncensored-HauhauCS-Balanced-MTP

Related

Total size
15.9 GB
Files
9
Quantizations
3
Registered
2026-08-22 13:56
Last updated on HF
2026-07-10 22:59

Files by quantization

Q4_K 1 file 15.6 GB
gemma4-26b-a4b-uncensored-1M-Q4_K_M.gguf 15.6 GB 6ac7a87c download
mmproj 1 file 1.11 GB
mmproj-gemma26b-hauhau.gguf 1.11 GB b5346e5b download
Auxiliary files 7 files 241 MB
mtp-gemma-4-26B-A4B-it.gguf 240 MB 62bd3af7 download
banner.jpeg 459 KB 3c608821 download
niah_heatmap.png 104 KB b7d275cb download
mtp_speedup.png 43.0 KB 04a1e68a download
README.md 6.82 KB 5daff5be download
results.jsonl 5.66 KB 39aa3735 download
.gitattributes 1.78 KB dd42e972 download

README current version from Hugging Face


license: gemma
pipeline_tag: text-generation
base_model: HauhauCS/Gemma4-26B-A4B-QAT-Uncensored-HauhauCS-Balanced-MTP
tags:

  • gguf
  • long-context
  • yarn
  • gemma4
  • uncensored
  • mtp
  • speculative-decoding
  • vision
  • llama.cpp
  • ollama

Gemma4-26B-A4B Uncensored: 1M Context + MTP + Vision

HauhauCS/Gemma4-26B-A4B-QAT-Uncensored-HauhauCS-Balanced-MTP (26B MoE, 4B active, Google QAT checkpoint) with a 1,048,576-token context baked in (4x the native 262,144), shipping with its MTP speculative-decoding draft head and vision tower. All numbers below were measured on these exact files.

Capability Status
1M context~91% mean recall across two seed sets (honesty note below)
MTP speculative decoding249.1 to 369.2 tok/s (+48%), acceptance 0.679 (measured on this trunk, RTX 5090)
VisionVerified July 6, 2026: reads image text and identifies objects
UncensoredHauhauCS Balanced abliteration; trunk weights bit-identical to the source release

Needle-in-a-haystack

Full transparency: across two complete seed sets this trunk averages ~91 percent needle recall, dropping roughly one needle per rung at random depths, including inside the native 262K range. The official censored trunk scored 10/10 at 393K under the same harness, so this is a small flat abliteration tax, not a context-length failure. If a rare retrieval miss is unacceptable for your workload, use the 12B, which is certified clean. Every run, including the imperfect ones, is in results.jsonl. Rungs above 524K exceed 32 GB VRAM and have not been run yet; they are queued for larger hardware and will publish here as measured, pass or fail.

MTP speculative decoding

The draft head predicts ahead and the trunk verifies every token, so output is identical to standard decoding, only faster. Measured speedup on this uncensored trunk beats the ~35 percent claimed upstream.

Files

File Size Role
gemma4-26b-a4b-uncensored-1M-Q4_K_M.gguf 16.8 GB Trunk, 1M baked, QAT 4-bit
mtp-gemma-4-26B-A4B-it.gguf 252 MB MTP draft head, pair with -md
mmproj-gemma26b-hauhau.gguf 1.2 GB Vision tower, pair with --mmproj
niah_heatmap.png, mtp_speedup.png, results.jsonl small Verification evidence

Every file, every mirror

Nothing was discontinued: every quant is one click away. Hugging Face carries the curated picks, ModelScope always carries everything, and Ollama serves ready-to-run tags.

On Ollama every tag ships with the vision tower bundled and the 1M rope metadata baked in.

File Size Hugging Face ModelScope Ollama
gemma4-26b-a4b-uncensored-1M-Q4_K_M.gguf 16.8 GB download download ollama run satgeze/gemma4-26b-uncensored-1m
mmproj-gemma26b-hauhau.gguf 1.2 GB download download bundled in every tag
mtp-gemma-4-26B-A4B-it.gguf 252 MB download download -

Run it

llama.cpp, everything on:

llama-server -m gemma4-26b-a4b-uncensored-1M-Q4_K_M.gguf \
  -c 1048576 -np 1 --jinja \
  -md mtp-gemma-4-26B-A4B-it.gguf --spec-type draft-mtp --spec-draft-n-max 3 \
  --mmproj mmproj-gemma26b-hauhau.gguf

Ollama (1M and vision work; Ollama has no speculative decoding yet, so the MTP head adds no speed there):

FROM ./gemma4-26b-a4b-uncensored-1M-Q4_K_M.gguf
RENDERER gemma4
PARSER gemma4
PARAMETER num_ctx 262144

The RENDERER and PARSER lines avoid imported-GGUF template bugs under tool-heavy use. Raise num_ctx as memory allows.

How this was built

YaRN rope-scaling metadata (factor 4.0 over native 262,144) baked into the GGUF header with gguf-py; weights are bit-identical to the HauhauCS release, no fine-tuning. Gemma 4's dual-rope design takes YaRN on its global-attention layers. Certification harness: 10 needles per rung at depths 5 to 95 percent, temperature 0, seeded prompts, f16 KV only. Method and tooling: github.com/satindergrewal/aviary-1m.

For base capability benchmarks see Google's official Gemma 4 cards; uncensoring quality versus the official trunk has not been independently benchmarked here.

How to actually use a 1M-context model

Habits that measurably help, from our RULER, hop and adherence testing across this fleet:

  1. Re-state standing instructions near the end of long prompts; recency beats depth.
  2. One big reference dump beats a long accumulated conversation. Fresh session per task.
  3. After any compaction or summarization, repeat your active rules yourself.
  4. Prefill at 500K+ takes real time on any hardware; stage your questions accordingly.
  5. Know your quant: the results tables on this card show what each quant actually holds at depth; pick the strongest one your memory allows.

Credits

Base model and QAT: Google (Gemma license; its terms flow down to these files). Uncensoring and packaging: HauhauCS. MTP head: Unsloth (via the HauhauCS repo). 1M YaRN extension, benchmarking, and certification: SatGeze.

Sister repos: 12B | 26B-A4B | 31B | Qwen3.6-35B

Mirrors: Hugging Face | ModelScope

README history 6 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-07-10card: add 'How to actually use a 1M-context model' section (fleet-wide)03719b06.8 KB
    Loading...
  2. 2026-07-09Upload README.md with huggingface_hub05cee606.2 KB
    Loading...
  3. 2026-07-06Upload README.md with huggingface_hubaeeee8f6.2 KB
    Loading...
  4. 2026-07-06Upload README.md with huggingface_hubee1f8d95.9 KB
    Loading...
  5. 2026-07-06Upload README.md with huggingface_hub25499c85.9 KB
    Loading...
  6. 2026-07-06Upload folder using huggingface_hub9aca4c54.7 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration