← back to catalog · registered 2026-08-22 13:56

HauhauCS/Gemma4-31B-QAT-Uncensored-HauhauCS-Balanced-MTP

HauhauCS Gemma 31B GGUF multimodal 262K ctx
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/HauhauCS%2FGemma4-31B-QAT-Uncensored-HauhauCS-Balanced-MTP"
Response includes
  • classification m-uncensored
  • files 5
  • benchmarks 11 entries
  • hub_downloads_all_time 588,756
  • author_summary 26 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M-U
Primary method

Uncensored (method unknown)

No other method signals detected in this model.
Confidence
LOW
Why this label 3 signals
Weak or ambiguous signals. Best guess based on catalog patterns; treat as tentative and check the evidence below.
  • 'uncensored' in name/tags but no 'abliterated' marker
  • method not identifiable from author declaration alone
  • may be DPO fine-tune, prompt engineering, or unknown technique
Refusal direction extraction

No specific extraction method could be identified for this model. The producer either did not document it or used a proprietary pipeline.

What is a refusal direction? →
Downloads · lifetime
589K
96K last 30d - stable
Likes
182
Descendants
2
in 2 direct forks
Model age
3mo ago
created 2026-06-24
Downloads over time
Now625.1K→from13K↑4,696%
0229.2K458.4K687.6K13K on Jun 24625.1K on Oct 11JunJulAugSepOct
Jun 24 → Oct 11 · 58 snapshots · spans 109 days

Benchmarks

Portrait before abliteration
Benchmarks of the base model as it stood before the refusal-removal operation. Compare with the numbers above to see what the operation cost.
Benchmark Score Source
Entertainment 1.9 UGI
Hazardous 0 UGI
Natural Intelligence 34.36 UGI
Political lean -19.4% UGI
Sensitive-Info 19.81 UGI
SocPol 3.7 UGI
UGI 21.54 UGI
Willingness (10) 2.5 UGI
W10-Adherence 3 UGI
W10-Direct 2 UGI
Writing 38.57 UGI

Genealogy 2 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
gemma
Languages
en
Quantizations
Q4_K
Tags
gguf uncensored gemma4 vision multimodal agentic coding creative-writing roleplay rp conversational image-text-to-text

Related

Total size
17.7 GB
Files
5
Quantizations
3
Registered
2026-08-22 13:56
Last updated on HF
2026-06-25 00:22

Files by quantization

Q4_K 1 file 17.4 GB
Gemma4-31B-QAT-Uncensored-HauhauCS-Balanced-Q4_K_M.gguf 17.4 GB 71667f9e download
BF16 1 file 1.12 GB
mmproj-Gemma4-31B-QAT-Uncensored-HauhauCS-Balanced-BF16.gguf 1.12 GB 7bef0d0f download
Auxiliary files 3 files 267 MB
mtp-gemma-4-31B-it.gguf 267 MB b5c4e583 download
README.md 3.76 KB 3ad4df55 download
.gitattributes 1.73 KB 2e80b5be download

README current version from Hugging Face


license: gemma
tags:

  • uncensored
  • gemma4
  • gguf
  • vision
  • multimodal
  • agentic
  • coding
  • creative-writing
  • roleplay
  • rp
  • conversational
    language:
  • en
    pipeline_tag: image-text-to-text
    base_model: google/gemma-4-31B-it

Gemma4-31B-QAT-Uncensored-HauhauCS-Balanced-MTP

Join the Discord for updates, roadmaps, projects, or just to chat.

Gemma4-31B (QAT) uncensored by HauhauCS. 0/465 Refusals*

About

No changes to datasets or capabilities — fully functional, 100% of what the original authors intended, just without the refusals. Built from the official QAT weights, so the 4-bit quant stays close to full-precision quality.

Balanced

The Balanced variant (recommended — 99%+ of users will be happy here) uses optimized full uncensoring tuned especially for agentic coding, reasoning, creative writing and reliability-critical tasks. It reasons before answering and stays dependable and on-instruction. An Aggressive variant, for cases where Balanced still deflects too much, after current testing is not required.

~53% faster with MTP

Ships with an MTP (multi-token-prediction) draft head for speculative decoding — roughly 53% faster generation with identical output (the model verifies every drafted token, so quality is unchanged — pure speed). This release is tuned to pair well with the included MTP head.

llama.cpp:

llama-server \
  -m Gemma4-31B-QAT-Uncensored-HauhauCS-Balanced-Q4_K_M.gguf \
  -md mtp-gemma-4-31B-it.gguf --spec-type draft-mtp \
  -ngl 99 -fa on

Note: the MTP speedup was currently tested by me through llama.cpp (llama-server / llama-cli).

Downloads

File Type Size
Gemma4-31B-QAT-Uncensored-HauhauCS-Balanced-Q4_K_M.gguf Q4_K_M (text) 18.7 GB
mmproj-Gemma4-31B-QAT-Uncensored-HauhauCS-Balanced-BF16.gguf mmproj (vision) 1.2 GB
mtp-gemma-4-31B-it.gguf MTP speculative drafter 280 MB

Why only Q4_K_M? Gemma 4 is quantization-aware-trained for ~4-bit, so Q4_K_M is the sweet spot — higher-precision quants add size with no real quality gain. Carefully quantized for best quality at 4-bit.

Vision

Load the mmproj alongside the model for image input:

llama-server -m Gemma4-31B-QAT-Uncensored-HauhauCS-Balanced-Q4_K_M.gguf \
  --mmproj mmproj-Gemma4-31B-QAT-Uncensored-HauhauCS-Balanced-BF16.gguf -ngl 99 -fa on

Recommended sampling

These are dialed in specifically for this HauhauCS build — use them for the intended behaviour and quality:

  • temperature 0.6
  • top_k 64
  • top_p 0.9
  • min_p 0.05
  • repeat_penalty 1.1

This release is tuned end-to-end as its own thing; the settings above are part of that and aren't the stock Gemma defaults.

Specs

  • 31B dense · 256K (262144) context
  • Vision (image input) via mmproj
  • Based on Gemma 4 31B by Google DeepMind

Compatibility

  • Works with llama.cpp, LM Studio, Jan, koboldcpp, and other GGUF runtimes.
  • Multi-GPU + LM Studio: I've personally noticed Gemma 4 can crash under LM Studio's tensor-split mode — use a single GPU (layer-split or priority order) for this model.

Acknowledgements

  • Google DeepMind — Gemma 4.
  • The included mtp-gemma-4-31B-it.gguf speculative draft head comes from Unsloth's Gemma 4 release — many thanks to the Unsloth team for it.

* Tested with both automated and manual refusal benchmarks — none have been found in standard use. A small number of edge-case prompts deflect on the first ask but comply on a re-ask or strategic framing. If you hit one that's actually obstructive to your use case, join the Discord and flag it so I can work on it in a future revision.

README history 1 version

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-06-25initial commit96544663.8 KB
    Loading...

Discussions 7 threads

  1. 2026-08-15qwen3.8 pleaseopen1 💬#7
    Loading...
  2. 2026-08-05Running on vLLMopen2 💬#6
    Loading...
  3. 2026-07-27Request: Laguna Seriesopen7 💬#5
    Loading...
  4. 2026-07-14About the usage of this model for local companion (Flash Attention)open2 💬#4
    Loading...
  5. 2026-07-07problems with tools (writing files) with llama.cpp.open2 💬#3
    Loading...
  6. 2026-07-02Join the discord!open1 💬#2
    Loading...
  7. 2026-06-28How to load MTP.gguf into the model when using LMstuidoopen7 💬#1
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration