← back to catalog · registered 2026-08-22 13:56

zaakirio/gemma-4-12b-it-uncensored-GGUF

zaakirio Gemma 12B GGUF multimodal second-order 131K ctx
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/zaakirio%2Fgemma-4-12b-it-uncensored-GGUF"
Response includes
  • classification m8
  • files 12
  • hub_downloads_all_time 247,816
  • author_summary 11 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M8
Primary method

Repackaging (quantization)

Applied on top of direct removal inherited from the base model.
Confidence
MEDIUM
Inherited from base model
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=1
  • assume M1 (base ablation) + M8 (GGUF quant) - default when producer unknown
Refusal direction extraction

No specific extraction method could be identified for this model. The producer either did not document it or used a proprietary pipeline.

What is a refusal direction? →
Downloads · lifetime
248K
78K last 30d - stable
Likes
68
Model age
4mo ago
created 2026-06-04
Downloads over time
Now276.2K→from12.3K↑2,152%
0100.9K201.8K302.6K12.3K on Jun 6276.2K on Oct 11JunJulAugSepOct
Jun 6 → Oct 11 · 59 snapshots · spans 127 days

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Variants by this author 2 formats · 78K downloads combined

The same weights this author released in different packaging. Pick the format that matches your runtime.

Metadata

License
gemma
Quantizations
F16 Q2_K Q3_K Q4_K Q5_K Q6_K Q8_0
Tags
gguf llama.cpp gemma gemma4 heretic abliterated uncensored decensored conversational multimodal audio image-text-to-text

Related

Total size
75.1 GB
Files
12
Quantizations
9
Registered
2026-08-22 13:56
Last updated on HF
2026-06-17 15:54

Files by quantization

F16 1 file 22.2 GB
gemma-4-12b-it-uncensored-f16.gguf 22.2 GB 0f5cbfa4 download
Q8_0 1 file 11.8 GB
gemma-4-12b-it-uncensored-Q8_0.gguf 11.8 GB 55e32e57 download
Q6_K 1 file 9.11 GB
gemma-4-12b-it-uncensored-Q6_K.gguf 9.11 GB 256c2874 download
Q5_K 1 file 7.96 GB
gemma-4-12b-it-uncensored-Q5_K_M.gguf 7.96 GB 54a6e309 download
Q4_K 2 files 13.4 GB
gemma-4-12b-it-uncensored-Q4_K_M.gguf 6.87 GB b2e1db5f download
gemma-4-12b-it-uncensored-Q4_K_S.gguf 6.54 GB b2b1ac0e download
Q3_K 1 file 5.67 GB
gemma-4-12b-it-uncensored-Q3_K_M.gguf 5.67 GB 54e5bfac download
Q2_K 1 file 4.50 GB
gemma-4-12b-it-uncensored-Q2_K.gguf 4.50 GB 8c55fa79 download
BF16 1 file 167 MB
mmproj-gemma-4-12B-it-bf16.gguf 167 MB 675ad6e6 download
Auxiliary files 3 files 444 MB
mtp-gemma-4-12b-it-uncensored.gguf 444 MB 145db909 download
README.md 10.8 KB c676d227 download
.gitattributes 2.19 KB 48fe27a8 download

README current version from Hugging Face


base_model: zaakirio/gemma-4-12b-it-uncensored
license: gemma
pipeline_tag: image-text-to-text
tags:

  • gguf
  • llama.cpp
  • gemma
  • gemma4
  • heretic
  • abliterated
  • uncensored
  • decensored
  • conversational
  • multimodal
  • audio

gemma-4-12b-it-uncensored — GGUF

GGUF quantizations of zaakirio/gemma-4-12b-it-uncensored, a decensored (Heretic-abliterated) version of google/gemma-4-12B-it.

These files run with llama.cpp.

Gemma 4 12B (the "Unified" release, June 2026) is an encoder-free unified multimodal model: text, image, audio, and video all project straight into a single decoder-only transformer, with a context window of up to 256K tokens. Abliteration touches only the language weights, so those capabilities carry over unchanged.

Files

Filenames follow gemma-4-12b-it-uncensored-<QUANT>.gguf.

Quant Size Notes
Q2_K 4.50 GB Smallest; lowest quality. Very tight memory only.
Q3_K_M 5.67 GB Small; usable on low RAM.
Q4_K_S 6.54 GB Compact 4-bit.
Q4_K_M 6.87 GB Recommended — best size/quality balance.
Q5_K_M 7.96 GB Higher quality, slightly larger.
Q6_K 9.11 GB Near-lossless.
Q8_0 11.80 GB Effectively lossless vs the BF16 source.
f16 22.20 GB Full precision; reference / re-quantizing.

Not sure which to pick? Start with Q4_K_M. Go up to Q5/Q6/Q8 if you have the memory and want maximum fidelity; drop to Q3/Q2 only if you're memory-constrained.

Multimodal projector (for image/audio input — see Multimodal):

File Size Notes
mmproj-gemma-4-12B-it-bf16.gguf 0.16 GB Multimodal projector — pair with any quant above.
mtp-gemma-4-12b-it-uncensored.gguf 0.47 GB Optional MTP speculative-decoding drafter — faster generation on GPUs. See Faster generation with MTP.

Usage

llama.cpp (auto-downloads the chosen quant from this repo):

# Interactive chat
llama-cli -hf zaakirio/gemma-4-12b-it-uncensored-GGUF:Q4_K_M --jinja

# OpenAI-compatible server with web UI
llama-server -hf zaakirio/gemma-4-12b-it-uncensored-GGUF:Q4_K_M --jinja -c 4096

-c sets the context window. Gemma 4 12B supports up to 256K tokens (-c 262144); raise it for long chats/documents as memory allows — 4096 is just a light default.

Or with a file you've already downloaded:

llama-cli -m gemma-4-12b-it-uncensored-Q4_K_M.gguf --jinja -p "Hello, who are you?"

Download a single file:

pip install -U "huggingface_hub[cli]"
hf download zaakirio/gemma-4-12b-it-uncensored-GGUF \
  --include "gemma-4-12b-it-uncensored-Q4_K_M.gguf" --local-dir ./

Prompt format & settings

The chat template is embedded in the GGUF — chat-aware tools apply it automatically (always pass --jinja with llama.cpp). For reference, Gemma 4 introduced new turn tokens — <|turn> / <turn|>, replacing Gemma 3's <start_of_turn> / <end_of_turn> — and now supports a dedicated system role:

<|turn>system
{system instructions — optional}<turn|>
<|turn>user
{prompt}<turn|>
<|turn>model

Image and audio inputs use the <|image|> and <|audio|> placeholder tokens, which the multimodal pipeline fills in for you (see Multimodal).

Recommended sampling (Google defaults): --temp 1.0 --top-p 0.95 --top-k 64.

Thinking mode: Gemma 4 has a reasoning channel. To disable it, pass --chat-template-kwargs '{"enable_thinking":false}' to llama-server.

Multimodal

In llama.cpp, multimodal input ships via a separate projector file; the language .gguf alone is text-only and will reject media. This repo includes mmproj-gemma-4-12B-it-bf16.gguf, which carries both a vision and an audio encoder (clip.has_vision_encoder and clip.has_audio_encoder).

In practice: image and audio input both work. Verified with llama-mtmd-cli: it accepts .wav/.mp3 directly and produced an accurate multilingual transcription of real speech. llama.cpp flags audio as experimental ("audio input is in experimental stage and may have reduced quality"), so treat ASR quality as best-effort. Video is handled as sampled frames (plus audio) and depends on your client doing the frame extraction.

When you load via -hf, llama.cpp auto-downloads the projector from this repo — images just work:

llama-server -hf zaakirio/gemma-4-12b-it-uncensored-GGUF:Q4_K_M --jinja

With local files, pass it explicitly with --mmproj:

llama-server -m gemma-4-12b-it-uncensored-Q4_K_M.gguf \
  --mmproj mmproj-gemma-4-12B-it-bf16.gguf --jinja

# download both files
hf download zaakirio/gemma-4-12b-it-uncensored-GGUF \
  --include "gemma-4-12b-it-uncensored-Q4_K_M.gguf" "mmproj-gemma-4-12B-it-bf16.gguf" --local-dir ./

One-shot image or audio from the CLI with llama-mtmd-cli:

# Image
llama-mtmd-cli -m gemma-4-12b-it-uncensored-Q4_K_M.gguf \
  --mmproj mmproj-gemma-4-12B-it-bf16.gguf --jinja \
  --image photo.png -p "What's in this image?"

# Audio (experimental in llama.cpp; 16 kHz WAV)
llama-mtmd-cli -m gemma-4-12b-it-uncensored-Q4_K_M.gguf \
  --mmproj mmproj-gemma-4-12B-it-bf16.gguf --jinja \
  --audio clip.wav -p "Transcribe this audio."

The projector is the unmodified Gemma 4 one (abliteration leaves it untouched). bf16 is preferred — it's small enough that quantizing it buys nothing.

Faster generation with MTP (speculative decoding)

Gemma 4 ships a small (~0.4 B) multi-token-prediction (MTP) drafter (google/gemma-4-12B-it-assistant) that speeds up generation through self-speculative decoding: the drafter proposes several tokens at once and the full model verifies them in a single forward pass. Output is identical to normal decoding — speculative decoding is mathematically exact, so it changes speed only, never quality, and the model stays exactly as uncensored: the main model verifies and decides every emitted token; the drafter merely proposes. (Google state the same — MTP "guarantee[s] the exact same quality as standard autoregressive generation" — and the implementation PR confirmed it by reproducing Gemma's published AIME-26 score.)

For that same reason the drafter does not need to be abliterated: this repo ships Google's unmodified upstream Gemma 4 drafter (the file is renamed mtp-gemma-4-12b-it-uncensored.gguf only so it sits beside the trunk — the weights are upstream), and your decensored trunk verifies every token. It's alignment-independent, exactly like the multimodal projector above.

MTP for Gemma 4 landed in llama.cpp PR #23398 (merged 2026-06-07; see also support discussion #22735), so it needs a llama.cpp build from that date or later.

# Grab the trunk + the MTP drafter
hf download zaakirio/gemma-4-12b-it-uncensored-GGUF \
  --include "gemma-4-12b-it-uncensored-Q4_K_M.gguf" "mtp-gemma-4-12b-it-uncensored.gguf" --local-dir ./

# Serve with MTP enabled
llama-server -m gemma-4-12b-it-uncensored-Q4_K_M.gguf \
  --model-draft mtp-gemma-4-12b-it-uncensored.gguf \
  --spec-type draft-mtp --spec-draft-n-max 4 \
  -ngl 99 --flash-attn auto --jinja

⚠️ MTP helps on GPUs but is slower on Apple Silicon

The benefit depends entirely on the backend.

On a GPU it's a win. The implementation PR measured >2× throughput on the dense Gemma 4 (31B) model — NVIDIA DGX Spark, ~0.59 mean draft acceptance at --spec-draft-n-max 4 — and Unsloth report 1.5–2.2× for Gemma 4 with MTP. (PR #23398, Unsloth MTP guide.)

On Apple Silicon (Metal) it's a regression. Measured here on this model — Q4_K_M, Apple M4 Pro, Gemma's default sampling (--temp 1.0 --top-p 0.95 --top-k 64), through llama-server:

M4 Pro / Metal · Q4_K_M Generation Draft acceptance
MTP off (baseline) ~27 tok/s —
MTP on (--spec-draft-n-max 4) ~12 tok/s (≈2× slower) ~0.45

The drafts are accepted (~45%), but on Metal the draft-evaluation overhead outweighs the decode steps it saves. This isn't specific to this model: an independent report (#23752 — Qwen3.5-9B on an M1 Max) found MTP slower at every setting, even at 100% acceptance, so no sampling tweak rescues it on Metal.

Recommendation: enable MTP when serving on an NVIDIA (CUDA) GPU; leave it off on a Mac. On a GPU, greedy / low-temperature sampling raises acceptance and widens the speedup.

About the base model

A decensored derivative produced with Heretic (automatic directional ablation). Compared with the original:

Metric Decensored Original
Refusals (/100 harmful prompts) 23 99
KL divergence (harmless prompts) 0.043 0 (by definition)

The refusal count is Heretic's keyword heuristic, which is known to over-count (it flags disclaimer-wrapped compliance as a refusal; ~11% precision per arXiv:2512.13655). We report only the measured marker figure and did not run a classifier-based eval on this model, so real compliance is likely higher. See the source model card for parameters and details.

Intended use & disclaimer

This model has had its refusal behaviour substantially removed and will comply with requests the original would have declined. Provided for research and unrestricted local use. You are responsible for how you use it and for complying with applicable law and the base model's Gemma license, which carries over to this derivative. Not for all audiences.

Provenance

README history 12 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-06-17Upload README.md with huggingface_hub328805610.8 KB
    Loading...
  2. 2026-06-17Upload README.md with huggingface_hube3301ce11.2 KB
    Loading...
  3. 2026-06-17Upload README.md with huggingface_hubcc8444d10 KB
    Loading...
  4. 2026-06-14Update README.md3ffb0417.2 KB
    Loading...
  5. 2026-06-14Update model card: Gemma 4 12B Unified context (256K), verified multimodal (i...80a56d87.3 KB
    Loading...
  6. 2026-06-06Update README.mdd70898b5.6 KB
    Loading...
  7. 2026-06-05Document multimodal image input via mmproj projector4ae38135.7 KB
    Loading...
  8. 2026-06-04Upload README.md with huggingface_hub946dfc54.6 KB
    Loading...
  9. 2026-06-04Upload README.md with huggingface_huba468deb4.9 KB
    Loading...
  10. 2026-06-04Upload README.md with huggingface_hub2a17def4.4 KB
    Loading...
  11. 2026-06-04Upload README.md with huggingface_hub21817af4.3 KB
    Loading...
  12. 2026-06-04Upload README.md with huggingface_hubedcf4a92.2 KB
    Loading...

Discussions 2 threads

  1. 2026-06-20Comprehensive benchmark & review: uncensored status, MTP speed, and limitationsopen1 💬#2
    Loading...
  2. 2026-06-04MTPclosed8 💬#1
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration