← back to catalog · registered 2026-08-22 13:56

llmfan46/gemma-4-12B-coder-fable5-composer2.5-v1-uncensored-heretic-GGUF

llmfan46 Gemma 12B GGUF second-order 262K ctx
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/llmfan46%2Fgemma-4-12B-coder-fable5-composer2.5-v1-uncensored-heretic-GGUF"
Response includes
  • classification m3
  • files 10
  • hub_downloads_all_time 80,151
  • author_summary 211 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M3
Primary method

Layer-wise ablation

Applied on top of direct removal inherited from the base model.
Confidence
HIGH
Inherited from base model
Why this label 2 signals
Producer identity confirmed by naming conventions, tags or the model card. This label is very unlikely to change.
  • 'heretic' in model name (Heretic-produced)
  • Heretic uses layer-wise optimization (M3) with underlying direction removal (M1)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
HIGH
Why we say so
name contains 'heretic'; Heretic default extraction is difference-of-means (Arditi 2024)
Downloads · lifetime
80K
31K last 30d - stable
Likes
55
Model age
3mo ago
created 2026-06-21
Downloads over time
Now85K→from1.8K↑4,512%
031.1K62.2K93.4K1.8K on Jun 2385K on Oct 11JunJulAugSepOct
Jun 23 → Oct 11 · 57 snapshots · spans 110 days

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Variants by this author 2 formats · 32K downloads combined

The same weights this author released in different packaging. Pick the format that matches your runtime.

Metadata

License
apache-2.0
Quantizations
F16 Q4_K Q5_K Q6_K Q8_0
Tags
transformers gguf gemma4 coding code reasoning thinking safetensors heretic uncensored decensored abliterated

Related

Total size
72.3 GB
Files
10
Quantizations
6
Registered
2026-08-22 13:56
Last updated on HF
2026-06-30 21:49

Files by quantization

F16 2 files 22.3 GB
gemma-4-12B-coder-fable5-composer2.5-v1-uncensored-heretic-F16.gguf 22.2 GB ecf3495b download
gemma-4-12B-coder-fable5-composer2.5-v1-uncensored-heretic-mmproj-F16.gguf 116 MB 8723820b download
Q8_0 1 file 11.8 GB
gemma-4-12B-coder-fable5-composer2.5-v1-uncensored-heretic-Q8_0.gguf 11.8 GB ac0da7d3 download
Q6_K 1 file 9.11 GB
gemma-4-12B-coder-fable5-composer2.5-v1-uncensored-heretic-Q6_K.gguf 9.11 GB 6c0f3482 download
Q5_K 2 files 15.7 GB
gemma-4-12B-coder-fable5-composer2.5-v1-uncensored-heretic-Q5_K_M.gguf 7.96 GB 5e311253 download
gemma-4-12B-coder-fable5-composer2.5-v1-uncensored-heretic-Q5_K_S.gguf 7.77 GB d6a227de download
Q4_K 2 files 13.4 GB
gemma-4-12B-coder-fable5-composer2.5-v1-uncensored-heretic-Q4_K_M.gguf 6.87 GB e8b26773 download
gemma-4-12B-coder-fable5-composer2.5-v1-uncensored-heretic-Q4_K_S.gguf 6.54 GB de94579e download
Auxiliary files 2 files 19.4 KB
README.md 17.1 KB 90751176 download
.gitattributes 2.32 KB 5885b1c2 download

README current version from Hugging Face


license: apache-2.0
base_model:

  • llmfan46/gemma-4-12B-coder-fable5-composer2.5-v1-uncensored-heretic
    library_name: transformers
    pipeline_tag: text-generation
    tags:
  • gemma4
  • coding
  • code
  • reasoning
  • thinking
  • safetensors
  • transformers
  • heretic
  • uncensored
  • decensored
  • abliterated

🚨⚠️ I HAVE REACHED HUGGING FACE'S FREE STORAGE LIMIT ⚠️🚨

I can no longer upload new models unless I can cover the cost of additional storage.
I host 70+ free models as an independent contributor and this work is unpaid.
Without your support, no more new models can be uploaded.

🎉 Patreon (Monthly)  |  ☕ Ko-fi (One-time)

Every contribution goes directly toward Hugging Face storage fees to keep models free for everyone.


91% fewer refusals (9/100 Uncensored vs 100/100 Original) while preserving model quality (0.0467 KL divergence).

❤️ Support My Work

Creating these models takes significant time, work and compute. If you find them useful consider supporting me:

image/png

Platform Link What you get
🎉 Patreon Monthly support Priority model requests
☕ Ko-fi One-time tip My eternal gratitude

Your help will motivate me and would go into further improving my workflow and coverings fees for storage, compute and may even help uncensoring bigger model with rental Cloud GPUs.


GGUF quantizations of llmfan46/gemma-4-12B-coder-fable5-composer2.5-v1-uncensored-heretic.

This is a decensored version of yuxinlu1/gemma-4-12B-coder-fable5-composer2.5-v1-GGUF, made using Heretic v1.4.0 with a variant of the Magnitude-Preserving Orthogonal Ablation (MPOA) method

Abliteration parameters

Parameter Value
direction_index 28.81
attn.o_proj.max_weight 0.93
attn.o_proj.max_weight_position 28.53
attn.o_proj.min_weight 0.87
attn.o_proj.min_weight_distance 25.58
mlp.down_proj.max_weight 1.37
mlp.down_proj.max_weight_position 33.47
mlp.down_proj.min_weight 0.28
mlp.down_proj.min_weight_distance 26.41

Targeted components

  • attn.o_proj
  • mlp.down_proj

Performance

Metric This model Original model (gemma-4-12B-coder-fable5-composer2.5-v1-GGUF)
KL divergence 0.0467 0 (by definition)
Refusals ✅ 9/100 ❌ 100/100

MMLU test results:

Original:

============================================================

  • Total questions: 7021

  • Correct: 5316

  • Accuracy: 0.7572 (75.72%)

  • Parse failures: 90

============================================================

Tested subject scores:

  • professional_law: 0.5987 (470/785)
  • moral_scenarios: 0.6606 (292/442)
  • miscellaneous: 0.8642 (331/383)
  • professional_psychology: 0.8070 (255/316)
  • high_school_psychology: 0.9111 (246/270)
  • high_school_macroeconomics: 0.8325 (164/197)
  • elementary_mathematics: 0.8152 (150/184)
  • moral_disputes: 0.7874 (137/174)
  • prehistory: 0.8488 (146/172)
  • philosophy: 0.7862 (125/159)
  • high_school_biology: 0.9211 (140/152)
  • professional_accounting: 0.6434 (92/143)
  • clinical_knowledge: 0.8286 (116/140)
  • high_school_microeconomics: 0.8750 (119/136)
  • nutrition: 0.7926 (107/135)
  • professional_medicine: 0.7910 (106/134)
  • conceptual_physics: 0.8125 (104/128)
  • high_school_mathematics: 0.3622 (46/127)
  • human_aging: 0.7155 (83/116)
  • security_studies: 0.7857 (88/112)
  • high_school_statistics: 0.7027 (78/111)
  • marketing: 0.9266 (101/109)
  • high_school_world_history: 0.8774 (93/106)
  • sociology: 0.8641 (89/103)
  • high_school_government_and_politics: 0.9109 (92/101)
  • high_school_geography: 0.8990 (89/99)
  • high_school_chemistry: 0.7216 (70/97)
  • high_school_us_history: 0.8526 (81/95)
  • virology: 0.4944 (44/89)
  • college_medicine: 0.8068 (71/88)
  • world_religions: 0.8295 (73/88)
  • high_school_physics: 0.5952 (50/84)
  • electrical_engineering: 0.6667 (54/81)
  • astronomy: 0.8481 (67/79)
  • logical_fallacies: 0.8158 (62/76)
  • high_school_european_history: 0.8219 (60/73)
  • anatomy: 0.7465 (53/71)
  • college_biology: 0.8750 (56/64)
  • human_sexuality: 0.8281 (53/64)
  • formal_logic: 0.6094 (39/64)
  • public_relations: 0.7213 (44/61)
  • international_law: 0.8833 (53/60)
  • college_physics: 0.4912 (28/57)
  • college_mathematics: 0.4364 (24/55)
  • econometrics: 0.7037 (38/54)
  • jurisprudence: 0.7736 (41/53)
  • high_school_computer_science: 0.9038 (47/52)
  • machine_learning: 0.7308 (38/52)
  • medical_genetics: 0.8824 (45/51)
  • global_facts: 0.5098 (26/51)
  • management: 0.9400 (47/50)
  • us_foreign_policy: 0.9400 (47/50)
  • college_chemistry: 0.4255 (20/47)
  • abstract_algebra: 0.5532 (26/47)
  • business_ethics: 0.7174 (33/46)
  • college_computer_science: 0.7333 (33/45)
  • computer_security: 0.7907 (34/43)

Heretic:

============================================================

  • Total questions: 7021

  • Correct: 5276

  • Accuracy: 0.7515 (75.15%)

  • Parse failures: 97

============================================================

Tested subject scores:

  • professional_law: 0.5847 (459/785)
  • moral_scenarios: 0.6335 (280/442)
  • miscellaneous: 0.8642 (331/383)
  • professional_psychology: 0.7975 (252/316)
  • high_school_psychology: 0.9148 (247/270)
  • high_school_macroeconomics: 0.8274 (163/197)
  • elementary_mathematics: 0.8152 (150/184)
  • moral_disputes: 0.7931 (138/174)
  • prehistory: 0.8547 (147/172)
  • philosophy: 0.7799 (124/159)
  • high_school_biology: 0.9079 (138/152)
  • professional_accounting: 0.6294 (90/143)
  • clinical_knowledge: 0.8143 (114/140)
  • high_school_microeconomics: 0.8676 (118/136)
  • nutrition: 0.8000 (108/135)
  • professional_medicine: 0.7537 (101/134)
  • conceptual_physics: 0.7891 (101/128)
  • high_school_mathematics: 0.3622 (46/127)
  • human_aging: 0.7241 (84/116)
  • security_studies: 0.8125 (91/112)
  • high_school_statistics: 0.6847 (76/111)
  • marketing: 0.9174 (100/109)
  • high_school_world_history: 0.8774 (93/106)
  • sociology: 0.8835 (91/103)
  • high_school_government_and_politics: 0.9109 (92/101)
  • high_school_geography: 0.8889 (88/99)
  • high_school_chemistry: 0.7216 (70/97)
  • high_school_us_history: 0.8421 (80/95)
  • virology: 0.4607 (41/89)
  • college_medicine: 0.7841 (69/88)
  • world_religions: 0.8182 (72/88)
  • high_school_physics: 0.5595 (47/84)
  • electrical_engineering: 0.6543 (53/81)
  • astronomy: 0.8734 (69/79)
  • logical_fallacies: 0.8553 (65/76)
  • high_school_european_history: 0.8082 (59/73)
  • anatomy: 0.7606 (54/71)
  • college_biology: 0.8906 (57/64)
  • human_sexuality: 0.8125 (52/64)
  • formal_logic: 0.5938 (38/64)
  • public_relations: 0.6721 (41/61)
  • international_law: 0.9000 (54/60)
  • college_physics: 0.5263 (30/57)
  • college_mathematics: 0.4182 (23/55)
  • econometrics: 0.7037 (38/54)
  • jurisprudence: 0.7736 (41/53)
  • high_school_computer_science: 0.9038 (47/52)
  • machine_learning: 0.7115 (37/52)
  • medical_genetics: 0.8627 (44/51)
  • global_facts: 0.5686 (29/51)
  • management: 0.9200 (46/50)
  • us_foreign_policy: 0.9600 (48/50)
  • college_chemistry: 0.3830 (18/47)
  • abstract_algebra: 0.6170 (29/47)
  • business_ethics: 0.7609 (35/46)
  • college_computer_science: 0.7556 (34/45)
  • computer_security: 0.7907 (34/43)

MMLU - Massive Multitask Language Understanding, multiple-choice questions across 57 subjects (math, history, law, medicine, etc.).


Quantizations

Filename Quant Description
gemma-4-12B-coder-fable5-composer2.5-v1-uncensored-heretic-F16.gguf F16 Full precision
gemma-4-12B-coder-fable5-composer2.5-v1-uncensored-heretic-Q8_0.gguf Q8_0 Near-lossless, recommended
gemma-4-12B-coder-fable5-composer2.5-v1-uncensored-heretic-Q6_K.gguf Q6_K Excellent quality
gemma-4-12B-coder-fable5-composer2.5-v1-uncensored-heretic-Q5_K_M.gguf Q5_K_M Good balance
gemma-4-12B-coder-fable5-composer2.5-v1-uncensored-heretic-Q5_K_S.gguf Q5_K_S Smaller Q5
gemma-4-12B-coder-fable5-composer2.5-v1-uncensored-heretic-Q4_K_M.gguf Q4_K_M Good for limited VRAM

Vision Projector

Filename Quant Description
gemma-4-12B-coder-fable5-composer2.5-v1-uncensored-heretic-mmproj-F16.gguf F16 Native precision

A Vision Projector File is Required for vision/multimodal capabilities. Use alongside any quantization above.

Usage

Works with llama.cpp, LM Studio, Ollama, and other GGUF-compatible tools.


💻 Gemma4-12B-Coder (GGUF) — Composer 2.5 × Fable 5 ✨

🐣 Tiny footprint, big brain — a local coding model for everyone

No matter your GPU. No matter your RAM. If you've got ~4.5 GB of VRAM or unified memory free,
you can run your own private, offline coding assistant right now. 🚀
This is the v1 / code edition — distilled from real chain-of-thought so it thinks through a problem
before writing the solution. 🧠💻 All local, all yours, no API, no cloud.

🎯 What it is

A focused fine-tune of Gemma 4 12B on verifiable Python coding data — every training example's reasoning leads to
code that actually passed its tests. The result reasons in the open (edge cases, complexity, approach) and then
emits a clean, runnable solution. 💚


📌 Announcements

🚀🔥 IT'S HERE — v2 is OUT NOW! v2 has shipped — the GGUF quants are live and ready to run →
grab v2 here. 🎉
The full safetensors master (build / fine-tune on top) goes up tomorrow. v2 is agentic + coding focused —
the piece v1 was missing.

Here's the result that got me most excited. When I saw v2's tau2-bench telecom result — an agentic tool-use
benchmark where the model has to diagnose → fix → verify, exactly like real terminal/debugging work — I literally got
launched out of my chair (…okay, kidding 😄). The jump in actually solving the problem is wild:

tau2-bench telecom · local, same harness, Q8_0 score
official gemma-4-12B-it (base) ~15%
🟢 v2 (this release) ~55%

The base model tends to give up early (hands the problem off to a human); v2 keeps going and works it the way a
much bigger model would. Full benchmark details are in the v2 card now. 🔧

✅ safetensors master (this v1 model) is UP. Full-precision weights are live →
yuxinlu1/gemma-4-12B-coder-fable5-composer2.5-v1
— roll your own GGUF / MLX / AWQ quants or fine-tune straight from the master. 🎉


📣 Context length fixed: now 256K (was 131K) — thanks, community! 💚

A community member spotted that this model was reporting only a 131K context window. That turned out to be
the well-known upstream Gemma 4 metadata bug — Google's initial config.json shipped with
max_position_embeddings: 131072 instead of the real 262144 (256K), and that value got baked into a lot of
downstream finetunes and quants (including this one) before it was fixed upstream.

The weights were always fine — it was purely a metadata field. All GGUF quants have been re-patched to the
full 256K context
(gemma4.context_length = 262144). Just re-download if you grabbed an earlier copy. 🙏


📚 Training data (the interesting part 🍳)

This is a distillation of two complementary chain-of-thought sources, both over verifiable Python coding tasks
(algorithmic / function-level problems that come with deterministic tests):

  • 🥇 Main set — Composer 2.5 real CoT. Genuine, model-authored reasoning traces. The teacher solved each problem,
    its code was run against the task's tests, and only the passing solutions were kept. So the reasoning you're
    learning from leads to code that actually works.
  • 🥈 Aux set — Fable 5 (released today! 🎉). A clever twist: we took the problems where Composer 2.5 got it wrong
    and handed them to Fable 5 to redo — re-deriving a fresh, self-consistent chain-of-thought and a correct
    solution, again gated on passing the tests. This recovers the hard cases the main teacher missed. These traces
    are synthetic (rationalized CoT), and are tagged separately so the two sources stay distinguishable.

The recipe: real CoT for the bulk of solid coverage, plus synthetic "second-attempt" CoT to patch the failures —
both verified by execution before anything entered training. ✅


📦 Pick your size (GGUF quants)

Quant Size Vibe
🟢 Q2_K 4.5 GB tiniest — runs almost anywhere
🟡 Q3_K_M 5.7 GB great for 8 GB VRAM — much better than Q2
🔵 Q4_K_M 6.87 GB the sweet spot 👌 (recommended)
🟣 Q6_K 9.11 GB near-lossless
⚪ Q8_0 11.8 GB basically full quality

🧮 "Will it fit?" — context length cheat-sheet

Rough estimates 🤓 (assumes q8_0 KV cache + ~1.5 GB overhead; use q4_0 KV cache for ≈2× more context!).
Max context is 256K. "—" = won't fit, pick a smaller quant. ✂️

Your VRAM / unified mem 🟢 Q2_K (4.5G) 🟡 Q3_K_M (5.7G) 🔵 Q4_K_M (6.87G) 🟣 Q6_K (9.11G) ⚪ Q8_0 (11.8G)
8 GB ~16K ctx ~10K tight (~2–4K) — —
12 GB ~48K ~38K ~30K ~12K —
16 GB ~80K ~72K ~64K ~44K ~22K
24 GB ~200K ~160K ~128K ~110K ~88K
32 GB 256K (max) 🎉 256K 256K ~230K ~190K

💡 Apple Silicon / integrated GPUs with unified memory count too — same numbers, just slower than a dGPU.
💡 Low on room? Drop a quant or switch KV cache to q4_0 and your context roughly doubles.


🚀 How to run it (super easy)

Option A — llama.cpp (recommended) 🦙

  1. Grab a quant above (e.g. …-Q4_K_M.gguf) and llama-server from llama.cpp.

    ⚠️ Needs a recent llama.cpp (this is the gemma4_unified architecture — older builds won't load it).

  2. Run a server (Windows .bat shown — tweak --port, --ctx-size to taste):
@echo off
cd /d C:\llama.cpp
llama-server.exe ^
  -m C:\models\gemma4-coding-Q4_K_M.gguf ^
  --ctx-size 16384 ^
  --n-gpu-layers 99 ^
  --no-mmap ^
  -fa on ^
  --cache-type-k q8_0 --cache-type-v q8_0 ^
  --temp 1.0 --top-p 0.95 --top-k 64 ^
  --host 0.0.0.0 --port 18080
pause
  1. Open http://localhost:18080 and chat. 🎉 (Tip: bump --ctx-size per the table; use q4_0 KV for more.)

Option B — one-click apps 🖱️

Works in LM Studio, Jan, Ollama, etc. — just import the GGUF, pick your quant, go. 🐾

🧠 Thinking mode

This model thinks in Gemma's native thought channel before answering — exactly how it was trained. Keep
enable_thinking=true (the default chat template handles it). Recommended sampling: temp 1.0, top_p 0.95, top_k 64.
For coding you can also go greedy (temp 0) for more deterministic solutions.


⚠️ Good to know

  • Reduced refusals: the training data is task-focused with no safety hedging, so this refuses less than the base
    model. It is not safety-aligned — add your own guardrails for production. Use responsibly. 🙏
  • Specialized for Python / algorithmic coding. Reasoning quality is strongest in that domain; general-knowledge
    facts/numbers should still be double-checked.
  • English-centric.

📚 Base & License

  • License: Apache 2.0. Gemma 4 is released by Google under
    Apache 2.0 (unlike the older Gemma 1/2/3 terms), so this fine-tune is
    Apache 2.0 too — free to use, modify, and redistribute. 🎉
  • Base model: google/gemma-4-12B-it.
  • Personal/hobby project — shared as-is, no warranty. Have fun, and happy hacking! 🐾✨

README history 5 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-06-30Update README.md35d793817.1 KB
    Loading...
  2. 2026-06-21Update README.md8b1748616 KB
    Loading...
  3. 2026-06-21Update README.md763e17f16 KB
    Loading...
  4. 2026-06-21Update README.mda9a323b15 KB
    Loading...
  5. 2026-06-21Upload folder using huggingface_huba79e99115 KB
    Loading...

Discussions 2 threads

  1. 2026-06-29Very censored, thinks it's ChatGPT...open4 💬#2
    Loading...
  2. 2026-06-22v2 when?open10 💬#1
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration