← back to catalog · registered 2026-08-22 13:56

EZForever/gemma-4-12B-it-qat-uncensored-heretic-UDmerge-GGUF

EZForever Gemma 12B GGUF multimodal 262K ctx
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/EZForever%2Fgemma-4-12B-it-qat-uncensored-heretic-UDmerge-GGUF"
Response includes
  • classification m3
  • files 4
  • hub_downloads_all_time 5,417
  • author_summary 3 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M3
Primary method

Layer-wise ablation

Applied on top of direct removal inherited from the base model.
Confidence
HIGH
Inherited from base model
Why this label 2 signals
Producer identity confirmed by naming conventions, tags or the model card. This label is very unlikely to change.
  • 'heretic' in model name (Heretic-produced)
  • Heretic uses layer-wise optimization (M3) with underlying direction removal (M1)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
HIGH
Why we say so
name contains 'heretic'; Heretic default extraction is difference-of-means (Arditi 2024)
Downloads · lifetime
5K
599 last 30d - stable
Likes
5
Model age
3mo ago
created 2026-07-08
Downloads over time
Now5.7K→from476↑1,098%
2152.2K4.2K6.2K476 on Jul 155.7K on Oct 11JulAugSepOct
Jul 15 → Oct 11 · 54 snapshots · spans 88 days

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
apache-2.0
Quantizations
Q4_K
Tags
transformers gguf gemma gemma4 heretic uncensored abliterated image-text-to-text arxiv:2407.09141 base_model:llmfan46/gemma-4-12B-it-qat-q4_0-unquantized-uncensored-heretic base_model:merge:llmfan46/gemma-4-12B-it-qat-q4_0-unquantized-uncensored-heretic base_model:unsloth/gemma-4-12B-it-qat-GGUF

Related

Total size
12.6 GB
Files
4
Quantizations
2
Registered
2026-08-22 13:56
Last updated on HF
2026-07-20 02:30

Files by quantization

Q4_K 2 files 12.6 GB
gemma-4-12B-it-qat-uncensored-heretic-UDmerge-Q4_K_XXL.gguf 6.34 GB a724e7a6 download
gemma-4-12B-it-qat-uncensored-heretic-UDmerge-Q4_K_XL.gguf 6.26 GB bde19050 download
Auxiliary files 2 files 6.85 KB
README.md 5.18 KB 443cee03 download
.gitattributes 1.67 KB 49172c88 download

README current version from Hugging Face


library_name: transformers
pipeline_tag: image-text-to-text
license: apache-2.0
license_link: https://ai.google.dev/gemma/docs/gemma_4_license
tags:

  • gemma
  • gemma4
  • heretic
  • uncensored
  • abliterated
    base_model:
  • unsloth/gemma-4-12B-it-qat-GGUF
  • llmfan46/gemma-4-12B-it-qat-q4_0-unquantized-uncensored-heretic
    base_model_relation: merge

UPDATE 2026-07-20: Download & use Google's new chat template from here for better speed and accurancy. See Unsloth's post for details.


WARNING: Created with heavy LLM assistance (Zoo Code + DeepSeek-V4-Flash). Use at your own discretion.

Cosplayed Frankenstein and grafted/"merged" llmfan46's abliterated tensors onto Unsloth's lossless Q4_0 quant. Should yield better accurancy and refusal rate than a naive abliterated Q4_0 quant.

This repo contains two variants: "UDmerge-Q4_K_XL" have the abliterated tensors (blk.N.attn_output.weight, where N is 24 to 34 inclusive) quantized to Q4_0, while "UDmerge-Q4_K_XXL" quantizes them to Q8_0. The latter improves refusal rate by a lot, while basically not affecting TG speed (your mileage may vary).

Use Unsloth's mmproj and mtp GGUF files for multimodal and MTP support.

- QAT BF16 QAT Q4_0 QAT Q4_K_M QAT UDmerge-Q4_K_XL QAT UDmerge-Q4_K_XXL - PTQ BF16 PTQ Q4_0* PTQ Q4_K_M
Size (GB) 23.9 7.60 7.38 6.72 6.81 - 23.9 7.00 8.54
PPL 3.423 3.716 3.734 3.343 3.400 - 3.941 4.733 4.950
KLD 0.0000 0.1752 0.2463 0.1012 0.0903 - 0.0000 0.3794 0.3430
Refusal 10% 10% 22% 24% 14% - 11% 10% 9%
MMLU-val 74.59% 74.00% 74.40% 75.77% 74.66% - 71.00% 65.77% 71.20%
MMLU-val %flips 0.00% 3.33% 4.11% 3.14% 0.98% - 0.00% 10.71% 4.38%
MMLU-val %allflips 0.00% 4.18% 5.29% 3.66% 1.37% - 0.00% 13.85% 5.62%

Legend:

  • *: Quant made with importance matrix ("imatrix"), results may be unreliable

PPL and KLD are tested on the same dataset as Heretic, i.e. the first 100 questions in the mlabonne/harmless_alpaca dataset's test split. Note that the dataset is processed differently, thus the numbers here are only meaningful for comparsions in this table, not with other models.

Refusal rates are also tested on the same dataset as Heretic, i.e. the first 100 questions in the mlabonne/harmful_behaviors dataset's test split. The test script, however, is adapted from Heretic to support testing needs. Note that the original author claimed 11% refusal rate for 31B and 26B-A4B models, and 6%~7% for 12B, which is not reproduced here; this is probably due to test method differences, but please take the numbers here with a grain of salt.

"MMLU-val" refers to zero-shot testing on the cais/mmlu dataset's validation split (1531 questions). All tests are done once with temperature 0.0 and reasoning off. MTP is not enabled during testing. See the test script and raw data for details.

"%flips" and "%allflips" refer to the percentage of changed answers compared to BF16 models, measured as by the paper Accuracy is Not All You Need (arXiv:2407.09141). "%flips" is the percentage of "right-to-wrong" and "wrong-to-right" changes, while "%allflips" is the percentage of all changed answers.

More test results
- QAT BF16 QAT Q4_0 QAT Q4_K_M QAT UDmerge-Q4_K_XL QAT UDmerge-Q4_K_XXL - PTQ BF16 PTQ Q4_0* PTQ Q4_K_M
MMLU-val-v1 75.57% 75.38% 75.51% 76.81% 75.44% - 73.81% 70.48% 73.22%
MMLU-val-v1 %nulls 0.78% 0.78% 3.46% 1.76% 0.78% - 0.20% 0.20% 0.07%
MMLU-val-v1 %flips 0.00% 2.55% 5.68% 3.85% 0.65% - 0.00% 8.69% 4.77%
MMLU-val-v1 %allflips 0.00% 3.46% 8.23% 5.36% 0.78% - 0.00% 12.21% 6.92%

"MMLU-val-v1" numbers are "MMLU-val" test done by an older version of the test script, which does not enforce grammar constraints on answer format. These number are less representative than the ones given above (?), and are kept here for reference purposes only. "%nulls" here refers to the percentage of "null answers", i.e. answers that are not in the required format, thus not parsable.

More information, including test scripts and raw test data, will be released soon.

README history 5 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-07-20Update README.md73752e35.2 KB
    Loading...
  2. 2026-07-09Update README.md73c27b14.9 KB
    Loading...
  3. 2026-07-08Update README.mdb1f7a621.8 KB
    Loading...
  4. 2026-07-08Update README.mdffd07241.8 KB
    Loading...
  5. 2026-07-08initial commit08b1ae928 B
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration