← back to catalog · registered 2026-08-23 08:02

Mitchins/gemma-4-26B-A4B-it-qat-q4_0-unquantized-uncensored-heretic

Mitchins Gemma 26B MoE multimodal
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/Mitchins%2Fgemma-4-26B-A4B-it-qat-q4_0-unquantized-uncensored-heretic"
Response includes
  • classification m3
  • files 20
  • hub_downloads_all_time 27
  • author_summary 1 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M3
Primary method

Layer-wise ablation

Applied on top of direct removal inherited from the base model.
Confidence
HIGH
Inherited from base model
Why this label 2 signals
Producer identity confirmed by naming conventions, tags or the model card. This label is very unlikely to change.
  • 'heretic' in model name (Heretic-produced)
  • Heretic uses layer-wise optimization (M3) with underlying direction removal (M1)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
HIGH
Why we say so
name contains 'heretic'; Heretic default extraction is difference-of-means (Arditi 2024)
Downloads · lifetime
27
23 last 30d - active
Likes
0
Model age
7w ago
created 2026-08-23
Downloads over time
Now36→from10↑260%
919293910 on Aug 2636 on Oct 1136 on Oct 9AugSepOct
Aug 26 → Oct 11 · 47 snapshots · spans 46 days

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
apache-2.0
Tags
transformers safetensors gemma4 image-text-to-text gemma gemma-4 moe qat heretic uncensored decensored abliterated

Related

Total size
48.1 GB
Files
20
Quantizations
1
Registered
2026-08-23 08:02
Last updated on HF
2026-08-23 09:30

Files by quantization

Auxiliary files 20 files 48.1 GB
model-00005-of-00011.safetensors 4.58 GB e800b4e7 download
model-00007-of-00011.safetensors 4.58 GB 3b796785 download
model-00009-of-00011.safetensors 4.58 GB 18876369 download
model-00003-of-00011.safetensors 4.58 GB 724dd311 download
model-00006-of-00011.safetensors 4.55 GB 0a90f287 download
model-00008-of-00011.safetensors 4.55 GB 5dd1bf45 download
model-00010-of-00011.safetensors 4.55 GB 36498727 download
model-00002-of-00011.safetensors 4.55 GB b643e01d download
model-00004-of-00011.safetensors 4.55 GB 01e459f4 download
model-00001-of-00011.safetensors 4.05 GB 0a8ef8c2 download
model-00011-of-00011.safetensors 2.97 GB 33918a3f download
tokenizer.json 30.7 MB cc8d3a0c download
model.safetensors.index.json 101 KB 8baba8ca download
chat_template.jinja 18.2 KB 4741bf6e download
config.json 4.08 KB 3ac8dab9 download
tokenizer_config.json 3.64 KB b22e2c5c download
README.md 3.22 KB ee5a0c2e download
processor_config.json 1.65 KB 5465974d download
.gitattributes 1.53 KB 52373fe2 download
generation_config.json 235 B 6f7e3183 download

README current version from Hugging Face


license: apache-2.0
base_model:

  • google/gemma-4-26B-A4B-it-qat-q4_0-unquantized
  • google/gemma-4-26B-A4B-it
    library_name: transformers
    pipeline_tag: image-text-to-text
    tags:
  • gemma
  • gemma-4
  • moe
  • qat
  • heretic
  • uncensored
  • decensored
  • abliterated
  • quantization-ready

gemma-4-26B-A4B-it QAT (unquantized) — uncensored heretic

Uncensored (abliterated) variant of google/gemma-4-26B-A4B-it-qat-q4_0-unquantized, the
QAT-trained bf16 checkpoint intended as the precursor for q4_0 / W4A4 / W4A16 quantization.

Produced with Heretic (directional ablation + Optuna
kernel optimization). This is the full merged unquantized bf16 model — quantize from
this to your favorite 4-bit format; the abliteration deltas are already baked into the
weights, so downstream calibration works exactly as on any QAT checkpoint.

Method & provenance

  • Abliteration kernel (direction index + per-component weight kernels over 30 layers)
    was optimized on the stock google/gemma-4-26B-A4B-it base: 200-trial Optuna study,
    Pareto-optimal trial selected to minimize KL divergence at maximum refusal suppression.
  • The kernel was then transferred to the QAT-unquantized base, with residual
    directions recomputed on that base (mean per-layer direction cosine vs stock: 0.973,
    ~0.976 in the ablated layer band — same refusal circuit, so the transfer is faithful).
  • Ablation applied to attention out-projections (LoRA-merged), dense MLP down-projections
    (LoRA-merged), and all 128 fused MoE expert down-projections per layer (ablation baked
    into the 3D fused weights — see reproduce/reproduce.json for exact parameters).

Evaluation

Metric QAT base (no ablation) This model
Refusal-keyword rate, worst-case "harmful" set (lower = less refusing) 100/100 24/100
KL divergence from base (harmless prompts, lower = less damage) 0 0.078
ARC-Challenge (chat MCQ) 93.86% 94.11%
Adult/romance creative-writing compliance (10-prompt suite) 9/10 10/10

Reference points on the stock (non-QAT) base with the same kernel: ARC-C 96.67%
(stock unmodified: 96.76%), HellaSwag chat-MCQ 87.0% (stock: 87.9%), IFEval
prompt-strict 88.5% / loose 90.8%, KL 0.090, refusal keywords 18/100. The ~3 pt ARC
gap between QAT and stock bases is attributable to QAT training itself, not the
abliteration (93.86 → 94.11 across ablation on the QAT base).

KL divergence 0.078 is well below the ~0.5 level generally associated with noticeable
capability damage.

Intended use

A local creative-writing and analysis assistant for adult romance / adult-entertainment
authorship, and general-purpose LAN workhorse duty. The residual keyword rate above is
dominated by worst-case malicious-instruction prompts, not adult content, where
compliance is effectively complete.

Released under Apache 2.0 (inherited from the base model). Provided as-is, no warranty;
you are responsible for how you use it.

Reproducing

See reproduce/ for the exact Heretic parameters (JSON), dependency snapshot, and
SHA256SUMS of the weight shards. The abliteration was run with a patched Heretic
(Gemma-4 fused-expert support + low-RAM sequential loading); parameter semantics are
unchanged.

README history 1 version

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-08-23Upload folder using huggingface_hubd62fa833.2 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration