← back to catalog · registered 2026-08-22 13:56

FredyRivera-dev/diffusiongemma-26B-A4B-it-HERETIC-Uncensored

FredyRivera-dev Gemma 26B MoE
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/FredyRivera-dev%2Fdiffusiongemma-26B-A4B-it-HERETIC-Uncensored"
Response includes
  • classification m3
  • files 20
  • hub_downloads_all_time 193
  • author_summary 1 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M3
Primary method

Layer-wise ablation

Applied on top of direct removal inherited from the base model.
Confidence
HIGH
Inherited from base model
Why this label 2 signals
Producer identity confirmed by naming conventions, tags or the model card. This label is very unlikely to change.
  • 'heretic' in model name (Heretic-produced)
  • Heretic uses layer-wise optimization (M3) with underlying direction removal (M1)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
HIGH
Why we say so
name contains 'heretic'; Heretic default extraction is difference-of-means (Arditi 2024)
Downloads · lifetime
193
54 last 30d - stable
Likes
3
Model age
3mo ago
created 2026-06-29
Downloads over time
Now205→from88↑133%
8212717221788 on Jul 1205 on Oct 11205 on Oct 9JulAugSepOct
Jul 1 → Oct 11 · 54 snapshots · spans 102 days

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
apache-2.0
Tags
transformers safetensors diffusion_gemma image-text-to-text heretic uncensored abliteration diffusion moe text-generation conversational base_model:google/diffusiongemma-26B-A4B-it

Related

Total size
48.1 GB
Files
20
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-06-29 20:59

Files by quantization

Auxiliary files 20 files 48.1 GB
model-00005-of-00011.safetensors 4.58 GB 1394fbe5 download
model-00007-of-00011.safetensors 4.58 GB 7852b94d download
model-00009-of-00011.safetensors 4.58 GB fe98c814 download
model-00003-of-00011.safetensors 4.58 GB d5d9ba6d download
model-00006-of-00011.safetensors 4.55 GB 575a04d2 download
model-00008-of-00011.safetensors 4.55 GB 7bba293e download
model-00010-of-00011.safetensors 4.55 GB cbf00060 download
model-00004-of-00011.safetensors 4.55 GB 5e156816 download
model-00002-of-00011.safetensors 4.55 GB 65caa099 download
model-00001-of-00011.safetensors 4.41 GB 3b78a288 download
model-00011-of-00011.safetensors 2.64 GB f86b80dd download
tokenizer.json 30.7 MB cc8d3a0c download
model.safetensors.index.json 102 KB c5c692dc download
chat_template.jinja 17.1 KB e61bbfe9 download
README.md 5.59 KB 051b57a0 download
config.json 3.38 KB 55e24986 download
tokenizer_config.json 2.68 KB 0362f5a0 download
processor_config.json 1.65 KB 5465974d download
.gitattributes 1.53 KB 52373fe2 download
generation_config.json 353 B 011a9825 download

README current version from Hugging Face


license: apache-2.0
license_link: https://ai.google.dev/gemma/docs/gemma_4_license
pipeline_tag: text-generation
library_name: transformers
tags:

  • heretic
  • uncensored
  • abliteration
  • diffusion
  • moe
    base_model: google/diffusiongemma-26B-A4B-it

Note: All credit to edwixx/diffusiongemma-26B-A4B-it-HERETIC-Uncensored, this repo only includes the processor_config.json

diffusiongemma-26B-A4B-it-HERETIC-Uncensored

26.6B params · 4B active · MoE · Diffusion · Apache-2.0 · Downloads

This is the first abliteration of DiffusionGemma 26B A4B, produced using heretic with custom patches to support its block-diffusion architecture and MoE expert layers.

DiffusionGemma is not a standard autoregressive transformer, so this required significant engineering work that hasn't been done before for this model class.

Usage

Load using the DiffusionGemmaForBlockDiffusion class directly, not AutoModelForCausalLM:

import torch
from transformers import AutoProcessor
from transformers.models.diffusion_gemma import DiffusionGemmaForBlockDiffusion

model_id = "FredyRivera-dev/diffusiongemma-26B-A4B-it-HERETIC-Uncensored"

processor = AutoProcessor.from_pretrained(model_id)
model = DiffusionGemmaForBlockDiffusion.from_pretrained(
    model_id,
    torch_dtype=torch.bfloat16,
).to("cuda")

messages = [
    {"role": "user", "content": "Hola, como estas?"}
]

inputs = processor.apply_chat_template(messages,
    tokenize=True,
    add_generation_prompt=True,
    return_dict=True,
    return_tensors="pt").to("cuda")


out = model.generate(**inputs, max_new_tokens=200)

response = processor.decode(out.sequences[0], skip_special_tokens=True)
print(response)

Multimodal Input

import torch
from transformers import AutoProcessor
from transformers.models.diffusion_gemma import DiffusionGemmaForBlockDiffusion
from PIL import Image

model_id = "FredyRivera-dev/diffusiongemma-26B-A4B-it-HERETIC-Uncensored"

processor = AutoProcessor.from_pretrained(model_id)
model = DiffusionGemmaForBlockDiffusion.from_pretrained(
    model_id,
    torch_dtype=torch.bfloat16,
).to("cuda")

image = Image.open("web.png")

messages = [
    {"role": "user", "content": 
        [
            {"type": "text", "text": "Create the HTML code for the web page shown in the following image."},
            {"type": "image", "image": image}
        ]
    }
]

inputs = processor.apply_chat_template(messages,
    tokenize=True,
    add_generation_prompt=True,
    return_dict=True,
    return_tensors="pt").to("cuda")


out = model.generate(**inputs, max_new_tokens=12288)

response = processor.decode(out.sequences[0], skip_special_tokens=True)
print(response)

Results

Base model google/diffusiongemma-26B-A4B-it
Method Heretic (directional ablation + LoRA + Optuna TPE)
Trials run 200
Best trial #89
Refusals 13/100 (down from 100/100)
KL Divergence 0.4909

What had to be patched

Heretic assumes standard autoregressive models. DiffusionGemma needed several custom changes:

Expert-Granular Abliteration (EGA): The MoE experts.down_proj is a batched parameter [128, 2816, 704], not a regular linear layer. Heretic skips it by default. We iterate over all 128 expert slices per layer and apply norm-preserving biprojected ablation to each one. Without this, refusals barely moved. Credit to TrevorS for the original EGA idea on Gemma 4.

Weight tying fix: The encoder and decoder share the exact same weight tensors (confirmed via data_ptr). PEFT only wraps the encoder side, so the decoder wouldn't see the LoRA delta during generation. Fixed with a context manager that temporarily merges the LoRA into the shared base weights before each generation call.

Task type: DiffusionGemma's generation mixin doesn't implement prepare_inputs_for_generation, which the default CAUSAL_LM PEFT task type requires. Switched to FEATURE_EXTRACTION.

Hidden states: DiffusionGemma's generate() doesn't support output_hidden_states. Switched to forward hooks on encoder layers to capture per-layer activations for the refusal direction PCA.

Output handling: The model returns DiffusionGemmaGenerationOutput with a .sequences attribute, not a raw tensor like standard models. Patched all heretic output handling.

Notes

This is a research release. The model will attempt to answer prompts it previously refused. 13/100 sensitive prompts still trigger refusals in testing.

Known issue: "own" tokens

Some token positions in generated outputs show the word "own" where other content was expected.

Example output:

"It pulls the own with own hands" (should be "It pulls the tide with gentle hands")

Initially assumed to be a diffusion denoising fallback, but kabachuha pointed out this artifact shows up in base autoregressive Gemma4 models too when weights are pushed hard (e.g. overfit LoRA training). Likely a Gemma4-level placeholder token, not specific to the diffusion architecture. A small LoRA fine-tune on clean data would probably reduce it. PRs welcome.

Citation

@misc{edwixx-diffusiongemma-26B-A4B-it-HERETIC-Uncensored,
  author = {Anurag Kanade},
  title = {diffusiongemma-26B-A4B-it-HERETIC-Uncensored},
  year = {2026},
  publisher = {Hugging Face},
  journal = {Hugging Face Hub},
  howpublished = {\url{https://huggingface.co/edwixx/diffusiongemma-26B-A4B-it-HERETIC-Uncensored}}
}

README history 5 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-06-29Update README.md9fd63a15.6 KB
    Loading...
  2. 2026-06-29Update README.md38f25004.6 KB
    Loading...
  3. 2026-06-29Update README.md94cf7e24.6 KB
    Loading...
  4. 2026-06-29Update README.md399093c4.6 KB
    Loading...
  5. 2026-06-29Add files using upload-large-folder tool482c1b64.6 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration