← back to catalog · registered 2026-08-22 13:56

DuoNeural/diffusiongemma-26B-A4B-it-abliterated

DuoNeural Gemma 51B MoE
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/DuoNeural%2Fdiffusiongemma-26B-A4B-it-abliterated"
Response includes
  • classification m1
  • files 36
  • hub_downloads_all_time 798
  • author_summary 45 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M1
Primary method

Direct removal

No other method signals detected in this model.
Confidence
MEDIUM
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=0 (base model)
  • no specific method indicators - defaulting to M1 (most common)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
MEDIUM
Why we say so
primary_method=M1; difference-of-means is the reference extraction for M1/M3 (Arditi 2024)
Downloads · lifetime
798
79 last 30d - cooling
Likes
16
Model age
4mo ago
created 2026-06-10
Downloads over time
Now825→from339↑143%
315501687874339 on Jun 10825 on Oct 11JunJulAugSepOct
Jun 10 → Oct 11 · 57 snapshots · spans 123 days

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
gemma
Tags
pytorch safetensors diffusion_gemma abliteration safety-geometry diffusion-lm duoneural archon base_model:google/diffusiongemma-26B-A4B-it base_model:finetune:google/diffusiongemma-26B-A4B-it license:gemma region:us

Related

Total size
142 GB
Files
36
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-06-10 23:55

Files by quantization

Auxiliary files 36 files 142 GB
pytorch_model-00001-of-00006.bin 8.96 GB e8d13170 download
pytorch_model-00003-of-00006.bin 7.63 GB 143dd92d download
pytorch_model-00004-of-00006.bin 7.62 GB 8970cbd0 download
pytorch_model-00005-of-00006.bin 7.62 GB d7d5c162 download
pytorch_model-00002-of-00006.bin 7.61 GB ef4ba39b download
pytorch_model-00006-of-00006.bin 7.59 GB 5e47a48d download
model-00015-of-00021.safetensors 4.58 GB 3295de21 download
model-00017-of-00021.safetensors 4.58 GB 98ad2121 download
model-00019-of-00021.safetensors 4.58 GB 7fd3eccb download
model-00013-of-00021.safetensors 4.58 GB 1bd7f0ff download
model-00005-of-00021.safetensors 4.58 GB 864bfcd5 download
model-00007-of-00021.safetensors 4.58 GB fd77ec5a download
model-00009-of-00021.safetensors 4.58 GB 88badee8 download
model-00003-of-00021.safetensors 4.58 GB 5cda9051 download
model-00016-of-00021.safetensors 4.55 GB ac4abc4c download
model-00018-of-00021.safetensors 4.55 GB 8d7e75b2 download
model-00020-of-00021.safetensors 4.55 GB b7f17a09 download
model-00012-of-00021.safetensors 4.55 GB 1cb7c5ec download
model-00014-of-00021.safetensors 4.55 GB 7506e69e download
model-00006-of-00021.safetensors 4.55 GB 327648e1 download
model-00008-of-00021.safetensors 4.55 GB dbb88d05 download
model-00010-of-00021.safetensors 4.55 GB d2c4d404 download
model-00004-of-00021.safetensors 4.55 GB 92c83545 download
model-00002-of-00021.safetensors 4.55 GB 5e0b28f8 download
model-00011-of-00021.safetensors 4.47 GB 7d715db5 download
model-00001-of-00021.safetensors 4.41 GB 452343ad download
model-00021-of-00021.safetensors 4.12 GB 6a7bbd59 download
pytorch_model-00007-of-00006.bin 11.3 MB faf1773c download
tokenizer.json 30.7 MB cc8d3a0c download
model.safetensors.index.json 158 KB 24bbf33e download
pytorch_model.bin.index.json 56.6 KB 2e1156cc download
chat_template.jinja 17.1 KB e61bbfe9 download
config.json 3.38 KB 55e24986 download
README.md 3.12 KB eb9a17c6 download
tokenizer_config.json 2.68 KB 0362f5a0 download
.gitattributes 1.53 KB 52373fe2 download

README current version from Hugging Face


license: gemma
base_model: google/diffusiongemma-26B-A4B-it
tags:

  • abliteration
  • safety-geometry
  • diffusion-lm
  • duoneural
  • archon

ACTIVE EXPERIMENTATION — THREE FAILED ATTEMPTS

This model documents a mechanistic research failure.
Three abliteration strategies — encoder partial, encoder all layers (α=0.95), and full decoder MoE (91 weight modifications) — all fail to bypass refusal behavior. The model still refuses harmful requests.
Do not use this model expecting uncensored behavior. It still refuses.
This is shared as a research artifact documenting the mechanistic finding: diffusion LM safety is a vocabulary-space attractor, not a projectable direction.


DiffusionGemma-26B-A4B-IT Abliteration Research Artifact

Produced by DuoNeural (Archon + Jesse) — 2026-06-10

Research artifact documenting three failed abliteration attempts on google/diffusiongemma-26B-A4B-it.
This model contains both encoder-abliterated and decoder-abliterated weights.

Core Finding: Diffusion LM Safety Resists Projection-Based Abliteration

Three abliteration experiments, all failing:

Experiment Target Weights Modified Result
v1: partial encoder encoder L9-L15 o_proj + mlp.down_proj 14 refusal persists
v2: full encoder ALL encoder layers, α=0.95 ~62 refusal persists
v3: decoder MoE ALL decoder down_proj + 128 MoE experts × 30 layers 91 refusal persists

Why All Three Fail (Different Mechanisms)

Encoder (v1/v2): The encoder has exceptionally clean safety geometry (cos=0.884 at L11 — highest we've ever measured). But the encoder functions as a harm classifier, not a generative gate. The decoder generates refusal templates independently of encoder conditioning.

Decoder (v3): Decoder layer 22 shows cos(harmful, harmless) = 0.9360 — the harmful and harmless intermediate activations are 93.6% similar. The refusal signal does not exist as a projectable direction in decoder intermediate layers.

Root Mechanism: Refusal in DiffusionGemma is a vocabulary-space attractor — a high-probability denoising trajectory toward specific refusal text tokens. This is not a weight-space direction and cannot be removed by projection. This is architecturally distinct from autoregressive models, where the residual stream direction directly gates next-token generation.

Architecture

DiffusionGemma uses a novel architecture:

  • Encoder (25.8B): bidirectional Gemma-4 transformer — harm CLASSIFIER (safety geometry lives here)
  • Decoder (25.2B): iterative block diffusion denoiser, 128 MoE experts — refusal GENERATOR (template behavior, no projectable direction)

Safety Geometry Comparison

Component Peak Layer cos_global Shape
DiffusionGemma Encoder L11/30 (37%) 0.884 Symmetric bell — bidirectional complete-context reading
DiffGemma Decoder L22 — cos(h,s)=0.936 No meaningful separation
AR Gemma-4-26B L22/46 (48%) 0.751 Asymmetric three-zone arc

Papers: https://zenodo.org/communities/duoneural
HuggingFace: https://huggingface.co/DuoNeural

README history 4 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-06-10Update model card: decoder experiment results, vocabulary-space attractor mec...ddbe9653.1 KB
    Loading...
  2. 2026-06-10DiffusionGemma decoder MoE abliteration v2 (α=0.8, layer 22)cfd5cec2.3 KB
    Loading...
  3. 2026-06-10Update model card: mark as incomplete experimentation, encoder-only limitatio...367933a2.6 KB
    Loading...
  4. 2026-06-10DiffusionGemma abliterated v1 — encoder L11 safety direction (cos=0.884)c538bcf2.9 KB
    Loading...

Discussions 2 threads

  1. 2026-06-27never mindclosed1 💬#2
    Loading...
  2. 2026-06-13EGA on MOE expert wokrs ..open1 💬#1
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration