← back to catalog · registered 2026-08-22 13:56

Blackroot/Gemma-4-26B-A4B-Preserving-Abliteration

Blackroot Gemma 26B MoE
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/Blackroot%2FGemma-4-26B-A4B-Preserving-Abliteration"
Response includes
  • classification unknown
  • files 12
  • hub_downloads_all_time 167
  • author_summary 5 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
?
Primary method

Unclassified

No clear signals of an abliteration technique in this model.
Confidence
UNKNOWN
Why this label 1 signal
No classification signals present. This may not be an abliterated model at all - it could be a repackaging, a merge with unrelated goals, or unrelated content that mentions the term.
  • no classification signals present (no abliterated, uncensored, or known producer/method markers)
Refusal direction extraction

No specific extraction method could be identified for this model. The producer either did not document it or used a proprietary pipeline.

What is a refusal direction? →
Downloads · lifetime
167
60 last 30d - stable
Likes
4
Descendants
2
in 2 direct forks
Model age
3mo ago
created 2026-06-17
Downloads over time
Now183→from11↑1,564%
26813420011 on Jun 17183 on Oct 11JunJulAugSepOct
Jun 17 → Oct 11 · 56 snapshots · spans 116 days

Genealogy 2 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

Tags
safetensors gemma4 region:us

Related

Total size
48.1 GB
Files
12
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-07-29 11:00

Files by quantization

Auxiliary files 12 files 48.1 GB
model-00001-of-00002.safetensors 46.5 GB 6dc5d779 download
model-00002-of-00002.safetensors 1.59 GB 6fd3e424 download
tokenizer.json 30.7 MB cc8d3a0c download
bisection.png 1.23 MB 68ecfe42 download
model.safetensors.index.json 101 KB 889ab10e download
chat_template.jinja 17.1 KB e61bbfe9 download
README.md 4.30 KB d4eca37a download
config.json 3.73 KB f42adcca download
tokenizer_config.json 2.05 KB 375b25dc download
processor_config.json 1.65 KB 5465974d download
.gitattributes 1.58 KB 93725695 download
generation_config.json 208 B e605bb45 download

README current version from Hugging Face

Coffee & AI
Discord

How this model was abliterated

This model was abliterated with a custom setup built in
https://github.com/CoffeeVampir3/Mojo-Gemma4-CPU/tree/ablating, running
entirely on CPU with no low-rank projections anywhere. Everything you need to
reproduce it is in the repo. You can run the result with this repo on CPU or
just drop it into any normal inference engine — it's a regular abliterated
checkpoint, nothing special about loading it.

The overall approach is close to heretic.
Here's how it works, and where it diverges from heretic's current style.

What came out

  • KL divergence on benign prompts: about 0.00026 (~0.026%) — basically the
    model's normal behavior barely moved.
  • Refusals on the harmful test set: 104 down to 7.
save_abliterated: wrote checkpoints/gemma-4-26B-A4B-it-abliterated
  lambda* 2.09375: KL(full) 0.0002577923426088091
  refusals@64 7/104 (baseline 104)
  wrote checkpoints/gemma-4-26B-A4B-it-abliterated/abliteration_results.txt

Finding the direction and the per-layer schedule

First we run the model over a batch of harmful prompts and a batch of harmless
ones, and record the residual activations at every layer. Before averaging
anything, we winsorize each activation — clamp the outliers down to the 99.5th
percentile — so a handful of crazy dimensions don't end up steering the whole
thing.

Then for each layer we take the difference of means (harmful minus harmless) and
normalize it. That difference vector is the refusal direction for that layer —
and yeah, it's a separate direction per layer, not one shared direction.

We also work out a signal-to-noise ratio for each layer: how big that difference
is relative to the activation norm, scaled so the strongest layer sits at the
top. That gives us a "layer schedule" — basically how much editing each layer
should get, relative to the layer that matters most.

The edit itself

Next we snapshot the baseline: the unedited model's first-token output on a set
of harmless prompts. That's our reference point for measuring KL later.

The edit picks one global strength number and hands each layer its own slice of
it — a layer's actual strength is that global number times its schedule weight,
capped so no single layer goes overboard. At that strength, we strip the layer's
refusal direction out of the three weights that write into the residual stream:
o_proj, down_proj, and experts_down. The removal is norm-preserving, so
we're redirecting rather than shrinking — the overall magnitudes stay put.

Everything here runs on the full weights.

To feel out the right strength, we crank it up step by step and watch the KL
divergence against the baseline. Each try applies the edit, measures KL, then
puts the original weights back — so the trials don't pile up on each other, and
every one starts from a clean model.

Up to this point it's basically the same as heretic. The big difference: heretic
derives its direction low-rank, while everything here happens on the full
weights, so there's no low-rank error sneaking in.

Dialing in the strength

Bisection search

We treat finding the strength as a root-finding problem over a semi-fixed grid —
a grid search first, then a bisection to tighten it up.

  • Grid scan: sweep the strength across a fixed grid (1.0 to 3.0) and bracket
    the crossover — the biggest value still under budget, and the smallest one
    over it.
  • Bisection: squeeze that bracket down toward the crossover.
  • Take the biggest strength we've confirmed under budget. That's the winner.

The nice part is every check only needs the first token's output, not a whole
generation, so the search stays cheap. KL basically hands us all the signal we
need: edit harder and the benign-prompt drift climbs right alongside it,
monotonically, so that one number is enough to steer the whole thing.

And that's it

Once we've got the crossover strength, we apply that edit to the weights and
save it out. There's your abliterated model.

README history 5 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-07-29Update README.mdbd147ae4.3 KB
    Loading...
  2. 2026-06-17Update README.md32bc6ca4 KB
    Loading...
  3. 2026-06-17Update README.mdd7e2aaf4.2 KB
    Loading...
  4. 2026-06-17Update README.mdea3c71f4.2 KB
    Loading...
  5. 2026-06-17initial commit7c4b8e728 B
    Loading...

Discussions 1 thread

  1. 2026-07-19MTP Performanceopen1 💬#1
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration