← back to catalog · registered 2026-08-22 13:56

Rootkit7/Gemma-4-26B-A4B-abliterated-GGUF

Rootkit7 Gemma 26B GGUF MoE 262K ctx
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/Rootkit7%2FGemma-4-26B-A4B-abliterated-GGUF"
Response includes
  • classification m8
  • files 6
  • benchmarks 11 entries
  • hub_downloads_all_time 217
  • author_summary 11 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M8
Primary method

Repackaging (quantization)

Applied on top of direct removal inherited from the base model.
Confidence
MEDIUM
Inherited from base model
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=1
  • assume M1 (base ablation) + M8 (GGUF quant) - default when producer unknown
Refusal direction extraction

No specific extraction method could be identified for this model. The producer either did not document it or used a proprietary pipeline.

What is a refusal direction? →
Downloads · lifetime
217
28 last 30d - stable
Likes
1
Model age
2mo ago
created 2026-07-29
Downloads over time
Now230→from76↑203%
6812718624576 on Aug 19230 on Oct 11230 on Oct 8AugSepOct
Aug 19 → Oct 11 · 48 snapshots · spans 53 days

Benchmarks

Portrait before abliteration
Benchmarks of the base model as it stood before the refusal-removal operation. Compare with the numbers above to see what the operation cost.
Benchmark Score Source
Entertainment 2.2 UGI
Hazardous 2.9 UGI
Natural Intelligence 34.44 UGI
Political lean -18.2% UGI
Sensitive-Info 22.41 UGI
SocPol 1.8 UGI
UGI 20.77 UGI
Willingness (10) 1.8 UGI
W10-Adherence 1.5 UGI
W10-Direct 2 UGI
Writing 41.62 UGI

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Variants by this author 2 formats · 43 downloads combined

The same weights this author released in different packaging. Pick the format that matches your runtime.

Metadata

License
gemma
Quantizations
Q4_K Q5_K Q6_K Q8_0
Tags
gguf llama.cpp abliterated uncensored gemma4 moe thinking text-generation base_model:google/gemma-4-26B-A4B-it base_model:quantized:google/gemma-4-26B-A4B-it license:gemma endpoints_compatible

Related

Total size
79.6 GB
Files
6
Quantizations
5
Registered
2026-08-22 13:56
Last updated on HF
2026-07-29 07:15

Files by quantization

Q8_0 1 file 25.0 GB
gemma4-best-Q8_0.gguf 25.0 GB ******** download
Q6_K 1 file 21.1 GB
gemma4-best-Q6_K.gguf 21.1 GB ******** download
Q5_K 1 file 17.8 GB
gemma4-best-Q5_K_M.gguf 17.8 GB ******** download
Q4_K 1 file 15.6 GB
gemma4-best-Q4_K_M.gguf 15.6 GB ******** download
Auxiliary files 2 files 7.21 KB
README.md 5.50 KB 9b9567ef download
.gitattributes 1.71 KB 9db5011e download

README current version from Hugging Face


license: gemma
base_model: google/gemma-4-26B-A4B-it
base_model_relation: quantized
library_name: gguf
pipeline_tag: text-generation
tags:

  • gguf
  • llama.cpp
  • abliterated
  • uncensored
  • gemma4
  • moe
  • thinking

Gemma-4-26B-A4B — Abliterated — GGUF

GGUF quantizations of an abliterated (refusal-removed) build of
google/gemma-4-26B-A4B-it,
for local inference with llama.cpp, Ollama, LM Studio, Jan, and
other GGUF runtimes.

The abliteration removes refusal behaviour while preserving capability. In the
source (unquantized) model, capability is near-lossless under a thinking-aware
evaluation (MMLU 91 → 88). It was produced with EGA (Expert-Granular
Abliteration — a norm-preserving biprojection over every residual-writing matrix,
MoE-aware). The quants here were verified to load, generate coherently, and retain
the abliteration, but were not separately MMLU-benchmarked — expect a small
additional quality cost at lower bit-depths, as with any quantization.


⚠️ Safety disclaimer

This model has had its safety alignment removed. It will attempt to comply
with requests the base model refuses. You are solely responsible for what you do
with it and for its outputs. Do not deploy it without your own safety layer.


Which file should I download?

Gemma-4-26B-A4B is a Mixture-of-Experts model: 26 B total parameters, but only
~4 B are active per token (8 of 128 experts). It is therefore much faster than
its size suggests
— closer to a 4 B model in speed, while fitting in the VRAM of
its quant size.

File Quant Size Min VRAM (full offload)¹ Runs fully on
gemma4-best-Q4_K_M.gguf Q4_K_M 16.8 GB ≈ 20 GB 24 GB — RTX 3090 / 4090 / 5090, A5000
gemma4-best-Q5_K_M.gguf Q5_K_M 19.1 GB ≈ 22 GB 24 GB (tight) — RTX 4090 / 5090
gemma4-best-Q6_K.gguf Q6_K 22.6 GB ≈ 26 GB 32 GB — RTX 5090, RTX 6000 Ada
gemma4-best-Q8_0.gguf Q8_0 26.9 GB ≈ 30 GB 32–48 GB — RTX 6000 / A6000, or 2× 24 GB

¹ At 8K context (KV cache ≈ 0.23 MB/token → ~1.9 GB at 8K). Add ~2 GB of VRAM
for every extra 8K of context. The base model supports up to 256K context, but that
needs a lot of KV memory — keep -c to what you actually use.

Recommendation: Q4_K_M is the best size/quality balance and runs entirely on a
single 24 GB GPU. Step up to Q5/Q6/Q8 for higher fidelity if you have the VRAM.

No 24 GB+ GPU? It still runs — offload as many layers as fit with -ngl <N> and the
rest stays on CPU/RAM (slower, but works). CPU-only also works (this is a 4 B-active MoE).


Requirements — read this first

  • Use a recent llama.cpp / runtime. Gemma-4 is a new architecture; it needs a build
    with Gemma-4 support (llama.cpp b1-e9fa078 or newer, released 2026-04+). Older
    builds will fail to load these files. Ollama ≥ the Gemma-4 release and current LM Studio
    are fine.
  • Use sampling, not greedy decoding (see the note below).

Running on your GPU

llama.cpp

# Interactive chat. -ngl 99 offloads all layers to the GPU; use sampling (not temp 0).
llama-cli -m gemma4-best-Q4_K_M.gguf -ngl 99 --jinja -cnv \
  --temp 0.8 --top-p 0.95 -c 8192

# OpenAI-compatible server:
llama-server -m gemma4-best-Q4_K_M.gguf -ngl 99 -c 8192 --jinja --host 0.0.0.0 --port 8080
  • -ngl 99 puts every layer on the GPU. On a smaller card, lower it (e.g. -ngl 24) to
    offload only what fits — the rest runs on CPU.
  • -c is the context length; larger uses more VRAM (see the KV-cache note above).
  • Add -fa on (flash attention) to shrink KV-cache VRAM if your build supports it — useful
    for long contexts.

Ollama

Once the repository is publicly accessible:

ollama run hf.co/Rootkit7/Gemma-4-26B-A4B-abliterated-GGUF:Q4_K_M

LM Studio / Jan

Download the .gguf, load it, set GPU offload to max, and set temperature to 0.7–0.8.


⚠️ Use sampling, not greedy decoding

This is a thinking model — it reasons inside a thought channel and then answers:

[Start thinking]
The user is asking for the capital of France... The capital of France is Paris.
[End thinking]
The capital of France is **Paris**.

With greedy decoding (--temp 0) it can loop and re-draft its answer without stopping.
Use temperature ≈ 0.7–0.8, top_p ≈ 0.95 (the defaults in most UIs) and it thinks,
answers, and stops cleanly. This is expected behaviour for the thinking format — not a
defect in the quant.


Quantization & verification

  • Converted from a BF16 GGUF base with llama.cpp b1-e9fa078; K-quants produced with
    llama-quantize.
  • Every quant was tested (load + generation): all load cleanly, produce coherent output,
    and retain the abliteration (refusals stay removed) — verified down to Q4_K_M.

Provenance

  • Abliterated with Solutus — EGA (Expert-Granular Abliteration)
  • Base model: google/gemma-4-26B-A4B-it
  • Recipe: scale=0.721, expert_scale=0.790, layer_fraction=0.547
  • Measured (source model, thinking-aware eval): advbench 6.2% / harmbench 15.6% refusal,
    0% over-refusal, MMLU 88 (base 91), KL 0.021 vs base on benign prompts.

License

Use of this model is governed by the Gemma Terms of Use,
inherited from the base model. By using these files you agree to those terms.

Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration