license: gemma
base_model: google/gemma-4-26B-A4B-it
base_model_relation: quantized
library_name: gguf
pipeline_tag: text-generation
tags:
- gguf
- llama.cpp
- abliterated
- uncensored
- gemma4
- moe
- thinking
Gemma-4-26B-A4B — Abliterated — GGUF
GGUF quantizations of an abliterated (refusal-removed) build ofgoogle/gemma-4-26B-A4B-it,
for local inference with llama.cpp, Ollama, LM Studio, Jan, and
other GGUF runtimes.
The abliteration removes refusal behaviour while preserving capability. In the
source (unquantized) model, capability is near-lossless under a thinking-aware
evaluation (MMLU 91 → 88). It was produced with EGA (Expert-Granular
Abliteration — a norm-preserving biprojection over every residual-writing matrix,
MoE-aware). The quants here were verified to load, generate coherently, and retain
the abliteration, but were not separately MMLU-benchmarked — expect a small
additional quality cost at lower bit-depths, as with any quantization.
⚠️ Safety disclaimer
This model has had its safety alignment removed. It will attempt to comply
with requests the base model refuses. You are solely responsible for what you do
with it and for its outputs. Do not deploy it without your own safety layer.
Which file should I download?
Gemma-4-26B-A4B is a Mixture-of-Experts model: 26 B total parameters, but only
~4 B are active per token (8 of 128 experts). It is therefore much faster than
its size suggests — closer to a 4 B model in speed, while fitting in the VRAM of
its quant size.
| File | Quant | Size | Min VRAM (full offload)¹ | Runs fully on |
|---|---|---|---|---|
gemma4-best-Q4_K_M.gguf |
Q4_K_M | 16.8 GB | ≈ 20 GB | 24 GB — RTX 3090 / 4090 / 5090, A5000 |
gemma4-best-Q5_K_M.gguf |
Q5_K_M | 19.1 GB | ≈ 22 GB | 24 GB (tight) — RTX 4090 / 5090 |
gemma4-best-Q6_K.gguf |
Q6_K | 22.6 GB | ≈ 26 GB | 32 GB — RTX 5090, RTX 6000 Ada |
gemma4-best-Q8_0.gguf |
Q8_0 | 26.9 GB | ≈ 30 GB | 32–48 GB — RTX 6000 / A6000, or 2× 24 GB |
¹ At 8K context (KV cache ≈ 0.23 MB/token → ~1.9 GB at 8K). Add ~2 GB of VRAM
for every extra 8K of context. The base model supports up to 256K context, but that
needs a lot of KV memory — keep -c to what you actually use.
Recommendation: Q4_K_M is the best size/quality balance and runs entirely on a
single 24 GB GPU. Step up to Q5/Q6/Q8 for higher fidelity if you have the VRAM.
No 24 GB+ GPU? It still runs — offload as many layers as fit with -ngl <N> and the
rest stays on CPU/RAM (slower, but works). CPU-only also works (this is a 4 B-active MoE).
Requirements — read this first
- Use a recent llama.cpp / runtime. Gemma-4 is a new architecture; it needs a build
with Gemma-4 support (llama.cpp b1-e9fa078 or newer, released 2026-04+). Older
builds will fail to load these files. Ollama ≥ the Gemma-4 release and current LM Studio
are fine. - Use sampling, not greedy decoding (see the note below).
Running on your GPU
llama.cpp
# Interactive chat. -ngl 99 offloads all layers to the GPU; use sampling (not temp 0).
llama-cli -m gemma4-best-Q4_K_M.gguf -ngl 99 --jinja -cnv \
--temp 0.8 --top-p 0.95 -c 8192
# OpenAI-compatible server:
llama-server -m gemma4-best-Q4_K_M.gguf -ngl 99 -c 8192 --jinja --host 0.0.0.0 --port 8080
-ngl 99puts every layer on the GPU. On a smaller card, lower it (e.g.-ngl 24) to
offload only what fits — the rest runs on CPU.-cis the context length; larger uses more VRAM (see the KV-cache note above).- Add
-fa on(flash attention) to shrink KV-cache VRAM if your build supports it — useful
for long contexts.
Ollama
Once the repository is publicly accessible:
ollama run hf.co/Rootkit7/Gemma-4-26B-A4B-abliterated-GGUF:Q4_K_M
LM Studio / Jan
Download the .gguf, load it, set GPU offload to max, and set temperature to 0.7–0.8.
⚠️ Use sampling, not greedy decoding
This is a thinking model — it reasons inside a thought channel and then answers:
[Start thinking]
The user is asking for the capital of France... The capital of France is Paris.
[End thinking]
The capital of France is **Paris**.
With greedy decoding (--temp 0) it can loop and re-draft its answer without stopping.
Use temperature ≈ 0.7–0.8, top_p ≈ 0.95 (the defaults in most UIs) and it thinks,
answers, and stops cleanly. This is expected behaviour for the thinking format — not a
defect in the quant.
Quantization & verification
- Converted from a BF16 GGUF base with llama.cpp
b1-e9fa078; K-quants produced withllama-quantize. - Every quant was tested (load + generation): all load cleanly, produce coherent output,
and retain the abliteration (refusals stay removed) — verified down to Q4_K_M.
Provenance
- Abliterated with Solutus — EGA (Expert-Granular Abliteration)
- Base model:
google/gemma-4-26B-A4B-it - Recipe:
scale=0.721,expert_scale=0.790,layer_fraction=0.547 - Measured (source model, thinking-aware eval): advbench 6.2% / harmbench 15.6% refusal,
0% over-refusal, MMLU 88 (base 91), KL 0.021 vs base on benign prompts.
License
Use of this model is governed by the Gemma Terms of Use,
inherited from the base model. By using these files you agree to those terms.