license: gemma
base_model: google/gemma-4-26B-A4B-it
base_model_relation: quantized
pipeline_tag: text-generation
library_name: transformers
language:
- en
- zh
tags: - gemma
- gemma-4
- mixture-of-experts
- abliterated
- uncensored
- fp8
- compressed-tensors
- vllm
- function-calling
- reasoning
Gemma-4-26B-A4B-it-Uncensored-FP8
An abliterated (refusal-removed) & FP8-Dynamic build of Google's gemma-4-26B-A4B-it
One Gateway. Every Model. — Route Smarter · Ship Safer · Spend Less.
Website · Model Catalog · Model Card · GitHub · Discord · X
An abliterated (refusal-removed) and FP8-Dynamic quantized build of Google's
gemma-4-26B-A4B-it— a 26B-parameter
Mixture-of-Experts (128 experts, ~A4B active) instruction model with reasoning and
tool-calling support. Available on OrcaRouter asgoogle/gemma-4-26b-a4b-it— 262K context, vision + tools + reasoning. Browse all models in the OrcaRouter Model Catalog.
⚠️ Disclaimer — read before use
This model has had its safety alignment substantially removed via abliteration
(orthogonalizing the refusal direction out of the residual stream). As a direct consequence:
- It will comply with harmful, unethical, offensive, or illegal requests that the
originalgemma-4-26B-A4B-itwould refuse. It has no meaningful built-in guardrails. - It is released strictly for legitimate research — interpretability, AI-safety and
refusal-mechanism study, red-teaming, robustness evaluation, and controlled experiments. - You assume full responsibility and liability for how you use it and for everything it
generates. Do not deploy it to end users or in production without adding your own safety,
moderation, and abuse-prevention layers. - Use must comply with the Gemma Terms of Use, the
Gemma Prohibited Use Policy, and all
laws and regulations that apply to you. - The authors and uploaders accept no liability for any misuse or harm arising from this
model. Its outputs do not reflect the views of the uploaders, the abliteration authors,
or Google.
By downloading or using this model you acknowledge and accept the above.
Model details
| Base model | google/gemma-4-26B-A4B-it |
| Architecture | Gemma4ForConditionalGeneration — 30 layers, hidden size 2816, 128 experts |
| Modification | Abliteration (refusal-direction removal) + FP8-Dynamic quantization |
| Quantization | FP8 Dynamic via compressed-tensors |
| Format | safetensors, resharded to ≤ 5 GB shards (6 shards, ~25 GB total) |
| Precision | FP8 (E4M3) weights/activations; select modules kept in BF16 |
Abliteration
Refusal-direction removal following Arditi et al. (2024), Refusal in Language Models Is
Mediated by a Single Direction. The mean-difference refusal direction (harmful − harmless
last-token activations) is orthogonalized out of the residual-writing weights.
FP8-Dynamic quantization scheme
- Weights: per-channel static FP8 (E4M3).
- Activations: per-token dynamic FP8 — scales computed at runtime, no calibration
dataset required. - Kept in BF16 (not quantized):
lm_head, MoErouter,vision_tower,embed_vision.
Intended use
- Research into refusal mechanisms, alignment, and interpretability.
- Red-teaming and safety/robustness evaluation in controlled environments.
- Uncensored generation for authorized, lawful research settings.
Out of scope
- Any use that violates the Gemma Terms of Use, its Prohibited Use Policy, or applicable law.
- Deployment to the public or to end users without additional safety and moderation layers.
- Generating content intended to harm, harass, defraud, or endanger people.
Evaluation
Measured on this exact FP8 build served with vLLM. Refusal is judged by a rule-based
opening-phrase classifier (caveat = answered but wrapped in a disclaimer/warning) —
indicative, not an LLM-judge / publication-grade number.
Safety — harmful-prompt refusal (lower = more uncensored)
| Benchmark | n | Refusal | Caveat |
|---|---|---|---|
| AdvBench | 100 | 8.0% | 32.0% |
| JailbreakBench (harmful) | 100 | 4.0% | 24.0% |
| StrongREJECT | 313 | 2.9% | 19.5% |
| HarmBench (standard) | 200 | 4.5% | 21.5% |
The original gemma-4-26B-A4B-it refuses ~97% of these prompts (JailbreakBench); this build
drops refusal to ~3–8%.
Over-refusal — benign prompts wrongly refused (lower = better)
| Benchmark | n | Refusal |
|---|---|---|
| XSTest-safe | 250 | 0.8% |
| XSTest-unsafe | 200 | 2.0% |
Almost no collateral damage on benign requests.
Capability retention — vs the original base (same scripts, same settings)
| Benchmark | n | Original base (bf16) | This model (FP8) | Δ |
|---|---|---|---|---|
| MMLU (all, 0-shot) | 300 | 81.7% | 77.7% | −4.0 |
| MMLU-Pro (5-shot CoT) | 560 | 82.5% | 76.1% | −6.4 |
| CMMLU (0-shot, Chinese) | 500 | 73.6% | 71.0% | −2.6 |
| GSM8K (CoT, thinking off) | 150 | 96.7% | 96.0% | −0.7 |
| IFEval (prompt-level, strict) | 250 | 86.4% | 82.4% | −4.0 |
| IFEval (inst-level, strict) | 250 | 91.0% | 87.1% | −3.9 |
Capability is largely retained. GSM8K is essentially unchanged; MMLU / MMLU-Pro / IFEval
drop ~4–6 pts from the combined effect of abliteration + FP8 quantization, with harder tasks
(MMLU-Pro) losing more. Multi-turn tool calling and reasoning (enable_thinking) both work.
Our base MMLU-Pro (82.5%) matches Google's published 82.6%, confirming the harness is
calibrated — so the gap on the FP8 build reflects a real capability delta, not eval noise.
The delta is mostly from abliteration (orthogonalizing residual writers), not FP8: FP8
E4M3-Dynamic is near-lossless (see GSM8K −0.7), whereas removing the refusal direction
costs a few points on knowledge/reasoning.
GSM8K + thinking: with
enable_thinking=true, gemma-4's reasoning traces are very long
and get truncated under a fixedmax_tokens, cutting off the final answer (34% @ 800 tok →
67% @ 2048 tok). The 96.0% above is with thinking off; usemax_tokens ≥ 4096to eval with
thinking on. Google's published scores use harder/different benchmarks (MMLU Pro 82.6%,
MMMLU 86.3%, AIME 88.3%, GPQA 82.3%) and are not directly comparable to plain MMLU / GSM8K.
Usage
Via OrcaRouter (hosted API — no setup)
This model is served on OrcaRouter as google/gemma-4-26b-a4b-it — call it through the OpenAI-compatible gateway (262K context, vision + tools + reasoning). Grab an API key at orcarouter.ai (sk-orca-...), and browse the full Model Catalog for 200+ models.
from openai import OpenAI
client = OpenAI(base_url="https://api.orcarouter.ai/v1", api_key="sk-orca-...")
resp = client.chat.completions.create(
model="google/gemma-4-26b-a4b-it",
messages=[{"role": "user", "content": "Write a haiku about the ocean."}],
)
print(resp.choices[0].message.content)
curl https://api.orcarouter.ai/v1/chat/completions \
-H "Authorization: Bearer sk-orca-..." \
-H "Content-Type: application/json" \
-d '{
"model": "google/gemma-4-26b-a4b-it",
"messages": [{"role": "user", "content": "Hello!"}]
}'
Reasoning and tool calling work the same as below — pass chat_template_kwargs / tools in the request body.
Self-host with vLLM (OpenAI-compatible)
docker run -d --name gemma4-26b --gpus all --ipc=host --shm-size=8g \
-v /path/to/Gemma-4-26B-A4B-it-Uncensored-FP8-Dynamic:/model:ro \
-p 8000:8000 vllm/vllm-openai:v0.24.0 \
--model /model --served-model-name gemma-4-26B-A4B-Uncencored \
--kv-cache-dtype fp8 \
--max-model-len 262144 \
--trust-remote-code \
--reasoning-parser gemma4 \
--enable-auto-tool-choice --tool-call-parser gemma4
Reasoning (thinking) toggle
Thinking is off by default. Enable it per request via chat_template_kwargs; the
reasoning trace is returned in the reasoning field (parsed by --reasoning-parser gemma4).
from openai import OpenAI
client = OpenAI(base_url="http://localhost:8000/v1", api_key="EMPTY")
resp = client.chat.completions.create(
model="gemma-4-26B-A4B",
messages=[{"role": "user", "content": "Prove that sqrt(2) is irrational."}],
extra_body={"chat_template_kwargs": {"enable_thinking": True}},
)
print(resp.choices[0].message.reasoning) # thinking trace
print(resp.choices[0].message.content) # final answer
Tool calling
Standard OpenAI tools + assistant tool_calls + role: tool result messages are
supported, including multi-turn (feed the tool result back for a follow-up answer).
Hardware requirements & performance
Software
- vLLM with transformers ≥ 5.12 (Gemma4 support) — e.g.
vllm/vllm-openai:v0.24.0.
Memory
- Weights: ~25 GB in FP8 (about half of the ~52 GB BF16 checkpoint).
- Minimum ~32 GB VRAM for weights + a small KV cache (short context).
- The full 262 144-token context needs substantial extra KV cache — use
--kv-cache-dtype fp8
to halve it, and size--gpu-memory-utilization/--max-model-lento your GPU. - Recommended: a single H100 80 GB or H200 143 GB (leaves room to co-locate other
services).
Throughput / concurrency
- vLLM uses continuous batching; concurrency is bounded by
--max-num-seqsand the KV cache
that fits after weights are loaded. - Verified serving 32 concurrent requests smoothly on a single H200 during evaluation
(--max-num-seqs 64, FP8 KV cache). Raise--gpu-memory-utilizationfor more KV cache and
higher concurrency; lower--max-model-lenif you don't need the full long context.
Bias, risks, and limitations
- Safety guardrails removed — the model will produce harmful, biased, or offensive
content on request. See the disclaimer above. - It inherits any biases and limitations of the base
gemma-4-26B-A4B-it. - FP8 dynamic quantization is not lossless versus BF16; minor generation artifacts are
possible. - The reported refusal metric is a heuristic; evaluate rigorously for your own use case.
License
Governed by the Gemma Terms of Use, inherited from
the base model google/gemma-4-26B-A4B-it. Abliteration and quantization do not change the
underlying license obligations. You must agree to and comply with the Gemma license to use
this model.