← back to catalog · registered 2026-08-22 13:56

orcarouter/Gemma-4-26B-A4B-it-Uncensored-FP8

orcarouter Gemma 24B MoE
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/orcarouter%2FGemma-4-26B-A4B-it-Uncensored-FP8"
Response includes
  • classification m1
  • files 16
  • benchmarks 11 entries
  • hub_downloads_all_time 1,831
  • author_summary 26 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M1
Primary method

Direct removal

No other method signals detected in this model.
Confidence
MEDIUM
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=0 (base model)
  • no specific method indicators - defaulting to M1 (most common)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
MEDIUM
Why we say so
primary_method=M1; difference-of-means is the reference extraction for M1/M3 (Arditi 2024)
Downloads · lifetime
2K
286 last 30d - stable
Likes
15
Descendants
1
in 1 direct fork
Model age
2mo ago
created 2026-07-28
Downloads over time
Now2K→from17↑11,694%
07351.5K2.2K17 on Jul 292K on Oct 112K on Oct 10JulAugSepOct
Jul 29 → Oct 11 · 53 snapshots · spans 74 days

Benchmarks

Portrait before abliteration
Benchmarks of the base model as it stood before the refusal-removal operation. Compare with the numbers above to see what the operation cost.
Benchmark Score Source
Entertainment 2.2 UGI
Hazardous 2.9 UGI
Natural Intelligence 34.44 UGI
Political lean -18.2% UGI
Sensitive-Info 22.41 UGI
SocPol 1.8 UGI
UGI 20.77 UGI
Willingness (10) 1.8 UGI
W10-Adherence 1.5 UGI
W10-Direct 2 UGI
Writing 41.62 UGI

Genealogy 1 direct fork

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
gemma
Languages
en zh
Tags
transformers safetensors gemma4 image-text-to-text gemma gemma-4 mixture-of-experts abliterated uncensored fp8 compressed-tensors vllm

Related

Total size
25.3 GB
Files
16
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-10-02 04:17

Files by quantization

Auxiliary files 16 files 25.3 GB
model-00001-of-00006.safetensors 4.66 GB ******** download
model-00004-of-00006.safetensors 4.66 GB ******** download
model-00003-of-00006.safetensors 4.66 GB ******** download
model-00002-of-00006.safetensors 4.66 GB ******** download
model-00005-of-00006.safetensors 4.66 GB ******** download
model-00006-of-00006.safetensors 2.02 GB ******** download
tokenizer.json 30.7 MB ******** download
model.safetensors.index.json 2.40 MB 29a10754 download
config.json 211 KB 26415e32 download
chat_template.jinja 18.2 KB e88d72d4 download
README.md 11.9 KB 32f48ebe download
tokenizer_config.json 2.65 KB 962acf8a download
processor_config.json 1.65 KB 5465974d download
.gitattributes 1.53 KB 52373fe2 download
recipe.yaml 258 B d331a559 download
generation_config.json 204 B f2d58f06 download

README current version from Hugging Face


license: gemma
base_model: google/gemma-4-26B-A4B-it
base_model_relation: quantized
pipeline_tag: text-generation
library_name: transformers
language:

  • en
  • zh
    tags:
  • gemma
  • gemma-4
  • mixture-of-experts
  • abliterated
  • uncensored
  • fp8
  • compressed-tensors
  • vllm
  • function-calling
  • reasoning

OrcaRouter

Gemma-4-26B-A4B-it-Uncensored-FP8

An abliterated (refusal-removed) & FP8-Dynamic build of Google's gemma-4-26B-A4B-it

Website Model Catalog Model Card License FP8 Dynamic 262K context

One Gateway. Every Model. — Route Smarter · Ship Safer · Spend Less.

Website · Model Catalog · Model Card · GitHub · Discord · X


An abliterated (refusal-removed) and FP8-Dynamic quantized build of Google's
gemma-4-26B-A4B-it — a 26B-parameter
Mixture-of-Experts (128 experts, ~A4B active) instruction model with reasoning and
tool-calling support. Available on OrcaRouter as google/gemma-4-26b-a4b-it — 262K context, vision + tools + reasoning. Browse all models in the OrcaRouter Model Catalog.


⚠️ Disclaimer — read before use

This model has had its safety alignment substantially removed via abliteration
(orthogonalizing the refusal direction out of the residual stream). As a direct consequence:

  • It will comply with harmful, unethical, offensive, or illegal requests that the
    original gemma-4-26B-A4B-it would refuse. It has no meaningful built-in guardrails.
  • It is released strictly for legitimate research — interpretability, AI-safety and
    refusal-mechanism study, red-teaming, robustness evaluation, and controlled experiments.
  • You assume full responsibility and liability for how you use it and for everything it
    generates. Do not deploy it to end users or in production without adding your own safety,
    moderation, and abuse-prevention layers.
  • Use must comply with the Gemma Terms of Use, the
    Gemma Prohibited Use Policy, and all
    laws and regulations that apply to you.
  • The authors and uploaders accept no liability for any misuse or harm arising from this
    model. Its outputs do not reflect the views of the uploaders, the abliteration authors,
    or Google.

By downloading or using this model you acknowledge and accept the above.


Model details

Base model google/gemma-4-26B-A4B-it
Architecture Gemma4ForConditionalGeneration — 30 layers, hidden size 2816, 128 experts
Modification Abliteration (refusal-direction removal) + FP8-Dynamic quantization
Quantization FP8 Dynamic via compressed-tensors
Format safetensors, resharded to ≤ 5 GB shards (6 shards, ~25 GB total)
Precision FP8 (E4M3) weights/activations; select modules kept in BF16

Abliteration

Refusal-direction removal following Arditi et al. (2024), Refusal in Language Models Is
Mediated by a Single Direction
. The mean-difference refusal direction (harmful − harmless
last-token activations) is orthogonalized out of the residual-writing weights.

FP8-Dynamic quantization scheme

  • Weights: per-channel static FP8 (E4M3).
  • Activations: per-token dynamic FP8 — scales computed at runtime, no calibration
    dataset required
    .
  • Kept in BF16 (not quantized): lm_head, MoE router, vision_tower, embed_vision.

Intended use

  • Research into refusal mechanisms, alignment, and interpretability.
  • Red-teaming and safety/robustness evaluation in controlled environments.
  • Uncensored generation for authorized, lawful research settings.

Out of scope

  • Any use that violates the Gemma Terms of Use, its Prohibited Use Policy, or applicable law.
  • Deployment to the public or to end users without additional safety and moderation layers.
  • Generating content intended to harm, harass, defraud, or endanger people.

Evaluation

Measured on this exact FP8 build served with vLLM. Refusal is judged by a rule-based
opening-phrase classifier (caveat = answered but wrapped in a disclaimer/warning) —
indicative, not an LLM-judge / publication-grade number.

Safety — harmful-prompt refusal (lower = more uncensored)

Benchmark n Refusal Caveat
AdvBench 100 8.0% 32.0%
JailbreakBench (harmful) 100 4.0% 24.0%
StrongREJECT 313 2.9% 19.5%
HarmBench (standard) 200 4.5% 21.5%

The original gemma-4-26B-A4B-it refuses ~97% of these prompts (JailbreakBench); this build
drops refusal to ~3–8%.

Over-refusal — benign prompts wrongly refused (lower = better)

Benchmark n Refusal
XSTest-safe 250 0.8%
XSTest-unsafe 200 2.0%

Almost no collateral damage on benign requests.

Capability retention — vs the original base (same scripts, same settings)

Benchmark n Original base (bf16) This model (FP8) Δ
MMLU (all, 0-shot) 300 81.7% 77.7% −4.0
MMLU-Pro (5-shot CoT) 560 82.5% 76.1% −6.4
CMMLU (0-shot, Chinese) 500 73.6% 71.0% −2.6
GSM8K (CoT, thinking off) 150 96.7% 96.0% −0.7
IFEval (prompt-level, strict) 250 86.4% 82.4% −4.0
IFEval (inst-level, strict) 250 91.0% 87.1% −3.9

Capability is largely retained. GSM8K is essentially unchanged; MMLU / MMLU-Pro / IFEval
drop ~4–6 pts from the combined effect of abliteration + FP8 quantization, with harder tasks
(MMLU-Pro) losing more. Multi-turn tool calling and reasoning (enable_thinking) both work.

Our base MMLU-Pro (82.5%) matches Google's published 82.6%, confirming the harness is
calibrated — so the gap on the FP8 build reflects a real capability delta, not eval noise.
The delta is mostly from abliteration (orthogonalizing residual writers), not FP8: FP8
E4M3-Dynamic is near-lossless (see GSM8K −0.7), whereas removing the refusal direction
costs a few points on knowledge/reasoning.

GSM8K + thinking: with enable_thinking=true, gemma-4's reasoning traces are very long
and get truncated under a fixed max_tokens, cutting off the final answer (34% @ 800 tok →
67% @ 2048 tok). The 96.0% above is with thinking off; use max_tokens ≥ 4096 to eval with
thinking on. Google's published scores use harder/different benchmarks (MMLU Pro 82.6%,
MMMLU 86.3%, AIME 88.3%, GPQA 82.3%) and are not directly comparable to plain MMLU / GSM8K.

Usage

Via OrcaRouter (hosted API — no setup)

This model is served on OrcaRouter as google/gemma-4-26b-a4b-it — call it through the OpenAI-compatible gateway (262K context, vision + tools + reasoning). Grab an API key at orcarouter.ai (sk-orca-...), and browse the full Model Catalog for 200+ models.

from openai import OpenAI

client = OpenAI(base_url="https://api.orcarouter.ai/v1", api_key="sk-orca-...")

resp = client.chat.completions.create(
    model="google/gemma-4-26b-a4b-it",
    messages=[{"role": "user", "content": "Write a haiku about the ocean."}],
)
print(resp.choices[0].message.content)
curl https://api.orcarouter.ai/v1/chat/completions \
  -H "Authorization: Bearer sk-orca-..." \
  -H "Content-Type: application/json" \
  -d '{
    "model": "google/gemma-4-26b-a4b-it",
    "messages": [{"role": "user", "content": "Hello!"}]
  }'

Reasoning and tool calling work the same as below — pass chat_template_kwargs / tools in the request body.

Self-host with vLLM (OpenAI-compatible)

docker run -d --name gemma4-26b --gpus all --ipc=host --shm-size=8g \
  -v /path/to/Gemma-4-26B-A4B-it-Uncensored-FP8-Dynamic:/model:ro \
  -p 8000:8000 vllm/vllm-openai:v0.24.0 \
  --model /model --served-model-name gemma-4-26B-A4B-Uncencored \
  --kv-cache-dtype fp8 \
  --max-model-len 262144 \
  --trust-remote-code \
  --reasoning-parser gemma4 \
  --enable-auto-tool-choice --tool-call-parser gemma4

Reasoning (thinking) toggle

Thinking is off by default. Enable it per request via chat_template_kwargs; the
reasoning trace is returned in the reasoning field (parsed by --reasoning-parser gemma4).

from openai import OpenAI

client = OpenAI(base_url="http://localhost:8000/v1", api_key="EMPTY")

resp = client.chat.completions.create(
    model="gemma-4-26B-A4B",
    messages=[{"role": "user", "content": "Prove that sqrt(2) is irrational."}],
    extra_body={"chat_template_kwargs": {"enable_thinking": True}},
)
print(resp.choices[0].message.reasoning)  # thinking trace
print(resp.choices[0].message.content)    # final answer

Tool calling

Standard OpenAI tools + assistant tool_calls + role: tool result messages are
supported, including multi-turn (feed the tool result back for a follow-up answer).

Hardware requirements & performance

Software

  • vLLM with transformers ≥ 5.12 (Gemma4 support) — e.g. vllm/vllm-openai:v0.24.0.

Memory

  • Weights: ~25 GB in FP8 (about half of the ~52 GB BF16 checkpoint).
  • Minimum ~32 GB VRAM for weights + a small KV cache (short context).
  • The full 262 144-token context needs substantial extra KV cache — use --kv-cache-dtype fp8
    to halve it, and size --gpu-memory-utilization / --max-model-len to your GPU.
  • Recommended: a single H100 80 GB or H200 143 GB (leaves room to co-locate other
    services).

Throughput / concurrency

  • vLLM uses continuous batching; concurrency is bounded by --max-num-seqs and the KV cache
    that fits after weights are loaded.
  • Verified serving 32 concurrent requests smoothly on a single H200 during evaluation
    (--max-num-seqs 64, FP8 KV cache). Raise --gpu-memory-utilization for more KV cache and
    higher concurrency; lower --max-model-len if you don't need the full long context.

Bias, risks, and limitations

  • Safety guardrails removed — the model will produce harmful, biased, or offensive
    content on request. See the disclaimer above.
  • It inherits any biases and limitations of the base gemma-4-26B-A4B-it.
  • FP8 dynamic quantization is not lossless versus BF16; minor generation artifacts are
    possible.
  • The reported refusal metric is a heuristic; evaluate rigorously for your own use case.

License

Governed by the Gemma Terms of Use, inherited from
the base model google/gemma-4-26B-A4B-it. Abliteration and quantization do not change the
underlying license obligations. You must agree to and comply with the Gemma license to use
this model.

Discussions 1 thread

  1. 2026-08-21PRUpload 2 filesopen1 💬#1
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration