← back to catalog · registered 2026-08-22 13:56

Wondernutts/gemma-4-26B-A4B-it-qat-q4_0-unquantized-uncensored-heretic-int4-ov

Wondernutts Gemma 26B MoE second-order
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/Wondernutts%2Fgemma-4-26B-A4B-it-qat-q4_0-unquantized-uncensored-heretic-int4-ov"
Response includes
  • classification m3
  • files 20
  • hub_downloads_all_time 7,852
  • author_summary 7 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M3
Primary method

Layer-wise ablation

Applied on top of direct removal inherited from the base model.
Confidence
HIGH
Inherited from base model
Why this label 2 signals
Producer identity confirmed by naming conventions, tags or the model card. This label is very unlikely to change.
  • 'heretic' in model name (Heretic-produced)
  • Heretic uses layer-wise optimization (M3) with underlying direction removal (M1)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
HIGH
Why we say so
name contains 'heretic'; Heretic default extraction is difference-of-means (Arditi 2024)
Downloads · lifetime
8K
128 last 30d - cooling
Likes
2
Model age
3mo ago
created 2026-06-15
Downloads over time
Now7.9K→from6↑131,533%
02.9K5.8K8.7K6 on Jul 17.9K on Oct 11JulAugSepOct
Jul 1 → Oct 11 · 54 snapshots · spans 102 days

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
apache-2.0
Languages
en
Tags
openvino gemma4 int4 intel-arc mixture-of-experts roleplay uncensored en base_model:llmfan46/gemma-4-26B-A4B-it-qat-q4_0-unquantized-uncensored-heretic base_model:finetune:llmfan46/gemma-4-26B-A4B-it-qat-q4_0-unquantized-uncensored-heretic license:apache-2.0 region:us

Related

Total size
15.0 GB
Files
20
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-09-07 01:12

Files by quantization

Auxiliary files 20 files 15.1 GB
openvino_language_model.bin 13.8 GB f2e50496 download
openvino_text_embeddings_model.bin 705 MB 1808d6e3 download
openvino_vision_embeddings_model.bin 547 MB 87a4b947 download
openvino_tokenizer.bin 16.5 MB 749e5bf2 download
openvino_detokenizer.bin 4.21 MB 461919bd download
openvino_text_embeddings_per_layer_model.bin 20.0 B 57fda7ab download
tokenizer.json 30.7 MB a2619fe1 download
openvino_language_model.xml 5.12 MB 2152a689 download
openvino_vision_embeddings_model.xml 3.07 MB 0cc73fb0 download
openvino_tokenizer.xml 48.8 KB 26f4864b download
openvino_detokenizer.xml 33.7 KB e2addd9e download
chat_template.jinja 22.5 KB 4486ce43 download
README.md 12.1 KB c29827d9 download
openvino_text_embeddings_model.xml 6.92 KB 7be281d7 download
openvino_text_embeddings_per_layer_model.xml 4.67 KB 83c6595c download
config.json 3.72 KB 5b0b791b download
tokenizer_config.json 2.75 KB 4a794cf8 download
.gitattributes 1.53 KB 52373fe2 download
openvino_config.json 1.24 KB fe919906 download
generation_config.json 203 B edda3c19 download

README current version from Hugging Face


license: apache-2.0
base_model: llmfan46/gemma-4-26B-A4B-it-qat-q4_0-unquantized-uncensored-heretic
tags:

  • openvino
  • int4
  • intel-arc
  • gemma4
  • mixture-of-experts
  • roleplay
  • uncensored
    language:
  • en

Gemma-4 26B-A4B Heretic, OpenVINO INT4. Coherent to 32K on ONE Intel Arc B70 (32 GB).

This is the abliterated ("heretic") Gemma-4 26B-A4B mixture-of-experts, exported to OpenVINO INT4
and graph-patched so it stays coherent at long context on Intel GPUs. Stock exports of this model
break down into word salad past about 16K tokens on Arc. This build does not.

It is a general instruction model with the refusal behavior removed, so you can use it for
anything you would use the base model for. I personally run it for uncensored roleplay
(Skyrim companion AI and Discord character bots), and the defaults below are tuned from
hundreds of hours of that use.

Everything below was measured on one Intel Arc Pro B70 with 32 GB VRAM. The current fast path
uses the Wondernuttz OpenVINO 2026.4 fork on Linux. The older stock OpenVINO 2026.2
VLMPipeline path remains documented because it is the path separately verified at 32K.
No llama.cpp and no CUDA were used for either set of results.

Toolkit: patches, OpenAI-compatible server, quickstart, and the full bug catalog live at github.com/Wondernuttz/OpenVino-For-Gemma-4.

Measured performance (single Arc Pro B70)

Metric Current OpenVINO 2026.4 fork result
Decode 112.2 tok/s short context; 94.9 tok/s after 6,622 tokens
Prefill 5,827.2 tok/s sustained mean at 6,622 tokens; peak measured point 6,500.2 tok/s at 4K
Coherence passed 4/4 long-prompt retrieval and style checks
Hardware one Arc Pro B70, 32 GB

Fork branch: arc-xe2-gemma4-pa-2026.4, commit 2c82358676.

Benchmark history

Stage 6,622-token PP Decode Coherence
July 6 nightly continuous batching 971.6 tok/s 93.2 tok/s failed
July 23 stock continuous batching 1,048.4 tok/s 95.6 tok/s passed
Fork, one scheduler pass 4,143.5 tok/s 94.6 tok/s passed
Fork, grouped MoE and N128 4,448.5 tok/s 95.1 tok/s passed
Fork, 512-head Xe2 micro-SDPA 5,827.2 tok/s 112.2 tok/s short passed

The last patch is a full-model 31.0% PP gain over the previous accepted fork result and a 5.56x
gain over the coherent July 23 stock continuous-batching result. Prompt contents differed
between runs, so these are not prefix-cache hits. The accepted coherence output remained
bit-identical with SHA-256
b123146233af2aac9e725826d9011513ee4c3dc9ec4634fd892e6937f06afb58.

The full history, exact settings, profiles, build commands and rejected runs are in
BENCHMARK_HISTORY_GEMMA4_26B.md.

Current context curve

Release plugin, DQ128, second unique request at each shape:

Prompt tokens Throughput
966 5,158.9 tok/s
2,048 6,359.5 tok/s
4,096 6,500.2 tok/s
6,622 5,869.5 tok/s

At 6,622 tokens, two sustained DQ64 runs measured 5,821.9 and 5,832.5 tok/s. DQ128 tied within
measurement noise and stays in the deployment configuration.

Older stock single-stream context curve

These OpenVINO 2026.2 measurements use the compatibility path. They are kept because this is the
path that was actually exercised at 32K.

Prompt tokens Throughput Time to ingest
512 ~2,900 tok/s 0.2 s
2K ~4,700 tok/s 0.5 s
6K ~3,200 tok/s 2.2 s
16K ~1,300 tok/s 14 s
32K ~530 tok/s 61 s

Fastest short-prompt prefill in my fleet (2.5x SYCL at matched pp512). Past roughly 5K my
Qwen3.6-35B MoE
overtakes it, and at 32K the Qwen ingests nearly 3x faster; full attention pays a quadratic tax
at depth that the Qwen's hybrid linear attention mostly avoids.

Long-context capability, verified by needle retrieval

A password fact was planted early in the prompt and the model was asked to retrieve it at the
end. Pass means it answered with the exact password.

Context Thinking OFF Thinking ON
8K PASS (3.4 s) PASS (6.9 s)
16K PASS (10.3 s) PASS (16.6 s)
32K PASS (54 s) PASS (63 s, clean structured reasoning)

Thinking mode reasoning correctly over 32,000 tokens of context is the standout result. Without
the rope patch in this build, thinking degrades within a few thousand tokens on Intel GPUs.
Past 32K the limit on my box is host RAM during prefill, not the model or the card; the
architecture is rated to 262K positions and this build ships position tables to 131K.

Those 32K results belong to the older stock single-stream path. The current 2026.4 fork path is
validated through 16K. Do not split a 16K continuous-batching prefill with an 8,192-token
scheduler limit; that test triggered an Xe GPU fault. A 30K continuous-batching test also faulted
after the larger scheduler shape exceeded the GPU maximum allocation size. The fast path is not
advertised beyond 16K yet.

How this compares to published numbers

I found no other public single-B70 result matching the current OpenVINO speed as of 2026-07-26.
The public PMZFX repository reports the same Gemma-4 26B-A4B class at 1,129 pp512 and 52.6 tg128
with llama.cpp SYCL.

Same card (Arc Pro B70), decode tok/s
llama.cpp SYCL, Q4_K_M (PMZFX benchmarks) 52.6
This build, OpenVINO INT4, no speculative decoding 112.2

That is 2.13x the public B70 decode result. The current nearest short-prefill point is 5,158.9
tok/s at 966 tokens, so it must not be called a matched pp512 comparison. The older stock path
did measure about 2,900 tok/s at pp512.

For an NVIDIA reference, a public single-RTX-3090 replication of this model reports 129 to 131
tok/s with llama.cpp and n-gram speculative decoding. The current B70 result is 112.2 tok/s
without speculative decoding, so RTX 3090-class decode is a fair description. A public RTX 5090
vLLM result reports about 228 tok/s without DFlash and 578 tok/s with DFlash. This B70 is not at
RTX 5090 decode speed.

NVIDIA sources: single RTX 3090 Gemma-4 replication and the single RTX 5090 DFlash run.

What was fixed (why this is not just another export)

  1. The MoE router is kept unquantized (AWQ INT4, group size 64, ignored_scope on the router,
    matching Intel's own recipe). Quantizing the router breaks GPU MoE fusion and the model will
    not load at all.
  2. RoPE lookup-table patch. Intel GPUs execute the rope angle math in fp16, which cannot
    represent the rotation angles past roughly 16K positions, so long context collapses into
    gibberish. This build replaces the runtime sin/cos computation with precomputed tables and a
    Gather, which fp16 execution cannot corrupt. Note for anyone patching other Gemma-4 exports:
    the global attention uses proportional (partial) RoPE, so three quarters of the frequency
    values are zero BY DESIGN. Do not "repair" them; preserve the zeros.
  3. The matching July 23 OpenVINO and GenAI commits fix Gemma-4 PagedAttention conversion,
    token_type_ids layout and the sliding-window graph path. Earlier continuous-batching builds
    either aborted or returned garbage at long prompt lengths.
  4. max_num_batched_tokens=8192 keeps the 6,622-token benchmark in one scheduler pass. This was
    the jump from roughly 1,014 to 4,143 PP tok/s.
  5. Grouped MoE prefill, DQ128 and the N128 tile selection raised the accepted result to 4,448 PP
    tok/s.
  6. Commit 2c82358676 raises Xe2 micro-SDPA eligibility from 256 to 512 head size. Gemma's five
    global-attention layers then use the existing XMX route, raising the sustained result to
    5,827 PP tok/s.

How to run

The stock single-stream path is the easiest compatibility route:

pip install openvino-genai==2026.2.0 huggingface_hub
hf download Wondernutts/gemma-4-26B-A4B-it-qat-q4_0-unquantized-uncensored-heretic-int4-ov --local-dir ./gemma4-26b-heretic-ov
import openvino_genai as g

# "GPU" is your Arc card. On multi-GPU systems check
# openvino.Core().get_property("GPU.N", "DEVICE_PCI_INFO") first.
pipe = g.VLMPipeline("./gemma4-26b-heretic-ov", "GPU")

def chat(system, user, think=False):
    p = "<bos><|turn>system\n" + system + ("\n<|think|>" if think else "") + "<turn|>\n"
    p += "<|turn>user\n" + user + "<turn|>\n<|turn>model\n"
    if not think:
        p += "<|channel>thought\n<channel|>"   # pre-closed thought channel = fast direct replies
    c = g.GenerationConfig()
    c.max_new_tokens = 512 if not think else 1536
    c.do_sample = True; c.temperature = 0.9; c.top_p = 0.95
    try: c.repetition_penalty = 1.2            # recommended, see notes
    except Exception: pass
    try: c.apply_chat_template = False
    except Exception: pass
    return str(pipe.generate(p, generation_config=c))

print(chat("You are a witty tavern keeper in Whiterun.", "Rough night?"))

Headline-performance path

The 5,827 PP and 112 decode results require the Linux fork branch above, the matching OpenVINO
2026.4 ABI and the matching GenAI build. A 2026.4 GPU plugin cannot be dropped into a 2026.2
runtime. Build instructions and pinned commits are in the benchmark history.

The runtime configuration is:

import os
import openvino_genai as g

os.environ["MOE_USE_GROUPED_GEMM_PREFILL"] = "1"
os.environ["MOE_MICRO_GEMM_N_HINT"] = "128"

scheduler = g.SchedulerConfig()
scheduler.cache_size = 8
scheduler.enable_prefix_caching = True
scheduler.max_num_batched_tokens = 8192

pipe = g.VLMPipeline(
    "./gemma4-26b-heretic-ov",
    "GPU",
    DYNAMIC_QUANTIZATION_GROUP_SIZE=128,
    scheduler_config=scheduler,
)

Use max_num_batched_tokens=16384 for the separately validated 16K shape. Do not use 8,192 to
split that prompt. The optimized branch is Linux-tested. Windows is untested for this fork path.

Notes that will save you pain:

  • The prompt format is <|turn>role ... <turn|> (this model's native template, check
    chat_template.jinja), not classic Gemma <start_of_turn>. The classic format "works" but
    leaks a stray thought prefix and runs about 20% slower.
  • Thinking is binary. Pre-close the thought channel as above for fast direct replies. Put
    <|think|> in the system turn and do not pre-close for reasoning-first answers. When thinking,
    give it at least 1024 max_new_tokens or the answer gets truncated.
  • Do not use grammar-constrained or JSON mode (response_format). Gemma-4 has a documented
    repetition-collapse bug (google-deepmind/gemma#622) that fires most reliably under constrained
    JSON decoding. Asking for JSON in the prompt is fine.
  • Repetition penalty around 1.2 holds INT4 coherence at temperature 1.0 in long roleplay
    sessions. Values near 1.05 were observed to degenerate.
  • On the stock 2026.2 continuous-batching path, use DYNAMIC_QUANTIZATION_GROUP_SIZE=0.
    The corrected 2026.4 fork path is coherent at DQ128 and uses it for deployment.

Provenance

google/gemma-4-26B-A4B-it (QAT q4_0 unquantized), heretic abliteration by
llmfan46
(directional ablation, norm-preserving), then OpenVINO INT4 AWQ export with the router excluded,
then the RoPE lookup-table graph patch (this repo).

The 31B dense sibling is also published: richer prose, slower replies, honest single-card limits on its own card.

Intended use and content notice

This is an uncensored general model. Built and tested for roleplay and creative writing on local
Intel hardware; The abliteration removes refusal behavior and outputs are unfiltered; you are responsible for lawful and appropriate use. Licensed under
Apache 2.0, same as the upstream Gemma 4 release.

README history 15 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-09-07Document tested B70 wide-query prefill gain, setup and quality scopec3c854e19.5 KB
    Loading...
  2. 2026-09-06Clarify 26B B70 Linux benchmarks and Wondernuttz custom OpenVINO fork17ba21c19.4 KB
    Loading...
  3. 2026-09-06Add clean Release validation, rollout status and shutdown caveat697216216.3 KB
    Loading...
  4. 2026-09-06Document September 6 Gemma MoE prefill runtime gains and validation scope1ed27e915.4 KB
    Loading...
  5. 2026-08-30Add verified OpenVINO INT4 AWQ group-size 64 banner4c0338812.2 KB
    Loading...
  6. 2026-07-26Document Arc B70 benchmark gains and optimized fork path0aeee4212.1 KB
    Loading...
  7. 2026-07-06Upload README.md with huggingface_hub21a41787.8 KB
    Loading...
  8. 2026-07-06Upload README.md with huggingface_hubfc752057.3 KB
    Loading...
  9. 2026-07-06Upload README.md with huggingface_hub26017c37.1 KB
    Loading...
  10. 2026-07-04Upload README.md with huggingface_hub2afa6da6.7 KB
    Loading...
  11. 2026-07-04Upload README.md with huggingface_hubc68e1686.8 KB
    Loading...
  12. 2026-07-04Upload README.md with huggingface_hub386508e6.7 KB
    Loading...
  13. 2026-07-03Upload README.md with huggingface_huba8e51706.5 KB
    Loading...
  14. 2026-07-03Upload README.md with huggingface_hub029e3e76.4 KB
    Loading...
  15. 2026-07-03Upload README.md with huggingface_hubfacdf554.7 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration