← back to catalog · registered 2026-08-22 13:56

Wondernutts/gemma-4-31B-it-qat-q4_0-unquantized-uncensored-heretic-int4-ov

Wondernutts Gemma 31B second-order
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/Wondernutts%2Fgemma-4-31B-it-qat-q4_0-unquantized-uncensored-heretic-int4-ov"
Response includes
  • classification m3
  • files 22
  • hub_downloads_all_time 127
  • author_summary 7 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M3
Primary method

Layer-wise ablation

Applied on top of direct removal inherited from the base model.
Confidence
HIGH
Inherited from base model
Why this label 2 signals
Producer identity confirmed by naming conventions, tags or the model card. This label is very unlikely to change.
  • 'heretic' in model name (Heretic-produced)
  • Heretic uses layer-wise optimization (M3) with underlying direction removal (M1)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
HIGH
Why we say so
name contains 'heretic'; Heretic default extraction is difference-of-means (Arditi 2024)
Downloads · lifetime
127
21 last 30d - stable
Likes
0
Model age
3mo ago
created 2026-06-15
Downloads over time
Now138→from2↑6,800%
0511011522 on Jul 1138 on Oct 11JulAugSepOct
Jul 1 → Oct 11 · 54 snapshots · spans 102 days

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
apache-2.0
Languages
en
Tags
openvino gemma4 int4 intel-arc roleplay uncensored en base_model:llmfan46/gemma-4-31B-it-qat-q4_0-unquantized-uncensored-heretic base_model:finetune:llmfan46/gemma-4-31B-it-qat-q4_0-unquantized-uncensored-heretic license:apache-2.0 region:us

Related

Total size
18.1 GB
Files
22
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-07-26 08:47

Files by quantization

Auxiliary files 22 files 18.1 GB
openvino_language_model.bin 16.2 GB df6e788d download
openvino_text_embeddings_model.bin 1.31 GB a861d72e download
openvino_vision_embeddings_model.bin 550 MB b7bf5c3d download
openvino_tokenizer.bin 16.5 MB d151f48d download
openvino_detokenizer.bin 4.21 MB 461919bd download
openvino_text_embeddings_per_layer_model.bin 20.0 B db06b2c7 download
tokenizer.json 30.7 MB a2619fe1 download
openvino_language_model.xml 6.89 MB e613cb6a download
openvino_vision_embeddings_model.xml 3.07 MB 6789ca2c download
openvino_tokenizer.xml 48.7 KB a6773b7f download
openvino_detokenizer.xml 33.7 KB fd228db6 download
chat_template.jinja 22.5 KB 4486ce43 download
openvino_text_embeddings_model.xml 6.93 KB 31827c46 download
README.md 5.95 KB 8b5175c7 download
openvino_text_embeddings_per_layer_model.xml 4.67 KB aa56e2fa download
config.json 4.50 KB 6a434fec download
tokenizer_config.json 2.75 KB 4a794cf8 download
processor_config.json 1.65 KB 5465974d download
.gitattributes 1.53 KB 52373fe2 download
openvino_config.json 1.12 KB 175154ed download
preprocessor_config.json 403 B 1b1350e0 download
generation_config.json 203 B edda3c19 download

README current version from Hugging Face


license: apache-2.0
base_model: llmfan46/gemma-4-31B-it-qat-q4_0-unquantized-uncensored-heretic
tags:

  • openvino
  • int4
  • intel-arc
  • gemma4
  • roleplay
  • uncensored
    language:
  • en

Gemma-4 31B Heretic, OpenVINO INT4. The smarter, slower sibling, on ONE Intel Arc B70 (32 GB).

This is the abliterated ("heretic") Gemma-4 31B dense model, exported to OpenVINO INT4 and
graph-patched with the same RoPE fix as its
26B-A4B MoE sibling.
Stock exports of this model produce word salad at long context on Intel GPUs. This build does not.

Pick this one when you want richer prose and deeper in-character writing and can accept slower
replies. Pick the 26B MoE when you want speed. Same uncensored base, same prompt format, same
serving stack. I personally run these for roleplay (Skyrim companion AI and Discord character
bots); this is a general instruction model with the refusal behavior removed.

Everything below was measured on a single Intel Arc Pro B70 (32 GB), OpenVINO 2026.2.

Toolkit: patches, OpenAI-compatible server, quickstart, and the full bug catalog live at github.com/Wondernuttz/OpenVino-For-Gemma-4.

Measured performance (single Arc Pro B70)

Metric Value
Decode, short context 27.5 tok/s
Decode at 6K context 19.4 tok/s
Prefill, 512 tokens 1,574 tok/s
Prefill, 2K tokens 1,656 tok/s
Prefill, 6K tokens 1,191 tok/s
Prefill, 6,620-token coherence prompt 1,123 tok/s
Model load ~23-29 s

These numbers use DYNAMIC_QUANTIZATION_GROUP_SIZE: 128. A matched fresh-process sweep against
DQGS 0 produced the following result:

Input DQGS 0 DQGS 128 Gain
512 1,079 tok/s 1,574 tok/s +45.8%
2,048 1,264 tok/s 1,656 tok/s +31.0%
6,144 987 tok/s 1,191 tok/s +20.7%
6,620 coherence 945 tok/s 1,123 tok/s +18.8%

Short decode remained 27.5 tok/s. Both settings passed all four long-prompt retrieval and style
checks and produced the exact same 110-token answer. An exact 16K generation was also attempted,
but the DQGS 0 baseline crossed the card's VRAM cliff and did not finish within ten minutes. It
was discarded rather than reported as a score.

Long-context capability and the honest single-card limit

Needle retrieval test: a password fact planted early in the prompt, retrieved at the end.

Context Thinking OFF Thinking ON
8K PASS (10.5 s) PASS (20.1 s)
16K PASS (119 s) not practical, see below

The dense 31B fills the whole card: 18.6 GB of weights plus KV cache saturates 32 GB near 16K
context with thinking generation on top. Retrieval stays CORRECT (the rope patch works), but the
card runs out of memory headroom and speed collapses. Practical guidance on one B70: treat this
as an 8K thinking model or a 16K no-think model.
The 26B MoE sibling is verified to 32K with
thinking on, because its weights leave twice the KV headroom. Dual-GPU or a larger-memory card
should extend this build further; the shipped position tables go to 131K.

How this compares to published numbers

Best published number for this model on the same card is 21.7 tok/s decode via llama.cpp SYCL
(PMZFX benchmarks); this build does ~27
tok/s decode, about 1.24x. Current DQGS 128 prefill at pp512 is 1,574 tok/s versus SYCL's 601,
about 2.6x (TTFT-based, so slightly conservative). For transparency: an
earlier version of this card quoted ~1,074 tok/s at a 6K prompt; re-measurement confirmed that
figure (~1,030 clean), it was accurate. Community numbers for a dense Gemma-4 31B at 4-bit on an RTX 3090 run
about 30-34 tok/s, so one B70 lands just under 3090 speed at far lower power.

What was fixed

Same two fixes as the 26B repo, plus one discovery worth repeating for anyone patching Gemma-4
exports: Intel GPUs execute the RoPE angle math in fp16, which cannot represent large rotation
angles, so long context collapses. This build replaces the runtime sin/cos with precomputed
lookup tables and a Gather. And Gemma-4 global attention uses proportional (partial) RoPE, so
three quarters of the global frequency values are zero BY DESIGN. Do not "repair" the zeros;
preserve them. Our first patch of this 31B filled them in with the standard geometric formula and
the model visibly degraded at 16K. The build in this repo preserves them and passes retrieval.

How to run

Identical to the 26B sibling: same <|turn> prompt format, same thinking pre-close trick, same
warnings (repetition penalty ~1.2 and never use JSON grammar mode). Use
DYNAMIC_QUANTIZATION_GROUP_SIZE: 128 with the ordinary VLMPipeline. Continuous batching at
the 16K memory edge was not part of this sweep. See the
26B model card
for the full runnable snippet; just point it at this folder.

pip install openvino-genai==2026.2.0 huggingface_hub
huggingface-cli download Wondernutts/gemma-4-31B-it-qat-q4_0-unquantized-uncensored-heretic-int4-ov --local-dir ./gemma4-31b-heretic-ov
import openvino_genai as genai

pipe = genai.VLMPipeline(
    "./gemma4-31b-heretic-ov",
    "GPU",
    **{"DYNAMIC_QUANTIZATION_GROUP_SIZE": 128},
)

Provenance

google/gemma-4-31B-it (QAT q4_0 unquantized), heretic abliteration by
llmfan46,
OpenVINO INT4 AWQ export, RoPE lookup-table graph patch (this repo).

Intended use and content notice

Uncensored general model, built and tested for roleplay and creative writing on local Intel hardware. The abliteration removes refusal behavior and outputs are unfiltered; you are responsible for lawful and appropriate use. Licensed under
Apache 2.0, same as the upstream Gemma 4 release.

README history 8 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-07-26Document Arc B70 DQGS=128 benchmarks0044be66 KB
    Loading...
  2. 2026-07-06Upload README.md with huggingface_hubb181ea25 KB
    Loading...
  3. 2026-07-06Upload README.md with huggingface_hub300401f4.9 KB
    Loading...
  4. 2026-07-06Upload README.md with huggingface_hub23076f04.8 KB
    Loading...
  5. 2026-07-04Upload README.md with huggingface_hub743c8574.6 KB
    Loading...
  6. 2026-07-04Upload README.md with huggingface_hub6a6306b4.7 KB
    Loading...
  7. 2026-07-04Upload README.md with huggingface_hub36b11634.5 KB
    Loading...
  8. 2026-06-15Create README.md0b3e98e105 B
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration