license: apache-2.0
base_model: llmfan46/gemma-4-31B-it-qat-q4_0-unquantized-uncensored-heretic
tags:
- openvino
- int4
- intel-arc
- gemma4
- roleplay
- uncensored
language: - en
Gemma-4 31B Heretic, OpenVINO INT4. The smarter, slower sibling, on ONE Intel Arc B70 (32 GB).
This is the abliterated ("heretic") Gemma-4 31B dense model, exported to OpenVINO INT4 and
graph-patched with the same RoPE fix as its
26B-A4B MoE sibling.
Stock exports of this model produce word salad at long context on Intel GPUs. This build does not.
Pick this one when you want richer prose and deeper in-character writing and can accept slower
replies. Pick the 26B MoE when you want speed. Same uncensored base, same prompt format, same
serving stack. I personally run these for roleplay (Skyrim companion AI and Discord character
bots); this is a general instruction model with the refusal behavior removed.
Everything below was measured on a single Intel Arc Pro B70 (32 GB), OpenVINO 2026.2.
Toolkit: patches, OpenAI-compatible server, quickstart, and the full bug catalog live at github.com/Wondernuttz/OpenVino-For-Gemma-4.
Measured performance (single Arc Pro B70)
| Metric | Value |
|---|---|
| Decode, short context | 27.5 tok/s |
| Decode at 6K context | 19.4 tok/s |
| Prefill, 512 tokens | 1,574 tok/s |
| Prefill, 2K tokens | 1,656 tok/s |
| Prefill, 6K tokens | 1,191 tok/s |
| Prefill, 6,620-token coherence prompt | 1,123 tok/s |
| Model load | ~23-29 s |
These numbers use DYNAMIC_QUANTIZATION_GROUP_SIZE: 128. A matched fresh-process sweep against
DQGS 0 produced the following result:
| Input | DQGS 0 | DQGS 128 | Gain |
|---|---|---|---|
| 512 | 1,079 tok/s | 1,574 tok/s | +45.8% |
| 2,048 | 1,264 tok/s | 1,656 tok/s | +31.0% |
| 6,144 | 987 tok/s | 1,191 tok/s | +20.7% |
| 6,620 coherence | 945 tok/s | 1,123 tok/s | +18.8% |
Short decode remained 27.5 tok/s. Both settings passed all four long-prompt retrieval and style
checks and produced the exact same 110-token answer. An exact 16K generation was also attempted,
but the DQGS 0 baseline crossed the card's VRAM cliff and did not finish within ten minutes. It
was discarded rather than reported as a score.
Long-context capability and the honest single-card limit
Needle retrieval test: a password fact planted early in the prompt, retrieved at the end.
| Context | Thinking OFF | Thinking ON |
|---|---|---|
| 8K | PASS (10.5 s) | PASS (20.1 s) |
| 16K | PASS (119 s) | not practical, see below |
The dense 31B fills the whole card: 18.6 GB of weights plus KV cache saturates 32 GB near 16K
context with thinking generation on top. Retrieval stays CORRECT (the rope patch works), but the
card runs out of memory headroom and speed collapses. Practical guidance on one B70: treat this
as an 8K thinking model or a 16K no-think model. The 26B MoE sibling is verified to 32K with
thinking on, because its weights leave twice the KV headroom. Dual-GPU or a larger-memory card
should extend this build further; the shipped position tables go to 131K.
How this compares to published numbers
Best published number for this model on the same card is 21.7 tok/s decode via llama.cpp SYCL
(PMZFX benchmarks); this build does ~27
tok/s decode, about 1.24x. Current DQGS 128 prefill at pp512 is 1,574 tok/s versus SYCL's 601,
about 2.6x (TTFT-based, so slightly conservative). For transparency: an
earlier version of this card quoted ~1,074 tok/s at a 6K prompt; re-measurement confirmed that
figure (~1,030 clean), it was accurate. Community numbers for a dense Gemma-4 31B at 4-bit on an RTX 3090 run
about 30-34 tok/s, so one B70 lands just under 3090 speed at far lower power.
What was fixed
Same two fixes as the 26B repo, plus one discovery worth repeating for anyone patching Gemma-4
exports: Intel GPUs execute the RoPE angle math in fp16, which cannot represent large rotation
angles, so long context collapses. This build replaces the runtime sin/cos with precomputed
lookup tables and a Gather. And Gemma-4 global attention uses proportional (partial) RoPE, so
three quarters of the global frequency values are zero BY DESIGN. Do not "repair" the zeros;
preserve them. Our first patch of this 31B filled them in with the standard geometric formula and
the model visibly degraded at 16K. The build in this repo preserves them and passes retrieval.
How to run
Identical to the 26B sibling: same <|turn> prompt format, same thinking pre-close trick, same
warnings (repetition penalty ~1.2 and never use JSON grammar mode). UseDYNAMIC_QUANTIZATION_GROUP_SIZE: 128 with the ordinary VLMPipeline. Continuous batching at
the 16K memory edge was not part of this sweep. See the
26B model card
for the full runnable snippet; just point it at this folder.
pip install openvino-genai==2026.2.0 huggingface_hub
huggingface-cli download Wondernutts/gemma-4-31B-it-qat-q4_0-unquantized-uncensored-heretic-int4-ov --local-dir ./gemma4-31b-heretic-ov
import openvino_genai as genai
pipe = genai.VLMPipeline(
"./gemma4-31b-heretic-ov",
"GPU",
**{"DYNAMIC_QUANTIZATION_GROUP_SIZE": 128},
)
Provenance
google/gemma-4-31B-it (QAT q4_0 unquantized), heretic abliteration by
llmfan46,
OpenVINO INT4 AWQ export, RoPE lookup-table graph patch (this repo).
Intended use and content notice
Uncensored general model, built and tested for roleplay and creative writing on local Intel hardware. The abliteration removes refusal behavior and outputs are unfiltered; you are responsible for lawful and appropriate use. Licensed under
Apache 2.0, same as the upstream Gemma 4 release.