← back to catalog · registered 2026-10-10 05:58

sjoe1244/gemma-4-31B-it-uncensored-heretic-exl3-4.00bpw-h4-vb4

sjoe1244 Gemma 31B multimodal second-order
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/sjoe1244%2Fgemma-4-31B-it-uncensored-heretic-exl3-4.00bpw-h4-vb4"
Response includes
  • classification m3
  • files 13
  • author_summary 8 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M3
Primary method

Layer-wise ablation

Applied on top of direct removal inherited from the base model.
Confidence
HIGH
Inherited from base model
Why this label 2 signals
Producer identity confirmed by naming conventions, tags or the model card. This label is very unlikely to change.
  • 'heretic' in model name (Heretic-produced)
  • Heretic uses layer-wise optimization (M3) with underlying direction removal (M1)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
HIGH
Why we say so
name contains 'heretic'; Heretic default extraction is difference-of-means (Arditi 2024)
Downloads · 30-day
0
Likes
0
Model age
today
created 2026-10-10

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
apache-2.0
Tags
exllamav3 safetensors gemma4 exl3 heretic uncensored vision quantized base_model:llmfan46/gemma-4-31B-it-uncensored-heretic base_model:quantized:llmfan46/gemma-4-31B-it-uncensored-heretic license:apache-2.0 4-bit

Related

Total size
17.2 GB
Files
13
Quantizations
1
Registered
2026-10-10 05:58
Last updated on HF
2026-10-10 05:27

Files by quantization

Auxiliary files 13 files 17.3 GB
model-00002-of-00003.safetensors 7.97 GB 8b8f3824 download
model-00001-of-00003.safetensors 7.84 GB aa72a290 download
model-00003-of-00003.safetensors 1.43 GB 940759e8 download
tokenizer.json 30.7 MB a2619fe1 download
quantization_config.json 648 KB 3074de6d download
model.safetensors.index.json 302 KB 0de11def download
chat_template.jinja 18.4 KB b0b68f32 download
config.json 5.64 KB 1e0569f4 download
README.md 5.43 KB 288deef7 download
tokenizer_config.json 3.01 KB 6068e357 download
processor_config.json 1.65 KB 5465974d download
.gitattributes 1.53 KB 52373fe2 download
generation_config.json 217 B ed42ae71 download

README current version from Hugging Face


license: apache-2.0
base_model: llmfan46/gemma-4-31B-it-uncensored-heretic
base_model_relation: quantized
library_name: exllamav3
tags:

  • exl3
  • exllamav3
  • gemma4
  • heretic
  • uncensored
  • vision
  • quantized
    inference: false

Gemma 4 31B Uncensored Heretic · EXL3 4.00bpw-h4-vb4

An ExLlamaV3 quant of
llmfan46/gemma-4-31B-it-uncensored-heretic.
It is built for long context with vision on a single 24 GB card. Weights, head, vision tower and KV cache are all 4-bit.

It runs in TabbyAPI or ExLlamaV3 only. It does not run in Transformers, vLLM or llama.cpp.

At a glance (one RTX 4090, 24 GB)

Download 18.5 GB
Context pool ~194K tokens with vision on (4-bit KV cache)
Writing speed 45 tok/s (short prompt), 42 tok/s at 32K
Quality (KL vs original, lower is better) 0.026, the same as our 4.00bpw-h6

Which one should I download?

Quant Size KL vs original Context on a 4090 Pick it if…
4.00bpw-h4-vb4 (this) 18.5 GB 0.026 ~194K, vision on you want the most context, with images
4.00bpw-h6 19.0 GB 0.026 ~139K, vision off you're on ExLlamaV3 1.5.x
4.15bpw-h6 19.6 GB 0.025 ~139K, vision off you want a hair more quality than 4.00
4.50bpw-h8 21.2 GB 0.018 70K–121K you want the best quality

Quick start (TabbyAPI, 24 GB)

model:
  max_seq_len: 64000
  cache_size: 194560
  cache_mode: 4,4
  chunk_size: 1024
  max_batch_size: 2
  vision: true
  prompt_template: gemma4

This exact setup survived two ~63K-token prompts running at the same time.
With vision: false you can raise cache_size to 206848.

Good to know

  • Made and tested with ExLlamaV3 1.6.0. It was not tried on older versions.
  • The 4-bit cache costs some accuracy. It adds about a quarter more drift than a full-precision cache.
    If you don't need the room, cache_mode: 6,6 or 8,8 is more accurate. Those modes were not fit-tested on this quant.
  • Vision loads and fits in the numbers above. Image quality was not benchmarked.
  • Not the coder3101 heretic. That one is
    sjoe1244/gemma-4-31B-it-heretic-exl3-4.00bpw-h6.
Full measurements and test setup

How it was made

Method EXL3, ExLlamaV3 1.6.0
Weights 4.00 bpw
Head (lm_head) 4 bits
Vision tower 4 bits (190 linear layers)
Embeddings BF16 (unquantized)
Codebook mul1, out_scales auto
Calibration 250 rows x 2048 cols (ExLlamaV3 default set)
Source llmfan46/gemma-4-31B-it-uncensored-heretic, BF16
File Bytes
model-00001-of-00003.safetensors 8,415,898,175
model-00002-of-00003.safetensors 8,557,177,089
model-00003-of-00003.safetensors 1,539,495,613
Total 18,512,570,877 (17.24 GiB)

Quality vs the BF16 source

The test is ExLlamaV3 eval/qbench.py on WikiText-2 test, chat-templated, 20 x 2048 rows.
The reference is BF16 logits from the unquantized source, with the same tokens for every model.

Quant KL mean KL median KL p90 Perplexity
BF16 — — — 16.586
4.00bpw-h4-vb4 (this) 0.0262 0.0045 0.0430 16.687
4.00bpw-h6 0.0262 0.0041 0.0407 16.655
4.50bpw-h8 0.0181 0.0029 0.0303 16.624

The mean is the same as 4.00bpw-h6. The median and tail are a little worse, because of the 4-bit head.

KV cache cost

The test is eval/model_diff.py on raw WikiText-2, 20 x 2048 rows. Only the comparison between rows is meaningful,
because raw-text scores are inflated for an instruct model.

Setup KL mean KL median KL p90
This quant, FP16 cache 0.400 0.041 0.896
This quant, 4,4 cache 0.497 0.068 1.212
4.50bpw-h8, 8,8 cache 0.328 0.028 0.673

Fit on one RTX 4090 (24,564 MiB)

The setup was TabbyAPI with ExLlamaV3 1.6.0, cache_mode 4,4, max_batch_size 2, chunk_size 1024 and max_seq_len 64000.

  • The largest pool that loads was found by bisecting in 2,048-token steps.
  • Stress-tested means the pool also survived two ~62.8K-token prompts prefilling at the same time.
  • VRAM is the whole-GPU nvidia-smi reading.
Setup Loads and stress-tested VRAM idle VRAM peak
This quant, vision on 194,560 22,263 MiB 22,947 MiB
This quant, vision off 206,848 — 22,879 MiB
4.50bpw-h8, vision off (same image, 4,4) 120,832 — 23,781 MiB

Each 8,192 tokens of pool costs about 180 MB at 4,4. Vision costs about 12K tokens of pool.

Speed (one request at a time)

Prompt Prefill Decode
2K tokens ~785 tok/s 45.0 tok/s
32K tokens ~1,620 tok/s 41.9 tok/s

For comparison, 4.50bpw-h8 decoded at 39.9 and 37.5 tok/s on the same setup.

Credit and license

Apache 2.0, the same as google/gemma-4-31B-it and the source model. The uncensoring is
llmfan46's work. This repo only adds the EXL3 export.

README history 1 version

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-10-10EXL3 4.00bpw, 4-bit head, 4-bit vision tower (ExLlamaV3 1.6.0)6aee44b5.4 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration