← back to catalog · registered 2026-09-29 09:57

sjoe1244/gemma-4-31B-it-uncensored-heretic-exl3-4.50bpw-h8

sjoe1244 Gemma 31B multimodal second-order
curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/sjoe1244%2Fgemma-4-31B-it-uncensored-heretic-exl3-4.50bpw-h8"
Response includes
  • classification m3
  • files 13
  • author_summary 7 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M3
Primary method

Layer-wise ablation

Applied on top of direct removal inherited from the base model.
Confidence
HIGH
Inherited from base model
Why this label 2 signals
Producer identity confirmed by naming conventions, tags or the model card. This label is very unlikely to change.
  • 'heretic' in model name (Heretic-produced)
  • Heretic uses layer-wise optimization (M3) with underlying direction removal (M1)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
HIGH
Why we say so
name contains 'heretic'; Heretic default extraction is difference-of-means (Arditi 2024)
Downloads · 30-day
0
Likes
0
Model age
today
created 2026-09-29

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
apache-2.0
Tags
safetensors gemma4 exl3 exllamav3 heretic uncensored vision quantized base_model:llmfan46/gemma-4-31B-it-uncensored-heretic base_model:quantized:llmfan46/gemma-4-31B-it-uncensored-heretic license:apache-2.0 region:us

Related

Total size
19.7 GB
Files
13
Quantizations
1
Registered
2026-09-29 09:57
Last updated on HF
2026-09-29 10:04

Files by quantization

Auxiliary files 13 files 19.8 GB
model-00001-of-00003.safetensors 7.86 GB 01261222 download
model-00002-of-00003.safetensors 7.83 GB 1ca66f30 download
model-00003-of-00003.safetensors 4.04 GB 91cabb5f download
tokenizer.json 30.7 MB a2619fe1 download
quantization_config.json 648 KB 8a84cb0b download
model.safetensors.index.json 302 KB 8e52bdb5 download
chat_template.jinja 22.5 KB 4486ce43 download
config.json 5.64 KB 1eb8bec9 download
README.md 2.37 KB 3be48bf0 download
tokenizer_config.json 2.05 KB 375b25dc download
processor_config.json 1.65 KB 5465974d download
.gitattributes 1.53 KB 52373fe2 download
generation_config.json 217 B ed42ae71 download

README current version from Hugging Face


license: apache-2.0
base_model: llmfan46/gemma-4-31B-it-uncensored-heretic
base_model_relation: quantized
tags:

  • exl3
  • exllamav3
  • gemma4
  • heretic
  • uncensored
  • vision
  • quantized
    inference: false

gemma-4-31B-it-uncensored-heretic-exl3-4.50bpw-h8

EXL3 export of llmfan46/gemma-4-31B-it-uncensored-heretic
for ExLlamaV3 / TabbyAPI. Vision tower quantized to 6-bit.

Load with ExLlamaV3 or TabbyAPI's exllamav3 loader. This will not load in
Transformers, vLLM or llama.cpp.

Not the coder3101 heretic — that daily quant lives at
sjoe1244/gemma-4-31B-it-heretic-exl3-4.00bpw-h6
(base_model: coder3101/gemma-4-31B-it-heretic). This repo is the uncensored-heretic lineage.

Quantization

Method EXL3
ExLlamaV3 1.5.3
Weights 4.5 bpw
Head 8 bits
Vision 6 bpw
Codebook mul1 (default)
Calibration 250 rows x 2048 cols

Converted on an RTX 4090 on 2026-09-29 from the local BF16 source
(/mnt/ssd2/models/llm/gemma-4-31B-it-uncensored-heretic-bf16, ~59 GB,
Gemma4ForConditionalGeneration).

Size

File Bytes
model-00001-of-00003.safetensors 8,437,700,726
model-00002-of-00003.safetensors 8,402,120,533
model-00003-of-00003.safetensors 4,342,120,961
Total 21,215,117,785 (19.758 GiB)

Measured on RTX 4090 (24 GB)

TabbyAPI config: max_seq_len / cache_size 139264, cache_mode 4,4, max_batch_size 2,
chunk_size 2048, vision: false (same as the daily gemma31-400-q4.yml lane).

Metric Value
Peak VRAM @ ~130K prompt n/a
Decode speed (wall, 128 tok) n/a tok/s
WikiText-2 PPL (8×2048, short) 975.521499
KL vs BF16 (model_diff, 8×2048) n/a

Recommended Tabby settings (24 GB)

max_seq_len: 139264
cache_size: 139264
cache_mode: 4,4
chunk_size: 2048
max_batch_size: 2
vision: false
prompt_template: gemma4

If peak VRAM exceeds ~23.5 GB, drop context or use a lower bpw / more aggressive KV.

Credit and license

Apache 2.0, matching google/gemma-4-31B-it and the source model.
Decensoring credit belongs to llmfan46 — this
repo is only the EXL3 export. Quantization by
ExLlamaV3.

Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.