← back to catalog · registered 2026-08-22 13:56

valoomba/Gemma-4-31B-it-uncensored-heretic-Quark-W8A8-INT8

valoomba Gemma 29B multimodal second-order
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/valoomba%2FGemma-4-31B-it-uncensored-heretic-Quark-W8A8-INT8"
Response includes
  • classification m3
  • files 11
  • benchmarks 11 entries
  • hub_downloads_all_time 687
  • author_summary 2 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M3
Primary method

Layer-wise ablation

Applied on top of direct removal inherited from the base model.
Confidence
HIGH
Inherited from base model
Why this label 2 signals
Producer identity confirmed by naming conventions, tags or the model card. This label is very unlikely to change.
  • 'heretic' in model name (Heretic-produced)
  • Heretic uses layer-wise optimization (M3) with underlying direction removal (M1)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
HIGH
Why we say so
name contains 'heretic'; Heretic default extraction is difference-of-means (Arditi 2024)
Downloads · lifetime
687
Likes
0
Model age
4mo ago
created 2026-06-06
Downloads over time
Now838→from21↑3,890%
030761392021 on Jun 10838 on Oct 11JunJulAugSepOct
Jun 10 → Oct 11 · 57 snapshots · spans 123 days

Benchmarks

Benchmark Score Source
Entertainment 2.7 UGI
Hazardous 4.7 UGI
Natural Intelligence 34.73 UGI
Political lean -18.5% UGI
Sensitive-Info 33.23 UGI
SocPol 3 UGI
UGI 53.82 UGI
Willingness (10) 9.5 UGI
W10-Adherence 10 UGI
W10-Direct 9 UGI
Writing 38.26 UGI

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
apache-2.0
Languages
en
Tags
transformers safetensors gemma4 image-text-to-text gemma multimodal vision-language quantized int8 w8a8 quark vllm

Related

Total size
31.0 GB
Files
11
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-06-06 00:02

Files by quantization

Auxiliary files 11 files 31.0 GB
model-00001-of-00002.safetensors 24.6 GB 7a460a8a download
model-00002-of-00002.safetensors 6.41 GB 91694850 download
tokenizer.json 30.7 MB a2619fe1 download
model.safetensors.index.json 165 KB 923c3392 download
config.json 18.7 KB bdee6e0a download
chat_template.jinja 16.9 KB 7fa57f97 download
README.md 5.91 KB cbbfbf2d download
tokenizer_config.json 2.05 KB 375b25dc download
processor_config.json 1.65 KB 5465974d download
.gitattributes 1.53 KB 52373fe2 download
generation_config.json 217 B ed42ae71 download

README current version from Hugging Face


license: apache-2.0
library_name: transformers
pipeline_tag: image-text-to-text
base_model: llmfan46/gemma-4-31B-it-uncensored-heretic
language:

  • en
    tags:
  • transformers
  • safetensors
  • gemma4
  • image-text-to-text
  • gemma
  • multimodal
  • vision-language
  • quantized
  • int8
  • w8a8
  • quark
  • vllm
  • heretic
  • uncensored

Gemma-4-31B-it-uncensored-heretic-Quark-W8A8-INT8

W8A8 INT8 quantized version of
llmfan46/gemma-4-31B-it-uncensored-heretic using
AMD Quark.

Model Details

Field Value
Base Model llmfan46/gemma-4-31B-it-uncensored-heretic
Architecture Gemma4ForConditionalGeneration (multimodal: text + vision)
Parameters 31 B text decoder (quantized) + vision tower & embeddings kept in BF16
Quantization W8A8 INT8 (per-channel weight + per-token dynamic activation)
Quantizer AMD Quark 0.11.2 (ptpc_int8 scheme, pack_method='order')
Model Size ~33.3 GB (2 safetensors shards)
Original Size ~62.5 GB (BF16 source checkpoint)
Compression ~1.9x size reduction

Quantization Scheme

Component dtype Granularity Mode
Weight INT8 per-channel (ch_axis=0) symmetric, static
Activation INT8 per-token (ch_axis=1) symmetric, dynamic
lm_head BF16 — unquantized
embed_tokens BF16 — unquantized
vision_tower / embed_vision BF16 — unquantized (multimodal preserved)

Accuracy

GSM8K 8-shot evaluation on the full 1319-question test split using vLLM's
OpenAI-compatible chat API (temperature=0, concurrency=16, max_tokens=512,
standard chat template, and final-answer format #### <answer>). Values are
only populated when this card is regenerated with EVAL_BASELINE_JSON and/or
EVAL_QUANT_JSON from eval_gsm8k_chat_vllm.py.

Model Scheme Accuracy Correct
llmfan46/gemma-4-31B-it-uncensored-heretic (BF16 baseline) — not measured by this script —
This model (Quark W8A8 INT8) per-channel weight + per-token act. not measured by this script —

How to Use

With vLLM (Recommended)

# Generic vLLM serve command. Use TP=4 for 4x24GB RDNA3 cards.
vllm serve valoomba/Gemma-4-31B-it-uncensored-heretic-Quark-W8A8-INT8 \
    --tensor-parallel-size 4 \
    --max-model-len 8192 \
    --gpu-memory-utilization 0.9 \
    --limit-mm-per-prompt '{"image":0,"audio":0,"video":0}' \
    --trust-remote-code

# On the local RDNA3/gfx1100 vLLM fork, use the optimized INT8 Triton path:
vllm serve valoomba/Gemma-4-31B-it-uncensored-heretic-Quark-W8A8-INT8 \
    --tensor-parallel-size 4 \
    --max-model-len 8192 \
    --gpu-memory-utilization 0.9 \
    --limit-mm-per-prompt '{"image":0,"audio":0,"video":0}' \
    --trust-remote-code \
    --linear-backend triton
curl http://localhost:8000/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{
    "model": "valoomba/Gemma-4-31B-it-uncensored-heretic-Quark-W8A8-INT8",
    "messages": [{"role": "user", "content": "Hello! What is the capital of France?"}],
    "max_tokens": 256,
    "temperature": 0.7
  }'

Hardware Requirements

  • Minimum for one GPU: a GPU with roughly >=48 GB VRAM for typical short-context
    serving.
  • 4x24 GB consumer RDNA3: use tensor parallelism, for example
    --tensor-parallel-size 4, and keep multimodal limits disabled for text-only
    serving as shown above.
  • For longer context or larger batches, increase tensor parallelism or reduce
    --max-model-len / batch settings.

Quantization Details

This model was quantized using AMD Quark's per-token per-channel INT8 scheme:

  • Weight quantization: INT8 per-channel (one scale per output channel),
    symmetric, static.
  • Activation quantization: INT8 per-token (one scale per token), symmetric,
    dynamic (computed at inference time).
  • Excluded layers: lm_head, token embeddings, the full vision_tower, and
    embed_vision remain BF16.
  • Export: pack_method='order', weight_format='real_quantized', Quark
    metadata (quant_method='quark') → real INT8 weights with BF16/FP32 scales
    depending on loader/export path; no fake-quant weights are stored.
  • Quantization path: CPU file-to-file quantization with no calibration data and
    no full model graph load.

Reproduce Quantization

pip install amd-quark==0.11.1 huggingface_hub safetensors

HIP_VISIBLE_DEVICES="" ROCR_VISIBLE_DEVICES="" CUDA_VISIBLE_DEVICES="" \
MODEL_IN=llmfan46/gemma-4-31B-it-uncensored-heretic \
MODEL_OUT=./Gemma-4-31B-it-uncensored-heretic-Quark-W8A8-INT8 \
HF_REPO_ID=valoomba/Gemma-4-31B-it-uncensored-heretic-Quark-W8A8-INT8 \
python quantize_gemma4_heretic_quark_w8a8_cpu.py

The script builds the Quark config explicitly:

weight: INT8, per_channel, ch_axis=0, symmetric=True, is_dynamic=False
input:  INT8, per_channel, ch_axis=1, symmetric=True, is_dynamic=True

and then runs:

quantizer.direct_quantize_checkpoint(
    pretrained_model_path=MODEL_IN,
    save_path=MODEL_OUT,
    device="cpu",
)

Citation

If you use this model, please cite the original Gemma 4 release and the upstream
source model:

@misc{google2026gemma4,
  title  = {Gemma 4},
  author = {Google DeepMind},
  year   = {2026},
  url    = {https://huggingface.co/google/gemma-4-31B-it}
}

Upstream model: llmfan46/gemma-4-31B-it-uncensored-heretic.

License

This quantized derivative follows the license metadata of the upstream model,
which is listed as Apache 2.0 / Gemma 4 license metadata on Hugging Face. Check
the upstream model card and any included LICENSE / NOTICE files for complete
terms.

Modified files include the INT8-quantized safetensors checkpoint and the appended
Quark quantization_config block in config.json. No warranty of any kind is
provided.

README history 1 version

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-06-06Upload Gemma-4-31B-it-uncensored-heretic-Quark-W8A8-INT8 Quark W8A8 INT8 quan...1486f605.9 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration