← back to catalog · registered 2026-08-22 13:56

alonsoko/gemma-4-31b-it-abliterated-heretic-AWQ-W4A16

alonsoko Gemma 29B multimodal second-order
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/alonsoko%2Fgemma-4-31b-it-abliterated-heretic-AWQ-W4A16"
Response includes
  • classification m3
  • files 10
  • hub_downloads_all_time 209,293
  • author_summary 2 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M3
Primary method

Layer-wise ablation

Applied on top of direct removal inherited from the base model.
Confidence
HIGH
Inherited from base model
Why this label 2 signals
Producer identity confirmed by naming conventions, tags or the model card. This label is very unlikely to change.
  • 'heretic' in model name (Heretic-produced)
  • Heretic uses layer-wise optimization (M3) with underlying direction removal (M1)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
HIGH
Why we say so
name contains 'heretic'; Heretic default extraction is difference-of-means (Arditi 2024)
Downloads · lifetime
209K
4K last 30d - cooling
Likes
15
Model age
6mo ago
created 2026-04-06
Downloads over time
Now210.7K→from11.5K↑1,732%
1.5K77.9K154.2K230.6K11.5K on May 13210.7K on Oct 11MayJunJulAugSepOct
May 13 → Oct 11 · 63 snapshots · spans 151 days

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
apache-2.0
Tags
transformers safetensors gemma4 image-text-to-text gemma gemma-4 multimodal vision-language abliterated heretic ara uncensored

Related

Total size
17.8 GB
Files
10
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-05-17 19:41

Files by quantization

Auxiliary files 10 files 17.8 GB
model.safetensors 17.8 GB 550f07aa download
tokenizer.json 30.7 MB 1cc9316e download
recipe.yaml 38.1 KB e5b29309 download
config.json 18.1 KB 8c950413 download
chat_template.jinja 11.8 KB 33c51c2d download
README.md 5.97 KB 951f249f download
tokenizer_config.json 2.70 KB 157943d1 download
processor_config.json 1.65 KB 5465974d download
.gitattributes 1.53 KB 52373fe2 download
generation_config.json 203 B ac6688a9 download

README current version from Hugging Face


license: apache-2.0
license_link: https://ai.google.dev/gemma/docs/gemma_4_license
pipeline_tag: image-text-to-text
library_name: transformers
base_model: trohrbaugh/gemma-4-31b-it-heretic-ara
base_model_relation: quantized
tags:

  • gemma
  • gemma-4
  • multimodal
  • vision-language
  • abliterated
  • heretic
  • ara
  • uncensored
  • decensored
  • awq
  • w4a16
  • quantized
  • compressed-tensors
  • vllm

Gemma-4-31B-it-abliterated-heretic-AWQ-W4A16

AWQ W4A16 (group_size 128, symmetric) quantization of
trohrbaugh/gemma-4-31b-it-heretic-ara
— a Heretic-ARA abliterated derivative of google/gemma-4-31b-it.

⚠️ Decensored model. Safety guardrails have been deliberately removed.
Research and experimentation only. See full disclaimer below.

Quantization Details

Parameter Value
Method AWQ (Activation-aware Weight Quantization)
Scheme W4A16 (symmetric)
Weight Bits 4
Activation Bits 16
Group Size 128
Format compressed-tensors
Calibration Dataset HuggingFaceH4/ultrachat_200k
Calibration Samples 256
Max Sequence Length 2048
Vision Tower Unquantized (full precision)
LM Head Unquantized (full precision)
Compatible Inference Engine vLLM (vllm/vllm-openai:gemma4)

Quantization Notes

  • All multimodal paths kept full precision: Vision tower, audio tower, video
    tower, multi-modal projector, and all modality-specific embedding and
    projection layers are excluded from quantization. Only language-model linear
    layers (attention Q/K/V/O and MLP gate/up/down) are quantized.
  • LM head unquantized: Standard practice to preserve output token
    distribution quality at negligible size cost.
  • v_proj → o_proj smoothing skipped: llm-compressor reports incompatible
    balance layer dimensions for v_proj → o_proj on this checkpoint across many
    of Gemma 4's decoder blocks. Per-channel smoothing is skipped for that path,
    matching the standard AWQ-for-GQA pattern.
  • Hybrid attention and context window unchanged: Quantization only touches
    Linear layer weights; Gemma 4's interleaved local/global attention pattern
    and 256K context capacity are structurally preserved. Actual long-context
    quality at 4-bit has not been benchmarked.
  • Multimodal input intact: Text + image input works as in the base model.
    Use the standard Gemma 4 chat template with image tokens placed before text.
  • Full quantization recipe is preserved in recipe.yaml in this repo for
    reproducibility. It records the exact AWQ mappings, ignore patterns, and
    scheme parameters applied.

Deployment

Recommended inference with vLLM (Gemma 4 requires the vllm/vllm-openai:gemma4 image):

vllm serve alonsoko/gemma-4-31b-it-abliterated-heretic-AWQ-W4A16 \
    --trust-remote-code \
    --tensor-parallel-size 1 \
    --max-model-len 32768

Recommended sampling (per upstream Gemma 4 guidance):
temperature=1.0, top_p=0.95, top_k=64

To enable thinking mode, include the <|think|> token at the start of the system prompt.

Hardware Requirements

Approximate VRAM for inference at this quantization (W4A16-G128):

  • Weights (quantized language model + unquantized vision tower, all in one safetensors): ~19 GB
  • KV cache (per request, grows with context length): ~1-4 GB at 32K context, more at longer contexts
  • Recommended: Single 24 GB GPU (RTX 3090/4090, A10G) for standard context up to ~32K,
    or single 48 GB GPU (L40S/A6000) for long context / batch serving
  • Long context (128K+): 48-80 GB recommended due to KV cache growth

⚠️ Disclaimer

This model is intended for research, experimentation, and testing purposes only.

  • This model may produce harmful, offensive, inappropriate, or otherwise objectionable content.
  • The abliteration process removes safety guardrails that were intentionally built into the original model.
  • Do not use this model in production systems, consumer-facing applications, or any context
    where harmful outputs could cause real-world harm.
  • The authors and contributors of this toolkit bear no responsibility for any misuse of this model
    or any harm caused by outputs generated by this model.
  • By using this model, you agree that you are solely responsible for ensuring its use complies
    with all applicable laws and ethical guidelines.

This model is shared purely for academic and technical exploration of model internals.

Abliteration

Performed with Heretic v1.2.0+custom using the
Arbitrary-Rank Ablation (ARA) method.

Parameter Value
start_layer_index 2
end_layer_index 60
preserve_good_behavior_weight 0.9920
steer_bad_behavior_weight 0.0001
overcorrect_relative_weight 0.4709
neighbor_count 10

Performance

Metric This model Original google/gemma-4-31b-it
KL divergence 0.0120 0 (by definition)
Refusals 5/100 98/100

Measured on the unquantized heretic-ara base; AWQ is expected to preserve these closely but has not been separately benchmarked.


About the Base Model

Original model: google/gemma-4-31b-it

Gemma 4 31B is a dense multimodal model (text + image input, text output) with
a 256K context window, native thinking-mode support, function calling, and
strong performance on reasoning, coding, and vision benchmarks. See the
base model card for architectural
details, benchmark results, training data, and Google's responsible-use guidance.

Gemma 4

README history 5 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-05-17Update README.md55e73ef6 KB
    Loading...
  2. 2026-04-17Update README.md9dfe0ae6 KB
    Loading...
  3. 2026-04-09Update README.md0d5e4cf28.5 KB
    Loading...
  4. 2026-04-08Update README.md68d671227.8 KB
    Loading...
  5. 2026-04-06Create README.md4afd4d627 KB
    Loading...

Discussions 1 thread

  1. 2026-04-06quant fails to load in VLLM due to vision tower?closed3 💬#1
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration