license: gemma
base_model: zaakirio/gemma-4-12b-it-uncensored
base_model_relation: quantized
tags:
- gemma4_unified
- quantization
- autoround
- w4a16
- vllm
gemma-4-12b-it-uncensored-W4A16-AutoRound
4-bit weight / 16-bit activation (W4A16) quantization ofzaakirio/gemma-4-12b-it-uncensored.
Attribution
- Original model:
google/gemma-4-12B-it— © Google, released under the Gemma license. - Immediate base (quantized here):
zaakirio/gemma-4-12b-it-uncensored - This repository only quantizes that base; all model capabilities and
weights originate upstream. License inherits from the original Gemma model.
Quantization
- Tool: Intel AutoRound (
v0.14.0) - Scheme: W4A16 — 4-bit weights, group size 128,
sym=True - Mode: RTN (
iters=0) — required for Gemma 4 numerical stability - Format: native
auto_round(quant_method=auto-round) - Multimodal projections (vision/audio embedders) kept unquantized
(quant_nontext_module=False) to preserve the encoder-free multimodal path.
Serving (vLLM)
Encoder-free Gemma 4 Unified; serve with a vLLM build that supportsGemma4UnifiedForConditionalGeneration. On Turing (Tesla T4) the Triton
attention TILE_SIZE patch (vllm-project/vllm#39018) is required.