license: gemma
library_name: transformers
base_model:
- Nabbers1999/Gemma-3-27B-it-NP-Abliterated
- google/gemma-3-27b-it
base_model_relation: quantized
pipeline_tag: image-text-to-text
tags: - gemma3
- int4
- gptq
- 4bit
- vllm
- compressed-tensors
- quantized
- abliterated
- multimodal
- vision
language: - en
datasets: - neuralmagic/calibration
Gemma 3 27B IT NP-Abliterated — GPTQ INT4 (128g)
GPTQ INT4 quantization of Nabbers1999/Gemma-3-27B-it-NP-Abliterated, a biprojected norm-preserving abliteration of Google's Gemma 3 27B IT.
Vision is fully preserved. Only the language model decoder layers are quantized to INT4. The SigLIP vision encoder and multimodal projector remain in BF16.
Quantization details
| Property | Value |
|---|---|
| Method | GPTQ (4-bit, symmetric, per-group) |
| Group size | 128 |
| Format | compressed-tensors |
| Calibration data | neuralmagic/calibration (1024 samples, seq_len 2048) |
| Dampening | 0.07 |
| Layers preserved in BF16 | lm_head, embed_tokens, vision_tower, multi_modal_projector |
| VRAM required | ~17 GB (fits RTX 4090 24GB) |
Evaluation
Benchmarked against the source BF16 model on the same hardware (A100 80GB) using lm-eval-harness v0.4.11 + vLLM v0.17.0.
| Benchmark | BF16 Abliterated | This (INT4) | Recovery |
|---|---|---|---|
| ARC-Challenge (acc_norm) | 0.6049 | 0.5973 | 98.7% |
| GSM8K (flexible-extract, 5-shot) | 0.9181 | 0.9212 | 100.3% |
| HellaSwag (acc_norm) | 0.8407 | 0.8326 | 99.0% |
| TruthfulQA MC2 | 0.5913 | 0.5748 | 97.2% |
| Winogrande | 0.7656 | 0.7545 | 98.6% |
| Average | 0.7441 | 0.7361 | 98.8% |
Usage — vLLM (recommended)
vllm serve suedegambit/Gemma-3-27B-it-NP-Abliterated-GPTQ-INT4-128g \
--max-model-len 8192 \
--gpu-memory-utilization 0.90 \
--enforce-eager
Requires SM80+ GPU (Ampere or newer) for GPTQ Marlin kernels.
Usage — transformers
Important: Load with
torch_dtype=torch.bfloat16. Gemma 3 overflows float16 range and will produce NaN outputs without this.
from transformers import Gemma3ForConditionalGeneration, AutoProcessor
import torch
model = Gemma3ForConditionalGeneration.from_pretrained(
"suedegambit/Gemma-3-27B-it-NP-Abliterated-GPTQ-INT4-128g",
device_map="auto",
torch_dtype=torch.bfloat16,
).eval()
processor = AutoProcessor.from_pretrained(
"suedegambit/Gemma-3-27B-it-NP-Abliterated-GPTQ-INT4-128g"
)
messages = [
{"role": "user", "content": [
{"type": "image", "image": "https://example.com/photo.jpg"},
{"type": "text", "text": "Describe this image."}
]}
]
inputs = processor.apply_chat_template(
messages, add_generation_prompt=True, tokenize=True,
return_dict=True, return_tensors="pt"
).to(model.device, dtype=torch.bfloat16)
with torch.inference_mode():
output = model.generate(**inputs, max_new_tokens=500, do_sample=False)
result = output[0][inputs["input_ids"].shape[-1]:]
print(processor.decode(result, skip_special_tokens=True))
Attribution
- Base model: google/gemma-3-27b-it by Google
- Abliteration: Nabbers1999/Gemma-3-27B-it-NP-Abliterated by Nabbers1999
- Quantization method: GPTQ via llm-compressor following RedHatAI's recipe
License
Gemma is provided under and subject to the Gemma Terms of Use.
Quantization script
import torch
import time
from datasets import load_dataset
from transformers import AutoProcessor, Gemma3ForConditionalGeneration
from llmcompressor.modifiers.quantization import GPTQModifier
from llmcompressor import oneshot
start = time.time()
MODEL_ID = "Nabbers1999/Gemma-3-27B-it-NP-Abliterated"
SAVE_DIR = "/workspace/Gemma-3-27B-it-NP-Abliterated-GPTQ-INT4-128g"
print("=== Loading model ===")
model = Gemma3ForConditionalGeneration.from_pretrained(
MODEL_ID,
device_map="auto",
torch_dtype="auto",
)
processor = AutoProcessor.from_pretrained(MODEL_ID, trust_remote_code=True)
print("=== Loading calibration data ===")
ds = load_dataset("neuralmagic/calibration", "LLM", split="train[:1024]")
ds = ds.shuffle(seed=42)
recipe = [
GPTQModifier(
scheme="W4A16",
targets="Linear",
ignore=[
"re:.*lm_head.*",
"re:.*embed_tokens.*",
"re:.*vision_tower.*",
"re:.*multi_modal_projector.*",
],
sequential_targets=["Gemma3DecoderLayer"],
dampening_frac=0.07,
block_size=128,
)
]
print("=== Starting GPTQ quantization ===")
print(f"GPU memory allocated: {torch.cuda.memory_allocated()/1e9:.1f} GB")
oneshot(
model=model,
tokenizer=MODEL_ID,
dataset=ds,
recipe=recipe,
max_seq_length=2048,
num_calibration_samples=1024,
trust_remote_code_model=True,
output_dir=SAVE_DIR,
)
elapsed = (time.time() - start) / 60
print(f"=== Quantization complete in {elapsed:.0f} minutes ==="
print(f"Output saved to: {SAVE_DIR}")