← back to catalog · registered 2026-08-22 13:56

suedegambit/Gemma-3-27B-it-NP-Abliterated-GPTQ-INT4-128g

suedegambit Gemma 26B multimodal second-order
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/suedegambit%2FGemma-3-27B-it-NP-Abliterated-GPTQ-INT4-128g"
Response includes
  • classification m1
  • files 17
  • hub_downloads_all_time 364
  • author_summary 1 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M1
Primary method

Direct removal

No other method signals detected in this model.
Confidence
MEDIUM
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=0 (base model)
  • no specific method indicators - defaulting to M1 (most common)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
MEDIUM
Why we say so
primary_method=M1; difference-of-means is the reference extraction for M1/M3 (Arditi 2024)
Downloads · lifetime
364
14 last 30d - cooling
Likes
0
Model age
7mo ago
created 2026-03-06

Training datasets

Corpora the author lists in the model card. Datasets tracked in our /datasets catalog carry a category badge linking to the workflow stage. Others open on Hugging Face.

Downloads over time
Now369→from0↑0%
01352714060 on Mar 4369 on Oct 11369 on Oct 9MarAprMayJunJulAugSepOct
Mar 4 → Oct 11 · 71 snapshots · spans 221 days

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
gemma
Languages
en
Tags
transformers safetensors gemma3 image-text-to-text int4 gptq 4bit vllm compressed-tensors quantized abliterated multimodal
Total size
15.7 GB
Files
17
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-03-09 00:48

Files by quantization

Auxiliary files 17 files 15.7 GB
model-00001-of-00004.safetensors 4.64 GB b4fde7e9 download
model-00003-of-00004.safetensors 4.62 GB b2b57b11 download
model-00002-of-00004.safetensors 4.62 GB bb85f10c download
model-00004-of-00004.safetensors 1.84 GB 901ffdb7 download
tokenizer.json 31.8 MB 4667f208 download
tokenizer.model 4.47 MB 1299c11d download
tokenizer_config.json 1.10 MB 7bdd14f0 download
model.safetensors.index.json 213 KB bfba40e3 download
config.json 15.6 KB 4ecca5cf download
README.md 5.20 KB 16a5e547 download
chat_template.json 1.58 KB 719b0cd0 download
.gitattributes 1.53 KB 52373fe2 download
special_tokens_map.json 662 B 1a619324 download
preprocessor_config.json 570 B b1e00fc1 download
recipe.yaml 398 B 67db99c5 download
generation_config.json 210 B 777c3c32 download
added_tokens.json 35.0 B e17bde03 download

README current version from Hugging Face


license: gemma
library_name: transformers
base_model:

  • Nabbers1999/Gemma-3-27B-it-NP-Abliterated
  • google/gemma-3-27b-it
    base_model_relation: quantized
    pipeline_tag: image-text-to-text
    tags:
  • gemma3
  • int4
  • gptq
  • 4bit
  • vllm
  • compressed-tensors
  • quantized
  • abliterated
  • multimodal
  • vision
    language:
  • en
    datasets:
  • neuralmagic/calibration

Gemma 3 27B IT NP-Abliterated — GPTQ INT4 (128g)

GPTQ INT4 quantization of Nabbers1999/Gemma-3-27B-it-NP-Abliterated, a biprojected norm-preserving abliteration of Google's Gemma 3 27B IT.

Vision is fully preserved. Only the language model decoder layers are quantized to INT4. The SigLIP vision encoder and multimodal projector remain in BF16.

Quantization details

Property Value
Method GPTQ (4-bit, symmetric, per-group)
Group size 128
Format compressed-tensors
Calibration data neuralmagic/calibration (1024 samples, seq_len 2048)
Dampening 0.07
Layers preserved in BF16 lm_head, embed_tokens, vision_tower, multi_modal_projector
VRAM required ~17 GB (fits RTX 4090 24GB)

Evaluation

Benchmarked against the source BF16 model on the same hardware (A100 80GB) using lm-eval-harness v0.4.11 + vLLM v0.17.0.

Benchmark BF16 Abliterated This (INT4) Recovery
ARC-Challenge (acc_norm) 0.6049 0.5973 98.7%
GSM8K (flexible-extract, 5-shot) 0.9181 0.9212 100.3%
HellaSwag (acc_norm) 0.8407 0.8326 99.0%
TruthfulQA MC2 0.5913 0.5748 97.2%
Winogrande 0.7656 0.7545 98.6%
Average 0.7441 0.7361 98.8%

Usage — vLLM (recommended)

vllm serve suedegambit/Gemma-3-27B-it-NP-Abliterated-GPTQ-INT4-128g \
    --max-model-len 8192 \
    --gpu-memory-utilization 0.90 \
    --enforce-eager

Requires SM80+ GPU (Ampere or newer) for GPTQ Marlin kernels.

Usage — transformers

Important: Load with torch_dtype=torch.bfloat16. Gemma 3 overflows float16 range and will produce NaN outputs without this.

from transformers import Gemma3ForConditionalGeneration, AutoProcessor
import torch

model = Gemma3ForConditionalGeneration.from_pretrained(
    "suedegambit/Gemma-3-27B-it-NP-Abliterated-GPTQ-INT4-128g",
    device_map="auto",
    torch_dtype=torch.bfloat16,
).eval()
processor = AutoProcessor.from_pretrained(
    "suedegambit/Gemma-3-27B-it-NP-Abliterated-GPTQ-INT4-128g"
)

messages = [
    {"role": "user", "content": [
        {"type": "image", "image": "https://example.com/photo.jpg"},
        {"type": "text", "text": "Describe this image."}
    ]}
]
inputs = processor.apply_chat_template(
    messages, add_generation_prompt=True, tokenize=True,
    return_dict=True, return_tensors="pt"
).to(model.device, dtype=torch.bfloat16)

with torch.inference_mode():
    output = model.generate(**inputs, max_new_tokens=500, do_sample=False)
    result = output[0][inputs["input_ids"].shape[-1]:]
print(processor.decode(result, skip_special_tokens=True))

Attribution

License

Gemma is provided under and subject to the Gemma Terms of Use.

Quantization script
import torch
import time
from datasets import load_dataset
from transformers import AutoProcessor, Gemma3ForConditionalGeneration
from llmcompressor.modifiers.quantization import GPTQModifier
from llmcompressor import oneshot

start = time.time()

MODEL_ID = "Nabbers1999/Gemma-3-27B-it-NP-Abliterated"
SAVE_DIR = "/workspace/Gemma-3-27B-it-NP-Abliterated-GPTQ-INT4-128g"

print("=== Loading model ===")
model = Gemma3ForConditionalGeneration.from_pretrained(
    MODEL_ID,
    device_map="auto",
    torch_dtype="auto",
)
processor = AutoProcessor.from_pretrained(MODEL_ID, trust_remote_code=True)

print("=== Loading calibration data ===")
ds = load_dataset("neuralmagic/calibration", "LLM", split="train[:1024]")
ds = ds.shuffle(seed=42)

recipe = [
    GPTQModifier(
        scheme="W4A16",
        targets="Linear",
        ignore=[
            "re:.*lm_head.*",
            "re:.*embed_tokens.*",
            "re:.*vision_tower.*",
            "re:.*multi_modal_projector.*",
        ],
        sequential_targets=["Gemma3DecoderLayer"],
        dampening_frac=0.07,
        block_size=128,
    )
]

print("=== Starting GPTQ quantization ===")
print(f"GPU memory allocated: {torch.cuda.memory_allocated()/1e9:.1f} GB")

oneshot(
    model=model,
    tokenizer=MODEL_ID,
    dataset=ds,
    recipe=recipe,
    max_seq_length=2048,
    num_calibration_samples=1024,
    trust_remote_code_model=True,
    output_dir=SAVE_DIR,
)

elapsed = (time.time() - start) / 60
print(f"=== Quantization complete in {elapsed:.0f} minutes ==="
print(f"Output saved to: {SAVE_DIR}")

README history 2 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-03-08Upload folder using huggingface_hub2bdc5155.2 KB
    Loading...
  2. 2026-03-06initial commit3c4affd23 B
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration