← back to catalog · registered 2026-08-22 13:56

groxaxo/Gemma-4-Abliterated-GPTQ-4bit-Pro-v1

groxaxo Gemma 29B second-order
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/groxaxo%2FGemma-4-Abliterated-GPTQ-4bit-Pro-v1"
Response includes
  • classification m1
  • files 15
  • hub_downloads_all_time 740
  • author_summary 27 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M1
Primary method

Direct removal

No other method signals detected in this model.
Confidence
MEDIUM
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=0 (base model)
  • no specific method indicators - defaulting to M1 (most common)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
MEDIUM
Why we say so
primary_method=M1; difference-of-means is the reference extraction for M1/M3 (Arditi 2024)
Downloads · lifetime
740
28 last 30d - cooling
Likes
1
Model age
5mo ago
created 2026-04-15
Downloads over time
Now750→from116↑547%
84327570813116 on Apr 22750 on Oct 11750 on Oct 10AprMayJunJulAugSepOct
Apr 22 → Oct 11 · 64 snapshots · spans 172 days

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
other
Tags
safetensors gemma4 gemma gptq gptq-pro 4bit quantized text-generation inference local-llm conversational base_model:wangzhang/gemma-4-31B-it-abliterated

Related

Total size
17.9 GB
Files
15
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-08-22 07:59

Files by quantization

Auxiliary files 15 files 17.9 GB
model-00002-of-00005.safetensors 4.00 GB a0aba9f3 download
model-00001-of-00005.safetensors 4.00 GB e36470b3 download
model-00004-of-00005.safetensors 3.96 GB 9731b743 download
model-00003-of-00005.safetensors 3.96 GB 122760ee download
model-00005-of-00005.safetensors 1.96 GB 3363677f download
tokenizer.json 30.7 MB 5d84efaa download
model.safetensors.index.json 233 KB 4b1eb3a6 download
quant_log.csv 18.5 KB 80db2d41 download
chat_template.jinja 16.1 KB 98da08eb download
README.md 10.5 KB 840112c8 download
config.json 5.73 KB 1167c9f3 download
tokenizer_config.json 2.72 KB 6f3337a6 download
.gitattributes 1.53 KB 52373fe2 download
quantize_config.json 1001 B 7f22f9ab download
generation_config.json 203 B edda3c19 download

README current version from Hugging Face


base_model:

  • wangzhang/gemma-4-31B-it-abliterated
    pipeline_tag: text-generation
    tags:
  • gemma
  • gemma4
  • gptq
  • gptq-pro
  • 4bit
  • quantized
  • safetensors
  • text-generation
  • inference
  • local-llm
    license: other

Gemma-4-Abliterated-GPTQ-4bit-Pro-v1

Overview

Gemma-4-Abliterated-GPTQ-4bit-Pro-v1 is a GPTQ-quantized checkpoint intended for efficient GPU inference, published by groxaxo.
It is intended for open-source evaluation, reproducible experimentation, and compatible local or
hosted inference workflows. The wording below is deliberately limited to what can be verified
from this repository's metadata and artifacts.

The repository name identifies a behavior-modified or reduced-filtering lineage. That label describes the source or conversion history; it is not a guarantee of unrestricted behavior in every prompt or runtime. Test outputs carefully before sharing or deploying them.

At a glance

Field Details
Format GPTQ
Source / base wangzhang/gemma-4-31B-it-abliterated
Intended task image-text-to-text
License other

What is included

  • *.safetensors (5 files)
  • config.json
  • generation_config.json
  • tokenizer.json
  • tokenizer_config.json
  • chat_template.jinja
  • quantize_config.json
  • Additional configuration, tokenizer, processor, or shard files (13 visible artifacts total)

Quick start

vLLM (documented configuration)

CUDA_VISIBLE_DEVICES=0,1 \
vllm serve groxaxo/Gemma-4-Abliterated-GPTQ-4bit-Pro-v1 \
  --tensor-parallel-size 2 \
  --gpu-memory-utilization 0.92 \
  --max-model-len 8192 \
  --max-num-batched-tokens 32768 \
  --trust-remote-code

This command is taken from the repository documentation. Adjust tensor parallelism, context
length, and cache settings to match your hardware and vLLM version.

Compatibility and responsible use

  • Use a runtime that explicitly supports this format, architecture, and modality.
  • Keep configuration, tokenizer, processor, projection, and weight files from the same revision together.
  • Review the source model card and license before redistribution or deployment.
  • Hardware needs depend on parameter count, context length, cache precision, quantization, and concurrency.
  • Report reproducible issues with the runtime version, hardware, launch command, and a minimal example.

Quantization or conversion changes numerical behavior, memory use, and throughput relative to the source checkpoint; validate quality on your own workload.

Generated outputs may be inaccurate or unsuitable for a given use case. Users are responsible for
testing behavior, applying appropriate safeguards, and complying with applicable licenses and laws.

Gemma-4-Abliterated-GPTQ-4bit-Pro-v1 banner

Gemma-4-Abliterated-GPTQ-4bit-Pro-v1

Big-model capability. 4-bit footprint. Built for local inference that does not melt your rig.

Base Model • GPTQ Pro • groxaxo


Overview

Gemma-4-Abliterated-GPTQ-4bit-Pro-v1 is a deployment-ready 4-bit GPTQ Pro quantization of wangzhang/gemma-4-31B-it-abliterated, built by groxaxo for efficient local inference, reduced VRAM usage, and practical 2-GPU serving.

This release targets users who want the behavior and capability of a 31B-class Gemma 4 model without the painful memory requirements of full-precision weights.

Quantized with GPTQ Pro by groxaxo.


Why This Quant?

Most large models are annoying to deploy locally: too much VRAM, slow loading, awkward sharding, and poor usability on real-world hardware.

This release focuses on what actually matters:

  • Lower VRAM pressure — 4-bit quantization makes a 31B-class model much easier to run locally.
  • Consumer GPU friendly — designed for practical inference on local GPU rigs, especially 2x GPU setups.
  • Deployment-first packaging — safetensors format, clean naming, and inference-oriented quantization.
  • Strong quality-to-size ratio — keeps the model useful while heavily reducing the weight footprint.
  • Better serving headroom — more VRAM remains available for KV cache, batching, and longer contexts.
  • Abliterated behavior profile — based on the abliterated Gemma 4 31B IT variant.
  • Built by groxaxo — quantized using the GPTQ Pro pipeline for practical local deployment.

Model Details

Field Value
Base model wangzhang/gemma-4-31B-it-abliterated
Quantization GPTQ Pro
Precision 4-bit
Format Safetensors
Model size ~19.2 GB
Quantized by groxaxo
GPTQ Pro repo groxaxo/GPTQ-Pro
Intended use Inference / serving

Key Benefits

Efficient Local Serving

Run a 31B-class model with a dramatically reduced memory footprint. This makes the model more practical for local labs, private inference servers, hobbyist clusters, and dual-GPU consumer setups.

Better Hardware Utilization

Because the weights are compressed to 4-bit, more VRAM remains available for the parts that matter during inference: KV cache, batching, longer contexts, and higher request concurrency.

Practical 2-GPU Deployment

This model is suitable for tensor-parallel serving across 2 GPUs. Homogeneous GPUs are strongly preferred for maximum stability and throughput.

Smaller Download, Faster Iteration

At roughly 19.2 GB, this release is easier to download, store, move, test, and redeploy than full-precision checkpoints.

Built for Builders

This quant is intended for people actually deploying models:

  • Local assistants
  • Private LLM APIs
  • vLLM serving
  • 2-GPU inference
  • Offline/private generation
  • Benchmarking
  • Rapid model comparison
  • Text and code generation experiments

Deployment

Install vLLM

pip install -U vllm

Recommended vLLM Serving Command

For a 2x GPU setup:

CUDA_VISIBLE_DEVICES=0,1 \
vllm serve groxaxo/Gemma-4-Abliterated-GPTQ-4bit-Pro-v1 \
  --tensor-parallel-size 2 \
  --gpu-memory-utilization 0.92 \
  --max-model-len 8192 \
  --trust-remote-code

For 2x RTX 3090, this is the clean setup. Keep the tensor-parallel group homogeneous; mixing a 3090 with a weaker card is only worth it when memory matters more than throughput.

Conservative 2-GPU Command

Use this if you hit OOM or want more KV-cache safety:

CUDA_VISIBLE_DEVICES=0,1 \
vllm serve groxaxo/Gemma-4-Abliterated-GPTQ-4bit-Pro-v1 \
  --tensor-parallel-size 2 \
  --gpu-memory-utilization 0.85 \
  --max-model-len 8192 \
  --trust-remote-code

Higher Throughput 2-GPU Command

Use this for API-style serving with batching:

CUDA_VISIBLE_DEVICES=0,1 \
vllm serve groxaxo/Gemma-4-Abliterated-GPTQ-4bit-Pro-v1 \
  --tensor-parallel-size 2 \
  --gpu-memory-utilization 0.92 \
  --max-model-len 8192 \
  --max-num-batched-tokens 32768 \
  --trust-remote-code

Useful knobs:

  • Lower --max-model-len if you do not need long context.
  • Increase --gpu-memory-utilization only if the server is stable.
  • Use two matching GPUs whenever possible.
  • Avoid mixing 3090 + 3060 unless memory is the priority over speed.
  • If throughput tanks, the weaker GPU is probably dragging the tensor-parallel group down.

Transformers / AutoGPTQ

pip install -U transformers accelerate auto-gptq
from transformers import AutoTokenizer
from auto_gptq import AutoGPTQForCausalLM

model_id = "groxaxo/Gemma-4-Abliterated-GPTQ-4bit-Pro-v1"

tokenizer = AutoTokenizer.from_pretrained(
    model_id,
    use_fast=True,
    trust_remote_code=True,
)

model = AutoGPTQForCausalLM.from_quantized(
    model_id,
    device="cuda:0",
    use_safetensors=True,
    trust_remote_code=True,
)

prompt = "Explain tensor parallelism in one paragraph."
inputs = tokenizer(prompt, return_tensors="pt").to(model.device)

outputs = model.generate(
    **inputs,
    max_new_tokens=256,
    temperature=0.7,
    top_p=0.9,
)

print(tokenizer.decode(outputs[0], skip_special_tokens=True))

Throughput Tuning

For better serving performance:

CUDA_VISIBLE_DEVICES=0,1 \
vllm serve groxaxo/Gemma-4-Abliterated-GPTQ-4bit-Pro-v1 \
  --tensor-parallel-size 2 \
  --gpu-memory-utilization 0.92 \
  --max-model-len 8192 \
  --max-num-batched-tokens 32768 \
  --trust-remote-code

Recommended tuning:

  • Lower --max-model-len for better throughput if long context is unnecessary.
  • Increase --max-num-batched-tokens for heavier concurrent workloads.
  • Keep tensor-parallel GPUs homogeneous.
  • Prefer 2x RTX 3090 or equivalent over mixed VRAM/bandwidth setups.
  • Watch VRAM during long prompts; KV cache grows fast.
  • Do not overbatch until latency becomes trash. Throughput without usable latency is just benchmark cosplay.

Recommended Use Cases

  • Local chat inference
  • Private LLM serving
  • vLLM deployment
  • 2-GPU inference
  • Research and experimentation
  • Instruction-following workloads
  • Text generation
  • Code generation
  • Offline/private assistant workflows

Limitations

This is a 4-bit quantized model. Some degradation compared to the original higher-precision checkpoint is expected, especially on:

  • highly sensitive reasoning tasks
  • exact numeric work
  • long-context edge cases
  • strict instruction-following edge cases
  • tasks requiring high precision in small probability differences

This model is intended for inference only. It is not suitable for training or fine-tuning.


Credits


Quantized for people who actually run models, not just collect checkpoints.

README history 7 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-08-22Polish model card overview and usage notesc282e5310.5 KB
    Loading...
  2. 2026-08-22Polish model card overview and usage notes1cbe05d10.4 KB
    Loading...
  3. 2026-08-22Polish model card overview and usage notes1a2a1ad10.4 KB
    Loading...
  4. 2026-08-22Polish model card overview and usage notes73e8a4e9.6 KB
    Loading...
  5. 2026-04-22Update README.md4d47cf97.9 KB
    Loading...
  6. 2026-04-22Update README.mde578a8b7.9 KB
    Loading...
  7. 2026-04-22Create README.md05f259b1.4 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration