base_model: shreyan35/gemma4-31b-abliterated-multimodal
library_name: vllm
pipeline_tag: image-text-to-text
tags:
- gemma4
- multimodal
- awq
- compressed-tensors
- w8a16
- vllm
Gemma4 31B Abliterated Multimodal AWQ8
Overview
gemma4-31b-abliterated-multimodal-awq8 is a weight-quantized checkpoint intended for efficient GPU inference, published by groxaxo.
It is intended for open-source evaluation, reproducible experimentation, and compatible local or
hosted inference workflows. The wording below is deliberately limited to what can be verified
from this repository's metadata and artifacts.
The repository name identifies a behavior-modified or reduced-filtering lineage. That label describes the source or conversion history; it is not a guarantee of unrestricted behavior in every prompt or runtime. Test outputs carefully before sharing or deploying them.
At a glance
| Field | Details |
|---|---|
| Format | AWQ / AutoRound |
| Source / base | shreyan35/gemma4-31b-abliterated-multimodal |
| Intended task | image-text-to-text |
| License | the license declared in the repository files |
What is included
*.safetensors(1 file)config.jsongeneration_config.jsontokenizer.jsontokenizer_config.jsonprocessor_config.jsonchat_template.jinja- Additional configuration, tokenizer, processor, or shard files (8 visible artifacts total)
Quick start
vLLM (AWQ-compatible runtimes)
vllm serve groxaxo/gemma4-31b-abliterated-multimodal-awq8 \
--quantization awq_marlin \
--dtype float16 \
--trust-remote-code
The exact kernel and flags depend on the quantizer and architecture. Check the files and source
model card before selecting a production serving configuration.
Compatibility and responsible use
- Use a runtime that explicitly supports this format, architecture, and modality.
- Keep configuration, tokenizer, processor, projection, and weight files from the same revision together.
- Review the source model card and license before redistribution or deployment.
- Hardware needs depend on parameter count, context length, cache precision, quantization, and concurrency.
- Report reproducible issues with the runtime version, hardware, launch command, and a minimal example.
Quantization or conversion changes numerical behavior, memory use, and throughput relative to the source checkpoint; validate quality on your own workload.
Generated outputs may be inaccurate or unsuitable for a given use case. Users are responsible for
testing behavior, applying appropriate safeguards, and complying with applicable licenses and laws.
This repository contains a compressed-tensors AWQ W8A16 checkpoint for shreyan35/gemma4-31b-abliterated-multimodal.
What is included
config.jsonwith the compressed-tensors quantization configmodel.safetensorstokenizer.jsonandtokenizer_config.jsonprocessor_config.jsonchat_template.jinjageneration_config.jsonrecipe.yaml
Quantization summary
- Format:
compressed-tensors - Method:
AWQ - Weight bits:
8 - Activation bits:
16 - Group size:
32 - Weights: symmetric
- Observer:
mse - Duo scaling: enabled
- Excluded from quantization: vision tower, multimodal projector/embed_vision, and
lm_head
Recommended serving
Use vLLM for inference.
Tested runtime:
torch 2.10.0+cu128transformers 5.5.1compressed-tensors 0.14.0.1vllm 0.19.0
pip install -U "torch==2.10.0" "transformers==5.5.1" "compressed-tensors==0.14.0.1" "vllm==0.19.0"
vllm serve groxaxo/gemma4-31b-abliterated-multimodal-awq8 \
--trust-remote-code \
--dtype auto \
--max-model-len 6144 \
--served-model-name gemma4-31b-abliterated-multimodal-awq8
For local image inputs, add:
--allowed-local-media-path /path/to/images --limit-mm-per-prompt '{"image":1}'
For text-only long-context serving:
--limit-mm-per-prompt '{"image":0,"video":0,"audio":0}' --skip-mm-profiling --mm-processor-cache-gb 0 --max-model-len 10240
Transformers fallback
The checkpoint also loads with trust_remote_code=True through Transformers:
from transformers import AutoProcessor, AutoModelForImageTextToText
repo_id = "groxaxo/gemma4-31b-abliterated-multimodal-awq8"
processor = AutoProcessor.from_pretrained(repo_id, trust_remote_code=True)
model = AutoModelForImageTextToText.from_pretrained(
repo_id,
trust_remote_code=True,
device_map="auto",
torch_dtype="auto",
)
Notes
- This checkpoint was built to preserve multimodal capability while avoiding quantization of the vision tower and projector modules.
- If you only need a local OpenAI-compatible endpoint, point clients at
http://127.0.0.1:1234/v1after starting vLLM.