language:
- en
license: other
base_model: - huihui-ai/Huihui-Qwen3.5-9B-abliterated
tags: - gptq
- 4bit
- quantized
- qwen
- text-generation
- gptq-pro
pipeline_tag: text-generation
library_name: transformers
Huihui-Qwen3.5-9B-abliterated GPTQ-Pro 4bit (g64)
Overview
Huihui-Qwen3.5-9B-abliterated-GPTQ-Pro-4bit-g64 is a GPTQ-quantized checkpoint intended for efficient GPU inference, published by groxaxo.
It is intended for open-source evaluation, reproducible experimentation, and compatible local or
hosted inference workflows. The wording below is deliberately limited to what can be verified
from this repository's metadata and artifacts.
The repository name identifies a behavior-modified or reduced-filtering lineage. That label describes the source or conversion history; it is not a guarantee of unrestricted behavior in every prompt or runtime. Test outputs carefully before sharing or deploying them.
At a glance
| Field | Details |
|---|---|
| Format | GPTQ |
| Source / base | huihui-ai/Huihui-Qwen3.5-9B-abliterated |
| Intended task | text-generation |
| License | other |
What is included
*.safetensors(2 files)config.jsongeneration_config.jsontokenizer.jsontokenizer_config.jsonchat_template.jinjaquantize_config.json- Additional configuration, tokenizer, processor, or shard files (10 visible artifacts total)
Quick start
vLLM (documented configuration)
vllm serve groxaxo/Huihui-Qwen3.5-9B-abliterated-GPTQ-Pro-4bit-g64 \
--quantization gptq_marlin \
--dtype float16 \
--trust-remote-code
This command is taken from the repository documentation. Adjust tensor parallelism, context
length, and cache settings to match your hardware and vLLM version.
Compatibility and responsible use
- Use a runtime that explicitly supports this format, architecture, and modality.
- Keep configuration, tokenizer, processor, projection, and weight files from the same revision together.
- Review the source model card and license before redistribution or deployment.
- Hardware needs depend on parameter count, context length, cache precision, quantization, and concurrency.
- Report reproducible issues with the runtime version, hardware, launch command, and a minimal example.
Quantization or conversion changes numerical behavior, memory use, and throughput relative to the source checkpoint; validate quality on your own workload.
Generated outputs may be inaccurate or unsuitable for a given use case. Users are responsible for
testing behavior, applying appropriate safeguards, and complying with applicable licenses and laws.
This is a GPTQ-Pro 4-bit quantization of huihui-ai/Huihui-Qwen3.5-9B-abliterated.
It was quantized with group size 64 and evaluated against the original model on Wikitext-2 using a strided perplexity setup, plus KL and token-agreement checks.
Highlights
- Base model:
huihui-ai/Huihui-Qwen3.5-9B-abliterated - Quantization: GPTQ-Pro, 4-bit, group size
64 - Calibration samples:
128 - Quantization time: about
11.1minutes - Quantized strided perplexity:
9.6579 - Original strided perplexity:
9.5234 - Perplexity degradation:
1.41% - Average KL divergence vs original:
0.03423 - Top-1 agreement vs original:
91.96% - Top-5 agreement vs original:
99.98%
Quality Notes
This quantized build stays very close to the source model in language modeling quality.
- Perplexity regression is small.
- KL divergence is low.
- Top-5 next-token agreement is effectively perfect.
- In practice, this should preserve most of the original model's behavior while reducing memory use substantially.
Files
model-00001-of-00002.safetensorsmodel-00002-of-00002.safetensorsquantize_config.json- tokenizer and config files
Load With Transformers / GPTQModel
from gptqmodel import GPTQModel
model = GPTQModel.load(
"groxaxo/Huihui-Qwen3.5-9B-abliterated-GPTQ-Pro-4bit-g64",
device_map="auto",
trust_remote_code=True,
)
Evaluation Summary
Measured locally:
- Quantized strided PPL:
9.6579304371 - Original strided PPL:
9.5233634665 - Quantized chunked PPL:
11.6689118281 - Original chunked PPL:
11.5080707440 - KL divergence:
0.0342324856 - Logit cosine similarity:
0.9935612157
Prompting
Use the same prompting and chat template behavior as the base model.
Disclaimer
This repo contains only the quantized checkpoint. Please review the base model card for intended use, limitations, and licensing details.