← back to catalog · registered 2026-08-22 13:56

groxaxo/Huihui-Qwen3.5-9B-abliterated-GPTQ-Pro-4bit-g64

groxaxo Qwen 6.9B second-order
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/groxaxo%2FHuihui-Qwen3.5-9B-abliterated-GPTQ-Pro-4bit-g64"
Response includes
  • classification m1
  • files 12
  • benchmarks 11 entries
  • hub_downloads_all_time 7,374
  • author_summary 27 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M1
Primary method

Direct removal

No other method signals detected in this model.
Confidence
MEDIUM
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=0 (base model)
  • no specific method indicators - defaulting to M1 (most common)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
MEDIUM
Why we say so
primary_method=M1; difference-of-means is the reference extraction for M1/M3 (Arditi 2024)
Downloads · lifetime
7K
464 last 30d - cooling
Likes
3
Model age
6mo ago
created 2026-03-28
Downloads over time
Now7.6K→from0↑0%
02.8K5.6K8.3K0 on Mar 257.6K on Oct 11MarAprMayJunJulAugSepOct
Mar 25 → Oct 11 · 68 snapshots · spans 200 days

Benchmarks

Benchmark Score Source
Entertainment 0.9 UGI
Hazardous 1.8 UGI
Natural Intelligence 14.88 UGI
Political lean -6.0% UGI
Sensitive-Info 11.44 UGI
SocPol 0.9 UGI
UGI 37.63 UGI
Willingness (10) 9 UGI
W10-Adherence 9 UGI
W10-Direct 9 UGI
Writing 29.12 UGI

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
other
Languages
en
Tags
transformers safetensors qwen3_5_text text-generation gptq 4bit quantized qwen gptq-pro conversational en base_model:huihui-ai/Huihui-Qwen3.5-9B-abliterated

Related

Total size
7.28 GB
Files
12
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-08-22 08:00

Files by quantization

Auxiliary files 12 files 7.30 GB
model-00001-of-00002.safetensors 4.00 GB 53cacd95 download
model-00002-of-00002.safetensors 3.28 GB 250dc484 download
tokenizer.json 19.1 MB 87a7830d download
model.safetensors.index.json 83.3 KB ec58d8aa download
quant_log.csv 9.38 KB 014ee33b download
chat_template.jinja 7.57 KB a585dec8 download
README.md 4.60 KB fa08efd7 download
config.json 3.28 KB 012adcc2 download
.gitattributes 1.53 KB 52373fe2 download
quantize_config.json 1.22 KB 7d528b6e download
tokenizer_config.json 1.10 KB 946a9ba7 download
generation_config.json 204 B bd1534f8 download

README current version from Hugging Face


language:

  • en
    license: other
    base_model:
  • huihui-ai/Huihui-Qwen3.5-9B-abliterated
    tags:
  • gptq
  • 4bit
  • quantized
  • qwen
  • text-generation
  • gptq-pro
    pipeline_tag: text-generation
    library_name: transformers

Huihui-Qwen3.5-9B-abliterated GPTQ-Pro 4bit (g64)

Overview

Huihui-Qwen3.5-9B-abliterated-GPTQ-Pro-4bit-g64 is a GPTQ-quantized checkpoint intended for efficient GPU inference, published by groxaxo.
It is intended for open-source evaluation, reproducible experimentation, and compatible local or
hosted inference workflows. The wording below is deliberately limited to what can be verified
from this repository's metadata and artifacts.

The repository name identifies a behavior-modified or reduced-filtering lineage. That label describes the source or conversion history; it is not a guarantee of unrestricted behavior in every prompt or runtime. Test outputs carefully before sharing or deploying them.

At a glance

Field Details
Format GPTQ
Source / base huihui-ai/Huihui-Qwen3.5-9B-abliterated
Intended task text-generation
License other

What is included

  • *.safetensors (2 files)
  • config.json
  • generation_config.json
  • tokenizer.json
  • tokenizer_config.json
  • chat_template.jinja
  • quantize_config.json
  • Additional configuration, tokenizer, processor, or shard files (10 visible artifacts total)

Quick start

vLLM (documented configuration)

vllm serve groxaxo/Huihui-Qwen3.5-9B-abliterated-GPTQ-Pro-4bit-g64 \
  --quantization gptq_marlin \
  --dtype float16 \
  --trust-remote-code

This command is taken from the repository documentation. Adjust tensor parallelism, context
length, and cache settings to match your hardware and vLLM version.

Compatibility and responsible use

  • Use a runtime that explicitly supports this format, architecture, and modality.
  • Keep configuration, tokenizer, processor, projection, and weight files from the same revision together.
  • Review the source model card and license before redistribution or deployment.
  • Hardware needs depend on parameter count, context length, cache precision, quantization, and concurrency.
  • Report reproducible issues with the runtime version, hardware, launch command, and a minimal example.

Quantization or conversion changes numerical behavior, memory use, and throughput relative to the source checkpoint; validate quality on your own workload.

Generated outputs may be inaccurate or unsuitable for a given use case. Users are responsible for
testing behavior, applying appropriate safeguards, and complying with applicable licenses and laws.

This is a GPTQ-Pro 4-bit quantization of huihui-ai/Huihui-Qwen3.5-9B-abliterated.

It was quantized with group size 64 and evaluated against the original model on Wikitext-2 using a strided perplexity setup, plus KL and token-agreement checks.

Highlights

  • Base model: huihui-ai/Huihui-Qwen3.5-9B-abliterated
  • Quantization: GPTQ-Pro, 4-bit, group size 64
  • Calibration samples: 128
  • Quantization time: about 11.1 minutes
  • Quantized strided perplexity: 9.6579
  • Original strided perplexity: 9.5234
  • Perplexity degradation: 1.41%
  • Average KL divergence vs original: 0.03423
  • Top-1 agreement vs original: 91.96%
  • Top-5 agreement vs original: 99.98%

Quality Notes

This quantized build stays very close to the source model in language modeling quality.

  • Perplexity regression is small.
  • KL divergence is low.
  • Top-5 next-token agreement is effectively perfect.
  • In practice, this should preserve most of the original model's behavior while reducing memory use substantially.

Files

  • model-00001-of-00002.safetensors
  • model-00002-of-00002.safetensors
  • quantize_config.json
  • tokenizer and config files

Load With Transformers / GPTQModel

from gptqmodel import GPTQModel

model = GPTQModel.load(
    "groxaxo/Huihui-Qwen3.5-9B-abliterated-GPTQ-Pro-4bit-g64",
    device_map="auto",
    trust_remote_code=True,
)

Evaluation Summary

Measured locally:

  • Quantized strided PPL: 9.6579304371
  • Original strided PPL: 9.5233634665
  • Quantized chunked PPL: 11.6689118281
  • Original chunked PPL: 11.5080707440
  • KL divergence: 0.0342324856
  • Logit cosine similarity: 0.9935612157

Prompting

Use the same prompting and chat template behavior as the base model.

Disclaimer

This repo contains only the quantized checkpoint. Please review the base model card for intended use, limitations, and licensing details.

README history 4 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-08-22Polish model card overview and usage notesd4c17254.6 KB
    Loading...
  2. 2026-08-22Polish model card overview and usage notes60248664.6 KB
    Loading...
  3. 2026-08-22Polish model card overview and usage notes89fe9a43.8 KB
    Loading...
  4. 2026-03-28Add files using upload-large-folder tool44f87142.1 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration