← back to catalog · registered 2026-08-22 13:56

groxaxo/Huihui-Qwen3.5-9B-abliterated-gptq-pro-w4g128

groxaxo Qwen 6.9B second-order
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/groxaxo%2FHuihui-Qwen3.5-9B-abliterated-gptq-pro-w4g128"
Response includes
  • classification m1
  • files 12
  • hub_downloads_all_time 137
  • author_summary 27 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M1
Primary method

Direct removal

No other method signals detected in this model.
Confidence
MEDIUM
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=0 (base model)
  • no specific method indicators - defaulting to M1 (most common)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
MEDIUM
Why we say so
primary_method=M1; difference-of-means is the reference extraction for M1/M3 (Arditi 2024)
Downloads · lifetime
137
47 last 30d - stable
Likes
1
Model age
6mo ago
created 2026-03-28
Downloads over time
Now159→from19↑737%
126611917319 on Apr 1159 on Oct 11AprMayJunJulAugSepOct
Apr 1 → Oct 11 · 67 snapshots · spans 193 days

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
other
Tags
safetensors qwen3_5_text exl3 gptq quantized qwen3.5 license:other 4-bit region:us

Related

Total size
7.15 GB
Files
12
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-08-22 07:58

Files by quantization

Auxiliary files 12 files 7.17 GB
model-00001-of-00002.safetensors 3.99 GB 4a035fa8 download
model-00002-of-00002.safetensors 3.16 GB c61d3da5 download
tokenizer.json 19.1 MB 87a7830d download
model.safetensors.index.json 83.3 KB 39fdcbc2 download
quant_log.csv 9.38 KB aaa5c1ac download
chat_template.jinja 7.57 KB a585dec8 download
config.json 3.31 KB ae0006e1 download
README.md 2.84 KB 52337e71 download
.gitattributes 1.53 KB 52373fe2 download
quantize_config.json 1.25 KB 4b7bd659 download
tokenizer_config.json 1.07 KB 6be6ce17 download
generation_config.json 199 B 1918645f download

README current version from Hugging Face


license: other
base_model: Huihui/Qwen3.5-9B-abliterated
tags:

  • exl3
  • gptq
  • quantized
  • qwen3.5

Huihui-Qwen3.5-9B-abliterated-gptq-pro-w4g128

Overview

Huihui-Qwen3.5-9B-abliterated-gptq-pro-w4g128 is a GPTQ-quantized checkpoint intended for efficient GPU inference, published by groxaxo.
It is intended for open-source evaluation, reproducible experimentation, and compatible local or
hosted inference workflows. The wording below is deliberately limited to what can be verified
from this repository's metadata and artifacts.

The repository name identifies a behavior-modified or reduced-filtering lineage. That label describes the source or conversion history; it is not a guarantee of unrestricted behavior in every prompt or runtime. Test outputs carefully before sharing or deploying them.

At a glance

Field Details
Format GPTQ
Source / base Huihui/Qwen3.5-9B-abliterated
Intended task text-generation
License other

What is included

  • *.safetensors (2 files)
  • config.json
  • generation_config.json
  • tokenizer.json
  • tokenizer_config.json
  • chat_template.jinja
  • quantize_config.json
  • Additional configuration, tokenizer, processor, or shard files (10 visible artifacts total)

Quick start

vLLM (documented configuration)

vllm serve groxaxo/Huihui-Qwen3.5-9B-abliterated-gptq-pro-w4g128 \
  --quantization gptq_marlin \
  --dtype float16 \
  --trust-remote-code

This command is taken from the repository documentation. Adjust tensor parallelism, context
length, and cache settings to match your hardware and vLLM version.

Compatibility and responsible use

  • Use a runtime that explicitly supports this format, architecture, and modality.
  • Keep configuration, tokenizer, processor, projection, and weight files from the same revision together.
  • Review the source model card and license before redistribution or deployment.
  • Hardware needs depend on parameter count, context length, cache precision, quantization, and concurrency.
  • Report reproducible issues with the runtime version, hardware, launch command, and a minimal example.

Quantization or conversion changes numerical behavior, memory use, and throughput relative to the source checkpoint; validate quality on your own workload.

Generated outputs may be inaccurate or unsuitable for a given use case. Users are responsible for
testing behavior, applying appropriate safeguards, and complying with applicable licenses and laws.

GPTQ-PRO W4A16 quantization (group size 128) of Huihui/Qwen3.5-9B-abliterated.

Quantization Details

  • Quantization method: GPTQ-PRO
  • Bits/Config: W4A16 (group 128)
  • Base model: Huihui/Qwen3.5-9B-abliterated

README history 5 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-08-22Polish model card overview and usage notes767b4702.8 KB
    Loading...
  2. 2026-08-22Polish model card overview and usage notes8edaa412.8 KB
    Loading...
  3. 2026-08-22Polish model card overview and usage notesf0c1f0b2.9 KB
    Loading...
  4. 2026-08-22Polish model card overview and usage notes40a39c71.9 KB
    Loading...
  5. 2026-03-29Upload folder using huggingface_hub4c2f078385 B
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration