← back to catalog · registered 2026-08-22 13:56

groxaxo/lukey03-Qwen3.5-9B-abliterated-gptq-pro-w4g128

groxaxo Qwen 6.9B second-order
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/groxaxo%2Flukey03-Qwen3.5-9B-abliterated-gptq-pro-w4g128"
Response includes
  • classification m1
  • files 12
  • hub_downloads_all_time 4,460
  • author_summary 27 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M1
Primary method

Direct removal

No other method signals detected in this model.
Confidence
MEDIUM
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=0 (base model)
  • no specific method indicators - defaulting to M1 (most common)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
MEDIUM
Why we say so
primary_method=M1; difference-of-means is the reference extraction for M1/M3 (Arditi 2024)
Downloads · lifetime
4K
143 last 30d - cooling
Likes
2
Model age
6mo ago
created 2026-03-23
Downloads over time
Now4.5K→from2.7K↑65%
2.7K3.3K4K4.7K2.7K on Mar 254.5K on Oct 11MarAprMayJunJulAugSepOct
Mar 25 → Oct 11 · 68 snapshots · spans 200 days

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

Tags
gptqmodel safetensors qwen3_5_text gptq quantized text-generation marlin vllm conversational base_model:lukey03/Qwen3.5-9B-abliterated base_model:quantized:lukey03/Qwen3.5-9B-abliterated 4-bit

Related

Total size
7.15 GB
Files
12
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-08-22 08:00

Files by quantization

Auxiliary files 12 files 7.17 GB
model-00001-of-00002.safetensors 3.99 GB 51dfbfb9 download
model-00002-of-00002.safetensors 3.16 GB bbb37c28 download
tokenizer.json 19.1 MB 6a0316e3 download
model.safetensors.index.json 83.3 KB 39fdcbc2 download
quant_log.csv 9.38 KB 7ae62693 download
chat_template.jinja 7.57 KB a585dec8 download
README.md 6.52 KB 2cdf2a0f download
config.json 3.28 KB 7e608de2 download
.gitattributes 1.53 KB 52373fe2 download
quantize_config.json 1.22 KB 301b2cfa download
tokenizer_config.json 1.18 KB d5471dee download
generation_config.json 136 B e0f0c5c1 download

README current version from Hugging Face


base_model: lukey03/Qwen3.5-9B-abliterated
library_name: gptqmodel
pipeline_tag: text-generation
tags:

  • gptq
  • gptqmodel
  • quantized
  • qwen3_5_text
  • text-generation
  • marlin
  • vllm

lukey03/Qwen3.5-9B-abliterated GPTQ-Pro 4-bit g128

Overview

lukey03-Qwen3.5-9B-abliterated-gptq-pro-w4g128 is a GPTQ-quantized checkpoint intended for efficient GPU inference, published by groxaxo.
It is intended for open-source evaluation, reproducible experimentation, and compatible local or
hosted inference workflows. The wording below is deliberately limited to what can be verified
from this repository's metadata and artifacts.

The repository name identifies a behavior-modified or reduced-filtering lineage. That label describes the source or conversion history; it is not a guarantee of unrestricted behavior in every prompt or runtime. Test outputs carefully before sharing or deploying them.

At a glance

Field Details
Format GPTQ
Source / base lukey03/Qwen3.5-9B-abliterated
Intended task text-generation
License the license declared in the repository files

What is included

  • *.safetensors (2 files)
  • config.json
  • generation_config.json
  • tokenizer.json
  • tokenizer_config.json
  • chat_template.jinja
  • quantize_config.json
  • Additional configuration, tokenizer, processor, or shard files (10 visible artifacts total)

Quick start

vLLM (documented configuration)

vllm serve groxaxo/lukey03-Qwen3.5-9B-abliterated-gptq-pro-w4g128 \
  --quantization gptq_marlin \
  --dtype float16 \
  --trust-remote-code

This command is taken from the repository documentation. Adjust tensor parallelism, context
length, and cache settings to match your hardware and vLLM version.

Compatibility and responsible use

  • Use a runtime that explicitly supports this format, architecture, and modality.
  • Keep configuration, tokenizer, processor, projection, and weight files from the same revision together.
  • Review the source model card and license before redistribution or deployment.
  • Hardware needs depend on parameter count, context length, cache precision, quantization, and concurrency.
  • Report reproducible issues with the runtime version, hardware, launch command, and a minimal example.

Quantization or conversion changes numerical behavior, memory use, and throughput relative to the source checkpoint; validate quality on your own workload.

Generated outputs may be inaccurate or unsuitable for a given use case. Users are responsible for
testing behavior, applying appropriate safeguards, and complying with applicable licenses and laws.

This is a GPTQ-Pro 4-bit / group-size-128 export of lukey03/Qwen3.5-9B-abliterated, produced with GPTQModel.

Why this checkpoint

I compared the local BF16 base model, this GPTQ-Pro export, symmetric AWQ GEMM, and asymmetric AWQ GEMM on 1x RTX 3090 with warmed generation runs and WikiText-2 perplexity:

Variant Perplexity Warm speed
BF16 base 10.2572 31.36 tok/s
GPTQ-Pro local runtime (BACKEND.GPTQ_PRO) 10.6271 4.95 tok/s
GPTQ-Pro deployed with Marlin quality preserved 31.18 tok/s
AWQ GEMM (sym=True) 37581.1150 14.35 tok/s
AWQ GEMM (sym=False) 75868.4507 37.14 tok/s

Result: GPTQ-Pro preserved quality, while AWQ was not quality-safe on this Qwen3.5 family.

Quantization settings

  • bits=4
  • group_size=128
  • sym=True
  • desc_act=False
  • format=gptq
  • quant_method=gptq
  • pack_dtype=int32

For Qwen3.5 text checkpoints, quantization should use batch_size=1.

Recommended runtime

For correctness or direct local validation, BACKEND.GPTQ_PRO works.

For actual deployment, the recommended path is:

  1. quantize with GPTQ-Pro
  2. serve with Marlin
  3. for highest throughput on this model family, use the patched Qwen3.5 vLLM wrapper from GPTQ-Pro

The repo-validated fast path is vLLM + gptq_marlin through:

  • scripts/serve_vllm_qwen35.py

That wrapper now auto-detects qwen3_5_text from either a local folder or a Hub repo ID.

Quickstart: GPTQModel + Marlin

from transformers import AutoTokenizer
from gptqmodel import GPTQModel, BACKEND

MODEL_ID = "groxaxo/lukey03-Qwen3.5-9B-abliterated-gptq-pro-w4g128"

tokenizer = AutoTokenizer.from_pretrained(MODEL_ID, trust_remote_code=True)
model = GPTQModel.load(
    MODEL_ID,
    backend=BACKEND.MARLIN,
    device="cuda:0",
    trust_remote_code=True,
)

prompt = "State one short fact about model quantization."
inputs = tokenizer(prompt, return_tensors="pt").to(model.device)
outputs = model.generate(**inputs, max_new_tokens=64)
print(tokenizer.decode(outputs[0], skip_special_tokens=True))

Quickstart: patched vLLM serve path

Clone the matching repo branch first so you get the Qwen3.5 wrapper and patches:

git clone https://github.com/groxaxo/GPTQ-Pro.git
cd GPTQ-Pro
git checkout gptq-pro-cuda-kernel

Then launch the model:

CUDA_VISIBLE_DEVICES=0 \
python scripts/serve_vllm_qwen35.py \
  --model groxaxo/lukey03-Qwen3.5-9B-abliterated-gptq-pro-w4g128 \
  --served-model-name qwen35-9b-gptq-pro \
  --host 0.0.0.0 \
  --port 8011 \
  --tensor-parallel-size 1

Two-GPU shared-host launch:

CUDA_VISIBLE_DEVICES=0,1 \
python scripts/serve_vllm_qwen35.py \
  --model groxaxo/lukey03-Qwen3.5-9B-abliterated-gptq-pro-w4g128 \
  --served-model-name qwen35-9b-gptq-pro \
  --host 0.0.0.0 \
  --port 8012 \
  --tensor-parallel-size 2 \
  --gpu-memory-utilization 0.4

When the fast path is active, the logs should contain:

Using MarlinLinearKernel for GPTQMarlinLinearMethod

Speed demon settings for future runs

If you want the same tradeoff in future Qwen3.5 quantization + deployment runs:

  • quantize with GPTQ-Pro, bits=4, group_size=128, sym=True, desc_act=False
  • keep batch_size=1 during Qwen3.5 quantization
  • deploy with BACKEND.MARLIN or the patched vLLM wrapper
  • prefer tensor_parallel_size=1 first, then scale to 2 GPUs if needed
  • on shared hosts, start with --gpu-memory-utilization 0.4
  • confirm the logs show MarlinLinearKernel

Notes

  • BACKEND.GPTQ_PRO is a functional runtime path, but it is not the throughput-optimized serving path.
  • I did not publish the AWQ exports for this model family as the recommended release because both symmetric and asymmetric AWQ failed the quality check on this benchmark setup.

README history 4 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-08-22Polish model card overview and usage notes1df9f746.5 KB
    Loading...
  2. 2026-08-22Polish model card overview and usage notes8d494726.5 KB
    Loading...
  3. 2026-08-22Polish model card overview and usage notese56055e5.7 KB
    Loading...
  4. 2026-03-23Add files using upload-large-folder toolacc56654 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration