← back to catalog · registered 2026-08-22 13:56

groxaxo/Qwen3.6-27B-abliterated-v2-GPTQ-Pro-FOEM-4bit-g128

groxaxo Qwen 24B multimodal
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/groxaxo%2FQwen3.6-27B-abliterated-v2-GPTQ-Pro-FOEM-4bit-g128"
Response includes
  • classification m1
  • files 15
  • hub_downloads_all_time 45,491
  • author_summary 27 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M1
Primary method

Direct removal

No other method signals detected in this model.
Confidence
MEDIUM
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=0 (base model)
  • no specific method indicators - defaulting to M1 (most common)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
MEDIUM
Why we say so
primary_method=M1; difference-of-means is the reference extraction for M1/M3 (Arditi 2024)
Downloads · lifetime
45K
268 last 30d - cooling
Likes
1
Model age
5mo ago
created 2026-04-29
Downloads over time
Now45.7K→from33↑138,252%
016.7K33.5K50.2K33 on Apr 2945.7K on Oct 11AprMayJunJulAugSepOct
Apr 29 → Oct 11 · 64 snapshots · spans 165 days

Metadata

Tags
transformers safetensors qwen3_5 image-text-to-text qwen qwen2-vl multimodal vision-language gptq gptq-pro 4bit foem

Related

Total size
17.4 GB
Files
15
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-08-22 07:59

Files by quantization

Auxiliary files 15 files 17.4 GB
model-00004-of-00005.safetensors 3.99 GB 70c5d89c download
model-00002-of-00005.safetensors 3.98 GB 2d281d2a download
model-00003-of-00005.safetensors 3.97 GB 2f19b75a download
model-00001-of-00005.safetensors 3.28 GB 028ad551 download
model-00005-of-00005.safetensors 2.21 GB c2a18f84 download
tokenizer.json 19.1 MB 87a7830d download
model.safetensors.index.json 223 KB d4bddac4 download
chat_template.jinja 7.87 KB f7a7d1b0 download
config.json 5.34 KB 86c5373e download
README.md 4.98 KB 61ebffaf download
.gitattributes 1.53 KB 52373fe2 download
quantize_config.json 1.42 KB 37dcb1fa download
processor_config.json 1.30 KB 32f01452 download
tokenizer_config.json 1.14 KB 3f49c3b4 download
generation_config.json 187 B 91511aaa download

README current version from Hugging Face


library_name: transformers
pipeline_tag: image-text-to-text
tags:

  • qwen
  • qwen2-vl
  • multimodal
  • vision-language
  • gptq
  • gptq-pro
  • 4bit
  • foem

Qwen3.6-27B-abliterated-v2-GPTQ-Pro-FOEM-4bit-g128

Overview

Qwen3.6-27B-abliterated-v2-GPTQ-Pro-FOEM-4bit-g128 is a GPTQ-quantized checkpoint intended for efficient GPU inference, published by groxaxo.
It is intended for open-source evaluation, reproducible experimentation, and compatible local or
hosted inference workflows. The wording below is deliberately limited to what can be verified
from this repository's metadata and artifacts.

The repository name identifies a behavior-modified or reduced-filtering lineage. That label describes the source or conversion history; it is not a guarantee of unrestricted behavior in every prompt or runtime. Test outputs carefully before sharing or deploying them.

At a glance

Field Details
Format GPTQ
Source / base the source checkpoint identified in the repository metadata
Intended task image-text-to-text
License the license declared in the repository files

What is included

  • *.safetensors (5 files)
  • config.json
  • generation_config.json
  • tokenizer.json
  • tokenizer_config.json
  • processor_config.json
  • chat_template.jinja
  • quantize_config.json
  • Additional configuration, tokenizer, processor, or shard files (13 visible artifacts total)

Quick start

vLLM (documented configuration)

vllm serve groxaxo/Qwen3.6-27B-abliterated-v2-GPTQ-Pro-FOEM-4bit-g128 \
  --quantization gptq_marlin \
  --dtype float16 \
  --trust-remote-code

This command is taken from the repository documentation. Adjust tensor parallelism, context
length, and cache settings to match your hardware and vLLM version.

Compatibility and responsible use

  • Use a runtime that explicitly supports this format, architecture, and modality.
  • Keep configuration, tokenizer, processor, projection, and weight files from the same revision together.
  • Review the source model card and license before redistribution or deployment.
  • Hardware needs depend on parameter count, context length, cache precision, quantization, and concurrency.
  • Report reproducible issues with the runtime version, hardware, launch command, and a minimal example.

Quantization or conversion changes numerical behavior, memory use, and throughput relative to the source checkpoint; validate quality on your own workload.

Generated outputs may be inaccurate or unsuitable for a given use case. Users are responsible for
testing behavior, applying appropriate safeguards, and complying with applicable licenses and laws.

FOEM-enhanced GPTQ-Pro W4G128 export of Qwen3.6-27B-abliterated-v2.

Quantization recipe

  • Base quant recipe: GPTQ-Pro W4G128
  • Enhancement: FOEM + activation-weighted MSE
  • bits: 4
  • group_size: 128
  • desc_act: false
  • sym: true
  • true_sequential: true
  • act_group_aware: true
  • lm_head: false
  • FOEM: alpha=0.25, beta=0.2

What stayed preserved

  • Vision tower tensors remain present and BF16
  • lm_head remains BF16
  • model.language_model.embed_tokens remains BF16
  • Processor and multimodal metadata are included

Verification notes

The local artifact was checked to confirm:

  • visual tensors > 0
  • quantized tensors are present
  • lm_head is not quantized
  • embed_tokens is not quantized

The source model was also validated locally with image input before quantization.

Local launch settings used to validate the source vision path

The stable local vLLM launch on this machine was:

setsid env \
  CUDA_VISIBLE_DEVICES=0,1,4 \
  CUDA_DEVICE_ORDER=PCI_BUS_ID \
  OMP_NUM_THREADS=1 \
  TOKENIZERS_PARALLELISM=false \
  NCCL_P2P_DISABLE=1 \
  NCCL_IB_DISABLE=1 \
  NCCL_NET_GDR_DISABLE=1 \
  NCCL_SHM_DISABLE=0 \
  NCCL_CUMEM_HANDLE_DISABLE=1 \
  PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True,max_split_size_mb:256 \
  /home/op/venvs/vllm-qwen36/bin/vllm serve "/home/op/models/Qwen3.6-27B-abliterated-v2" \
    --host 0.0.0.0 \
    --port 8000 \
    --tensor-parallel-size 1 \
    --pipeline-parallel-size 3 \
    --max-model-len 4096 \
    --kv-cache-dtype fp8 \
    --gpu-memory-utilization 0.98 \
    --max-num-seqs 1 \
    --reasoning-parser qwen3 \
    --enable-auto-tool-choice \
    --tool-call-parser qwen3_coder \
    --trust-remote-code \
    --served-model-name qwen36-27b-abliterated-v2 \
    --disable-custom-all-reduce \
    --generation-config vllm \
    --enforce-eager \
    --limit-mm-per-prompt '{"image":1}'

Notes:

  • --max-model-len 32144 did not fit KV cache on this host.
  • For direct chat requests, chat_template_kwargs.enable_thinking=false was used to avoid spending the visible token budget in the reasoning channel.

Files

  • quantize_config.json records the GPTQ-Pro + FOEM config
  • processor_config.json keeps the multimodal processor config
  • model.safetensors.index.json and shard files contain the final export

README history 6 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-08-22Polish model card overview and usage notes8d853055 KB
    Loading...
  2. 2026-08-22Polish model card overview and usage notes586afb25 KB
    Loading...
  3. 2026-08-22Polish model card overview and usage notese8a83555 KB
    Loading...
  4. 2026-08-22Polish model card overview and usage notes2912b8b4 KB
    Loading...
  5. 2026-04-29Upload README.md with huggingface_hubcef087d2.5 KB
    Loading...
  6. 2026-04-29Add files using upload-large-folder toolc6a37312.3 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration