← back to catalog · registered 2026-10-04 19:58

ericmey/Qwen3.8-27B-abliterated-GPTQ-Int4-MTP

ericmey 27B multimodal second-order
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/ericmey%2FQwen3.8-27B-abliterated-GPTQ-Int4-MTP"
Response includes
  • classification m-uncensored
  • files 18
  • author_summary 1 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M-U
Primary method

Uncensored (method unknown)

No other method signals detected in this model.
Confidence
LOW
Why this label 3 signals
Weak or ambiguous signals. Best guess based on catalog patterns; treat as tentative and check the evidence below.
  • 'uncensored' in name/tags but no 'abliterated' marker
  • method not identifiable from author declaration alone
  • may be DPO fine-tune, prompt engineering, or unknown technique
Refusal direction extraction

No specific extraction method could be identified for this model. The producer either did not document it or used a proprietary pipeline.

What is a refusal direction? →
Downloads · 30-day
0
Likes
1
Model age
today
created 2026-10-04

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
apache-2.0
Languages
en zh
Tags
transformers safetensors qwen3_5 image-text-to-text gptq int4 w4a16 vllm xpu intel-arc mtp vision

Related

Total size
18.2 GB
Files
18
Quantizations
1
Registered
2026-10-04 19:58
Last updated on HF
2026-10-04 19:52

Files by quantization

Auxiliary files 18 files 18.2 GB
model-00004-of-00005.safetensors 3.99 GB 0e4c1737 download
model-00002-of-00005.safetensors 3.98 GB 28613e3d download
model-00003-of-00005.safetensors 3.97 GB 25a37658 download
model-00001-of-00005.safetensors 3.28 GB 16cc514d download
model-00005-of-00005.safetensors 3.00 GB 24e5c14f download
tokenizer.json 19.1 MB a5cd9732 download
model.safetensors.index.json 224 KB 4580b435 download
quant_log.csv 18.8 KB 0b53c93e download
chat_template.jinja 8.74 KB c0c686f9 download
README.md 7.99 KB 78430664 download
config.json 5.07 KB a17b1869 download
.gitattributes 1.53 KB 52373fe2 download
tokenizer_config.json 1.46 KB 24bac0b1 download
processor_config.json 1.19 KB 43c4343e download
quantize_config.json 1.16 KB c76a363c download
preprocessor_config.json 390 B 2ea84a43 download
video_preprocessor_config.json 385 B 3ba673a5 download
generation_config.json 214 B a6d9d5ca download

README current version from Hugging Face


library_name: transformers
license: apache-2.0
pipeline_tag: image-text-to-text
language:

  • en
  • zh
    base_model: huihui-ai/Huihui-Qwen3.8-27B-abliterated
    base_model_relation: quantized
    quantized_by: ericmey
    tags:
  • gptq
  • int4
  • w4a16
  • vllm
  • xpu
  • intel-arc
  • mtp
  • vision
  • abliterated
  • uncensored
  • qwen3

Qwen3.8-27B-abliterated — GPTQ INT4 (sym, G128) + BF16 MTP

license quant engine mtp vision context uncensored

GPTQ-Int4 quantization of
huihui-ai/Huihui-Qwen3.8-27B-abliterated,
packaged for vLLM serving on Intel Arc (Xe2 / B-series) with native MTP speculative
decoding
and the vision tower preserved. The 15 mtp.* draft tensors are deliberately
excluded from quantization and kept in BF16, so speculative decoding loads and runs out of
the box (the Qwen3_5ForConditionalGeneration class drops mtp.* on load, so many quantizers
silently lose it — this build keeps it).

Qwen/Qwen3.8-27B                       (base, Qwen team)
  └─ huihui-ai/Huihui-Qwen3.8-27B-abliterated   (refusal-direction ablation, BF16)
       └─ this repo                    (GPTQ-Int4 + BF16 MTP + F16 vision, for vLLM XPU)

Scope of this repo. I did the quantization and Intel-Arc packaging only. The weights are
huihui-ai's abliterated model, re-quantized — not a new fine-tune or a different abliteration.
Credit to the Qwen team (base) and huihui-ai (abliteration).

TL;DR

Base huihui-ai/Huihui-Qwen3.8-27B-abliterated (abliterated Qwen3.8-27B)
Format GPTQ INT4, symmetric, group size 128, desc_act=false, int32 pack
Preserved 15 mtp.* draft tensors (BF16) · vision tower (F16)
Modality Text + vision (image), tool-calling, thinking toggle
Context 262,144 native (served/tested at 131,072)
Size on disk ~19 GB (5 shards)
Target engine vLLM (Intel Arc / XPU W4A16; also loads on CUDA)
Censorship Abliterated / uncensored (see Usage Warnings)

Files

File Size
model-00001-of-00005.safetensors 3.3 GB
model-0000{2,3,4}-of-00005.safetensors 4.0 GB each
model-00005-of-00005.safetensors 3.0 GB
config.json, quantize_config.json, generation_config.json configs
tokenizer*.json, chat_template.jinja tokenizer + chat template
preprocessor_config.json, video_preprocessor_config.json, processor_config.json vision/video processor

Quantization recipe

Produced with gptqmodel 7.3.2:

from gptqmodel import GPTQModel, QuantizeConfig
qcfg = QuantizeConfig(
    bits=4, group_size=128, sym=True, desc_act=False,
    lm_head=False, true_sequential=True,
    dynamic={"-:.*mtp.*": {}},   # keep the 15 mtp.* draft tensors in BF16
)
model = GPTQModel.load("huihui-ai/Huihui-Qwen3.8-27B-abliterated", qcfg)
model.quantize(calibration, batch_size=1)   # ~256 general-text samples, seqlen 2048
model.save("Qwen3.8-27B-abliterated-GPTQ-Int4-MTP")

Result: 400 quantized weight tensors (int32) + 15 BF16 mtp.* tensors + the vision tower (F16).
The vision tower is left unquantized automatically (gptqmodel only quantizes the decoder Linears).

Usage (vLLM, Intel Arc / XPU)

vllm serve ericmey/Qwen3.8-27B-abliterated-GPTQ-Int4-MTP \
  --quantization gptq --dtype float16 --max-model-len 131072 \
  --tensor-parallel-size 2 --gpu-memory-utilization 0.92 \
  --kv-cache-dtype auto --enable-prefix-caching \
  --mamba-cache-mode align --mamba-ssm-cache-dtype float16 \
  --limit-mm-per-prompt '{"image":4,"video":0}' \
  --mm-processor-kwargs '{"max_pixels":1048576}' \
  --speculative-config '{"method":"mtp","num_speculative_tokens":3}' \
  --enable-auto-tool-choice --tool-call-parser qwen3_coder --reasoning-parser qwen3

Minimal (any backend — GPTQ kernels are hardware-independent, so this also loads on CUDA):

from vllm import LLM
llm = LLM("ericmey/Qwen3.8-27B-abliterated-GPTQ-Int4-MTP", tensor_parallel_size=2)

Speculative decoding (MTP)

The 15 mtp.* draft tensors are shipped in BF16 and excluded from quantization via
dynamic={"-:.*mtp.*": {}}. Under vLLM the model resolves as Qwen3_5MTP and accepts
--speculative-config '{"method":"mtp","num_speculative_tokens":3}'. Load- and generation-verified;
it is what produces the decode speeds below.

Prompt template & sampling

Uses the Qwen3 chat template (shipped chat_template.jinja) with thinking control and tool-calling.
Recommended sampling from the Qwen3.8-27B card:

  • Non-thinking: temperature 0.7, top_p 0.80, top_k 20
  • Thinking: temperature 1.0, top_p 0.95, top_k 20

Benchmarks

Measured on 2× Intel Arc Pro B60 (24 GB), TP=2, vLLM XPU, MTP3, client-side timing;
compared against the non-abliterated GPTQ build of the same base, same recipe/args:

Check This model Non-abliterated baseline
Decode — code / explain / agent (tok/s) 47.3 / 48.4 / 50.7 47.2 / 42.7 / 49.2
Decode mean (tok/s) 48.8 46.4
Exact-answer sanity (20 items, greedy) 20 / 20 20 / 20
Tool-call well-formedness (60 ×2 temps) 60/60 · 60/60 60/60 · 60/60
Vision smoke test (read text + shape) PASS PASS

Capabilities and speed match the non-abliterated build — abliteration changes no tensor shapes
and does not touch the MTP head, so there is no quantization or speed penalty.

Limitations

  • Abliterated base. Refusal directions were removed upstream by huihui-ai; that behavior is
    inherited here and is not something quantization changes.
  • Benchmarks above are a smoke-test suite (speed, exact-answer sanity, tool well-formedness, one
    vision probe), not broad downstream task benchmarks (MMLU/GSM8K/etc. not run here).
  • Tuned and measured for the Intel Arc W4A16 (vLLM XPU) path; it loads on CUDA but the
    recipe/args are Arc-oriented.
  • MTP acceptance rate not separately reported (speculative decoding is load- and speed-verified).

Usage Warnings

This is an uncensored / abliterated model: safety refusals have been removed, and outputs are
unfiltered and may be offensive, inaccurate, or otherwise objectionable. It is intended for
adults, for research, creative writing, and local/private use.

  • You are solely responsible for how you use it and for compliance with all applicable laws.
  • Do not use it to generate illegal content — including, unconditionally, any sexual content
    involving minors — or to facilitate harm to others.
  • No guardrails and no warranty. Validate outputs before relying on them.

Credits & license

Licensed Apache-2.0, following the base model.

Citation

@misc{qwen38_27b_abliterated_gptq_int4_mtp,
  title  = {Qwen3.8-27B-abliterated-GPTQ-Int4-MTP},
  author = {ericmey},
  year   = {2026},
  note   = {GPTQ-Int4 (sym, g128) + BF16 MTP quantization of huihui-ai/Huihui-Qwen3.8-27B-abliterated for vLLM XPU},
  howpublished = {\url{https://huggingface.co/ericmey/Qwen3.8-27B-abliterated-GPTQ-Int4-MTP}}
}
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration