← back to catalog · registered 2026-10-08 17:58

davetha/Huihui-Qwen3.8-27B-abliterated-AutoRound-Int4-baked-embed-int8

davetha 27B multimodal second-order
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/davetha%2FHuihui-Qwen3.8-27B-abliterated-AutoRound-Int4-baked-embed-int8"
Response includes
  • classification m-uncensored
  • files 25
  • author_summary 5 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M-U
Primary method

Uncensored (method unknown)

No other method signals detected in this model.
Confidence
LOW
Why this label 3 signals
Weak or ambiguous signals. Best guess based on catalog patterns; treat as tentative and check the evidence below.
  • 'uncensored' in name/tags but no 'abliterated' marker
  • method not identifiable from author declaration alone
  • may be DPO fine-tune, prompt engineering, or unknown technique
Refusal direction extraction

No specific extraction method could be identified for this model. The producer either did not document it or used a proprietary pipeline.

What is a refusal direction? →
Downloads · 30-day
0
Likes
0
Model age
today
created 2026-10-08

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
apache-2.0
Tags
transformers safetensors qwen3_5 image-text-to-text qwen3.8 abliterated uncensored autoround gptq int4 int8-embedding mtp

Related

Total size
17.1 GB
Files
25
Quantizations
1
Registered
2026-10-08 17:58
Last updated on HF
2026-10-08 17:34

Files by quantization

Auxiliary files 25 files 17.1 GB
model-00002-of-00004.safetensors 5.00 GB 04b766e0 download
model-00001-of-00004.safetensors 4.99 GB bcf7ed6b download
model-00003-of-00004.safetensors 2.69 GB 07383a74 download
model-00004-of-00004.safetensors 2.37 GB 55a14ee7 download
model-embed-int8.safetensors 1.18 GB 1722f347 download
model-bake-int4.safetensors 841 MB 26d3fb4a download
compat_g_idx.safetensors 11.1 MB ca16bd7e download
model_extra_tensors.safetensors 51.7 KB 852b2283 download
tokenizer.json 19.1 MB 06b95093 download
model.safetensors.index.json 223 KB e8ad7b4a download
LICENSE 11.3 KB f938136e download
chat_template.jinja 8.74 KB c0c686f9 download
chat_template_low.jinja 8.74 KB 3b904292 download
README.md 8.05 KB 42707df1 download
config.json 4.36 KB cdf48e89 download
BAKE.json 3.58 KB a5b7b057 download
.gitattributes 1.53 KB 52373fe2 download
processor_config.json 1.19 KB 43c4343e download
tokenizer_config.json 1.14 KB 1d134cd2 download
B65-LICENSE 1.06 KB 62b94ee7 download
preprocessor_config.json 443 B 8ed39680 download
COMPATIBILITY.json 360 B 90e7765a download
quantization_config.json 215 B 184c2bc1 download
quantize_config.json 215 B 184c2bc1 download
generation_config.json 214 B 1f489fac download

README current version from Hugging Face


license: apache-2.0
library_name: transformers
pipeline_tag: image-text-to-text
base_model:

  • huihui-ai/Huihui-Qwen3.8-27B-abliterated
    tags:
  • qwen3.8
  • abliterated
  • uncensored
  • autoround
  • gptq
  • int4
  • int8-embedding
  • mtp
  • intel-arc
  • arc-pro-b70
  • xpu
  • vllm

Huihui-Qwen3.8-27B-abliterated — AutoRound INT4, baked head/MTP, INT8 embedding

A 4-bit build of huihui-ai/Huihui-Qwen3.8-27B-abliterated (an abliterated, i.e. refusal-removed, version of Qwen/Qwen3.8-27B) made for a single Intel Arc Pro B70 (32 GB) under vLLM XPU. It follows the recipe of ukisai/Swift-1.5-Qwen3.8-27b-W4A16-AutoRound for the body and of Wrapzii's Swift 1.5 GPTQ-Int4-baked-v1-embed-int8 for the head, MTP and embedding — applied to the uncensored Huihui weights.

Uncensored model. Abliteration removes most refusal behaviour. It will follow requests the original model refuses. You are responsible for how you use it.

What is in the files

part form how
400 text body linears INT4, symmetric, group 128 (GPTQ-compatible packing) AutoRound / SignRound from the bf16 Huihui weights
lm_head INT4 GPTQ, group 128 Hessian GPTQ on this model's own activations
MTP block (fc, q/k/v/o, gate/up/down) INT4 GPTQ, group 128 same
token embedding per-row symmetric INT8 + FP16 scales (model-embed-int8.safetensors) absmax/127
vision tower, MTP norms, GDN in_proj_a / in_proj_b, norms BF16 (unchanged) —

Weights resident on the GPU: 14.75 GiB (vs 18.28 GiB for a GPTQ build that keeps lm_head / MTP dense), which leaves room for the full 262,144-token context on one 32 GB card (vLLM reports a 319K-token KV pool with fp8 KV).

The dense embedding is still in the shards, so stock loaders work; the INT8 embedding only takes effect in a runtime that reads embed_tokens_quant (the patched image below).

Quality

WikiText-2 test perplexity, 40 chunks x 2,048 tokens (81,880 scored tokens), the same token chunks for both, computed from vLLM prompt_logprobs on the B70:

build perplexity
this repo (AutoRound body from bf16 + baked head/MTP + INT8 embedding) 6.266
zrlu's GPTQ-Int4 Huihui export + the same head/MTP/embedding bake 6.346

Head / MTP bake on held-out activations (relative output L2 error; GPTQ vs round-to-nearest):

linear GPTQ round-to-nearest
lm_head 3.39% 7.61%
mtp.fc 5.67% 11.86%
MTP q / k / v / o 2.11% / 4.99% / 4.06% / 5.04% 4.42% / 10.46% / 8.33% / 12.12%
MTP gate / up / down 2.93% / 4.59% / 5.74% 6.46% / 9.75% / 11.30%
embedding (INT8, weight L2) 0.90% —

These are matrix and perplexity checks, not a full benchmark suite; upstream evaluation scores have not been re-established for this derivative. A quick functional check (arithmetic, thinking on/off, tool-call template, an uncensored prompt) passed.

Speed on one Arc Pro B70

Image vllm-openai-xpu:v0.30.0-k8v4-tp1 (vLLM 0.30 XPU + Wrapzii/k8v4-xpu patches incl. the INT8-embedding loader), TP1, fp8 KV, MTP with 3 speculative tokens, decode cudagraph mode FULL_DECODE_ONLY, max_num_seqs=4.

BetterBench (temperature 0.7, top-p 0.95, thinking on, 20 runs per category):

category decode t/s (median) tokens per update
code 73.6 2.53
reasoning 62.8 2.29
prose 64.3 2.27
json 92.6 3.36
file_edit 86.5 3.06
summarization 88.6 3.15
math 88.9 3.17
chat 68.4 2.37
weighted combined 75.7

TTFT p50 ~160 ms; update gap p99 37.4 ms.

concurrent requests aggregate t/s per-request decode t/s
1 72.0 81.7
2 129.7 74.4
4 196.0 61.6
8 / 16 ~200 (capped by max_num_seqs=4) ~62
prompt length prefill t/s (median)
1.6K 1,848
6.0K 1,888
11.8K 1,772
23.6K 1,566
47.1K 1,285

Greedy, thinking off (single runs): prose ~72 t/s, JSON ~104 t/s; MTP draft acceptance ~60%.

For reference, the Swift 1.5 bake (a different fine-tune, run with 6 speculative tokens) measured 12–34% faster per BetterBench category on the same card, except prose, where this build was 5% faster. This build has not been tuned for speculative depth yet.

Serving (vLLM XPU, one Arc Pro B70)

docker run -d --name qwen27-b70 \
  --device /dev/dri/card1 --device /dev/dri/renderD128 --group-add render \
  --ipc=host --shm-size=8g -p 8200:8000 \
  -v /path/to/this/repo:/model:ro \
  -e VLLM_TARGET_DEVICE=xpu -e VLLM_XPU_ENABLE_XPU_GRAPH=1 \
  vllm-openai-xpu:v0.30.0-k8v4-tp1 /model \
  --host 0.0.0.0 --port 8000 --served-model-name qwen27-huihui-unc \
  --gpu-memory-utilization 0.92 --dtype bfloat16 \
  --max-model-len 262144 --kv-cache-dtype fp8 --tensor-parallel-size 1 \
  --max-num-seqs 4 --max-num-batched-tokens 4224 \
  --enable-auto-tool-choice --tool-call-parser qwen3_xml --reasoning-parser qwen3 \
  --enable-prefix-caching --trust-remote-code \
  --limit-mm-per-prompt '{"image":4,"video":0}' --mm-processor-kwargs '{"max_pixels":1048576}' \
  --speculative-config '{"method":"mtp","num_speculative_tokens":3}' \
  --compilation-config '{"cudagraph_mode":"FULL_DECODE_ONLY","cudagraph_capture_sizes":[4,8,12,16]}'

Build the image from Wrapzii/k8v4-xpu (k8v4_v030/, plus Dockerfile.swift-bake for the INT8-embedding loader). A stock vLLM XPU image loads the checkpoint too, with the embedding dense.

Thinking: chat_template_kwargs: {"enable_thinking": false} turns it off; reasoning_effort takes low / medium / xhigh (the bundled chat_template_low.jinja defaults to low). Tool calls use the qwen3 XML format.

How it was made

  1. Body — AutoRound 0.16.0 on the bf16 checkpoint, on the B70 (2.5 h, 27.5 GB peak VRAM):
    auto-round --model huihui-ai/Huihui-Qwen3.8-27B-abliterated --scheme W4A16 --bits 4 --group_size 128 \
      --iters 200 --batch_size 8 --nsamples 256 --seqlen 2048 --seed 42 \
      --dataset HuggingFaceH4/ultrachat_200k \
      --ignore_layers "lm_head,.*visual.*,.*mtp.*,.*in_proj_a,.*in_proj_b" \
      --format auto_round --low_gpu_mem_usage
    
    AutoRound writes the 15 MTP tensors (dropped by the Transformers loader) back from the source into model_extra_tensors.safetensors.
  2. GPTQ compatibility — tools/prepare_swift_autoround.py from k8v4-xpu: verifies the 400 packed linears and zero points, adds deterministic g_idx, switches the config to gptq. No weight values change.
  3. Calibration — tools/calibrate_swift_bake.py (run at TP1 with fp8 KV on the single card): 128 WikiText-2 validation prompts of 384 tokens, 64 generated tokens each, MTP on; Hessians for lm_head and the MTP inputs, every 64th row held out.
  4. Bake — tools/bake_gptq_resume.py: symmetric GPTQ INT4, group 128, 1% damping, no activation ordering.
  5. Assemble — tools/assemble_swift_bake.py: rewritten shards without the dense replaced tensors, the 36 GPTQ arrays in model-bake-int4.safetensors, the INT8 embedding side file. Settings and per-linear errors are in BAKE.json.

Credits and licenses

Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration