← back to catalog · registered 2026-08-22 13:56

johannrplaster/Qwen3.8-27B-Uncensored-int4-AutoRound

johannrplaster Qwen 25B multimodal
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/johannrplaster%2FQwen3.8-27B-Uncensored-int4-AutoRound"
Response includes
  • classification m-uncensored
  • files 18
  • hub_downloads_all_time 2,308
  • author_summary 1 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M-U
Primary method

Uncensored (method unknown)

No other method signals detected in this model.
Confidence
LOW
Why this label 3 signals
Weak or ambiguous signals. Best guess based on catalog patterns; treat as tentative and check the evidence below.
  • 'uncensored' in name/tags but no 'abliterated' marker
  • method not identifiable from author declaration alone
  • may be DPO fine-tune, prompt engineering, or unknown technique
Refusal direction extraction

No specific extraction method could be identified for this model. The producer either did not document it or used a proprietary pipeline.

What is a refusal direction? →
Downloads · lifetime
2K
1K last 30d - active
Likes
1
Model age
7w ago
created 2026-08-20
Downloads over time
Now2.5K→from101↑2,391%
09191.8K2.8K101 on Aug 192.5K on Oct 11AugSepOct
Aug 19 → Oct 11 · 48 snapshots · spans 53 days

Metadata

License
apache-2.0
Tags
transformers safetensors qwen3_5 image-text-to-text vllm intel-xpu int4 autoround gptq mtp speculative-decoding qwen
Total size
17.7 GB
Files
18
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-08-20 01:19

Files by quantization

Auxiliary files 18 files 17.7 GB
model-00003-of-00007.safetensors 3.00 GB 48c480ac download
model-00004-of-00007.safetensors 3.00 GB c4c935c8 download
model-00001-of-00007.safetensors 3.00 GB 90308620 download
model-00002-of-00007.safetensors 2.97 GB b0d50d3b download
model-00007-of-00007.safetensors 2.65 GB f95c8268 download
model-00006-of-00007.safetensors 2.37 GB 55a14ee7 download
model-00005-of-00007.safetensors 734 MB f0cd6396 download
tokenizer.json 19.1 MB 06b95093 download
model.safetensors.index.json 188 KB 9482719b download
config.json 15.4 KB f5bbc68c download
quantization_config.json 10.9 KB a6d67570 download
chat_template.jinja 8.74 KB c0c686f9 download
README.md 4.56 KB 8defa5ff download
.gitattributes 1.53 KB 52373fe2 download
processor_config.json 1.16 KB 33818c7f download
tokenizer_config.json 1.14 KB 1d134cd2 download
preprocessor_config.json 443 B 8ed39680 download
generation_config.json 214 B 8b9f95da download

README current version from Hugging Face


license: apache-2.0
library_name: transformers
tags:

  • vllm
  • intel-xpu
  • int4
  • autoround
  • gptq
  • mtp
  • speculative-decoding
  • qwen
  • uncensored

Qwen3.8-27B-Uncensored INT4 AutoRound (MTP-preserved)

An INT4 (W4A16) quantization of an uncensored/abliterated 27B dense Qwen3.8
variant, built for fast local serving on Intel Arc Pro B70 GPUs with vLLM XPU.

Two deliberate choices distinguish this checkpoint from a stock INT4 quant:

  1. The MTP (multi-token prediction) head is preserved and quantized, not
    stripped. Speculative decoding works out of the box with vLLM's
    qwen3_5_mtp method — no separate draft model or weights surgery needed.
  2. The linear-attention in_proj_a / in_proj_b projections are kept in
    FP16.
    Quantizing those two projection families to INT4 g128 measurably
    degrades exact-answer behavior (e.g. deterministic code-evaluation prompts
    return wrong values). Keeping them FP16 costs a few percent of decode
    bandwidth and preserves target-model quality. Verify any other INT4 build
    of this model against a deterministic canary before trusting it.

Quantization

Setting Value
Method AutoRound 0.14.2, default tuning recipe
Format W4A16, symmetric, group_size 128, auto_round:auto_gptq packing
Calibration batch_size 4, gradient_accumulate_steps 2
Quantized all decoder layers + the MTP head
Kept FP16/BF16 GDN linear_attn.in_proj_a/b, norms, mtp.fc, vision tower
Size on disk ~18 GB (7 shards, MTP tensors included in shard 7)

The checkpoint is directly servable as GPTQ INT4 (Marlin /
XPUwNa16LinearKernel path in vLLM).

Measured performance (Intel Arc Pro B70)

Cold-response measurement discipline: fixed realistic prompt suite, each
prompt run once, prefix/prompt caching disabled, cached_tokens=0 on every
request; primary metric is the median inter-token rate over tokens 1-100
after TTFT.

Serving config Decode (tok/s, median)
2x B70 (TP2), MTP3, FP16 KV 81.8
2x B70 (TP2), MTP3, FP16 KV, experimental INT8 LM head 94.6
1x B70, MTP3, FP16 KV 57.5
1x B70, MTP3, FP16 KV, experimental INT8 LM head 66.8

Quality gates on this exact checkpoint: arithmetic, deterministic code
evaluation, copy, factual, JSON-schema, long-context needle, and
repeat-stability checks all pass; all exact canaries are identical to the
FP16-head reference.

Notes on interpretation: warmed, repeated-prompt, same-shape benchmarks of
this stack read much higher (80+ on a single card); those numbers describe
repeat traffic, not fresh requests.

Serving (vLLM XPU)

vllm serve /path/to/this/model \
  --quantization gptq \
  --dtype float16 \
  --tensor-parallel-size 2 \
  --max-model-len 8192 \
  --gpu-memory-utilization 0.9 \
  --enable-prompt-tokens-details \
  --speculative-config '{"method":"qwen3_5_mtp","num_speculative_tokens":3}'
  • Single-GPU serving needs no tensor parallelism flags; the model fits in one
    32 GB B70 with FP16 KV at moderate context.
  • FP8 KV cache (--kv-cache-dtype fp8) works and is the capacity option for
    128K context on one card; FP16 KV decodes slightly faster.
  • The MTP head accepts num_speculative_tokens 1-4. Depth 3 was the best
    measured trade-off on fresh-response workloads; depth 4 can win on highly
    predictable content (acceptance is workload-dependent — check
    vllm:spec_decode_num_accepted_tokens_total against
    spec_decode_num_draft_tokens_total on your own traffic).
  • Optional experimental lever: an INT8-quantized LM head (W8A8) served via a
    patched vllm-xpu-kernels build adds ~15% decode on this model class; the
    stock FP16 head path is what the 81.8/57.5 rows above use.

Chat template

Ships the stock Qwen3.8 chat template (chat_template.jinja; reasoning and
tool-call blocks follow the base model's conventions). Note for serving:
vLLM's thinking_token_budget (or an equivalent hard cap) is the reliable
way to bound reasoning length; prompt-level effort instructions are a hint
the model can ignore on hard prompts.

Limitations

  • This is an uncensored model: it will answer prompts that safety-tuned
    models refuse. You are responsible for how you use it; review generated
    content before relying on it.
  • INT4 quantization trades some quality for speed and size. The preserved
    FP16 projection families mitigate the worst regression we measured, but
    outputs are not bit-identical to the BF16 original, and stylistic casing
    differences on short answers have been observed.
  • Benchmarks above are single-stream; concurrent serving aggregates higher.

README history 1 version

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-08-20Upload folder using huggingface_hub74f84ec4.6 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration