license: apache-2.0
license_link: LICENSE
base_model: orcarouter/Qwen3.8-27B-Uncensored
base_model_relation: quantized
pipeline_tag: image-text-to-text
library_name: transformers
language:
- en
- zh
tags: - qwen3.8
- abliterated
- uncensored
- w4a16
- 4-bit
- int8
- compressed-tensors
- autoround
- vllm
- marlin
- mtp
- speculative-decoding
- reasoning
- function-calling
- vision-language
- rtx-3090
Orca-Qwen3.8-27B-Uncensored-W4A16
A 4-bit (W4A16) quantization of
orcarouter/Qwen3.8-27B-Uncensored,
the abliterated (refusal-removed) BF16 build of Qwen/Qwen3.8-27B.
It serves on one 24 GB GPU (RTX 3090 class) with vLLM and the
HyperQwen stack, with the vision tower, the MTP
speculative-decoding head and tool calling kept.
A faster variant with int4 heads is
Orca-Qwen3.8-27B-Uncensored-W4A16-fast.
Lineage
Qwen/Qwen3.8-27B Apache 2.0
└─ orcarouter/Qwen3.8-27B-Uncensored BF16, refusal direction removed (abliteration)
└─ this model W4A16 g128 body, int8 lm_head / embed_tokens / MTP
└─ ...-W4A16-fast int4 GPTQ lm_head / MTP, own-output draft vocabulary
What is in this checkpoint
| Part | Format |
|---|---|
| Decoder linear layers (64 layers) | int4, group 128, symmetric (Intel AutoRound), compressed-tensors pack-quantized, Marlin kernels on Ampere |
lm_head |
int8, group 128, symmetric |
embed_tokens |
int8, group 128, symmetric |
| MTP module (8 linear layers) | int8, group 128, symmetric |
MTP draft head (mtp.draft_lm_head) |
40,960 rows sliced from the int8 lm_head; ids in mtp_draft_vocab_ids.pt |
Vision tower (333 tensors), GatedDeltaNet in_proj_a / in_proj_b, norms |
BF16, unchanged |
2,022 tensors, 67 safetensors files, 16.7 GB on disk. No group has zero points.
The draft vocabulary is the generic 40,960-id list that the officialQwen3.8-27B-W4A16-AutoRound checkpoint also uses. On this model's own outputs it
covers 98.39 % of held-out tokens; a list counted over the model's own outputs covers
97.90 % (see the fast variant for how that was measured).
How it was made
run_quant.shfrom HyperQwen: Intel AutoRound,--scheme W4A16 --bits 4 --group_size 128, 128 calibration samples x 2,048 tokens,lm_headnot quantized,auto_round:llm_compressorexport.prepare/quant_lm_head.py,prepare/quant_embed.py,prepare/quant_mtp.py: int8
round-to-nearest heads. Relative error against the BF16 source weights
(||dequant - w|| / ||w||):lm_head0.0064, MTP linears 0.0066–0.0153.prepare/build_draft_vocab.py --ids prepare/draft_vocab_ids.json: the draft head.
The abliteration itself is a property of the BF16 source and is carried through
unchanged. orcarouter reports (on their own block-FP8 build, rule-based classifier):
AdvBench refusal 0.0 % vs 99.0 % for the base model, MMLU 84.7 % vs 84.3 %, GSM8K CoT
88.7 % vs 90.0 %. See the source card
for the method and the full tables.
Measured results (single RTX 3090, 24 GB)
All numbers use the HyperQwen stack on vLLM 0.28.0, WSL2 + Docker, thinking off
(protocol v2). The official Qwen3.8-27B-W4A16-AutoRound checkpoint is the reference.
Single-user decode (bench/run_benchmarks.sh single), 2026-09-23, card at 250 W,SPEC=dflash2 (DFlash2 drafter, 7 draft tokens), PREFIX_CACHE=1, KV_MEM=4529848320
(4.22 GiB), MAX_LEN=49152. End-to-end throughput in tok/s, sampling T=default / T=0.
One run per checkpoint after a warm-up request.
| Concurrency | This model | Fast variant | Official AutoRound base |
|---|---|---|---|
| C1 | 116.1 / 121.1 | 124.9 / 127.8 | 116.6 / 118.9 |
| C2 | 172.6 / 189.2 | 180.3 / 191.2 | 175.7 / 169.6 |
| C4 | 224.7 / 245.5 | 238.1 / 260.3 | 229.4 / 237.6 |
| C8 | 209.8 / 256.7 | 224.7 / 233.5 | 223.6 / 207.4 |
C1 accepts 3.71 / 3.97 tokens per verify step, mean time to first token 186 ms. Expect
3–5 % variation between sessions.
Batch serving (bench/run_benchmarks.sh batch --prefill --long), 2026-09-21, card at
350 W, KV=fp8, no speculation, second of two runs.
| Row | This model | Fast variant |
|---|---|---|
| 64 concurrent, 128 in / 512 out | 947.2 tok/s | 1,044.8 tok/s |
| 64 concurrent, 256 in / 256 out | 700.4 tok/s | 752.4 tok/s |
| Prefill 1,024 tokens | 1,804 tok/s | 1,818 tok/s |
| Prefill 102,400 tokens | 1,063 tok/s | 1,064 tok/s |
| 1 x 100k prompt: TTFT / TPOT | 93.0 s / 25.4 ms | 92.9 s / 24.6 ms |
| 4 x 60k prompts, 1,024 out | 17.1 tok/s | 17.4 tok/s |
Quality (bench/quality_battery.py against the served model): perplexity over
~300-token windows of wikitext-2 (en), fineweb-2 Danish (da) and Python source (code);
GSM8K exact match, first 200 test questions, greedy, thinking off.
| Checkpoint | PPL all (en / da / code) | GSM8K | Mean answer tokens |
|---|---|---|---|
| This model | 8.216 (10.80 / 10.90 / 3.27) | 94.5 % (95.5 % in an earlier run) | 384 |
| Fast variant | 8.260 (10.83 / 10.98 / 3.29) | 96.0 % (96.5 %) | 384 |
| Official AutoRound base | 8.186 (10.68 / 10.85 / 3.30) | 94.5 % | 379 |
With n=200 the GSM8K standard error is about 1.6 points.
How to serve
HyperQwen (Linux or WSL2, one 24 GB GPU), in .env:
MODEL=/app/models/Orca-Qwen3.8-27B-Uncensored-W4A16
SPEC=dflash2 # or SPEC=mtp to draft with the built-in MTP head
PREFIX_CACHE=1
then docker compose --profile single up -d. bash verify.sh --no-server checks the
directory before you serve it.
Plain vLLM (0.28 or later) loads the body and heads through compressed-tensors:
vllm serve TyroneNel/Orca-Qwen3.8-27B-Uncensored-W4A16 \
--max-model-len 32768 --reasoning-parser qwen3 \
--enable-auto-tool-choice --tool-call-parser qwen3_coder
The truncated MTP draft head (mtp.draft_lm_head.* with mtp_draft_vocab_ids.pt) is
read by HyperQwen's patches/qwen3_5-mtp-draft-vocab.patch. It is not tested on
unpatched vLLM.
Sampling defaults (generation_config.json): temperature 1.0, top_p 0.95, top_k 20.
The chat template accepts reasoning_effort low, medium and xhigh (default), and
maps minimal to low and high / max to xhigh.
Files
| File | Content |
|---|---|
model-000NN-of-00066.safetensors, model_extra_tensors.safetensors |
weights (extras: MTP module, draft head; lm_head is in file 66) |
model.safetensors.index.json |
tensor-to-file map |
config.json, quantization_config.json |
architecture and quantization groups |
mtp_draft_vocab_ids.pt, draft_vocab_ids.json |
the draft head's 40,960 token ids (same list, two formats) |
chat_template.jinja |
Qwen3.8 template with reasoning-effort aliases and string tool arguments |
tokenizer.json |
the Qwen3.8 tokenizer, byte-identical to Qwen/Qwen3.8-27B |
tokenizer_config.json, generation_config.json, preprocessor_config.json, processor_config.json |
tokenizer, sampling and vision configuration |
LICENSE |
Apache License 2.0 |
Disclaimer
This model has its safety alignment removed. It will follow harmful, unethical or
illegal requests that the original model refuses. It is for research: refusal
mechanisms, interpretability, red-teaming and robustness evaluation. Do not put it in
front of end users without your own moderation layer. You are responsible for how you
use it and for what it generates. Its outputs do not represent the views of the
uploader, OrcaRouter, or Qwen / Alibaba.
License
Apache License 2.0, inherited from Qwen/Qwen3.8-27B
through the OrcaRouter release. Changes made in this work: the weights are quantized as
described above, the chat template is extended, and this README is replaced.