license: apache-2.0
license_link: LICENSE
base_model: orcarouter/Qwen3.8-27B-Uncensored
base_model_relation: quantized
pipeline_tag: image-text-to-text
library_name: transformers
language:
- en
- zh
tags: - qwen3.8
- abliterated
- uncensored
- w4a16
- 4-bit
- int4
- int8
- gptq
- compressed-tensors
- autoround
- vllm
- marlin
- mtp
- speculative-decoding
- reasoning
- function-calling
- vision-language
- rtx-3090
Orca-Qwen3.8-27B-Uncensored-W4A16-fast
The "fast variant" of
Orca-Qwen3.8-27B-Uncensored-W4A16:
the same 4-bit body of orcarouter/Qwen3.8-27B-Uncensored
(the abliterated BF16 build of Qwen/Qwen3.8-27B),
with the lm_head and the MTP module requantized to int4 GPTQ and a draft vocabulary
counted over this model's own outputs. It serves on one 24 GB GPU (RTX 3090 class) with
the HyperQwen stack. The int4 lm_head makes every
decode step read about 0.65 GB less; that is where the gain over the g128 build comes from.
Lineage
Qwen/Qwen3.8-27B Apache 2.0
└─ orcarouter/Qwen3.8-27B-Uncensored BF16, refusal direction removed (abliteration)
└─ Orca-...-W4A16 W4A16 g128 body, int8 lm_head / embed_tokens / MTP
└─ this model int4 GPTQ lm_head / MTP, own-output draft vocabulary
What changed against the g128 build
| Part | g128 | This model |
|---|---|---|
| Decoder linear layers | int4 g128 symmetric (AutoRound) | unchanged (65 of 67 files byte-identical) |
lm_head |
int8 RTN, rel. error 0.0064 | int4 GPTQ g128 symmetric, rel. error 0.1408 |
| MTP module (8 linear layers) | int8 RTN, rel. error 0.0066–0.0153 | int4 GPTQ g128 symmetric, rel. error 0.141–0.170 |
| MTP draft head | 40,960 rows of the int8 lm_head, generic id list |
40,960 rows of the int4 lm_head, id list counted over this model's own outputs |
embed_tokens, vision tower, norms |
int8 / BF16 | unchanged |
| Size on disk | 16.7 GB | 15.8 GB |
Relative error is ||dequant - w|| / ||w|| against the BF16 source weights. GPTQ uses
hidden states captured from this model's own generations, so it minimizes the output
error on real activations rather than the weight error that this number shows. All
groups are symmetric (no zero points).
The own-output draft vocabulary is the top 40,960 ids of 6,761 generations by this model
(4.85 M output tokens, 3,119 of them with thinking on; prompts: English chat, code,
Danish instructions and reasoning, GSM8K train). On the held-out 10 % of those generations it covers
97.90 % of tokens; the generic list of the g128 build covers 98.39 %. The draft head
only matters for SPEC=mtp.
How it was made
drafter/collect_prompts.py,drafter/gen_data.py,drafter/capture.py: generate
with this model and capture its final hidden states.drafter/gptq_lm_head.py --bits 4 --calib-rows 300000: GPTQ int4lm_head.drafter/train_mtp.py --eval-only 1 --dump-hessians, thendrafter/requant_mtp_gptq.py --bits 4: GPTQ int4 MTP module.prepare/build_draft_vocab.py: the draft head from the own-output id list.
The tools are in HyperQwen (drafter/, prepare/).
The abliteration is a property of the BF16 source and is carried through unchanged; see
the g128 card and the
source card.
Measured results (single RTX 3090, 24 GB)
All numbers use the HyperQwen stack on vLLM 0.28.0, WSL2 + Docker, thinking off
(protocol v2).
Single-user decode (bench/run_benchmarks.sh single), 2026-09-23, card at 250 W,SPEC=dflash2 (DFlash2 drafter, 7 draft tokens), PREFIX_CACHE=1, KV_MEM=4529848320
(4.22 GiB), MAX_LEN=49152. End-to-end throughput in tok/s, sampling T=default / T=0.
One run per checkpoint after a warm-up request.
| Concurrency | This model | g128 build | Official AutoRound base |
|---|---|---|---|
| C1 | 124.9 / 127.8 | 116.1 / 121.1 | 116.6 / 118.9 |
| C2 | 180.3 / 191.2 | 172.6 / 189.2 | 175.7 / 169.6 |
| C4 | 238.1 / 260.3 | 224.7 / 245.5 | 229.4 / 237.6 |
| C8 | 224.7 / 233.5 | 209.8 / 256.7 | 223.6 / 207.4 |
C1 accepts 3.70 / 3.93 tokens per verify step, mean time to first token 175 ms. Expect
3–5 % variation between sessions.
Batch serving (bench/run_benchmarks.sh batch --prefill --long), 2026-09-21, card at
350 W, KV=fp8, no speculation, second of two runs.
| Row | This model | g128 build |
|---|---|---|
| 64 concurrent, 128 in / 512 out | 1,044.8 tok/s | 947.2 tok/s (+10 %) |
| 64 concurrent, 256 in / 256 out | 752.4 tok/s | 700.4 tok/s (+7 %) |
| Prefill 1,024 tokens | 1,818 tok/s | 1,804 tok/s |
| Prefill 102,400 tokens | 1,064 tok/s | 1,063 tok/s |
| 1 x 100k prompt: TTFT / TPOT | 92.9 s / 24.6 ms | 93.0 s / 25.4 ms |
| 4 x 60k prompts, 1,024 out | 17.4 tok/s | 17.1 tok/s |
Quality (bench/quality_battery.py against the served model): perplexity over
~300-token windows of wikitext-2 (en), fineweb-2 Danish (da) and Python source (code);
GSM8K exact match, first 200 test questions, greedy, thinking off.
| Checkpoint | PPL all (en / da / code) | GSM8K | Mean answer tokens |
|---|---|---|---|
| This model | 8.260 (10.83 / 10.98 / 3.29) | 96.0 % (96.5 % in an earlier run) | 384 |
| g128 build | 8.216 (10.80 / 10.90 / 3.27) | 94.5 % (95.5 %) | 384 |
| Official AutoRound base | 8.186 (10.68 / 10.85 / 3.30) | 94.5 % | 379 |
The int4 heads cost 0.5 % perplexity. With n=200 the GSM8K standard error is about
1.6 points, so the GSM8K difference is not significant.
How to serve
HyperQwen (Linux or WSL2, one 24 GB GPU), in .env:
MODEL=/app/models/Orca-Qwen3.8-27B-Uncensored-W4A16-fast
SPEC=dflash2 # or SPEC=mtp to draft with the int4 MTP head
PREFIX_CACHE=1
then docker compose --profile single up -d. bash verify.sh --no-server checks the
directory before you serve it.
Plain vLLM (0.28 or later) loads the body and heads through compressed-tensors:
vllm serve TyroneNel/Orca-Qwen3.8-27B-Uncensored-W4A16-fast \
--max-model-len 32768 --reasoning-parser qwen3 \
--enable-auto-tool-choice --tool-call-parser qwen3_coder
The truncated MTP draft head (mtp.draft_lm_head.* with mtp_draft_vocab_ids.pt) is
read by HyperQwen's patches/qwen3_5-mtp-draft-vocab.patch. It is not tested on
unpatched vLLM.
Sampling defaults (generation_config.json): temperature 1.0, top_p 0.95, top_k 20.
The chat template accepts reasoning_effort low, medium and xhigh (default), and
maps minimal to low and high / max to xhigh.
Files
| File | Content |
|---|---|
model-000NN-of-00066.safetensors, model_extra_tensors.safetensors |
weights (extras: MTP module, draft head; lm_head is in file 66) |
model.safetensors.index.json |
tensor-to-file map |
config.json, quantization_config.json |
architecture and quantization groups (lm_head int4, embed_tokens int8, mtp.* int4) |
mtp_draft_vocab_ids.pt, draft_vocab_ids.json |
the draft head's 40,960 token ids (same list, two formats) |
chat_template.jinja |
Qwen3.8 template with reasoning-effort aliases and string tool arguments |
tokenizer.json |
the Qwen3.8 tokenizer, byte-identical to Qwen/Qwen3.8-27B |
tokenizer_config.json, generation_config.json, preprocessor_config.json, processor_config.json |
tokenizer, sampling and vision configuration |
LICENSE |
Apache License 2.0 |
Disclaimer
This model has its safety alignment removed. It will follow harmful, unethical or
illegal requests that the original model refuses. It is for research: refusal
mechanisms, interpretability, red-teaming and robustness evaluation. Do not put it in
front of end users without your own moderation layer. You are responsible for how you
use it and for what it generates. Its outputs do not represent the views of the
uploader, OrcaRouter, or Qwen / Alibaba. Fine-tuning should start from the BF16 source,
not from these 4-bit weights.
License
Apache License 2.0, inherited from Qwen/Qwen3.8-27B
through the OrcaRouter release. Changes made in this work: the weights are quantized as
described above, the chat template is extended, and this README is replaced.