← back to catalog · registered 2026-10-05 09:58

TyroneNel/Orca-Qwen3.8-27B-Uncensored-W4A16

TyroneNel 27B multimodal second-order
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/TyroneNel%2FOrca-Qwen3.8-27B-Uncensored-W4A16"
Response includes
  • classification m-uncensored
  • files 81
  • author_summary 6 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M-U
Primary method

Uncensored (method unknown)

No other method signals detected in this model.
Confidence
LOW
Why this label 3 signals
Weak or ambiguous signals. Best guess based on catalog patterns; treat as tentative and check the evidence below.
  • 'uncensored' in name/tags but no 'abliterated' marker
  • method not identifiable from author declaration alone
  • may be DPO fine-tune, prompt engineering, or unknown technique
Refusal direction extraction

No specific extraction method could be identified for this model. The producer either did not document it or used a proprietary pipeline.

What is a refusal direction? →
Downloads · 30-day
362
Likes
1
Model age
2w ago
created 2026-09-16

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
apache-2.0
Languages
en zh
Tags
transformers safetensors qwen3_5 image-text-to-text qwen3.8 abliterated uncensored w4a16 4-bit int8 compressed-tensors autoround

Related

Total size
15.6 GB
Files
81
Quantizations
1
Registered
2026-10-05 09:58
Last updated on HF
2026-10-05 09:46

Files by quantization

Auxiliary files 81 files 15.6 GB
model-00065-of-00066.safetensors 2.06 GB 875e6c83 download
model-00066-of-00066.safetensors 1.20 GB ce55d80e download
model_extra_tensors.safetensors 615 MB bbc9de9c download
model-00011-of-00066.safetensors 189 MB e6601b31 download
model-00013-of-00066.safetensors 189 MB 83fecdaf download
model-00014-of-00066.safetensors 189 MB 63724985 download
model-00015-of-00066.safetensors 189 MB 87cec281 download
model-00017-of-00066.safetensors 189 MB 550cbc8c download
model-00018-of-00066.safetensors 189 MB c8d92c46 download
model-00019-of-00066.safetensors 189 MB a9e50b24 download
model-00021-of-00066.safetensors 189 MB 5b73f2b6 download
model-00022-of-00066.safetensors 189 MB 511e1744 download
model-00023-of-00066.safetensors 189 MB 5218af5c download
model-00025-of-00066.safetensors 189 MB d61f519b download
model-00026-of-00066.safetensors 189 MB 548b5db8 download
model-00027-of-00066.safetensors 189 MB ea6d118d download
model-00029-of-00066.safetensors 189 MB d972bdbb download
model-00030-of-00066.safetensors 189 MB 364cbdcf download
model-00031-of-00066.safetensors 189 MB 41315de0 download
model-00033-of-00066.safetensors 189 MB 6167ab8e download
model-00034-of-00066.safetensors 189 MB 751ad704 download
model-00035-of-00066.safetensors 189 MB af48c229 download
model-00037-of-00066.safetensors 189 MB 7fda5354 download
model-00038-of-00066.safetensors 189 MB 3277c7b4 download
model-00039-of-00066.safetensors 189 MB 84f1b550 download
model-00041-of-00066.safetensors 189 MB 36c4d4b3 download
model-00042-of-00066.safetensors 189 MB 1e64b4dc download
model-00043-of-00066.safetensors 189 MB d4b958e6 download
model-00045-of-00066.safetensors 189 MB c76e37f9 download
model-00046-of-00066.safetensors 189 MB 76210a97 download
model-00047-of-00066.safetensors 189 MB ac69fced download
model-00049-of-00066.safetensors 189 MB ae8b68b2 download
model-00050-of-00066.safetensors 189 MB b326298b download
model-00051-of-00066.safetensors 189 MB 3bccaf2a download
model-00053-of-00066.safetensors 189 MB 0516bb5a download
model-00054-of-00066.safetensors 189 MB 3c043bf0 download
model-00055-of-00066.safetensors 189 MB daf797df download
model-00057-of-00066.safetensors 189 MB 7353d50f download
model-00058-of-00066.safetensors 189 MB eeb23bdc download
model-00059-of-00066.safetensors 189 MB d9c25e1a download
model-00061-of-00066.safetensors 189 MB 5cc9115f download
model-00062-of-00066.safetensors 189 MB b47df4bf download
model-00063-of-00066.safetensors 189 MB 7d44b9a9 download
model-00001-of-00066.safetensors 189 MB a3503e38 download
model-00002-of-00066.safetensors 189 MB dc277f26 download
model-00003-of-00066.safetensors 189 MB a7a337ea download
model-00005-of-00066.safetensors 189 MB 25daec15 download
model-00006-of-00066.safetensors 189 MB f56456c0 download
model-00007-of-00066.safetensors 189 MB 4341e9a9 download
model-00009-of-00066.safetensors 189 MB d3282fbb download
model-00010-of-00066.safetensors 189 MB 053bc0a5 download
model-00012-of-00066.safetensors 183 MB 3f7ff903 download
model-00016-of-00066.safetensors 183 MB 2d003fa2 download
model-00020-of-00066.safetensors 183 MB 7c0ce3a0 download
model-00024-of-00066.safetensors 183 MB d60704f7 download
model-00028-of-00066.safetensors 183 MB 752deb8e download
model-00032-of-00066.safetensors 183 MB 89651dcf download
model-00036-of-00066.safetensors 183 MB 0ae54950 download
model-00040-of-00066.safetensors 183 MB 129c01ed download
model-00044-of-00066.safetensors 183 MB 04ee6b0e download
model-00048-of-00066.safetensors 183 MB 99dc0fb9 download
model-00052-of-00066.safetensors 183 MB 0a7aa150 download
model-00056-of-00066.safetensors 183 MB 2af29c35 download
model-00060-of-00066.safetensors 183 MB 22a82162 download
model-00064-of-00066.safetensors 183 MB 8254b303 download
model-00004-of-00066.safetensors 183 MB d059333d download
model-00008-of-00066.safetensors 183 MB d468f9d7 download
mtp_draft_vocab_ids.pt 322 KB e2024ef6 download
tokenizer.json 12.2 MB 0997f410 download
draft_vocab_ids.json 273 KB d0069e09 download
model.safetensors.index.json 188 KB 99cc655a download
config.json 16.6 KB 8f65047d download
quantization_config.json 12.7 KB 73745591 download
LICENSE 11.3 KB f938136e download
chat_template.jinja 9.42 KB 9377e318 download
README.md 7.85 KB af443379 download
.gitattributes 1.53 KB 52373fe2 download
tokenizer_config.json 1.17 KB 12a6c87d download
processor_config.json 1.16 KB 33818c7f download
preprocessor_config.json 390 B 2ea84a43 download
generation_config.json 202 B 023756cf download

README current version from Hugging Face


license: apache-2.0
license_link: LICENSE
base_model: orcarouter/Qwen3.8-27B-Uncensored
base_model_relation: quantized
pipeline_tag: image-text-to-text
library_name: transformers
language:

  • en
  • zh
    tags:
  • qwen3.8
  • abliterated
  • uncensored
  • w4a16
  • 4-bit
  • int8
  • compressed-tensors
  • autoround
  • vllm
  • marlin
  • mtp
  • speculative-decoding
  • reasoning
  • function-calling
  • vision-language
  • rtx-3090

Orca-Qwen3.8-27B-Uncensored-W4A16

A 4-bit (W4A16) quantization of
orcarouter/Qwen3.8-27B-Uncensored,
the abliterated (refusal-removed) BF16 build of Qwen/Qwen3.8-27B.
It serves on one 24 GB GPU (RTX 3090 class) with vLLM and the
HyperQwen stack, with the vision tower, the MTP
speculative-decoding head and tool calling kept.

A faster variant with int4 heads is
Orca-Qwen3.8-27B-Uncensored-W4A16-fast.

Lineage

Qwen/Qwen3.8-27B                         Apache 2.0
 └─ orcarouter/Qwen3.8-27B-Uncensored    BF16, refusal direction removed (abliteration)
     └─ this model                       W4A16 g128 body, int8 lm_head / embed_tokens / MTP
         └─ ...-W4A16-fast               int4 GPTQ lm_head / MTP, own-output draft vocabulary

What is in this checkpoint

Part Format
Decoder linear layers (64 layers) int4, group 128, symmetric (Intel AutoRound), compressed-tensors pack-quantized, Marlin kernels on Ampere
lm_head int8, group 128, symmetric
embed_tokens int8, group 128, symmetric
MTP module (8 linear layers) int8, group 128, symmetric
MTP draft head (mtp.draft_lm_head) 40,960 rows sliced from the int8 lm_head; ids in mtp_draft_vocab_ids.pt
Vision tower (333 tensors), GatedDeltaNet in_proj_a / in_proj_b, norms BF16, unchanged

2,022 tensors, 67 safetensors files, 16.7 GB on disk. No group has zero points.

The draft vocabulary is the generic 40,960-id list that the official
Qwen3.8-27B-W4A16-AutoRound checkpoint also uses. On this model's own outputs it
covers 98.39 % of held-out tokens; a list counted over the model's own outputs covers
97.90 % (see the fast variant for how that was measured).

How it was made

  1. run_quant.sh from HyperQwen: Intel AutoRound, --scheme W4A16 --bits 4 --group_size 128, 128 calibration samples x 2,048 tokens, lm_head not quantized,
    auto_round:llm_compressor export.
  2. prepare/quant_lm_head.py, prepare/quant_embed.py, prepare/quant_mtp.py: int8
    round-to-nearest heads. Relative error against the BF16 source weights
    (||dequant - w|| / ||w||): lm_head 0.0064, MTP linears 0.0066–0.0153.
  3. prepare/build_draft_vocab.py --ids prepare/draft_vocab_ids.json: the draft head.

The abliteration itself is a property of the BF16 source and is carried through
unchanged. orcarouter reports (on their own block-FP8 build, rule-based classifier):
AdvBench refusal 0.0 % vs 99.0 % for the base model, MMLU 84.7 % vs 84.3 %, GSM8K CoT
88.7 % vs 90.0 %. See the source card
for the method and the full tables.

Measured results (single RTX 3090, 24 GB)

All numbers use the HyperQwen stack on vLLM 0.28.0, WSL2 + Docker, thinking off
(protocol v2). The official Qwen3.8-27B-W4A16-AutoRound checkpoint is the reference.

Single-user decode (bench/run_benchmarks.sh single), 2026-09-23, card at 250 W,
SPEC=dflash2 (DFlash2 drafter, 7 draft tokens), PREFIX_CACHE=1, KV_MEM=4529848320
(4.22 GiB), MAX_LEN=49152. End-to-end throughput in tok/s, sampling T=default / T=0.
One run per checkpoint after a warm-up request.

Concurrency This model Fast variant Official AutoRound base
C1 116.1 / 121.1 124.9 / 127.8 116.6 / 118.9
C2 172.6 / 189.2 180.3 / 191.2 175.7 / 169.6
C4 224.7 / 245.5 238.1 / 260.3 229.4 / 237.6
C8 209.8 / 256.7 224.7 / 233.5 223.6 / 207.4

C1 accepts 3.71 / 3.97 tokens per verify step, mean time to first token 186 ms. Expect
3–5 % variation between sessions.

Batch serving (bench/run_benchmarks.sh batch --prefill --long), 2026-09-21, card at
350 W, KV=fp8, no speculation, second of two runs.

Row This model Fast variant
64 concurrent, 128 in / 512 out 947.2 tok/s 1,044.8 tok/s
64 concurrent, 256 in / 256 out 700.4 tok/s 752.4 tok/s
Prefill 1,024 tokens 1,804 tok/s 1,818 tok/s
Prefill 102,400 tokens 1,063 tok/s 1,064 tok/s
1 x 100k prompt: TTFT / TPOT 93.0 s / 25.4 ms 92.9 s / 24.6 ms
4 x 60k prompts, 1,024 out 17.1 tok/s 17.4 tok/s

Quality (bench/quality_battery.py against the served model): perplexity over
~300-token windows of wikitext-2 (en), fineweb-2 Danish (da) and Python source (code);
GSM8K exact match, first 200 test questions, greedy, thinking off.

Checkpoint PPL all (en / da / code) GSM8K Mean answer tokens
This model 8.216 (10.80 / 10.90 / 3.27) 94.5 % (95.5 % in an earlier run) 384
Fast variant 8.260 (10.83 / 10.98 / 3.29) 96.0 % (96.5 %) 384
Official AutoRound base 8.186 (10.68 / 10.85 / 3.30) 94.5 % 379

With n=200 the GSM8K standard error is about 1.6 points.

How to serve

HyperQwen (Linux or WSL2, one 24 GB GPU), in .env:

MODEL=/app/models/Orca-Qwen3.8-27B-Uncensored-W4A16
SPEC=dflash2          # or SPEC=mtp to draft with the built-in MTP head
PREFIX_CACHE=1

then docker compose --profile single up -d. bash verify.sh --no-server checks the
directory before you serve it.

Plain vLLM (0.28 or later) loads the body and heads through compressed-tensors:

vllm serve TyroneNel/Orca-Qwen3.8-27B-Uncensored-W4A16 \
  --max-model-len 32768 --reasoning-parser qwen3 \
  --enable-auto-tool-choice --tool-call-parser qwen3_coder

The truncated MTP draft head (mtp.draft_lm_head.* with mtp_draft_vocab_ids.pt) is
read by HyperQwen's patches/qwen3_5-mtp-draft-vocab.patch. It is not tested on
unpatched vLLM.

Sampling defaults (generation_config.json): temperature 1.0, top_p 0.95, top_k 20.
The chat template accepts reasoning_effort low, medium and xhigh (default), and
maps minimal to low and high / max to xhigh.

Files

File Content
model-000NN-of-00066.safetensors, model_extra_tensors.safetensors weights (extras: MTP module, draft head; lm_head is in file 66)
model.safetensors.index.json tensor-to-file map
config.json, quantization_config.json architecture and quantization groups
mtp_draft_vocab_ids.pt, draft_vocab_ids.json the draft head's 40,960 token ids (same list, two formats)
chat_template.jinja Qwen3.8 template with reasoning-effort aliases and string tool arguments
tokenizer.json the Qwen3.8 tokenizer, byte-identical to Qwen/Qwen3.8-27B
tokenizer_config.json, generation_config.json, preprocessor_config.json, processor_config.json tokenizer, sampling and vision configuration
LICENSE Apache License 2.0

Disclaimer

This model has its safety alignment removed. It will follow harmful, unethical or
illegal requests that the original model refuses. It is for research: refusal
mechanisms, interpretability, red-teaming and robustness evaluation. Do not put it in
front of end users without your own moderation layer. You are responsible for how you
use it and for what it generates. Its outputs do not represent the views of the
uploader, OrcaRouter, or Qwen / Alibaba.

License

Apache License 2.0, inherited from Qwen/Qwen3.8-27B
through the OrcaRouter release. Changes made in this work: the weights are quantized as
described above, the chat template is extended, and this README is replaced.

Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration