license: apache-2.0
base_model: orcarouter/Qwen3.8-27B-Uncensored
base_model_relation: quantized
tags:
- exl3
- exllamav3
- qwen3.8
- abliterated
- uncensored
- mtp
- quantized
language: - en
- fr
- zh
pipeline_tag: text-generation
Qwen3.8-27B-Uncensored - EXL3 5.0 bpw
EXL3 quantization of
orcarouter/Qwen3.8-27B-Uncensored,
an abliterated (refusal-removed) build of Qwen/Qwen3.8-27B.
Built with ExLlamaV3 1.4.6.
20 GB on disk. The middle step between 4.0 bpw (16 GB) and 6.0 bpw
(22 GB), aimed at 24-28 GB cards.
The MTP head is preserved at 8 bpw and works. GGUF conversions of this
family drop the mtp.* tensors, so they have no self-drafter. TabbyAPI logsUsing main model MTP component for drafting on load.
Sibling repos
| Variant | Size | Repo |
|---|---|---|
| 4.0 bpw | 16 GB | Qwen3.8-27B-Uncensored-exl3-4bpw |
| 5.0 bpw | 20 GB | this one |
| 6.0 bpw | 22 GB | Qwen3.8-27B-Uncensored-exl3-6bpw |
| 8.0 bpw | 28 GB | Qwen3.8-27B-Uncensored-exl3-8bpw |
Build
| Language model | 5.00 bpw |
lm_head |
8 bpw |
| MTP layers | 8 bpw |
| Vision tower | unquantized (16-bit) |
| Embeddings | unquantized (16-bit) |
The lm_head is at 8 bpw here (6 bpw on the 4.0 bpw build). It costs about
0.3 GB and removes any doubt on the layer that picks every token.
python convert.py \
-i Qwen3.8-27B-Uncensored \
-o Qwen3.8-27B-Uncensored-exl3-5bpw \
-w /tmp/exl3-work \
-b 5.0 -hb 8 -mb 8 -vb 16 \
-d 0,1,2,3 -v
Default bundled calibration corpus (wiki 50, C4 20, code 20, random tokens
20, technical 10, multilingual 10, tiny 5), 250 rows x 2048 columns.
Usage - TabbyAPI
model:
model_dir: /path/to/models
model_name: Qwen3.8-27B-Uncensored-exl3-5bpw
cache_size: 32768
cache_mode: "8,8"
tensor_parallel: true
draft_model:
draft_mode: mtp
Benchmarks
Not run at 5.0 bpw. On the 4 / 6 / 8 bpw siblings (4x RTX 4000 Ada, PCIe,
no NVLink, tensor-parallel), going from 4 to 8 bpw cost only 6% latency, so
expect 5.0 bpw to land between the 4.0 and 6.0 figures. Full tables on the
4.0 bpw card.
Quality
Not benchmarked at 5.0 bpw. For reference on the same abliterated weights,
the FP8 build scores 88.0% on MMLU-Pro (business subset, 100 questions, 2
unparsed) vs 89.0% for official Qwen3.8-27B-FP8, within noise. How 5.0 bpw
compares has not been verified.
Safety
Inherited from the base model: safety alignment has been substantially
removed via abliteration. It will comply with requests the original
Qwen3.8-27B refuses, and has no meaningful built-in guardrails. Released for
research and controlled experimentation. Add your own moderation layer
before any deployment. You are responsible for what you do with it.
License
Apache 2.0, inherited from Qwen/Qwen3.8-27B.
Credits
- Qwen: base model
- orcarouter: abliterated build
- turboderp: ExLlamaV3 / EXL3
- theroyallab: TabbyAPI