license: apache-2.0
license_link: https://huggingface.co/Qwen/Qwen3.8-27B/blob/main/LICENSE
base_model: llmfan46/Qwen3.8-27B-Uncensored-Heretic-Native-MTP-Preserved
base_model_relation: quantized
quantized_by: sjoe1244
pipeline_tag: image-text-to-text
library_name: exllamav3
tags:
- exl3
- exllamav3
- tabbyapi
- qwen3_5
- heretic
- uncensored
- abliterated
- mtp
- vision
- conversational
- quantized
inference: false
Qwen3.8 27B Uncensored Heretic · EXL3 4.00bpw-h4-vb4
An ExLlamaV3 quant ofllmfan46/Qwen3.8-27B-Uncensored-Heretic-Native-MTP-Preserved.
It is built for the full 262K context, with vision and MTP speculative decoding, on a single 24 GB card.
Weights, head, vision tower and MTP head are all 4-bit.
It runs in TabbyAPI or ExLlamaV3 only. It does not run in Transformers, vLLM or llama.cpp.
At a glance (one RTX 4090, 24 GB)
| This quant | |
|---|---|
| Download | 16.1 GB |
| Context | the full 262,144 tokens loads with vision and MTP on (4-bit KV cache) |
| VRAM after load | 21.0 GB |
| Writing speed with MTP | 155–167 tok/s without thinking, 142–144 tok/s with thinking |
Quick start (TabbyAPI, 24 GB)
model:
max_seq_len: 262144
cache_size: 266240
cache_mode: 4,4
chunk_size: 1024
max_batch_size: 2
vision: true
tool_format: auto
prompt_template:
draft_model:
draft_mode: mtp
- Leave
prompt_templateempty. The model then uses its own chat template. draft_mode: mtpdrafts tokens with the model's own MTP head. No second model is needed.
Samplers
These are Qwen's recommended values.
| Mode | temperature | top_p | top_k |
|---|---|---|---|
| Thinking | 1.0 | 0.95 | 20 |
| No thinking | 0.7 | 0.8 | 20 |
- Keep
min_pat 0, repetition penalty at 1.0 and DRY off. - Turn thinking off with
chat_template_kwargs: {"enable_thinking": false}. reasoning_effortdefaults toxhigh, which can think for thousands of tokens. Uselowormediumfor everyday work.
Good to know
- Made and tested with ExLlamaV3 1.6.0. It was not tried on older versions.
- Only short tests were run on this quant. The longest prompt was 9K tokens. The full window loads, but it was not filled.
For reference, turboderp's Huihui 4.00bpw quant of the same model family (6-bit head, 5-bit vision)
peaked at 22.7 GB with a 232K-token prompt on this setup. - Quality against the BF16 source was not measured.
- MTP is where the speed comes from. Without it, that Huihui quant wrote 54 tok/s on this setup.
MTP off was not timed on this quant. - Vision loads and fits in the numbers above. No image test was run.
- The "params" count in the Hugging Face sidebar is wrong for EXL3 files. This is the 27B model.
Full measurements and test setup
How it was made
| Setting | Value |
|---|---|
| Method | EXL3, ExLlamaV3 1.6.0 |
| Weights | 4.00 bpw |
Head (lm_head) |
4 bits |
| Vision tower | 4 bits |
| MTP head | 4 bits |
| Embeddings | BF16 (unquantized) |
| Codebook | mul1, out_scales auto |
| Calibration | 250 rows x 2048 cols (ExLlamaV3 default set) |
| Source | llmfan46/Qwen3.8-27B-Uncensored-Heretic-Native-MTP-Preserved, BF16 |
python convert.py -i <source> -o <out> -w <work> -b 4.0 -hb 4 -vb 4 -mb 4
The source repo has no preprocessor_config.json, and ExLlamaV3 needs it for vision.
This repo adds preprocessor_config.json and video_preprocessor_config.json from Qwen/Qwen3.8-27B.
Their image settings are the same as in the source's processor_config.json.
| File | Bytes |
|---|---|
model-00001-of-00002.safetensors |
8,471,336,006 |
model-00002-of-00002.safetensors |
7,605,654,924 |
| Total | 16,076,990,930 (14.97 GiB) |
Test run
The setup was TabbyAPI with ExLlamaV3 1.6.0 on one RTX 4090 (24,564 MiB), using the Quick start config above.
VRAM is the whole-GPU nvidia-smi reading. Each test ran once.
| Check | Result |
|---|---|
| Load | 35 s, 21,027 MiB after load |
| VRAM after all tests | 21,595 MiB |
| Short coding tasks, checked by hidden unit tests | 6 of 6 passed |
| Tool-call loops (read a file, fix it, write it, run the tests) | 2 of 2 finished, 4 calls each, no bad arguments |
| Lookups in 4K and 9K-token prompts (exact function copies, exact tool arguments) | 9 of 9 passed |
| Speed without thinking | 155–167 tok/s |
Speed with thinking (reasoning_effort: low) |
142–144 tok/s |
| Speed on the 4K and 9K-token prompts | 173–193 tok/s |
| MTP drafts accepted | 81% (6,297 of 7,788 tokens) |
- The first request after loading ran at 48 tok/s.
- One long-program test hit its 3,000-token output limit before finishing. It is not counted above.
Credit and license
Apache 2.0, the same as Qwen/Qwen3.8-27B and the source model. The uncensoring isllmfan46's work. This repo only adds the EXL3 export.