← back to catalog · registered 2026-10-10 18:58

sjoe1244/Qwen3.8-27B-Uncensored-Heretic-exl3-4.00bpw-h4-vb4

sjoe1244 27B multimodal second-order
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/sjoe1244%2FQwen3.8-27B-Uncensored-Heretic-exl3-4.00bpw-h4-vb4"
Response includes
  • classification m3
  • files 14
  • author_summary 9 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M3
Primary method

Layer-wise ablation

Applied on top of direct removal inherited from the base model.
Confidence
HIGH
Inherited from base model
Why this label 2 signals
Producer identity confirmed by naming conventions, tags or the model card. This label is very unlikely to change.
  • 'heretic' in model name (Heretic-produced)
  • Heretic uses layer-wise optimization (M3) with underlying direction removal (M1)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
HIGH
Why we say so
name contains 'heretic'; Heretic default extraction is difference-of-means (Arditi 2024)
Downloads · 30-day
0
Likes
0
Model age
today
created 2026-10-10

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
apache-2.0
Tags
exllamav3 safetensors qwen3_5 exl3 tabbyapi heretic uncensored abliterated mtp vision conversational quantized

Related

Total size
15.0 GB
Files
14
Quantizations
1
Registered
2026-10-10 18:58
Last updated on HF
2026-10-10 18:48

Files by quantization

Auxiliary files 14 files 15.0 GB
model-00001-of-00002.safetensors 7.89 GB ef243b6a download
model-00002-of-00002.safetensors 7.08 GB eb397386 download
tokenizer.json 19.1 MB 06b95093 download
quantization_config.json 614 KB 82d6feea download
model.safetensors.index.json 289 KB 4e243edf download
chat_template.jinja 28.2 KB 0478ac3f download
README.md 4.94 KB f06ad3ea download
config.json 4.59 KB 4790fdd7 download
.gitattributes 1.53 KB 52373fe2 download
processor_config.json 1.25 KB 47f77d30 download
tokenizer_config.json 1.17 KB 7b0f9106 download
preprocessor_config.json 390 B 2ea84a43 download
video_preprocessor_config.json 385 B 3ba673a5 download
generation_config.json 227 B a5baa6a6 download

README current version from Hugging Face


license: apache-2.0
license_link: https://huggingface.co/Qwen/Qwen3.8-27B/blob/main/LICENSE
base_model: llmfan46/Qwen3.8-27B-Uncensored-Heretic-Native-MTP-Preserved
base_model_relation: quantized
quantized_by: sjoe1244
pipeline_tag: image-text-to-text
library_name: exllamav3
tags:

  • exl3
  • exllamav3
  • tabbyapi
  • qwen3_5
  • heretic
  • uncensored
  • abliterated
  • mtp
  • vision
  • conversational
  • quantized
    inference: false

Qwen3.8 27B Uncensored Heretic · EXL3 4.00bpw-h4-vb4

An ExLlamaV3 quant of
llmfan46/Qwen3.8-27B-Uncensored-Heretic-Native-MTP-Preserved.
It is built for the full 262K context, with vision and MTP speculative decoding, on a single 24 GB card.
Weights, head, vision tower and MTP head are all 4-bit.

It runs in TabbyAPI or ExLlamaV3 only. It does not run in Transformers, vLLM or llama.cpp.

At a glance (one RTX 4090, 24 GB)

This quant
Download 16.1 GB
Context the full 262,144 tokens loads with vision and MTP on (4-bit KV cache)
VRAM after load 21.0 GB
Writing speed with MTP 155–167 tok/s without thinking, 142–144 tok/s with thinking

Quick start (TabbyAPI, 24 GB)

model:
  max_seq_len: 262144
  cache_size: 266240
  cache_mode: 4,4
  chunk_size: 1024
  max_batch_size: 2
  vision: true
  tool_format: auto
  prompt_template:

draft_model:
  draft_mode: mtp
  • Leave prompt_template empty. The model then uses its own chat template.
  • draft_mode: mtp drafts tokens with the model's own MTP head. No second model is needed.

Samplers

These are Qwen's recommended values.

Mode temperature top_p top_k
Thinking 1.0 0.95 20
No thinking 0.7 0.8 20
  • Keep min_p at 0, repetition penalty at 1.0 and DRY off.
  • Turn thinking off with chat_template_kwargs: {"enable_thinking": false}.
  • reasoning_effort defaults to xhigh, which can think for thousands of tokens. Use low or medium for everyday work.

Good to know

  • Made and tested with ExLlamaV3 1.6.0. It was not tried on older versions.
  • Only short tests were run on this quant. The longest prompt was 9K tokens. The full window loads, but it was not filled.
    For reference, turboderp's Huihui 4.00bpw quant of the same model family (6-bit head, 5-bit vision)
    peaked at 22.7 GB with a 232K-token prompt on this setup.
  • Quality against the BF16 source was not measured.
  • MTP is where the speed comes from. Without it, that Huihui quant wrote 54 tok/s on this setup.
    MTP off was not timed on this quant.
  • Vision loads and fits in the numbers above. No image test was run.
  • The "params" count in the Hugging Face sidebar is wrong for EXL3 files. This is the 27B model.
Full measurements and test setup

How it was made

Setting Value
Method EXL3, ExLlamaV3 1.6.0
Weights 4.00 bpw
Head (lm_head) 4 bits
Vision tower 4 bits
MTP head 4 bits
Embeddings BF16 (unquantized)
Codebook mul1, out_scales auto
Calibration 250 rows x 2048 cols (ExLlamaV3 default set)
Source llmfan46/Qwen3.8-27B-Uncensored-Heretic-Native-MTP-Preserved, BF16
python convert.py -i <source> -o <out> -w <work> -b 4.0 -hb 4 -vb 4 -mb 4

The source repo has no preprocessor_config.json, and ExLlamaV3 needs it for vision.
This repo adds preprocessor_config.json and video_preprocessor_config.json from Qwen/Qwen3.8-27B.
Their image settings are the same as in the source's processor_config.json.

File Bytes
model-00001-of-00002.safetensors 8,471,336,006
model-00002-of-00002.safetensors 7,605,654,924
Total 16,076,990,930 (14.97 GiB)

Test run

The setup was TabbyAPI with ExLlamaV3 1.6.0 on one RTX 4090 (24,564 MiB), using the Quick start config above.
VRAM is the whole-GPU nvidia-smi reading. Each test ran once.

Check Result
Load 35 s, 21,027 MiB after load
VRAM after all tests 21,595 MiB
Short coding tasks, checked by hidden unit tests 6 of 6 passed
Tool-call loops (read a file, fix it, write it, run the tests) 2 of 2 finished, 4 calls each, no bad arguments
Lookups in 4K and 9K-token prompts (exact function copies, exact tool arguments) 9 of 9 passed
Speed without thinking 155–167 tok/s
Speed with thinking (reasoning_effort: low) 142–144 tok/s
Speed on the 4K and 9K-token prompts 173–193 tok/s
MTP drafts accepted 81% (6,297 of 7,788 tokens)
  • The first request after loading ran at 48 tok/s.
  • One long-program test hit its 3,000-token output limit before finishing. It is not counted above.

Credit and license

Apache 2.0, the same as Qwen/Qwen3.8-27B and the source model. The uncensoring is
llmfan46's work. This repo only adds the EXL3 export.

Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration