← back to catalog · registered 2026-10-06 19:58

Lygodactylus/Qwen3.8-27B-Uncensored-exl3-5bpw

Lygodactylus 27B second-order
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/Lygodactylus%2FQwen3.8-27B-Uncensored-exl3-5bpw"
Response includes
  • classification m-uncensored
  • files 17
  • author_summary 6 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M-U
Primary method

Uncensored (method unknown)

No other method signals detected in this model.
Confidence
LOW
Why this label 3 signals
Weak or ambiguous signals. Best guess based on catalog patterns; treat as tentative and check the evidence below.
  • 'uncensored' in name/tags but no 'abliterated' marker
  • method not identifiable from author declaration alone
  • may be DPO fine-tune, prompt engineering, or unknown technique
Refusal direction extraction

No specific extraction method could be identified for this model. The producer either did not document it or used a proprietary pipeline.

What is a refusal direction? →
Downloads · 30-day
0
Likes
0
Model age
today
created 2026-10-06

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
apache-2.0
Languages
en fr zh
Tags
safetensors qwen3_5 exl3 exllamav3 qwen3.8 abliterated uncensored mtp quantized text-generation conversational en

Related

Total size
19.0 GB
Files
17
Quantizations
1
Registered
2026-10-06 19:58
Last updated on HF
2026-10-06 19:18

Files by quantization

Auxiliary files 17 files 19.1 GB
model-00002-of-00003.safetensors 8.00 GB 84b24ecd download
model-00001-of-00003.safetensors 7.93 GB 6123ad43 download
model-00003-of-00003.safetensors 3.10 GB 05bbd3a3 download
tokenizer.json 12.2 MB 0997f410 download
vocab.json 6.41 MB 0aa0ce06 download
merges.txt 3.20 MB a494e019 download
quantization_config.json 614 KB 6f760823 download
model.safetensors.index.json 233 KB 21e608bd download
tokenizer_config.json 17.5 KB 5de744b3 download
LICENSE 11.3 KB f938136e download
chat_template.jinja 8.74 KB c0c686f9 download
config.json 4.51 KB a037ad6d download
README.md 3.62 KB 4a2f1e26 download
.gitattributes 1.53 KB 52373fe2 download
preprocessor_config.json 390 B 2ea84a43 download
video_preprocessor_config.json 385 B 3ba673a5 download
generation_config.json 202 B 023756cf download

README current version from Hugging Face


license: apache-2.0
base_model: orcarouter/Qwen3.8-27B-Uncensored
base_model_relation: quantized
tags:

  • exl3
  • exllamav3
  • qwen3.8
  • abliterated
  • uncensored
  • mtp
  • quantized
    language:
  • en
  • fr
  • zh
    pipeline_tag: text-generation

Qwen3.8-27B-Uncensored EXL3

Qwen3.8-27B-Uncensored - EXL3 5.0 bpw

EXL3 quantization of
orcarouter/Qwen3.8-27B-Uncensored,
an abliterated (refusal-removed) build of Qwen/Qwen3.8-27B.

Built with ExLlamaV3 1.4.6.

20 GB on disk. The middle step between 4.0 bpw (16 GB) and 6.0 bpw
(22 GB), aimed at 24-28 GB cards.

The MTP head is preserved at 8 bpw and works. GGUF conversions of this
family drop the mtp.* tensors, so they have no self-drafter. TabbyAPI logs
Using main model MTP component for drafting on load.

Sibling repos

Variant Size Repo
4.0 bpw 16 GB Qwen3.8-27B-Uncensored-exl3-4bpw
5.0 bpw 20 GB this one
6.0 bpw 22 GB Qwen3.8-27B-Uncensored-exl3-6bpw
8.0 bpw 28 GB Qwen3.8-27B-Uncensored-exl3-8bpw

Build

Language model 5.00 bpw
lm_head 8 bpw
MTP layers 8 bpw
Vision tower unquantized (16-bit)
Embeddings unquantized (16-bit)

The lm_head is at 8 bpw here (6 bpw on the 4.0 bpw build). It costs about
0.3 GB and removes any doubt on the layer that picks every token.

python convert.py \
  -i Qwen3.8-27B-Uncensored \
  -o Qwen3.8-27B-Uncensored-exl3-5bpw \
  -w /tmp/exl3-work \
  -b 5.0 -hb 8 -mb 8 -vb 16 \
  -d 0,1,2,3 -v

Default bundled calibration corpus (wiki 50, C4 20, code 20, random tokens
20, technical 10, multilingual 10, tiny 5), 250 rows x 2048 columns.

Usage - TabbyAPI

model:
  model_dir: /path/to/models
  model_name: Qwen3.8-27B-Uncensored-exl3-5bpw
  cache_size: 32768
  cache_mode: "8,8"
  tensor_parallel: true

draft_model:
  draft_mode: mtp

Benchmarks

Not run at 5.0 bpw. On the 4 / 6 / 8 bpw siblings (4x RTX 4000 Ada, PCIe,
no NVLink, tensor-parallel), going from 4 to 8 bpw cost only 6% latency, so
expect 5.0 bpw to land between the 4.0 and 6.0 figures. Full tables on the
4.0 bpw card.

Quality

Not benchmarked at 5.0 bpw. For reference on the same abliterated weights,
the FP8 build scores 88.0% on MMLU-Pro (business subset, 100 questions, 2
unparsed) vs 89.0% for official Qwen3.8-27B-FP8, within noise. How 5.0 bpw
compares has not been verified.

Safety

Inherited from the base model: safety alignment has been substantially
removed via abliteration. It will comply with requests the original
Qwen3.8-27B refuses, and has no meaningful built-in guardrails. Released for
research and controlled experimentation. Add your own moderation layer
before any deployment. You are responsible for what you do with it.

License

Apache 2.0, inherited from Qwen/Qwen3.8-27B.

Credits

Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration