← back to catalog · registered 2026-08-22 13:56

bowmanslayer/Qwen3.8-27B-Uncensored-GGUF

bowmanslayer Qwen 27B GGUF second-order 262K ctx
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/bowmanslayer%2FQwen3.8-27B-Uncensored-GGUF"
Response includes
  • classification m8
  • files 10
  • hub_downloads_all_time 1,776
  • author_summary 9 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M8
Primary method

Repackaging (quantization)

Applied on top of direct removal inherited from the base model.
Confidence
MEDIUM
Inherited from base model
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=1
  • assume M1 (base ablation) + M8 (GGUF quant) - default when producer unknown
Refusal direction extraction

No specific extraction method could be identified for this model. The producer either did not document it or used a proprietary pipeline.

What is a refusal direction? →
Downloads · lifetime
2K
147 last 30d - cooling
Likes
6
Model age
7w ago
created 2026-08-20
Downloads over time
Now1.8K→from481↑271%
4169161.4K1.9K481 on Aug 191.8K on Oct 111.8K on Oct 5AugSepOct
Aug 19 → Oct 11 · 48 snapshots · spans 53 days

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
apache-2.0
Quantizations
IQ4 Q3_K Q4_K Q5_K Q6_K Q8_0
Tags
gguf qwen3.5 llama.cpp imatrix quantized speculative-decoding mtp not-for-all-audiences uncensored abliterated text-generation base_model:JonathanColetti/Qwen3.8-27B-Uncensored

Related

Total size
109 GB
Files
10
Quantizations
8
Registered
2026-08-22 13:56
Last updated on HF
2026-08-21 13:47

Files by quantization

Q8_0 1 file 26.6 GB
Qwen3.8-27B-Uncensored-Text-Only-Q8_0.gguf 26.6 GB ******** download
Q6_K 1 file 20.6 GB
Qwen3.8-27B-Uncensored-Text-Only-Q6_K.gguf 20.6 GB ******** download
Q5_K 1 file 17.9 GB
Qwen3.8-27B-Uncensored-Text-Only-Q5_K_M.gguf 17.9 GB ******** download
Q4_K 2 files 17.3 GB
Qwen3.8-27B-Uncensored-Text-Only-Q4_K_M.gguf 15.4 GB ******** download
mtp-Qwen3.8-27B-Uncensored-Q4_K_M.gguf 1.89 GB ******** download
IQ4 1 file 14.0 GB
Qwen3.8-27B-Uncensored-Text-Only-IQ4_XS.gguf 14.0 GB ******** download
Q3_K 1 file 12.4 GB
Qwen3.8-27B-Uncensored-Text-Only-Q3_K_M.gguf 12.4 GB ******** download
F16 1 file 885 MB
mmproj-Qwen3.8-27B-Uncensored-f16.gguf 885 MB ******** download
Auxiliary files 2 files 15.9 KB
README.md 13.8 KB aac26161 download
.gitattributes 2.10 KB 2e821dec download

README current version from Hugging Face


license: apache-2.0
base_model: JonathanColetti/Qwen3.8-27B-Uncensored
base_model_relation: quantized
library_name: gguf
pipeline_tag: text-generation
tags:

  • qwen3.5
  • gguf
  • llama.cpp
  • imatrix
  • quantized
  • speculative-decoding
  • mtp
  • not-for-all-audiences
  • uncensored
  • abliterated
    extra_gated_prompt: >-
    This model has had its safety alignment removed. It will comply with requests
    that the original model refuses. By requesting access you confirm that you are
    of legal age in your jurisdiction, that you will not deploy it to third parties
    without your own safety layer, and that you accept sole responsibility for its
    outputs and for compliance with applicable law.
    extra_gated_fields:
    I am of legal age in my jurisdiction: checkbox
    I will not deploy this to third parties without my own safety layer: checkbox
    I accept sole responsibility for outputs and legal compliance: checkbox

Qwen3.8-27B-Uncensored · GGUF (imatrix)

Qwen3.8-27B for llama.cpp — importance-matrix quantized from the original BF16.
Runs from 16 GB VRAM up.

The main weight files carry the language model. Vision ships as a separate
mmproj file and the model's multi-token-prediction head as a separate
speculative draft — llama.cpp keeps both outside the main file by design, not
because anything was cut down. Take either, both, or neither.

Renamed 2026-08-21 from Qwen3.8-27B-Uncensored-Text-Only-GGUF. The old name described the layout of one
file but read as "this model cannot see", which was never true — the vision projector
has been in this repo since it was first published. Old links redirect automatically.

The file names still say Text-Only. Hugging Face redirects a renamed repo
but not individual file URLs, so renaming the files would 404 every existing direct
link and download script. They stay as they are.

Community quantization. Not an official Qwen release; not endorsed by or
affiliated with the Qwen team or Alibaba Cloud. "Qwen3.8" identifies the upstream
model this artifact derives from (Apache-2.0 §6).


Which file do I download?

File Size pp512 tg128 Max context on one 3090
Q8_0 26.6 GiB — — does not fit Near-lossless. Needs 32 GB, or partial offload on 24 GB
Q6_K 20.6 GiB 1088 32.3 32,512 Best quality on one 24 GB card — but look at that context
Q5_K_M 17.9 GiB 1110 36.9 73,472
Q4_K_M 15.4 GiB 1183 41.4 111,872 Fastest, and 3.4× the context of Q6_K. The default pick
IQ4_XS 14.0 GiB 1042 31.9 133,632 Pick for a 16 GB card — see the caveat below
Q3_K_M 12.4 GiB 992 38.5 158,976 Quality drops noticeably
mmproj-* 0.9 GiB — — — Optional vision, pairs with any of the above
mtp-*-Q4_K_M 1.9 GiB — — — Optional MTP speculative draft, +36 % on CUDA — see below

All three columns measured on one RTX 3090 (24 GB), Vulkan, full offload. Context
figures come from llama-fit-params, i.e. what actually fits — not a calculation.

The size/context trade is steeper than the size/quality trade. Dropping from Q6_K
to Q4_K_M costs a little quality and buys 3.4× the context — on the same card.

The counter-intuitive part: IQ4_XS is 9 % smaller than Q4_K_M but ~23 % slower.
IQ-family quants cost more compute to dequantize. Take IQ4_XS because you need the
size, not because you want speed. If Q4_K_M fits your card, it is both faster and
higher quality.

16 GB card: IQ4_XS. Q4_K_M technically loads but leaves almost nothing for KV.
24 GB card: Q4_K_M for speed, Q6_K for quality.
Apple Silicon: unified memory is the budget — 32 GB → Q5_K_M/Q6_K, 64 GB → Q8_0.

Vulkan numbers. CUDA builds are typically faster.


Why the KV cache is unusually small

Qwen3.8 is a hybrid-attention model. Of its 64 layers only 16 are full attention —
the other 48 are Gated DeltaNet linear attention and hold no KV cache. With
num_key_value_heads = 4, head_dim = 256:

KV dtype Per token 32K ctx 128K ctx
f16 64 KiB 2.0 GiB 8.0 GiB
q8 32 KiB 1.0 GiB 4.0 GiB

A comparable dense-attention 27B needs roughly four times this — which is why the
context figures in the table above are as large as they are.

The table is measured on a 24 GB card. For a 16 GB card with IQ4_XS (14.0 GiB) the
same arithmetic gives roughly 32K at f16 KV or 64K at q8 — estimated, not measured,
as the author has no 16 GB card to test on.


Vision

llama-mtmd-cli -m Qwen3.8-27B-Uncensored-Text-Only-Q4_K_M.gguf \
               --mmproj mmproj-Qwen3.8-27B-Uncensored-f16.gguf \
               --image photo.jpg -p "Describe this image."

The projector is the general-purpose tower from Qwen3.8-27B, unchanged. Dedicated
Qwen3-VL-* models will still do better on dense OCR and small-object counting.


Speculative decoding (MTP)

Qwen3.8 was trained with a multi-token-prediction head. llama.cpp's converter publishes
the target model and the MTP draft as two files by design, so the head is here as its
own optional download rather than inside the main weights.

llama-cli -m Qwen3.8-27B-Uncensored-Text-Only-Q4_K_M.gguf \
          -md mtp-Qwen3.8-27B-Uncensored-Q4_K_M.gguf \
          --spec-type draft-mtp --spec-draft-n-max 1 \
          -p "..."

Works with llama-cli and llama-server. llama-completion does not accept -md.
Requires build b10502 or newer — draft-mtp does not exist before it.

The numbers

One RTX 3090, Q4_K_M target, 512 tokens, greedy, single stream. no draft and
n-max 1 are the mean of three runs; the rest are single runs.

CUDA

--spec-draft-n-max tok/s Draft acceptance
no draft 41.2 —
1 56.1 75 %
2 54.6 63 %
3 (default) 48.6 50 %
4 45.1 41 %

+36 % at n-max 1. Every setting helps on CUDA, but the default of 3 leaves a third
of the gain on the table.

The backend matters more than anything else here

The same card, same model, same prompt, on the Vulkan build:

--spec-draft-n-max tok/s Draft acceptance
no draft 39.1 —
1 40.4 77 %
2 33.7 62 %
4 28.3 44 %

+3 % at best, and negative at the default. Acceptance is the same 77 % — the draft
is doing its job either way. What differs is the cost of batched verification, which
Vulkan does not make cheap enough to pay for the draft's own forward pass.

Metal (M1 Max, 64 GB unified, same files, same method):

--spec-draft-n-max tok/s Draft acceptance
no draft 11.17 —
1 9.96 77 %
2 8.80 66 %

Negative. −11 % at n-max 1, −21 % at 2.

So: +36 % on CUDA, +3 % on Vulkan, −11 % on Metal — from identical files, at the same
75–77 % acceptance on all three.
The draft does its job everywhere; what differs is
what batched verification costs on each backend. On Apple Silicon, do not attach the
draft.
If you are on ROCm, SYCL, or anything else, measure before assuming any of these
numbers applies to you.

Apple Silicon users: llama.cpp is not the fastest path for this model anyway. An MLX
build of the same weights ran 15.9 tok/s on the same machine — 42 % faster than
llama.cpp Metal, and no draft involved.

Only one draft file, and why

Q8_0 (2.9 GiB) was built and tested too: 55.0 tok/s at 73.8 % acceptance, versus
56.1 at 75.0 % for this Q4_K_M (1.9 GiB) — and the two produced byte-identical
output
. The bigger draft was slower and no more accurate, so it is not published. A
draft cannot change what the target accepts, so there is no quality argument for it.

It is not bit-identical to running without a draft

Speculative decoding preserves the output distribution — every drafted token is
verified against the target, so nothing gets through that the target would not have
produced. It does not reproduce the same string. Batched verification changes
floating-point reduction order, so a near-tie between two candidate tokens can land on
the other one.

Measured: each configuration is perfectly reproducible with itself (two greedy runs
byte-identical), but greedy output with the draft diverged from greedy output without it
after ~338 characters on CUDA and ~121 on Vulkan, continuing as an equally coherent
answer. If you need byte-reproducible output, do not attach a draft.

The draft is not abliterated

The MTP head in JonathanColetti/Qwen3.8-27B-Uncensored is byte-identical to the
one in the base model — verified by SHA-256 over every mtp.* tensor. Abliteration
only touched part of the language model. So this draft is the stock one.

That cannot censor anything: the target verifies every drafted token, and the target
here is the abliterated model. The only consequence is that acceptance may run lower
on content the stock model would have declined — costing speed, never refusals.


How these were made

Quantized from the original BF16 weights, not re-quantized from an existing INT4
release — so no compounding loss.

  1. Vision tower and MTP block split off at the safetensors level;
    model.language_model.* promoted to model.*
  2. convert_hf_to_gguf.py --outtype bf16 --no-mtp → BF16 GGUF (lossless from source).
    The mmproj and mtp-* files come from second passes over the unmodified
    source with --mmproj and --mtp — the pairing llama.cpp's converter documents
  3. Importance matrix over 300 chunks of the same calibration corpus used for this
    author's W4A16 releases (512 passages, pile-val news text), on 2×RTX 3090
  4. Every level quantized with that imatrix, K-quants included

Every file in this repo was loaded and generated with before publishing — including
both mmproj files, which were checked against a synthetic image with known content.
Not just checksum-verified.


Requirements

A llama.cpp build that knows the qwen35 architecture. Build b10502 or newer works.

Older builds fail on the main files with check_tensor_dims: tensor 'blk.64...' not found. That is a converter that counted an MTP block the main files do not carry —
not a corrupt download. The mtp-* files are the opposite case: they are block 64,
and --spec-type draft-mtp only exists in b10502 and newer.

① Provenance and attribution

The abliterated weights are not my work.

Layer Author
Base model Qwen/Qwen3.8-27B — Qwen team, Alibaba Cloud
Abliteration JonathanColetti/Qwen3.8-27B-Uncensored — via Heretic, 200-trial Pareto search
This repo GGUF conversion + imatrix quantization only. No weight modification beyond quantization.

Not an official Qwen release; not endorsed by or affiliated with the Qwen team
or Alibaba Cloud. "Qwen3.8" identifies the upstream model this artifact derives
from (Apache-2.0 §6).


② Safety alignment has been removed

This is the point of the model, and you should read this before downloading.

The upstream abliteration removes the refusal behaviour trained into
Qwen/Qwen3.8-27B. On the author's own W4A16 build of the same weights,
automated refusal testing measured 0 refusals out of 100 adversarial prompts,
against 12/100 for the upstream BF16. Quantization does not restore refusals.

Consequences you are accepting:

  • It will produce content the original model declines to produce, including
    content that is offensive, dangerous, or illegal in your jurisdiction.
  • It has no content filter. There is no safe-completion path, no refusal
    fallback, and no guardrail to fail back to.
  • Refusal removal is not proven exhaustive — absence of observed refusals
    in testing is not proof that none remain, and equally not proof that no
    harmful behaviour was introduced.
  • TruthfulQA drops ~5 pp versus the base model. "Do not refuse" and
    "prefer the truthful answer" are partly aligned optimization targets;
    pulling on one moves the other.

Not intended for:

  • Deployment to third parties, end users, or any public-facing service
    without your own safety layer
  • Anyone under the legal age in their jurisdiction
  • Any use prohibited by Qwen's acceptable use policy,
    which applies to this derivative exactly as it does to the base model

Intended for: local inference and research, by people who understand the
above and take responsibility for it.


③ No warranty; responsibility rests with the user

This model is provided "AS IS", without warranty of any kind, express or
implied, including but not limited to warranties of merchantability, fitness
for a particular purpose, and non-infringement.

  • I do not endorse, recommend, or condone any particular use of this model.
  • I make no representation that its outputs are accurate, lawful, or fit for
    any purpose.
  • You are solely responsible for what you generate with it, for how you
    deploy it, and for compliance with all laws and regulations applicable to
    you — including but not limited to laws on illegal content, data protection,
    export control, and AI-specific regulation in your jurisdiction.
  • To the maximum extent permitted by law, I accept no liability for any claim,
    damage, or other liability arising from the model or its use.

Downloading these files means you accept the above. If you do not, do not
download them.

License

Apache 2.0, inherited through the chain above. The Apache-2.0 grant covers the
weights; it does not grant permission for uses that are unlawful where you are.

Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration