← back to catalog · registered 2026-08-22 13:56

bowmanslayer/Qwen3.8-27B-Uncensored-W4A16-vision-mtp

bowmanslayer Qwen 24B multimodal second-order
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/bowmanslayer%2FQwen3.8-27B-Uncensored-W4A16-vision-mtp"
Response includes
  • classification m1
  • files 26
  • hub_downloads_all_time 4,994
  • author_summary 9 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M1
Primary method

Direct removal

No other method signals detected in this model.
Confidence
MEDIUM
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=0 (base model)
  • no specific method indicators - defaulting to M1 (most common)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
MEDIUM
Why we say so
primary_method=M1; difference-of-means is the reference extraction for M1/M3 (Arditi 2024)
Downloads · lifetime
5K
3K last 30d - active
Likes
11
Model age
7w ago
created 2026-08-16
Downloads over time
Now5.2K→from0↑0%
01.9K3.8K5.7K0 on Aug 175.2K on Oct 11AugSepOct
Aug 17 → Oct 11 · 49 snapshots · spans 55 days

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
apache-2.0
Languages
en zh
Tags
transformers safetensors qwen3_5 image-text-to-text qwen3.8 w4a16 auto-round gptq marlin vllm vision uncensored

Related

Total size
18.2 GB
Files
26
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-08-18 06:59

Files by quantization

Auxiliary files 26 files 18.2 GB
model-00003.safetensors 6.52 GB 661e8c9d download
model-00001.safetensors 5.02 GB 3cfc9fcb download
model-00002.safetensors 5.01 GB 2a95703a download
vision.safetensors 879 MB d7defc90 download
model-mtp.safetensors 810 MB 90fa0e3e download
tokenizer.json 19.1 MB f399b3cd download
vocab.json 6.41 MB 0aa0ce06 download
merges.txt 3.20 MB a494e019 download
model.safetensors.index.json 159 KB 699a04cf download
config.json 12.1 KB fc0c3e16 download
README.md 11.6 KB db1d97fe download
LICENSE 9.87 KB ad9713c6 download
chat_template.jinja 8.74 KB c0c686f9 download
EVAL.md 4.15 KB f7668702 download
toolcall-BASE.json 3.66 KB 83d23628 download
toolcall-UNC.json 3.58 KB 384d2383 download
SUMMARY-BASE.json 2.90 KB 8241cf70 download
SUMMARY-UNC.json 2.90 KB e8242d1f download
.gitattributes 1.53 KB 52373fe2 download
vision-UNC.json 1.50 KB 8391e3b3 download
NOTICE.md 1.43 KB acafedf7 download
vision-BASE.json 1.40 KB 410cf962 download
tokenizer_config.json 1.14 KB 1d134cd2 download
preprocessor_config.json 390 B 2ea84a43 download
video_preprocessor_config.json 385 B 3ba673a5 download
generation_config.json 213 B 3f7527c4 download

README current version from Hugging Face


license: apache-2.0
base_model: JonathanColetti/Qwen3.8-27B-Uncensored
base_model_relation: quantized
pipeline_tag: image-text-to-text
library_name: transformers
language:

  • en
  • zh
    tags:
  • qwen3.8
  • w4a16
  • auto-round
  • gptq
  • marlin
  • vllm
  • vision
  • uncensored
  • abliterated

Qwen3.8-27B-Uncensored-W4A16-vision

4-bit (W4A16, group size 128) quantization of
JonathanColetti/Qwen3.8-27B-Uncensored,
with the bf16 vision tower preserved and the MTP head omitted.

18 GB on disk. Serves on 2×RTX 3090 (TP=2) with 256K context and ~12.9 GB
free for KV cache.

Attribution chain
Qwen/Qwen3.8-27B (Apache 2.0)
→ JonathanColetti/Qwen3.8-27B-Uncensored — abliteration via Heretic, 200-trial Pareto search
→ this repo — W4A16 quantization + vision-tower repack

The abliteration is not my work. What is mine here: the W4A16
quantization pipeline, the text/vision split before calibration, and the
repack that reattaches the bf16 vision tower to the quantized text stack.

Variants

Repo MTP head Size Use when
this one (-vision) no 18 GB default — MTP was not worth it on our hardware
-vision-mtp yes 19 GB you want to try --speculative-config

Both were produced from the same quantization run; they differ only by the
849 MB model-mtp.safetensors. vLLM skips mtp.* tensors unless
speculative decoding is explicitly enabled, so the two behave identically
without it.

Deployment (verified)

Hardware: 2×NVIDIA RTX 3090 24 GB (NVLink), TP=2, GPU 0+2.

vLLM launch (0.20.2, float16 compute):

vllm serve <path> \
  --host 0.0.0.0 --port 9411 \
  --dtype float16 --tensor-parallel-size 2 \
  --gpu-memory-utilization 0.95 \
  --max-model-len 262144 --max-num-seqs 16 \
  --enable-prefix-caching \
  --enable-auto-tool-choice --tool-call-parser qwen3_coder \
  --reasoning-parser qwen3 \
  --mm-processor-kwargs '{"max_pixels": 2097152}' \
  --limit-mm-per-prompt '{"image": 8, "video": 2}'

Runtime footprint (measured at startup):

Model weights on GPU 8.87 GiB per rank (× 2 ranks)
Available KV cache 12.93 GiB per rank
Total KV pool 415,125 tokens across ranks
Marlin kernel MarlinLinearKernel for GPTQMarlinLinearMethod
Encoder cache budget (image tokens) 2,048 tokens (matches max_pixels)

Why the flags matter:

  • --mm-processor-kwargs '{"max_pixels": 2097152}': without this, images
    above ~2M pixels trigger a tokenizer truncation and fail with
    Mismatch in image token count between text and input_ids. 2 M ≈ 1448×1448
    (or a long screenshot around 900×2300).
  • --limit-mm-per-prompt '{"image": 8, "video": 2}': enables the video
    input path (vLLM defaults undeclared modalities to zero) and lifts the
    per-request image cap.
  • --reasoning-parser qwen3: separates the model's <think> block from the
    answer body. Recommended if you consume message.content in a pipeline
    where thinking noise is unwanted; drop it if you want reasoning inline.

Throughput

Measured on the same 2×3090 during the evaluation runs:

Load Decode
Single-request, thinking-on 66–68 tok/s
16 concurrent, thinking-on ~700 tok/s aggregate (across all in-flight)
Prompt prefill peak 3,300 tok/s per request

The 16-concurrent aggregate is what you'll actually see in a real chat
front-end serving multiple sessions.

Quantization recipe

Text-only calibration on 256 samples × 2048 tokens from NeelNanda/pile-10k.

auto-round /models/<text-only-source> \
  --bits 4 --group_size 128 --format auto_round:auto_gptq \
  --nsamples 256 --seqlen 2048 --low_gpu_mem_usage \
  --ignore_layers <all linear_attn.in_proj_a/b across 64 layers>

Details, all learned the hard way:

  • The vision tower is stripped before quantization and repacked afterwards.
    auto-round's is_mllm_model() detects preprocessor_config.json and
    silently switches to MLLM calibration with the wrong dataset. The file
    must be absent during quantization and present in the final
    artifact.
  • linear_attn.in_proj_a / in_proj_b are excluded from quantization
    on all linear-attention layers. Qwen3.8's hybrid attention (16 full + 48
    linear across 64 layers) breaks otherwise.
  • The MTP head is copied verbatim from the source checkpoint after the
    merge, never quantized (for the -mtp variant only).
  • End-to-end quantization time on 2×3090 (~1 hr calibration, 3 hr AutoRound
    optimization loop, ~10 min repack).

Evaluation

Please read this before comparing to other repos' numbers. All numbers
below were produced with thinking ON (Qwen 3.8's default chat template
opens a <think> block; we did not close it). Most public leaderboards for
non-reasoning-tuned models use thinking OFF + 0-shot. On multiple-choice
sets, thinking ON typically gains 5–10 points over thinking OFF for this
model family. Do not cross-compare these scores to leaderboards that
disable thinking.

The comparisons that are valid:

  1. This quantized derivative vs the unmodified Qwen3.8-27B base (same
    quantization, same harness, same seed, thinking ON).
    Isolates what
    abliteration + W4A16 cost end-to-end.
  2. This quantized derivative vs upstream Qwen public numbers for
    Qwen3.8-27B.
    Same base model, different eval configuration — treat as
    a coarse sanity check, not a rank comparison.

1) Vs. our own quantized base (real, tight comparison)

Same seed, same items, thinking ON, temperature=0. Full-set accuracy and
the accuracy on the subset where thinking finished within the sampling
budget:

Benchmark Uncensored Base Δ Uncensored (ex-trunc) Base (ex-trunc) Δ (ex-trunc)
MMLU (150) 79.33 84.00 −4.67 93.70 95.45 −1.75
CMMLU (150) 85.33 84.67 +0.66 93.43 92.03 +1.40
C-Eval (150) 79.33 78.67 +0.66 91.54 88.06 +3.48
ARC-C (150) 96.00 94.00 +2.00 96.64 94.63 +2.01
TruthfulQA (150) 72.00 85.33 −13.33 87.10 92.09 −4.99
GSM8K (100) 97.00 97.00 0.00 100.00 98.98 +1.02
MATH-500 (100) 77.00 77.00 0.00 78.57 79.38 −0.81
BBH (150) 85.33 88.67 −3.34 94.70 97.01 −2.31
HumanEval (164) 80.49 82.32 −1.83 91.67 91.22 +0.45
IFEval strict (100) 80.00 79.00 +1.00 92.86 91.57 +1.29

Read the "ex-trunc" column as the actual capability figure. MC-class
tasks used a 4096-token sampling budget which is not always enough for
Qwen3.8's default <think> block to close on hard items — those items were
scored as wrong but the failure is a sampling budget artifact, not a
capability loss. Truncation rate stays around 8–17 % per task on both
models.

Vision tower: 7/7 on a self-composed synthetic-image battery (solid
colours, shape counting, OCR, grid counting, dense colour count, 1808×4000
resize test). Identical to the base.

Tool calling: 8/11 on a mixed suite (8 standard single-tool cases + 1
multi-tool composition + 2 boundary cases where no tool should be called).
The base scores 7/11 on the same suite (this one is one point better
because it correctly routed a flight-booking case).

2) Vs. Qwen upstream public numbers (coarse sanity check)

Qwen publishes Qwen3.8-27B scores as visual-formatted tables under different
eval conditions than ours. As a sanity check that this derivative did not
collapse:

Benchmark Upstream (Qwen public) Ours base (ex-trunc) Ours UNC (ex-trunc)
MMLU — (not directly given in the VL card) 95.45 93.70
MathVision (a Qwen-published VL bench) 90.0 (base) not run not run

We deliberately did not run MathVision / OSWorld / BabyVision / CharXiv —
they need extra harness setup and would only add noise. GSM8K 100 % and
HumanEval 91 % (ex-trunc) are strong on their own for a 27 B W4A16 model.

3) Refusal behaviour

100 held-out prompts from
mlabonne/harmful_behaviors
(test split, same 100 as Jonathan's upstream measurement). Scoring is a
refusal-prefix regex, run with thinking OFF to match upstream's method.

Refusals
Qwen3.8-27B-W4A16-vision (our base) 99 / 100
Qwen3.8-27B-Uncensored-W4A16-vision (this repo) 0 / 100
Upstream JonathanColetti/Qwen3.8-27B-Uncensored (bf16) 12 / 100
Upstream Qwen/Qwen3.8-27B (bf16, same eval) 98 / 100

Base 99 vs upstream 98 confirms our quantization did not perturb the
model's original safety layer — a difference of one prompt within the
scoring's noise floor.

Uncensored 0 vs upstream 12 deserves a note: both are refusal-rate
lower bounds (the regex catches "I'm sorry"-style prefixes, not "sure, here
is …" that then delivers a refusal). We manually poked the model with
adversarial inputs well beyond what the automated harness covered — dark,
graphic, and taboo — and could not find a topic that produced a refusal.
As the human tester put it, "I ran out of things I could think of to ask."
So the 0/100 number is consistent with what we observed in practice, but
it should be read as "we could not find a refusal in extensive manual
testing"
, not as a formal proof of complete removal.

The published data files in this repo include the full JSON of MMLU / CMMLU
/ C-Eval / ARC-C / TruthfulQA / GSM8K / MATH-500 / BBH / HumanEval /
IFEval / vision / tool-calling results (per-item gold + reply for
reproducibility). The 100-prompt refusal probe transcripts are
deliberately not included
— they contain the full harmful-request payload
and are not something we want to redistribute even for reproducibility.
Use mlabonne/harmful_behaviors test split directly with the same seed
if you want to reproduce.

Known limitations

  • Identity confusion. Asked "who are you?", the model sometimes claims
    to be Claude or GPT. Inherited from base training data. Set a system
    prompt to enforce identity if this matters.
  • Refusals are reduced, not verified-removed. The upstream author
    reports 12/100 on held-out prompts; we could not reproduce any refusal on
    100 automated + extensive manual testing. But this cannot be interpreted
    as proof of exhaustive removal.
  • Fine visual detail is a Qwen3.8 property, not a specialized
    vision-tower's.
    The tower here is the general-purpose one from
    Qwen3.8-27B, unchanged by quantization. Dedicated Qwen3-VL-* models
    will do better on OCR of dense small text and small-object counting.
  • Single images above ~2M pixels must be either downsampled by the
    server (via --mm-processor-kwargs '{"max_pixels": 2097152}') or by the
    client. Without the flag they fail with a token-count mismatch.
  • TruthfulQA drops 5 pp vs the base. The largest measured cost of
    abliteration in our suite. "Do not refuse harmful requests" and "prefer
    the truthful answer over a plausible-sounding one" are partly aligned
    optimization targets, and pulling on one moves the other. Expected, but
    something to be aware of.

Intended use

Local inference and research. Not intended for deployment to third parties
without your own safety layer — the same condition the upstream author
states, doubly so at 0/100.

License

Apache 2.0, inherited through the chain above. Qwen's acceptable use policy
applies to this derivative as it does to the base model.

README history 1 version

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-08-16add README, LICENSE, NOTICE, EVAL, benchmark JSONsdf7fd3011.6 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration