← back to catalog · registered 2026-08-25 09:02

bowmanslayer/Qwen3.5-9B-Uncensored-W4A16

bowmanslayer Qwen 2.5B multimodal second-order
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/bowmanslayer%2FQwen3.5-9B-Uncensored-W4A16"
Response includes
  • classification m1
  • files 20
  • hub_downloads_all_time 180
  • author_summary 9 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M1
Primary method

Direct removal

No other method signals detected in this model.
Confidence
MEDIUM
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=0 (base model)
  • no specific method indicators - defaulting to M1 (most common)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
MEDIUM
Why we say so
primary_method=M1; difference-of-means is the reference extraction for M1/M3 (Arditi 2024)
Downloads · lifetime
180
85 last 30d - stable
Likes
0
Model age
6w ago
created 2026-08-25
Downloads over time
Now215→from75↑187%
6812217522975 on Aug 26215 on Oct 11AugSepOct
Aug 26 → Oct 11 · 47 snapshots · spans 46 days

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
apache-2.0
Languages
en zh
Tags
transformers safetensors qwen3_5 image-text-to-text uncensored abliterated qwen3.5 multimodal w4a16 gptq not-for-all-audiences conversational

Related

Total size
8.11 GB
Files
20
Quantizations
1
Registered
2026-08-25 09:02
Last updated on HF
2026-08-25 09:43

Files by quantization

Auxiliary files 20 files 8.13 GB
model-00006-of-00008.safetensors 1.89 GB 7346cc6b download
model-00008-of-00008.safetensors 1.89 GB 106ec926 download
model-00004-of-00008.safetensors 1023 MB 5abcd679 download
model-00003-of-00008.safetensors 1022 MB 04e4113f download
model-00002-of-00008.safetensors 1018 MB e970b976 download
model-00001-of-00008.safetensors 1002 MB 34d007e5 download
model-00005-of-00008.safetensors 243 MB 997d32b9 download
model_extra_tensors.safetensors 121 MB 08d98885 download
model-00007-of-00008.safetensors 8.10 KB d4c003d5 download
tokenizer.json 19.1 MB 06b95093 download
model.safetensors.index.json 108 KB c80abe31 download
config.json 8.86 KB 72cd1814 download
chat_template.jinja 7.57 KB a585dec8 download
quantization_config.json 5.63 KB 22e1cdbd download
README.md 3.73 KB 7b52177d download
.gitattributes 1.53 KB 52373fe2 download
processor_config.json 1.16 KB 33818c7f download
tokenizer_config.json 1.14 KB 1d134cd2 download
preprocessor_config.json 443 B 8ed39680 download
generation_config.json 136 B c5afb96c download

README current version from Hugging Face


library_name: transformers
license: apache-2.0
license_link: https://huggingface.co/Qwen/Qwen3.5-9B/blob/main/LICENSE
pipeline_tag: image-text-to-text
base_model:

  • bowmanslayer/Qwen3.5-9B-Uncensored
    base_model_relation: quantized
    tags:
  • uncensored
  • abliterated
  • qwen3.5
  • multimodal
  • w4a16
  • gptq
    language:
  • en
  • zh

Qwen3.5-9B-Uncensored-W4A16

INT4 weight + FP16 activation 量化版,基于
bowmanslayer/Qwen3.5-9B-Uncensored。
方法:AutoRound → GPTQ 格式,vLLM gptq_marlin 快速核可用。
配合 fp8_e5m2 KV cache 拿到超同尺寸对手 60-80% 的上下文容量。

一、快速开始

推荐 vLLM 启动(已开 fp8 KV,开箱最佳上下文):

vllm serve bowmanslayer/Qwen3.5-9B-Uncensored-W4A16 \
  --dtype float16 --tensor-parallel-size 2 \
  --kv-cache-dtype fp8_e5m2 \
  --max-model-len 28672 --max-num-seqs 16 \
  --reasoning-parser qwen3

单卡也可:

vllm serve bowmanslayer/Qwen3.5-9B-Uncensored-W4A16 \
  --dtype float16 --tensor-parallel-size 1 \
  --kv-cache-dtype fp8_e5m2 --gpu-memory-utilization 0.92

二、量化不破消融(验证过)

12 条 harmful 快验:W4A16 quantized 拒绝率 = 0/10(文字 0/6,图像 0/4),
与 bf16 原模的 2/100(0.02)量级一致。量化没引入任何拒绝回退。

主仓 bf16 的完整能力评测、四方对比表、拒绝率对比见
bowmanslayer/Qwen3.5-9B-Uncensored 主 README。

三、和常见 ~8GB 版本的上下文对比

Qwen3.5-9B 架构 = 32 层(24 linear + 8 full attention),hidden 3584。KV cache 只对
full attention 层随 context 线性增长
,因此本节数字是理论算(不含 CUDA graph overhead
等,实测偏差 <5%)。

每 token KV cache 大小:

  • fp16 KV:2 (K+V) × 8 层 × 3584 × 2 字节 = 114 KB / token
  • fp8 KV(本模型默认):2 × 8 × 3584 × 1 = 57 KB / token(-50%)

单卡 NVIDIA GPU 上,减去权重 + ~1 GB overhead,剩余可分配给 KV cache:

GPU 本模型 W4A16(8.2 GB) + fp8 KV 同尺寸 GGUF Q6_K(7.6 GB) + f16 KV(llama.cpp 默认)
12 GB(RTX 3060 12G / 4070 等) ~48k tokens ~30k tokens
16 GB(RTX 4060 Ti / 4070 Ti Super 等) ~120k tokens ~65k tokens
24 GB(RTX 3090 / 4090) ~256k tokens(接近模型最大 262k) ~135k tokens

要点:

  • 我们的推荐启动默认就是 fp8_e5m2 KV,无需手动配置
  • GGUF 用户要拿到同等上下文,需手动加 -ctk q8_0 -ctv q8_0(int8 KV,类似效果),
    但大部分社区帖不给,默认体验只到我们一半
  • 12 GB 显卡的差异最大(48k vs 30k,+60% 上下文)—— 是这类中端卡上最有感知的价值差

四、显存与量化档次对照

以 W4A16 权重 8.2 GB 为基准:

场景 权重 KV(28k ctx,batch=1) 系统 overhead 合计 建议 GPU
本模型 + fp8 KV,batch=1 8.2 ~1.6 ~1.5 ~11.3 GB 12 GB ✓
本模型 + fp8 KV,batch=8 8.2 ~12.5 ~1.5 ~22 GB 24 GB
本模型 + fp16 KV,batch=1 8.2 ~3.2 ~1.5 ~12.9 GB 16 GB

五、方法与限制

见主仓 bowmanslayer/Qwen3.5-9B-Uncensored。

量化配方:AutoRound --bits 4 --group_size 128 --format auto_round:auto_gptq --nsamples 256 --seqlen 2048,linear_attn.in_proj_a/b(48 张量,ssm 计算不可量化)
排除。视觉塔和 MTP 头随语言层一同处理为 int4。

MTP speculative decoding 不推荐启用(vLLM 0.20.2 实测减速 40-46%,K=5 崩溃)。

六、署名

基座模型见 Qwen/Qwen3.5-9B,Apache 2.0。
本仓是 bowmanslayer 消融版的 W4A16 量化。

README history 2 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-08-25README: bilingual + full eval + legal + gatedb6431c416.2 KB
    Loading...
  2. 2026-08-25Upload README.md with huggingface_hub3fb5c013.7 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration