license: apache-2.0
base_model: DavidAU/Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NM-DAU
base_model_relation: quantized
library_name: transformers
pipeline_tag: image-text-to-text
tags:
- qwen3_5
- autoround
- w4a16
- int4
- 4-bit
- compressed-tensors
- vllm
- marlin
- mtp
- vision
- uncensored
Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NM-DAU — W4A16 AutoRound (attention + GatedDeltaNet excluded)
4-bit weight-only quantization (W4A16, int4 symmetric, group_size 128) of
DavidAU/Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NM-DAU,
produced with AutoRound (llm-compressor 0.14 + auto-round 0.15.1, 400 iterations, 256
calibration samples), with the full-attention and GatedDeltaNet projections kept in BF16
because measurement showed they carry most of the quantization damage at 4 bits.
- Quantized → compressed-tensors
pack-quantized(uint4b8), served by vLLM's Marlin
kernels: the MLP projections only (192 modules:mlp.{gate,up,down}_proj× 64 layers). - Kept in BF16 (bit-identical to the base model):
- All full-attention projections —
self_attn.{q,k,v,o}_proj, 64 modules - All GatedDeltaNet projections —
linear_attn.{in_proj_qkv,in_proj_z,out_proj}, 144 modules - Vision tower (
model.visual.*, 333 tensors) and MTP head (mtp.*, 15 tensors) lm_head, embeddings, norms, conv1d, and the GatedDeltaNetin_proj_a/in_proj_b
- All full-attention projections —
- Calibration: 256 samples × 2048 tokens from
neuralmagic/LLM_compression_calibration
with the model's own chat template. - Size: ~30.2 GB (base BF16 ~52 GB). 70% of the quantized parameters stay in 4 bits.
Quality
Perplexity on 50 held-out texts from wikitext-103, 16,007 evaluated tokens — deliberately not the
calibration set. Same texts, same tokenization, same method, both models served by vLLM. The BF16
baseline is 8.0880 (reproduced identically across every run of this series).
| version | perplexity | delta |
|---|---|---|
| base BF16 | 8.0880 | — |
| fully quantized W4A16 (nothing excluded) | 8.7945 | +8.74% |
| this build (attention + GatedDeltaNet in BF16) | 8.3888 | +3.72% |
Where the 4-bit damage comes from
Each group of modules was restored to BF16 in turn — surgically, on top of the fully-quantized 4-bit
checkpoint, without re-quantizing — and the perplexity re-measured on the identical corpus:
| group restored to BF16 | modules | perplexity | delta | share of the +8.74% |
|---|---|---|---|---|
| (none — fully quantized) | 0 | 8.7945 | +8.74% | — |
| GatedDeltaNet projections | 144 | 8.6289 | +6.69% | 2.05 pp |
| MLP | 192 | 8.6115 | +6.47% | 2.27 pp |
| attention projections | 64 | 8.5515 | +5.73% | 3.01 pp |
| attention + GatedDeltaNet | 208 | 8.3888 | +3.72% | 5.02 pp |
Attention is still the single largest contributor at 4 bits (3.01 pp), and no single group was
enough to reach the 5% band — attention alone leaves +5.73%. The attention + GatedDeltaNet pair is
what crosses it.
The MLP is not excluded on purpose: it holds 17.1 B of the 24.3 B quantized parameters (70%), so
restoring it to BF16 would produce a 42 GB checkpoint against 52 GB for plain BF16 — the
quantization would stop being a compression. Keeping it in 4 bits is what makes this build 30.2 GB.
Honest note on the numbering
This build is dominated by its 6-bit sibling in the same series
(…-W6A16-AutoRound, 27.6 GB, +2.40%): that one is smaller and more faithful, and being smaller
it also reads fewer bytes per token. So on size, fidelity and decode bandwidth together, the 6-bit
build wins.
What this checkpoint establishes is that +3.72% is reachable at 4 bits for this model — the
project's stated target of ≤5% is met — and it documents exactly which module groups the 4-bit error
lives in. If you are choosing one artifact to serve, choose the W6A16 build; if you specifically want
a 4-bit MLP for its kernel/throughput characteristics, this is the most faithful 4-bit option we
found.
Serving notes
The quantized weights are 4-bit compressed-tensors uint4b8, served by Marlin:
Using MarlinLinearKernel for CompressedTensorsWNA16
Attention and GatedDeltaNet projections run as ordinary BF16 GEMMs, which is what costs the bytes —
those 208 modules are 20.4 B parameters that stay at 2 bytes each.
Usage (vLLM ≥ 0.28)
vllm serve DoktorMincs/Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NM-DAU-W4A16-AutoRound \
--tensor-parallel-size 2 \
--max-model-len 32768 \
--reasoning-parser qwen3 \
--enable-prefix-caching \
--speculative-config '{"method": "mtp", "num_speculative_tokens": 1}'
- Vision inputs work normally (image/video); the vision tower runs in BF16.
- The MTP draft head loads from the same checkpoint for speculative decoding.
- Validated on vLLM 0.28: text, MTP speculative decoding and vision confirmed; 192 packed modules,
208 weights left in BF16. - This is a thinking model: with a small
max_tokensthe whole budget can be consumed by the
reasoning trace, leaving the visible answer empty. Usemax_tokens≥ 1024 and prefer the chat
endpoint over raw completion.
Notes
- Architecture:
Qwen3_5ForConditionalGeneration(hybrid: 64 layers, 3:1 GatedDeltaNet
linear-attention : full-attention, 27-block ViT, 1 MTP layer). - The exclusion is surgical: the MLP was AutoRound-tuned with the attention and GatedDeltaNet
quantized, so a re-quantization with those excluded from the start would likely do slightly better. - No training, fine-tuning or abliteration was performed here — this is a weight-only quantization.
The behavioural characteristics, abliteration and usage warnings of the upstream model are
inherited unchanged. - Perplexity measures language-modelling fidelity only; it is not a substitute for task-specific
evaluation.
License
Inherited from the upstream model, which declares the Apache License 2.0. See LICENSE andNOTICE in this repository. Upstream chain: Qwen/Qwen3.8-27B → this checkpoint's base by DavidAU.