license: apache-2.0
base_model: HauhauCS/Qwen3.6-27B-Uncensored-HauhauCS-Aggressive
tags:
- qwen3.6
- safetensors
- uncensored
- vision
- not-for-all-audiences
- fp8
- compressed-tensors
- vllm
Qwen3.6-27B Uncensored HauhauCS Aggressive — FP8-Dynamic
vLLM-native FP8 quantization (compressed-tensors, data-free dynamic scheme) of the HauhauCS uncensored finetune.
Provenance & fidelity: HauhauCS released this finetune only as GGUF — no safetensors
source exists publicly. These weights were dequantized from the highest-fidelity public
artifact (Q8_K_P GGUF, which internally is Q8_0/F16/F32 tensors) back into the officialQwen/Qwen3.6-27B HF layout, with every tensor validated against the stock checkpoint's
exact shape and dtype (1199/1199). Quality ceiling is therefore Q8 (near-lossless vs the
author's private BF16), not true BF16.
- Vision tower included (rebuilt from the author's f16 mmproj; qkv re-fused, Conv3D
patch-embed re-stacked) — full image/video input support. - MTP head included (
mtp.*, taken from stockQwen/Qwen3.6-27Bsince the finetune
never shipped one) — enables speculative decoding in engines that support Qwen3.6 MTP. - Tokenizer/configs from stock
Qwen/Qwen3.6-27B(the finetune does not modify them).
Usage (vLLM)
vllm serve zeeksa/Qwen3.6-27B-Uncensored-HauhauCS-Aggressive-FP8-Dynamic
compressed-tensors FP8_DYNAMIC (float-quantized weights, dynamic per-token activations).
Linear layers FP8; lm_head and the vision tower kept BF16. ~31 GB.
Related
- GGUF + DFlash speculative decoding bundle: zeeksa/Qwen3.6-27B-Uncensored-HauhauCS-Aggressive-DFlash-GGUF
- Original finetune (GGUF only): HauhauCS/Qwen3.6-27B-Uncensored-HauhauCS-Aggressive (Apache-2.0)
Note: Qwen3.6's gated attention is new — use a recent transformers (>=5.10) / recent vLLM;
older engine versions may not support the architecture.