license: apache-2.0
base_model:
- croll83/Qwopus3.6-27B-v2-Abliterated-NVFP4
tags: - qwen3.5
- vlm
- multimodal
- nvfp4
- vllm
- abliterated
pipeline_tag: image-text-to-text
language: - en
- de
Qwopus3.6-27B-v2-Abliterated-NVFP4 — vision tower in bf16
Drop-in re-export of croll83/Qwopus3.6-27B-v2-Abliterated-NVFP4
with the vision tower kept in bf16 instead of NVFP4. The NVFP4 text body and the MTP head are unchanged.
Why this exists
The upstream checkpoint NVFP4-quantized the vision tower (model.visual.*, 110 linear layers → packed FP4)
and did not list model.visual* in hf_quant_config.json → exclude_modules.
Under vLLM, the NVFP4 GEMM path mis-runs the quantized vision tower → the image embeddings come out as NaN →
the model emits an endless stream of !!!!! for any request containing an image. Text-only requests are unaffected,
which is what makes this easy to miss.
A known-good reference is sakamakismile/Huihui-Qwen3.6-* (same Qwen3_5ForConditionalGeneration architecture,
same vLLM, same flags): it serves images correctly because it excludes the vision tower from quantization.
This export does the same. The upstream model card also intends an unquantized vision projector
("F16 for the vision projector") — this just makes the safetensors checkpoint match that intent for vLLM.
What changed vs upstream
model.visual.*(333 tensors): NVFP4 packed FP4 → bf16 (the original, untouched Qwen3.6 vision weights —
abliteration and quantization never touched the vision tower, so this is lossless, not a downgrade).hf_quant_config.json:model.visual*added toexclude_modules.- The stale packed-FP4 vision tensors were physically removed from
model.safetensors— vLLM iterates every
tensor present in a shard file, not just the entries inmodel.safetensors.index.json, so leftover FP4 vision
weights (e.g. shape[1152, 576]) would otherwise collide with the new bf16 params. - Text body (NVFP4) and MTP draft head: byte-identical to upstream.
Serving (vLLM)
vllm serve SolbachLeads/Qwopus3.6-27B-v2-Abliterated-NVFP4-vision-bf16 \
--quantization modelopt --dtype bfloat16 --trust-remote-code \
--limit-mm-per-prompt '{"image":20}' \
--reasoning-parser qwen3 --enable-auto-tool-choice --tool-call-parser qwen3_xml
Images now return correct output instead of !!!! (verified on RTX PRO 6000 Blackwell and B200).
Provenance
Built deterministically from public inputs with build_checkpoint.py (SolbachLeads infrastructure):
NVFP4 text body from croll83/Qwopus3.6-27B-v2-Abliterated-NVFP4, bf16 vision tower from the BF16 base.
Lineage: croll83/Qwopus3.6-27B-v2-Abliterated-NVFP4 ← Jackrong/Qwopus3.6-27B-v2 ← Unsloth/Qwen3.6-27B (Qwen3.5).
License: Apache-2.0.