license: apache-2.0
base_model: trohrbaugh/Qwen3.8-27B-heretic-ara
base_model_relation: quantized
pipeline_tag: image-text-to-text
tags:
- qwen3_5
- qwen3.8
- abliterated
- uncensored
- nvfp4
- mtp
- multimodal
- compressed-tensors
Qwen3.8-27B ARA Abliterated / Uncensored NVFP4 + MTP
This is a multimodal, refusal-ablated (commonly described as "uncensored") W4A4
NVFP4 derivative of the reproducibletrohrbaugh/Qwen3.8-27B-heretic-ara
BF16 checkpoint. It is intended for native Blackwell NVFP4 inference.
What is preserved
- The vision tower, recurrent convolutions, language head, and all 15 native MTP
tensors remain BF16. - The MTP tensors were grafted from the hash-verified source after Transformers
serialization and verified bit-exact. - All 333 vision tensors were verified bit-exact against the BF16 source.
- The language-model linear layers use compressed-tensors NVFP4 W4A4 group-16
quantization.
Validation
This artifact passed its text-capability, image-vision, benign refusal-surface,
native MTP-acceptance, integrity, and clean-load gates on vLLM 0.23. SeeBUILD_MANIFEST.json, VALIDATION_REPORT.json, and SHA256SUMS for exact
provenance and results. Video tensors/processors are preserved, but video input
was not part of the live runtime gate.
The live gate used Qwen3_5ForConditionalGeneration, native three-token MTP,
the FlashInfer CUTLASS NVFP4 kernel, an 8,192-token context, and an RTX PRO 6000
Blackwell GPU. Loaded model memory was approximately 19.53 GiB.
vLLM example
vllm serve aday777/Qwen3.8-27B-ARA-abliterated-NVFP4-MTP \
--served-model-name Qwen3.8-27B-ARA-NVFP4-MTP \
--max-model-len 8192 \
--speculative-config '{"method":"mtp","num_speculative_tokens":3}'
Use a recent vLLM build with Qwen3.5 multimodal and compressed-tensors NVFP4
support. Native NVFP4 execution requires compatible Blackwell hardware and CUDA
runtime support.
Notes
"Abliterated" or "uncensored" describes the source checkpoint's refusal-ablation
process; it is not a guarantee that every prompt will receive a particular answer.
Users remain responsible for evaluating outputs and applying safeguards appropriate
to their deployment.