base_model:
- huihui-ai/Huihui-Qwen3.5-35B-A3B-Claude-4.6-Opus-abliterated
pipeline_tag: image-text-to-text
This is a quantization of the model Huihui-Qwen3.5-35B-A3B-Claude-4.6-Opus-abliterated
Setup Instructions:
- Download the latest llama.cpp.
- Use the F32 version of the mmproj file for optimal results. Recommended quality ranking: F32 > BF16 > F16.
For configuration tips, follow the Unsloth Qwen3.5 local run guide
About this Variant:
The typical standard uses MXFP4 for MoE tensors and Q8 for everything else. In this build, I switched the non-MoE tensors to BF16. While this may result in slower inference on certain architectures, it offers the highest possible quality, effectively keeping the tensor values unquantized (or closest to the originals).