library_name: mlx-vlm
pipeline_tag: image-text-to-text
base_model: huihui-ai/Huihui-Qwen3.6-27B-abliterated
base_model_relation: quantized
tags:
- mlx
- mlx-vlm
- qwen3.6
- multimodal
- vision-language
- abliterated
Huihui-Qwen3.6-27B-abliterated-mlx-nvfp4
Local MLX-VLM conversion of huihui-ai/Huihui-Qwen3.6-27B-abliterated.
Overview
- Format: MLX-VLM
- Precision: nvfp4
- Size: 15G
- q_mode: nvfp4
- q_bits: 4
- Source model type: Qwen3_5ForConditionalGeneration
- Source pipeline: image-text-to-text
- Intended runtime: mlx-vlm, LM Studio
Validation
- text generation smoke test: passed
- image generation path smoke test: passed
Quantization
- Quantization mode: nvfp4
- Quantization bits: 4
Test Results (2026-04-23)
Test runner: mlx_vlm.generate
Text prompt:你好,请用一句自然中文回应。
Image prompt:请用一句中文描述这张图片。
Image asset:
64x64 pure red square
Observed results:
| Case | Prompt TPS | Generation TPS | Peak Memory |
|---|---|---|---|
| Text | 44.390 | 20.134 | 16.219 GB |
| Image | 100.428 | 20.164 | 16.394 GB |
Notes:
- Both text and image tests completed successfully.
- This build gives the lowest peak memory and the highest measured generation speed among the three local variants.
- The current
mlx_vlmruntime still emits<think>content even with--processor-kwargs '{"enable_thinking": false}'.