license: apache-2.0
library_name: mlx
base_model: huihui-ai/Huihui-ThinkingCap-Qwen3.6-27B-abliterated
pipeline_tag: image-text-to-text
tags:
- mlx
- omlx
- qwen3.6
- mtp
- speculative-decoding
- 4-bit
Huihui ThinkingCap Qwen3.6 27B — oMLX 4-bit with MTP
This repository is packaged specifically for oMLX native MTP. The target checkpoint's main safetensors index includes all 15 language_model.mtp.* tensors, which bind directly to oMLX's Qwen3.5/3.6 VLM MTP model tree.
Recommended MTP runtime: use this oMLX artifact for MTP. oMLX provides the faster, more mature integrated path for this model and requires no separate drafter.
oMLX
Add this model repository to oMLX and enable Native MTP in the model settings. No separate draft-model repository is required for this oMLX artifact.
Compatibility
| Runtime | Recommended artifact |
|---|---|
| oMLX | This repository: embedded/indexed language_model.mtp.* target |
Direct mlx-vlm |
Use the normal 4-bit target plus the standalone direct mlx-vlm MTP drafter |
| LM Studio MLX | Use the normal 4-bit target without MTP; runtime 1.10.1 does not support draft models for this batched VLM |
| MTPLX | Use the MTPLX sidecar target |
Do not use this embedded-MTP target as an LM Studio target: LM Studio's current MLX target loader and oMLX use different MTP packaging contracts.
Verified runtime
Verified end-to-end on Apple Silicon with oMLX 0.4.4rc1: the model loaded directly through VLMBatchedEngine, oMLX reported its native MTP patch active, and a bounded chat generation completed with the MTP path active. That smoke accepted 2 of 5 drafted tokens (40%); acceptance depends on the prompt.
Technical details
- Target trunk: MLX affine 4-bit, group size 64
- MTP tensors: 15
language_model.mtp.*entries in the mainmodel.safetensors.index.json - MTP precision: BF16
- Source revision:
44f63da8141407af529405c1e4b83fa39b70abe0 - The three previously validated target trunk shards are unchanged; the MTP payload is an additional indexed shard.
Upstream
- Model: huihui-ai/Huihui-ThinkingCap-Qwen3.6-27B-abliterated
- Base architecture: Qwen/Qwen3.6-27B
- Runtime: oMLX