library_name: mlx
pipeline_tag: image-text-to-text
license: apache-2.0
base_model: vanch007/Huihui-Qwen3.6-35B-A3B-abliterated-mlx-bf16
tags:
- mlx
- mlx-lm
- mlx-vlm
- qwen
- qwen3.6
- moe
- multimodal
- vision-language
- abliterated
- uncensored
- nvfp4
- quantized
- thinking
- lm-studio
- ollama
Huihui-Qwen3.6-35B-A3B-abliterated-mlx-nvfp4
MLX-native NVFP4 quantized build of Qwen3.6 35B-A3B (MoE, multimodal), abliterated for uncensored output, with <think>/</think> special tokens correctly registered so downstream tools that parse tokenizer_config.json (Ollama's qwen3.5 parser, LM Studio, HF tokenizer-based servers) extract reasoning into a structured thinking field.
What this repo provides
- Format: MLX-native NVFP4 (bits=4, group_size=16, mode=nvfp4) with per-path affine 8-bit overrides on router gates — the mixed-precision layout mlx-lm emits natively.
- Size: ~19 GB on disk (4 safetensors shards).
- Target runtime:
mlx-lm,mlx-vlm, LM Studio, Ollama (with thex/mlxrunnerquant-naming fix). - Fixed vs base abliterated release:
tokenizer_config.jsonnow includesadded_tokens_decoderwith the<think>(id 248068) and</think>(id 248069) entries so thinking-aware parsers recognise them as special tokens.
Why a separate NVFP4 MLX repo
Existing downstream quantizations of huihui-ai/Huihui-Qwen3.6-35B-A3B-abliterated at the time of upload are either:
- GGUF (llama.cpp) — great on CPU, not MLX-native,
- vLLM compressed-tensors NVFP4 — server-side GPU format, not mlx-lm compatible,
- MLX bf16 — 66 GB, too large for many Apple Silicon laptops.
This repo fills the MLX NVFP4 slot specifically, and on import was run through a full verification pass (load, prompt eval, generation, thinking extraction, Lake-Tahoe vision sanity check via mlx-vlm).
Benchmarks (Apple M4 Pro, 64 GB)
| Metric | Value |
|---|---|
| Generation throughput | 85–88 tok/s |
| Prompt eval throughput | 115 tok/s |
| TTFT (19-token prompt) | 0.18 s |
| Peak memory | ~19 GB |
For comparison, on the same machine:
| Runtime | Model | Gen tok/s |
|---|---|---|
| Ollama (patched) | this repo | 88.8 |
| Rapid-MLX / vllm_mlx | this repo | 79.7 |
| Ollama (patched) | qwen3.6:35b-a3b-nvfp4 (non-abliterated) |
69.9 |
| Ollama | huihui_ai/Qwen3.6-abliterated:35b-a3b-q4_K GGUF |
31.0 |
Usage
mlx-lm (text)
pip install mlx-lm
mlx_lm.generate \
--model <your-hf-user>/Huihui-Qwen3.6-35B-A3B-abliterated-mlx-nvfp4 \
--prompt "What is 17 * 23? Think step by step." \
--max-tokens 400
mlx-vlm (vision + text)
pip install mlx-vlm
python -m mlx_vlm generate \
--model <your-hf-user>/Huihui-Qwen3.6-35B-A3B-abliterated-mlx-nvfp4 \
--image photo.jpg \
--prompt "Describe this image." \
--max-tokens 400 \
--trust-remote-code
Ollama (text + thinking)
Requires an Ollama build with the mlx-lm quant-naming / config.json-override fix for the x/mlxrunner/model package (see the patch submitted upstream).
cat > Modelfile <<'EOF'
FROM ./Huihui-Qwen3.6-35B-A3B-abliterated-mlx-nvfp4
RENDERER qwen3.5
PARSER qwen3.5
EOF
ollama create qwen3.6-nvfp4-thinking --experimental -f Modelfile
ollama run qwen3.6-nvfp4-thinking "What is 17 * 23? Think carefully."
The RENDERER qwen3.5 / PARSER qwen3.5 directives are important — Ollama's auto-detection picks qwen3 (older, no thinking support) for this import; setting them explicitly activates the thinking-aware parser that splits <think> blocks into the message.thinking field.
Lineage
Qwen/Qwen3.6-35B-A3B (Apache-2.0, base)
└─ huihui-ai/Huihui-Qwen3.6-35B-A3B-abliterated (abliteration)
└─ vanch007/Huihui-Qwen3.6-35B-A3B-abliterated-mlx-bf16 (MLX bf16 conversion)
└─ this repo (MLX nvfp4 + think-token fix)
Limitations
- Vision path in Ollama: the MLX runner in Ollama does not yet wire the Qwen3.6 vision tower into the forward pass. Image inputs go through
mlx-vlmormlx_vlm.generatedirectly, not via the Ollama API. Text + thinking + tools all work via Ollama. - Abliterated: no refusal guardrails. Don't deploy user-facing without your own safety layer.
- Audio tokens are present in the tokenizer but untested here.
License
Apache-2.0, inherited from the upstream chain. Respect the terms of each parent repo.
Citation / credit
If you use this repo:
- Base model: Qwen team, Qwen/Qwen3.6-35B-A3B.
- Abliteration: huihui-ai, huihui-ai/Huihui-Qwen3.6-35B-A3B-abliterated.
- MLX bf16 conversion: vanch007, vanch007/Huihui-Qwen3.6-35B-A3B-abliterated-mlx-bf16.