license: apache-2.0
base_model: huihui-ai/Huihui-Qwen3.6-35B-A3B-abliterated
base_model_relation: quantized
library_name: mlx
pipeline_tag: text-generation
tags:
- mlx
- 4-bit
- tensorfold
- dgx-spark
- mtp
- abliterated
- uncensored
Huihui Qwen3.6 35B-A3B abliterated — MLX 4-bit + MTP (TensorFold)
MLX affine 4-bit weights of
huihui-ai/Huihui-Qwen3.6-35B-A3B-abliterated,
with the checkpoint's own MTP head quantized into mtp-4bit.safetensors, laid out for
TensorFold's Qwen3.6 MoE CUDA family. It is the format
Spark Studio converts to locally; download it here to
skip the 67 GB BF16 download and the 20–40 minute conversion.
Speed (one NVIDIA DGX Spark, GB10, one request at a time)
TensorFold 0.6.6, --mtp-drafts 2 --prefill-fp8 --parallel 1:
| Prompt | Prefill | Decode | Time to first token |
|---|---|---|---|
| 8k tokens | 5,700 – 7,200 tok/s | 98 – 128 tok/s | 1.1 – 1.4 s |
| 21k tokens | 5,600 – 6,100 tok/s | prose ~128, code 98 – 147 tok/s | 3.8 s |
For reference, the BF16 checkpoint on vLLM 0.29 with FP8 online quantization + MTP 2 measured about
53 / 69 tok/s decode and 3,000 – 3,600 tok/s prefill on the same machine.
Run
With Spark Studio: open the Models screen and launch Huihui Qwen3.6 35B-A3B abliterated.
By hand, with TensorFold Python 0.6.6 on CUDA (see Spark Studio's docker/tensorfold.Dockerfile):
hf download swdq/Huihui-Qwen3.6-35B-A3B-abliterated-MLX-4bit-MTP --local-dir ~/models/huihui-qwen36-mlx4-mtp
tensorfold serve ~/models/huihui-qwen36-mlx4-mtp --backend cuda --port 8110 \
--name huihui-qwen3.6-abliterated --context 262144 --parallel 1 --mtp-drafts 2 --prefill-fp8
The server speaks OpenAI Chat Completions / Responses and Anthropic Messages. The weights also load
with mlx-lm on Apple Silicon (text only; mtp-4bit.safetensors is used by TensorFold).
How it was made
- Source:
huihui-ai/Huihui-Qwen3.6-35B-A3B-abliterated@8f0ee727aff5e771ea72466d64d13ecd851d2cc7(BF16). mlx_lm.convert -q --q-bits 4 --q-group-size 64 --dtype bfloat16(mlx 0.32.3, mlx-lm 0.31.3, CPU):
4-bit/group-64 affine weights, 8-bit router and shared-expert gates, 4.503 bits per weight.
mlx-lm's text conversion drops the vision tower and the MTP layer.- The source's
mtp.*tensors were kept: expertgate_up_projsplit intogate_proj/up_proj,
RMSNorm weights shifted by +1 (MLX convention), linear weights quantized withmlx.core.quantize
(4-bit/group-64; 8-bit for router gates) and saved asmtp-4bit.safetensors. No external draft
model is involved.
The exact script is in Spark Studio's src/tensorfold.rs (CONVERT_MTP).
Notes
- Text only (no vision).
- Abliterated models have reduced refusals. You are responsible for how you deploy and use them;
add your own safeguards where needed. - License: Apache-2.0, inherited from Qwen3.6 and the huihui-ai checkpoint. Credit to the Qwen team,
huihui-ai, the MLX team and TensorFold's authors.