← back to catalog · registered 2026-10-10 12:58

swdq/Huihui-Qwen3.6-35B-A3B-abliterated-MLX-4bit-MTP

swdq 35B MoE second-order
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/swdq%2FHuihui-Qwen3.6-35B-A3B-abliterated-MLX-4bit-MTP"
Response includes
  • classification m-uncensored
  • files 2
  • author_summary 1 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M-U
Primary method

Uncensored (method unknown)

No other method signals detected in this model.
Confidence
LOW
Why this label 3 signals
Weak or ambiguous signals. Best guess based on catalog patterns; treat as tentative and check the evidence below.
  • 'uncensored' in name/tags but no 'abliterated' marker
  • method not identifiable from author declaration alone
  • may be DPO fine-tune, prompt engineering, or unknown technique
Refusal direction extraction

No specific extraction method could be identified for this model. The producer either did not document it or used a proprietary pipeline.

What is a refusal direction? →
Downloads · 30-day
0
Likes
0
Model age
today
created 2026-10-10

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
apache-2.0
Tags
mlx safetensors qwen3_5_moe 4-bit tensorfold dgx-spark mtp abliterated uncensored text-generation conversational base_model:huihui-ai/Huihui-Qwen3.6-35B-A3B-abliterated

Related

Total size
0 B
Files
2
Quantizations
1
Registered
2026-10-10 12:58
Last updated on HF
2026-10-10 13:09

Files by quantization

Auxiliary files 2 files 4.58 KB
README.md 3.10 KB d346591c download
.gitattributes 1.48 KB a6344aac download

README current version from Hugging Face


license: apache-2.0
base_model: huihui-ai/Huihui-Qwen3.6-35B-A3B-abliterated
base_model_relation: quantized
library_name: mlx
pipeline_tag: text-generation
tags:

  • mlx
  • 4-bit
  • tensorfold
  • dgx-spark
  • mtp
  • abliterated
  • uncensored

Huihui Qwen3.6 35B-A3B abliterated — MLX 4-bit + MTP (TensorFold)

MLX affine 4-bit weights of
huihui-ai/Huihui-Qwen3.6-35B-A3B-abliterated,
with the checkpoint's own MTP head quantized into mtp-4bit.safetensors, laid out for
TensorFold's Qwen3.6 MoE CUDA family. It is the format
Spark Studio converts to locally; download it here to
skip the 67 GB BF16 download and the 20–40 minute conversion.

Speed (one NVIDIA DGX Spark, GB10, one request at a time)

TensorFold 0.6.6, --mtp-drafts 2 --prefill-fp8 --parallel 1:

Prompt Prefill Decode Time to first token
8k tokens 5,700 – 7,200 tok/s 98 – 128 tok/s 1.1 – 1.4 s
21k tokens 5,600 – 6,100 tok/s prose ~128, code 98 – 147 tok/s 3.8 s

For reference, the BF16 checkpoint on vLLM 0.29 with FP8 online quantization + MTP 2 measured about
53 / 69 tok/s decode and 3,000 – 3,600 tok/s prefill on the same machine.

Run

With Spark Studio: open the Models screen and launch Huihui Qwen3.6 35B-A3B abliterated.

By hand, with TensorFold Python 0.6.6 on CUDA (see Spark Studio's docker/tensorfold.Dockerfile):

hf download swdq/Huihui-Qwen3.6-35B-A3B-abliterated-MLX-4bit-MTP --local-dir ~/models/huihui-qwen36-mlx4-mtp
tensorfold serve ~/models/huihui-qwen36-mlx4-mtp --backend cuda --port 8110 \
  --name huihui-qwen3.6-abliterated --context 262144 --parallel 1 --mtp-drafts 2 --prefill-fp8

The server speaks OpenAI Chat Completions / Responses and Anthropic Messages. The weights also load
with mlx-lm on Apple Silicon (text only; mtp-4bit.safetensors is used by TensorFold).

How it was made

  • Source: huihui-ai/Huihui-Qwen3.6-35B-A3B-abliterated @ 8f0ee727aff5e771ea72466d64d13ecd851d2cc7 (BF16).
  • mlx_lm.convert -q --q-bits 4 --q-group-size 64 --dtype bfloat16 (mlx 0.32.3, mlx-lm 0.31.3, CPU):
    4-bit/group-64 affine weights, 8-bit router and shared-expert gates, 4.503 bits per weight.
    mlx-lm's text conversion drops the vision tower and the MTP layer.
  • The source's mtp.* tensors were kept: expert gate_up_proj split into gate_proj / up_proj,
    RMSNorm weights shifted by +1 (MLX convention), linear weights quantized with mlx.core.quantize
    (4-bit/group-64; 8-bit for router gates) and saved as mtp-4bit.safetensors. No external draft
    model is involved.

The exact script is in Spark Studio's src/tensorfold.rs (CONVERT_MTP).

Notes

  • Text only (no vision).
  • Abliterated models have reduced refusals. You are responsible for how you deploy and use them;
    add your own safeguards where needed.
  • License: Apache-2.0, inherited from Qwen3.6 and the huihui-ai checkpoint. Credit to the Qwen team,
    huihui-ai, the MLX team and TensorFold's authors.
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration