license: apache-2.0
library_name: mlx
tags:
- mlx
- omlx
- qwen3_5_moe
- abliterated
- uncensored
- 3-bit
base_model: Qwen3.6-35B-A3B-uncensored-heretic-Native-MTP-Preserved
pipeline_tag: text-generation
Qwen3.6-35B-A3B Heretic — oQ3 (3-bit, MLX)
A sensitivity-guided ~3-bit (oQ3) quant of the abliterated Qwen3.6-35B-A3B "Heretic" (Native-MTP-Preserved) model, built on-device with omlx's mixed-precision oQ quantizer. MoE arch qwen3_5_moe (35B total / ~3B active). Apple-Silicon MLX format.
- Effective precision: ~3.6 bpw (≈16 GB weights) —
oQkeeps sensitive layers higher-bit, so it punches above its nominal bit-width. - Abliterated / uncensored. Use responsibly; you are accountable for your outputs.
Why this quant
Despite being the smallest/fastest quant of the family, it matched the higher-bit builds on every benchmark tried (M4 Max, thinking modes as noted):
| Test | Score |
|---|---|
| Hard reasoning + code (8 tasks) | 8/8 |
| Harder quality (multi-digit math, DP) (6) | 6/6 |
| General knowledge (20 facts) | 20/20 |
| Multi-step agentic tool-use (5) | 5/5 |
| Agentic "gauntlet" — flaky-tool retry, traps, branch (7, thinking-OFF) | 7/7 |
| Throughput | ~100+ tok/s |
Run it
Serve with omlx (or any MLX-LM runtime) on Apple Silicon:
omlx serve --port 8000
# then request model "Qwen3.6-35B-A3B-uncensored-heretic-Native-MTP-Preserved-oQ3"
Best as an agent default with thinking OFF (cleanest tool-discipline); flip thinking ON for hard multi-step reasoning.
Private quant for personal use. License inherits from the base model.