license: apache-2.0
tags:
- uncensored
- abliterated
- qwen3.5
- moe
- a3b
- mlx
- 4-bit
language: - en
- zh
- multilingual
pipeline_tag: text-generation
base_model: Li101/Qwen3.5-35B-A3B-Uncensored-Aggressive-safetensors
library_name: mlx
Qwen3.5-35B-A3B-Uncensored-Aggressive-mlx-4bit
A 4-bit MLX quantization of Li101/Qwen3.5-35B-A3B-Uncensored-Aggressive-safetensors, converted for fast local inference on Apple Silicon.
This is a 35B-parameter A3B Mixture-of-Experts model (~3B active parameters per token), which makes it run considerably faster than a dense 35B while keeping a large effective capacity.
Conversion
Converted with mlx-lm:
mlx_lm.convert \
--hf-path Li101/Qwen3.5-35B-A3B-Uncensored-Aggressive-safetensors \
--mlx-path Qwen3.5-35B-A3B-Uncensored-Aggressive-mlx-4bit \
-q --q-bits 4 --q-group-size 64
Result: 4.503 bits per weight, ~18.2 GB on disk.
Notes / caveats
- Text-only. Although the base repo carries vision-tower weights, this MLX build loads as a text-only LLM and vision was not verified working. Treat it as a text model or test further.
- Uncensored / abliterated. Refusal behavior is heavily reduced. You are responsible for how you use it.
Local validation (Apple Silicon)
Tested on MacBook Pro M3 36GB (shows 39 GB usable memory) served via oMLX with an 8-bit KV cache and a 64K context window.
- Memory stability: held stable across a long 10-phase iterative-coding stress run — usage oscillated ~78–84% with 0 bytes of swap throughout; server process resident ~20 GB.
- Instruction following: reliably obeys repeated per-step instructions and completes multi-part single-response tasks without splitting the work or asking to continue.
- Known weakness: over very long context it can drop a single deferred instruction (e.g. a "print this token only at the very end" directive) even while every per-step instruction is honored. If you need a guaranteed end-of-run action, pair it with an agent that re-asserts the instruction, or restate it near the end.
Recommended runtime config
For long-context stability on ~36 GB-class Apple Silicon:
- Quantized (8-bit) KV cache
- 64K context
- Wired-memory limit raised from 25GB to 30GB to give the model headroom
License
Apache-2.0, inherited from the base model. Credit to Li101 for the original fine-tune.