library_name: mlx
base_model: empero-ai/Qwen3.8-35B-A3B-Distill
pipeline_tag: text-generation
tags: [mlx, quantized, moe, qwen3.8, abliterated, apex, nanoplus]
Qwen3.8 35B A3B Distill MLX APEX NanoPlus Abliterated
Text-only native MLX build of empero-ai/Qwen3.8-35B-A3B-Distill, using IsValorum's deterministic Abliterated recipe and the compact MLX APEX NanoPlus mixed-precision map.
NanoPlus is the smaller member of the MLX APEX Abliterated pair. Its Safetensors weights are 18.8156 GB and the repository artifact is approximately 18.836 GB.
Quick Navigation Index
- Model profile
- MLX APEX NanoPlus tensor map
- WikiText-2 PPL estimate
- Abliteration recipe
- Validation and compatibility
- Usage
- Reproducibility
- Sibling MiniPlus build
- Voluntary support
Model profile
| Architecture | Qwen3.8 35B A3B Distill (MoE) |
| Format | MLX / Safetensors |
| Variant | Abliterated |
| Recipe | MLX APEX NanoPlus |
| Quantization | Mixed-precision MLX affine |
| Base quantization | 3-bit G32 |
| Group sizes | G32 / G64 |
| Safetensors size | 18.8156 GB |
| Modality | Text only |
NanoPlus uses a 3-bit G32 base selectively, with higher precision retained on attention, shared experts, embeddings, the LM head and protected control paths. The label shown by the Hub reflects the base MLX quantization; it does not mean every tensor is 3-bit.
MLX APEX NanoPlus tensor map
| Component | MLX precision |
|---|---|
| LM head | 6-bit affine, G64 |
| Token embeddings | 4-bit affine, G64 |
| Norms | FP32 |
| MoE routers | FP32 |
| Linear-attention in_proj_z | 8-bit affine, G64 |
| Shared experts | 5-bit affine, G64 |
| Full-attention Q/K/V | 4-bit affine, G64 |
| Full-attention o_proj | 6-bit affine, G64 |
| in_proj_a, A_log, conv1d, dt_bias | FP32 |
| Linear-attention in_proj_qkv / in_proj_b / out_proj | 3-bit affine, G32 |
| Routed experts, layers 0-9 and 30-39 | 3-bit affine, G32 |
| Routed experts, layers 10-29 | 4-bit affine, G64 |
MLX affine is a different quantization family from GGUF K-quants/IQ quants. This is an MLX-native mixed-precision map, not a GGUF quant copied bit-for-bit.
WikiText-2 PPL estimate
The BF16 reference and NanoPlus were measured in the same MLX-LM runtime on the same full WikiText-2 raw test split using non-overlapping 512-token blocks. PPL is an estimate and does not fully predict real-world quality.
| Model | PPL |
|---|---|
| Abliterated BF16 reference | 8.822815 |
| MLX APEX NanoPlus | 9.649103 |
| Delta | +0.826288 (+9.365%) |
Scored tokens: 296,380. Dataset revision: b08601e04326c79dfdd32d625aee71d232d685c3.
The BF16 reference uses the same FP32-protected norms, MoE routers, and linear-attention state/control tensors as NanoPlus; all other unquantized floating weights remain at checkpoint precision.
Abliteration recipe
The TPE search was not repeated. The previously selected deterministic recipe was applied directly.
| Parameter | Value |
|---|---|
| Target | attn.o_proj |
| Direction scope | per layer |
| max_weight | 2.6323508125163686 |
| max_weight_position | 21.9566207717385 |
| min_weight | 1.643626880852864 |
| min_weight_distance | 21.506408702671163 |
| Orthogonalization | enabled |
| Row normalization | full |
| Calibration | 25 harmless + 25 harmful prompts |
Source revision: bcc2dbe2f21b213625df2dc1a5a690212373af07.
Validation and compatibility
Conversion and PPL validation ran with MLX 0.32.3 on NVIDIA CUDA using an RTX PRO 6000. The saved artifact was reloaded, its native source chat template was rendered for user-only, system+user and multi-turn cases, and a finite-logit forward pass was verified.
This is not a macOS performance benchmark. It is a native MLX affine model intended for MLX-LM on Apple Silicon. MLX-LM's Qwen3.8 text path excludes vision-tower and MTP weights from this artifact.
MLX-LM conversion commit: 53b9af378278d6dd5dacac447edeb0f523b16253.
Usage
pip install -U mlx-lm
mlx_lm.generate --model IsValorum/Qwen3.8-35B-A3B-Distill-MLX-APEX-NanoPlus-Abliterated --prompt "Hello"
Server:
mlx_lm.server --model IsValorum/Qwen3.8-35B-A3B-Distill-MLX-APEX-NanoPlus-Abliterated
Reproducibility
- evaluation.json contains the full WikiText-2 BF16 vs NanoPlus comparison.
- quant_recipe.json contains the exact mixed-precision tensor map.
Sibling MiniPlus build
For the higher-precision build under the 20 GB target, see MLX APEX MiniPlus Abliterated.
Voluntary support
If this release or my MiniPlus/NanoPlus work has been useful to you and you would like to support it, you can do so voluntarily through Ko-fi. Your contribution helps with evaluation, hosting, and future handcrafted quantizations. Every release will always remain free to download and use; there are no paywalled files, updates, or features.
