library_name: mlx
base_model: empero-ai/Qwen3.8-35B-A3B-Distill
pipeline_tag: text-generation
tags: [mlx, quantized, moe, qwen3.8, abliterated, apex, miniplus]
Qwen3.8 35B A3B Distill MLX APEX MiniPlus Abliterated
Text-only native MLX build of empero-ai/Qwen3.8-35B-A3B-Distill, using IsValorum's deterministic Abliterated recipe and the MLX APEX MiniPlus mixed-precision map.
This repository contains the MLX APEX MiniPlus Abliterated variant.
⚡ Quick Navigation Index
- Model profile
- MLX APEX MiniPlus tensor map
- WikiText-2 PPL estimate
- Abliteration recipe
- Validation and compatibility
- Usage
- Reproducibility
- Voluntary support
Model profile
| Architecture | Qwen3.8 35B A3B Distill (MoE) |
| Format | MLX / Safetensors |
| Variant | Abliterated |
| Recipe | MLX APEX MiniPlus |
| Quality class | Q5 |
| Quantization | Mixed-precision MLX affine |
| Group sizes | G32 / G64 |
| Modality | Text only |
MiniPlus uses precision selectively rather than applying one bit-width to every tensor. High-sensitivity paths are promoted or kept in floating point while less sensitive expert and linear-attention weights carry the size reduction.
MLX APEX MiniPlus tensor map
| Component | MLX precision |
|---|---|
| LM head | 6-bit affine, G64 |
| Token embeddings | 4-bit affine, G64 |
| Norms | FP32 |
| MoE routers | FP32 |
| Linear-attention in_proj_z | 8-bit affine, G64 |
| Shared experts | 5-bit affine, G64 |
| Full-attention Q/K/V | 4-bit affine, G64 |
| Full-attention o_proj | 6-bit affine, G64 |
| in_proj_a, A_log, conv1d, dt_bias | FP32 |
| Linear-attention in_proj_qkv / in_proj_b / out_proj | 3-bit affine, G32 |
| Routed experts, layers 0-9 and 30-39 | 3-bit affine, G32 |
| Routed experts, layers 10-29 | 4-bit affine, G64 |
The middle routed-expert band is promoted to 4-bit affine G64. MLX affine is a different quantization family from GGUF K-quants/IQ quants, so the tensor map is an MLX-native translation rather than a bit-for-bit copy of the GGUF recipe.
WikiText-2 PPL estimate
Both values were measured in the same MLX-LM runtime on the same full WikiText-2 raw test split using non-overlapping 512-token blocks. PPL is an estimate and does not fully predict real-world quality.
For an apples-to-apples quantization comparison, the BF16 reference uses the same FP32-protected norms, MoE routers, and linear-attention state/control tensors as MiniPlus; all other unquantized floating weights remain at checkpoint precision.
| Model | PPL |
|---|---|
| Abliterated BF16 reference | 8.822815 |
| MLX APEX MiniPlus | 9.649103 |
| Delta | +0.826288 (+9.365%) |
Scored tokens: 296,380. Dataset revision: b08601e04326c79dfdd32d625aee71d232d685c3.
Abliteration recipe
The TPE search was not repeated. The previously selected deterministic recipe was applied directly.
| Parameter | Value |
|---|---|
| Target | attn.o_proj |
| Direction scope | per layer |
| max_weight | 2.6323508125163686 |
| max_weight_position | 21.9566207717385 |
| min_weight | 1.643626880852864 |
| min_weight_distance | 21.506408702671163 |
| Orthogonalization | enabled |
| Row normalization | full |
| Calibration | 25 harmless + 25 harmful prompts |
Source revision: bcc2dbe2f21b213625df2dc1a5a690212373af07.
Validation and compatibility
Conversion and PPL validation ran with MLX 0.32.3 on NVIDIA CUDA using an RTX PRO 6000. The saved artifact was reloaded, its native source chat template was rendered for user-only, system+user and multi-turn cases, and a finite-logit forward pass was verified.
This is not a macOS performance benchmark. It is a native MLX affine model intended for MLX-LM on Apple Silicon. MLX-LM's Qwen3.8 text path excludes vision-tower and MTP weights from this artifact.
MLX-LM conversion commit: 53b9af378278d6dd5dacac447edeb0f523b16253.
Usage
pip install -U mlx-lm
mlx_lm.generate \
--model IsValorum/Qwen3.8-35B-A3B-Distill-MLX-APEX-MiniPlus-Abliterated \
--prompt "Hello"
Server:
mlx_lm.server --model IsValorum/Qwen3.8-35B-A3B-Distill-MLX-APEX-MiniPlus-Abliterated
Reproducibility
evaluation.jsoncontains the full WikiText-2 BF16 vs MiniPlus comparison.quant_recipe.jsoncontains the exact mixed-precision tensor map.
Voluntary support
If this release or my MiniPlus/NanoPlus work has been useful to you and you would like to support it, you can do so voluntarily through Ko-fi. Your contribution helps with evaluation, hosting, and future handcrafted quantizations. Every release will always remain free to download and use; there are no paywalled files, updates, or features.
