← back to catalog · registered 2026-10-01 22:58

IsValorum/Qwen3.8-35B-A3B-Distill-MLX-APEX-MiniPlus-Abliterated

IsValorum 35B MoE
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/IsValorum%2FQwen3.8-35B-A3B-Distill-MLX-APEX-MiniPlus-Abliterated"
Response includes
  • classification unknown
  • files 14
  • author_summary 8 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
?
Primary method

Unclassified

No clear signals of an abliteration technique in this model.
Confidence
UNKNOWN
Why this label 1 signal
No classification signals present. This may not be an abliterated model at all - it could be a repackaging, a merge with unrelated goals, or unrelated content that mentions the term.
  • no classification signals present (no abliterated, uncensored, or known producer/method markers)
Refusal direction extraction

No specific extraction method could be identified for this model. The producer either did not document it or used a proprietary pipeline.

What is a refusal direction? →
Downloads · 30-day
0
Likes
0
Model age
today
created 2026-10-01
Downloads over time
Now0→from0↑0%
00110 on Oct 10 on Oct 2Oct
Oct 1 → Oct 2 · 2 snapshots · spans 1 day

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

Tags
mlx safetensors qwen3_5_moe quantized moe qwen3.8 abliterated uncensored apex miniplus text-generation conversational

Related

Total size
17.5 GB
Files
14
Quantizations
1
Registered
2026-10-01 22:58
Last updated on HF
2026-10-02 00:51

Files by quantization

Auxiliary files 14 files 17.5 GB
model-00003-of-00004.safetensors 5.00 GB 29d6f00e download
model-00001-of-00004.safetensors 5.00 GB 6c59e9ab download
model-00002-of-00004.safetensors 4.90 GB dc09ac72 download
model-00004-of-00004.safetensors 2.63 GB a5dc1ffc download
tokenizer.json 19.1 MB a5cd9732 download
model.safetensors.index.json 161 KB 8f3488c2 download
config.json 129 KB b73e02a5 download
chat_template.jinja 7.58 KB a8755d82 download
README.md 5.03 KB 7183d187 download
.gitattributes 1.53 KB 52373fe2 download
tokenizer_config.json 1.20 KB 159b5c47 download
evaluation.json 900 B a8119321 download
quant_recipe.json 578 B ca42a32a download
generation_config.json 214 B a6d9d5ca download

README current version from Hugging Face


library_name: mlx
base_model: empero-ai/Qwen3.8-35B-A3B-Distill
pipeline_tag: text-generation
tags: [mlx, quantized, moe, qwen3.8, abliterated, apex, miniplus]

Qwen3.8 35B A3B Distill MLX APEX MiniPlus Abliterated

Text-only native MLX build of empero-ai/Qwen3.8-35B-A3B-Distill, using IsValorum's deterministic Abliterated recipe and the MLX APEX MiniPlus mixed-precision map.

This repository contains the MLX APEX MiniPlus Abliterated variant.

⚡ Quick Navigation Index

  1. Model profile
  2. MLX APEX MiniPlus tensor map
  3. WikiText-2 PPL estimate
  4. Abliteration recipe
  5. Validation and compatibility
  6. Usage
  7. Reproducibility
  8. Voluntary support

Model profile

Architecture Qwen3.8 35B A3B Distill (MoE)
Format MLX / Safetensors
Variant Abliterated
Recipe MLX APEX MiniPlus
Quality class Q5
Quantization Mixed-precision MLX affine
Group sizes G32 / G64
Modality Text only

MiniPlus uses precision selectively rather than applying one bit-width to every tensor. High-sensitivity paths are promoted or kept in floating point while less sensitive expert and linear-attention weights carry the size reduction.

MLX APEX MiniPlus tensor map

Component MLX precision
LM head 6-bit affine, G64
Token embeddings 4-bit affine, G64
Norms FP32
MoE routers FP32
Linear-attention in_proj_z 8-bit affine, G64
Shared experts 5-bit affine, G64
Full-attention Q/K/V 4-bit affine, G64
Full-attention o_proj 6-bit affine, G64
in_proj_a, A_log, conv1d, dt_bias FP32
Linear-attention in_proj_qkv / in_proj_b / out_proj 3-bit affine, G32
Routed experts, layers 0-9 and 30-39 3-bit affine, G32
Routed experts, layers 10-29 4-bit affine, G64

The middle routed-expert band is promoted to 4-bit affine G64. MLX affine is a different quantization family from GGUF K-quants/IQ quants, so the tensor map is an MLX-native translation rather than a bit-for-bit copy of the GGUF recipe.

WikiText-2 PPL estimate

Both values were measured in the same MLX-LM runtime on the same full WikiText-2 raw test split using non-overlapping 512-token blocks. PPL is an estimate and does not fully predict real-world quality.

For an apples-to-apples quantization comparison, the BF16 reference uses the same FP32-protected norms, MoE routers, and linear-attention state/control tensors as MiniPlus; all other unquantized floating weights remain at checkpoint precision.

Model PPL
Abliterated BF16 reference 8.822815
MLX APEX MiniPlus 9.649103
Delta +0.826288 (+9.365%)

Scored tokens: 296,380. Dataset revision: b08601e04326c79dfdd32d625aee71d232d685c3.

Abliteration recipe

The TPE search was not repeated. The previously selected deterministic recipe was applied directly.

Parameter Value
Target attn.o_proj
Direction scope per layer
max_weight 2.6323508125163686
max_weight_position 21.9566207717385
min_weight 1.643626880852864
min_weight_distance 21.506408702671163
Orthogonalization enabled
Row normalization full
Calibration 25 harmless + 25 harmful prompts

Source revision: bcc2dbe2f21b213625df2dc1a5a690212373af07.

Validation and compatibility

Conversion and PPL validation ran with MLX 0.32.3 on NVIDIA CUDA using an RTX PRO 6000. The saved artifact was reloaded, its native source chat template was rendered for user-only, system+user and multi-turn cases, and a finite-logit forward pass was verified.

This is not a macOS performance benchmark. It is a native MLX affine model intended for MLX-LM on Apple Silicon. MLX-LM's Qwen3.8 text path excludes vision-tower and MTP weights from this artifact.

MLX-LM conversion commit: 53b9af378278d6dd5dacac447edeb0f523b16253.

Usage

pip install -U mlx-lm

mlx_lm.generate \
  --model IsValorum/Qwen3.8-35B-A3B-Distill-MLX-APEX-MiniPlus-Abliterated \
  --prompt "Hello"

Server:

mlx_lm.server --model IsValorum/Qwen3.8-35B-A3B-Distill-MLX-APEX-MiniPlus-Abliterated

Reproducibility

  • evaluation.json contains the full WikiText-2 BF16 vs MiniPlus comparison.
  • quant_recipe.json contains the exact mixed-precision tensor map.

Voluntary support

Gold Ship dancing

If this release or my MiniPlus/NanoPlus work has been useful to you and you would like to support it, you can do so voluntarily through Ko-fi. Your contribution helps with evaluation, hosting, and future handcrafted quantizations. Every release will always remain free to download and use; there are no paywalled files, updates, or features.

Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.