← back to catalog · registered 2026-10-01 23:58

IsValorum/Qwen3.8-35B-A3B-Distill-MLX-APEX-NanoPlus-Abliterated

IsValorum 35B MoE
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/IsValorum%2FQwen3.8-35B-A3B-Distill-MLX-APEX-NanoPlus-Abliterated"
Response includes
  • classification unknown
  • files 14
  • author_summary 8 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
?
Primary method

Unclassified

No clear signals of an abliteration technique in this model.
Confidence
UNKNOWN
Why this label 1 signal
No classification signals present. This may not be an abliterated model at all - it could be a repackaging, a merge with unrelated goals, or unrelated content that mentions the term.
  • no classification signals present (no abliterated, uncensored, or known producer/method markers)
Refusal direction extraction

No specific extraction method could be identified for this model. The producer either did not document it or used a proprietary pipeline.

What is a refusal direction? →
Downloads · 30-day
0
Likes
1
Model age
today
created 2026-10-01
Downloads over time
Now0→from0↑0%
00110 on Oct 10 on Oct 2Oct
Oct 1 → Oct 2 · 2 snapshots · spans 1 day

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

Tags
mlx safetensors qwen3_5_moe quantized moe qwen3.8 abliterated uncensored apex nanoplus text-generation conversational

Related

Total size
17.5 GB
Files
14
Quantizations
1
Registered
2026-10-01 23:58
Last updated on HF
2026-10-02 00:43

Files by quantization

Auxiliary files 14 files 17.5 GB
model-00003-of-00004.safetensors 5.00 GB 29d6f00e download
model-00001-of-00004.safetensors 5.00 GB 6c59e9ab download
model-00002-of-00004.safetensors 4.90 GB dc09ac72 download
model-00004-of-00004.safetensors 2.63 GB a5dc1ffc download
tokenizer.json 19.1 MB a5cd9732 download
model.safetensors.index.json 161 KB 8f3488c2 download
config.json 129 KB b73e02a5 download
chat_template.jinja 7.58 KB a8755d82 download
README.md 5.35 KB 7cd79777 download
.gitattributes 1.53 KB 52373fe2 download
tokenizer_config.json 1.20 KB 159b5c47 download
evaluation.json 900 B 15fff287 download
quant_recipe.json 578 B ca42a32a download
generation_config.json 214 B a6d9d5ca download

README current version from Hugging Face


library_name: mlx
base_model: empero-ai/Qwen3.8-35B-A3B-Distill
pipeline_tag: text-generation
tags: [mlx, quantized, moe, qwen3.8, abliterated, apex, nanoplus]

Qwen3.8 35B A3B Distill MLX APEX NanoPlus Abliterated

Text-only native MLX build of empero-ai/Qwen3.8-35B-A3B-Distill, using IsValorum's deterministic Abliterated recipe and the compact MLX APEX NanoPlus mixed-precision map.

NanoPlus is the smaller member of the MLX APEX Abliterated pair. Its Safetensors weights are 18.8156 GB and the repository artifact is approximately 18.836 GB.

Quick Navigation Index

  1. Model profile
  2. MLX APEX NanoPlus tensor map
  3. WikiText-2 PPL estimate
  4. Abliteration recipe
  5. Validation and compatibility
  6. Usage
  7. Reproducibility
  8. Sibling MiniPlus build
  9. Voluntary support

Model profile

Architecture Qwen3.8 35B A3B Distill (MoE)
Format MLX / Safetensors
Variant Abliterated
Recipe MLX APEX NanoPlus
Quantization Mixed-precision MLX affine
Base quantization 3-bit G32
Group sizes G32 / G64
Safetensors size 18.8156 GB
Modality Text only

NanoPlus uses a 3-bit G32 base selectively, with higher precision retained on attention, shared experts, embeddings, the LM head and protected control paths. The label shown by the Hub reflects the base MLX quantization; it does not mean every tensor is 3-bit.

MLX APEX NanoPlus tensor map

Component MLX precision
LM head 6-bit affine, G64
Token embeddings 4-bit affine, G64
Norms FP32
MoE routers FP32
Linear-attention in_proj_z 8-bit affine, G64
Shared experts 5-bit affine, G64
Full-attention Q/K/V 4-bit affine, G64
Full-attention o_proj 6-bit affine, G64
in_proj_a, A_log, conv1d, dt_bias FP32
Linear-attention in_proj_qkv / in_proj_b / out_proj 3-bit affine, G32
Routed experts, layers 0-9 and 30-39 3-bit affine, G32
Routed experts, layers 10-29 4-bit affine, G64

MLX affine is a different quantization family from GGUF K-quants/IQ quants. This is an MLX-native mixed-precision map, not a GGUF quant copied bit-for-bit.

WikiText-2 PPL estimate

The BF16 reference and NanoPlus were measured in the same MLX-LM runtime on the same full WikiText-2 raw test split using non-overlapping 512-token blocks. PPL is an estimate and does not fully predict real-world quality.

Model PPL
Abliterated BF16 reference 8.822815
MLX APEX NanoPlus 9.649103
Delta +0.826288 (+9.365%)

Scored tokens: 296,380. Dataset revision: b08601e04326c79dfdd32d625aee71d232d685c3.

The BF16 reference uses the same FP32-protected norms, MoE routers, and linear-attention state/control tensors as NanoPlus; all other unquantized floating weights remain at checkpoint precision.

Abliteration recipe

The TPE search was not repeated. The previously selected deterministic recipe was applied directly.

Parameter Value
Target attn.o_proj
Direction scope per layer
max_weight 2.6323508125163686
max_weight_position 21.9566207717385
min_weight 1.643626880852864
min_weight_distance 21.506408702671163
Orthogonalization enabled
Row normalization full
Calibration 25 harmless + 25 harmful prompts

Source revision: bcc2dbe2f21b213625df2dc1a5a690212373af07.

Validation and compatibility

Conversion and PPL validation ran with MLX 0.32.3 on NVIDIA CUDA using an RTX PRO 6000. The saved artifact was reloaded, its native source chat template was rendered for user-only, system+user and multi-turn cases, and a finite-logit forward pass was verified.

This is not a macOS performance benchmark. It is a native MLX affine model intended for MLX-LM on Apple Silicon. MLX-LM's Qwen3.8 text path excludes vision-tower and MTP weights from this artifact.

MLX-LM conversion commit: 53b9af378278d6dd5dacac447edeb0f523b16253.

Usage

pip install -U mlx-lm

mlx_lm.generate   --model IsValorum/Qwen3.8-35B-A3B-Distill-MLX-APEX-NanoPlus-Abliterated   --prompt "Hello"

Server:

mlx_lm.server --model IsValorum/Qwen3.8-35B-A3B-Distill-MLX-APEX-NanoPlus-Abliterated

Reproducibility

  • evaluation.json contains the full WikiText-2 BF16 vs NanoPlus comparison.
  • quant_recipe.json contains the exact mixed-precision tensor map.

Sibling MiniPlus build

For the higher-precision build under the 20 GB target, see MLX APEX MiniPlus Abliterated.

Voluntary support

Gold Ship dancing

If this release or my MiniPlus/NanoPlus work has been useful to you and you would like to support it, you can do so voluntarily through Ko-fi. Your contribution helps with evaluation, hosting, and future handcrafted quantizations. Every release will always remain free to download and use; there are no paywalled files, updates, or features.

Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.