← back to catalog · registered 2026-08-22 13:56

leonsarmiento/Huihui-Nex-N2-mini-abliterated-5bit-XL-mlx

leonsarmiento 34B MoE multimodal second-order
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/leonsarmiento%2FHuihui-Nex-N2-mini-abliterated-5bit-XL-mlx"
Response includes
  • classification m1
  • files 14
  • hub_downloads_all_time 458
  • author_summary 7 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M1
Primary method

Direct removal

No other method signals detected in this model.
Confidence
MEDIUM
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=0 (base model)
  • no specific method indicators - defaulting to M1 (most common)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
MEDIUM
Why we say so
primary_method=M1; difference-of-means is the reference extraction for M1/M3 (Arditi 2024)
Downloads · lifetime
458
42 last 30d - cooling
Likes
0
Model age
3mo ago
created 2026-07-03
Downloads over time
Now481→from0↑0%
01763535290 on Jul 1481 on Oct 11481 on Oct 9JulAugSepOct
Jul 1 → Oct 11 · 54 snapshots · spans 102 days

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

Tags
mlx safetensors qwen3_5_moe mlx-vlm moe multimodal vision agentic coding abliterated uncensored basequant-xl

Related

Total size
24.0 GB
Files
14
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-08-04 13:52

Files by quantization

Auxiliary files 14 files 24.1 GB
model-00004-of-00005.safetensors 4.98 GB fc729253 download
model-00003-of-00005.safetensors 4.98 GB 6276f195 download
model-00002-of-00005.safetensors 4.98 GB 411ca29f download
model-00001-of-00005.safetensors 4.86 GB c56d503e download
model-00005-of-00005.safetensors 4.23 GB 417d23ce download
tokenizer.json 19.1 MB 87a7830d download
model.safetensors.index.json 192 KB 1bb003a6 download
config.json 136 KB c6845dfe download
tokenizer_config.json 8.94 KB 0d652c24 download
chat_template.jinja 7.57 KB fa6e2772 download
README.md 4.00 KB fd95c453 download
.gitattributes 1.53 KB 52373fe2 download
processor_config.json 1.27 KB 331fb318 download
preprocessor_config.json 390 B 2ea84a43 download

README current version from Hugging Face


library_name: mlx
base_model: huihui-ai/Huihui-Nex-N2-mini-abliterated
tags:

  • mlx
  • mlx-vlm
  • moe
  • multimodal
  • vision
  • agentic
  • coding
  • abliterated
  • uncensored
  • basequant-xl
    pipeline_tag: image-text-to-text

leonsarmiento/Huihui-Nex-N2-mini-abliterated-5bit-XL-mlx

This model was converted to MLX format from huihui-ai/Huihui-Nex-N2-mini-abliterated using BaseQuant_XL 5/8-bit mixed quantization optimized for Apple Silicon. The vision encoder is preserved and quantized at 5-bit, making this a full multimodal model.

BaseQuant_XL keeps the most routing-critical layers in full bf16 precision — the MoE router gate, shared expert gate, shared expert, and lm_head — while applying aggressive quantization to the bulk parameters. This preserves routing accuracy and output quality where it matters most.

This is the abliterated (uncensored) version of Nex-N2-mini, a 35B-parameter MoE (Mixture of Experts) model fine-tuned from Qwen3.5-35B-A3B-Base by Nex-AGI, featuring 256 experts (8 active per token + 1 shared expert), hybrid full + linear (Gated DeltaNet) attention, an "Agentic Thinking" framework (Adaptive Thinking + Coherent Thinking), and a vision encoder for multimodal input. Despite 35B total parameters, only ~3B are activated per token for efficient inference.

About XL Quantization

BaseQuant_XL is a fully data-agnostic, static quantization. No calibration dataset, no sensitivity analysis, no importance matrix. Precision is allocated purely by architectural role — routing-critical layers get higher precision, bulk expert parameters get lower precision. The result is a transparent, faithful capture of the source model.

Data-dependent calibration quantizations (iMatrix, AWQ, GPTQ, oQ, oQ4e, etc.) use a calibration set to guide bit allocation. This can produce a skewed representation of the model: domains well-represented in the calibration data (English, popular topics, public or leaked benchmarks) are preserved better, while underrepresented domains (non-English languages, niche use cases, your own data) are preserved worse. XL avoids this trade-off entirely — it generalizes honestly because it is never fit to any particular data distribution.

Use with mlx

pip install -U mlx-vlm
python -m mlx_vlm.generate --model leonsarmiento/Huihui-Nex-N2-mini-abliterated-5bit-XL-mlx --max-tokens 256 --temperature 0.7 --top-p 0.95 --prompt "Hello"

BaseQuant_XL Quantization Strategy

Bit Depth Layers Rationale
bf16 (unquantized) mlp.gate (router), shared_expert_gate, lm_head, shared_expert Routing decisions and shared computation path — errors here are qualitatively different from precision loss
8-bit embed_tokens, self_attn (full attention), linear_attn (DeltaNet) Every-token layers with moderate sensitivity — 8-bit is near-lossless
5-bit vision_tower, switch_mlp (routed experts) Bulk of parameters, only 8 of 256 experts active per token — natural redundancy tolerates lower precision

Quantization Details

Layer Bits Group Size
mlp.gate (router) bf16 —
shared_expert_gate bf16 —
lm_head bf16 —
shared_expert bf16 —
embed_tokens 8 64
self_attn (full attention) 8 64
linear_attn (DeltaNet) 8 64
vision_tower 5 64
switch_mlp (routed experts) 5 64
Default fallback 8 64
  • Quantization type: BaseQuant_XL mixed (multimodal, vision preserved)
  • Bits per weight: 5.881
  • Total size: ~24 GB (5 shards)
  • Group size: 64
  • Method: Custom quant_predicate via mlx_vlm

Recommended Inference Parameters

Parameter Value
temperature 0.7
top_p 0.95
top_k 40
min_p 0.01
repeat_penalty 1.0

Note: This is a Qwen3.5-based model — preserve_thinking is not applicable.

README history 3 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-08-04Add standardized 'About XL Quantization' paragraph06702394 KB
    Loading...
  2. 2026-07-22Upload README.md with huggingface_hub5f6cbe93.1 KB
    Loading...
  3. 2026-07-03Upload folder using huggingface_hubc1978f03.1 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration