← back to catalog · registered 2026-08-22 13:56

andrevp/Huihui-Nex-N2-mini-abliterated-APEX-MTP-GGUF

andrevp Qwen GGUF MoE multimodal 262K ctx
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/andrevp%2FHuihui-Nex-N2-mini-abliterated-APEX-MTP-GGUF"
Response includes
  • classification m8
  • files 6
  • hub_downloads_all_time 5,135
  • author_summary 2 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M8
Primary method

Repackaging (quantization)

Applied on top of direct removal inherited from the base model.
Confidence
MEDIUM
Inherited from base model
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=1
  • assume M1 (base ablation) + M8 (GGUF quant) - default when producer unknown
Refusal direction extraction

No specific extraction method could be identified for this model. The producer either did not document it or used a proprietary pipeline.

What is a refusal direction? →
Downloads · lifetime
5K
241 last 30d - cooling
Likes
3
Model age
3mo ago
created 2026-06-16
Downloads over time
Now5.2K→from3.3K↑60%
3.2K3.9K4.7K5.4K3.3K on Jun 175.2K on Oct 115.2K on Oct 10JunJulAugSepOct
Jun 17 → Oct 11 · 56 snapshots · spans 116 days

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
apache-2.0
Tags
gguf apex apex-quant mtp nextn qwen3.5 moE abliterated uncensored quantized image-text-to-text base_model:huihui-ai/Huihui-Nex-N2-mini-abliterated

Related

Total size
88.2 GB
Files
6
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-06-17 05:35

Files by quantization

Auxiliary files 6 files 88.2 GB
Huihui-Nex-N2-mini-APEX-I-Quality-MTP.gguf 22.8 GB 5645a27e download
Huihui-Nex-N2-mini-APEX-Quality-MTP.gguf 22.8 GB 451fac3c download
Huihui-Nex-N2-mini-APEX-I-Quality-noMTP.gguf 21.3 GB 4fab4e9b download
Huihui-Nex-N2-mini-APEX-Quality-noMTP.gguf 21.3 GB 56bba297 download
README.md 5.96 KB 3c0d842d download
.gitattributes 1.79 KB 6a7de5c3 download

README current version from Hugging Face


license: apache-2.0
base_model:

  • nex-agi/Nex-N2-mini
  • huihui-ai/Huihui-Nex-N2-mini-abliterated
    tags:
  • gguf
  • apex
  • apex-quant
  • mtp
  • nextn
  • qwen3.5
  • moE
  • abliterated
  • uncensored
  • quantized
    pipeline_tag: image-text-to-text

Huihui-Nex-N2-mini-abliterated — APEX I-Quality + MTP (GGUF)

GGUF quantizations of huihui-ai/Huihui-Nex-N2-mini-abliterated using the APEX (Adaptive Precision for EXpert Models) MoE-aware mixed-precision technique, with the built-in MTP (Multi-Token-Prediction) draft head retained for speculative decoding.

  • Base arch: qwen3_5_moe (Qwen3.5-style hybrid MoE: 40 layers, 256 experts / 8 active, linear-attention SSM + full-attention, MRoPE, multimodal vision projector)
  • Parameters: ~35.1B total (A3B active)
  • MTP: built-in 1-layer nextn draft head (kept embedded in the GGUF → no separate draft model needed)

Files

Both MTP (built-in Multi-Token-Prediction draft head retained for speculative decoding) and non-MTP (clean, max compatibility) variants are provided:

File Tier MTP Size BPW imatrix Notes
Huihui-Nex-N2-mini-APEX-I-Quality-MTP.gguf APEX I-Quality ✅ ~24.5 GB 5.52 ✅ diverse Best accuracy + speculative decoding (needs recent llama.cpp w/ bundled-MTP support)
Huihui-Nex-N2-mini-APEX-Quality-MTP.gguf APEX Quality ✅ ~24.5 GB 5.52 ❌ Lowest perplexity + speculative decoding
Huihui-Nex-N2-mini-APEX-I-Quality-noMTP.gguf APEX I-Quality ❌ ~22.8 GB 5.52 ✅ diverse Best accuracy, loads on any llama.cpp build
Huihui-Nex-N2-mini-APEX-Quality-noMTP.gguf APEX Quality ❌ ~22.8 GB 5.52 ❌ Lowest perplexity, max compatibility

Which to pick?

  • MTP variants → faster tok/s via speculative decoding, but need a recent llama.cpp (bundled-MTP loader support, PR #22673 or newer; c1304d7 confirmed working).
  • non-MTP variants → load on any llama.cpp build (incl. older), slightly smaller. Drop-in if your tool reports missing tensor 'blk.40.ssm_conv1d.weight' on the MTP file.

The MTP variants include the full draft block as blk.40.* + blk.40.nextn.{eh_proj,enorm,hnorm,shared_head_norm} tensors, with the MTP projection (eh_proj) kept at F16 and norms at F32 (per the APEX-MTP methodology: the draft head must stay high-precision or speculative-decoding acceptance drops). Non-MTP variants have block_count=40 (no blk.40.* / no nextn_predict_layers metadata).

APEX tier details (Quality profile)

APEX assigns precision per tensor role + per layer:

  • Edge layers (L0-4, L35-39): routed experts Q6_K
  • Near-edge (L5-9, L30-34): routed experts Q5_K
  • Middle (L10-29): routed experts IQ4_XS
  • Shared expert: Q8_0 (heavy-tailed, kurtosis ~13 → needs high precision)
  • Attention: Q6_K
  • MTP draft head (blk.40.*): F16 (override)

I-variant (imatrix)

APEX-I-Quality is calibrated with a diverse importance matrix (chat + code + reasoning + tool-calling, no Wikipedia) rather than encyclopedic text. This trades a tiny wikitext perplexity increase for better real-world (assistant/coding/tool-use) accuracy and lower KL divergence. See the APEX technical report.

Note: the imatrix here is a locally-built diverse calibration set (~560 documents sampled from UltraChat, Alpaca, GSM8K chain-of-thought, Hermes function-calling, HumanEval code), not the private APEX calibration_v1.2.txt. Methodology matches; results are not byte-identical to upstream APEX I-variants.

Loading / Inference

  • Built-in MTP is load-bearing: use a recent llama.cpp build that supports bundled MTP / nextn (llama + spec: MTP Support PR #22673 or newer — c1304d7 confirmed working). Older builds fail with missing tensor 'blk.40.ssm_conv1d.weight'.
  • MTP speculative decoding activates automatically on load → faster tok/s, no -md draft flag needed.
# MTP variant (needs recent llama.cpp with bundled-MTP support)
llama-server -m Huihui-Nex-N2-mini-APEX-I-Quality-MTP.gguf -ngl 12 -c 8192 --port 8080

# non-MTP variant (any llama.cpp build)
llama-server -m Huihui-Nex-N2-mini-APEX-I-Quality-noMTP.gguf -ngl 12 -c 8192 --port 8080

On a 12 GB VRAM GPU (e.g. RTX 3060), -ngl around 8-12 fits the ~23 GB model with partial CPU offload.

⚠️ Usage warnings (abliterated / uncensored)

This model is an abliterated (uncensored) derivative — its refusal direction has been removed via direction-ablation. The original huihui-ai warnings apply:

  • Risk of sensitive/controversial outputs — safety filtering is significantly reduced.
  • Not suitable for all audiences — outputs may be inappropriate for public settings, underage users, or high-security applications.
  • Legal & ethical responsibility — ensure your usage complies with local laws. You are solely responsible for any consequences.
  • Research / experimental use recommended — avoid unmonitored production or public-facing deployment.
  • No default safety guarantees — this model has not undergone rigorous safety optimization. The uploader bears no responsibility for any consequences arising from its use.

Credits

Donation

If this is useful, donations are appreciated:

BTC: bc1q6xxf0j3e7zn52cqrprc6gplql225wj8mnq75yw

README history 2 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-06-16Upload README.md with huggingface_hub2f0ea6e6 KB
    Loading...
  2. 2026-06-16Upload README.md with huggingface_hub956b82e4.8 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration