license: apache-2.0
base_model:
- nex-agi/Nex-N2-mini
- huihui-ai/Huihui-Nex-N2-mini-abliterated
tags: - gguf
- apex
- apex-quant
- mtp
- nextn
- qwen3.5
- moE
- abliterated
- uncensored
- quantized
pipeline_tag: image-text-to-text
Huihui-Nex-N2-mini-abliterated — APEX I-Quality + MTP (GGUF)
GGUF quantizations of huihui-ai/Huihui-Nex-N2-mini-abliterated using the APEX (Adaptive Precision for EXpert Models) MoE-aware mixed-precision technique, with the built-in MTP (Multi-Token-Prediction) draft head retained for speculative decoding.
- Base arch:
qwen3_5_moe(Qwen3.5-style hybrid MoE: 40 layers, 256 experts / 8 active, linear-attention SSM + full-attention, MRoPE, multimodal vision projector) - Parameters: ~35.1B total (A3B active)
- MTP: built-in 1-layer nextn draft head (kept embedded in the GGUF → no separate draft model needed)
Files
Both MTP (built-in Multi-Token-Prediction draft head retained for speculative decoding) and non-MTP (clean, max compatibility) variants are provided:
| File | Tier | MTP | Size | BPW | imatrix | Notes |
|---|---|---|---|---|---|---|
Huihui-Nex-N2-mini-APEX-I-Quality-MTP.gguf |
APEX I-Quality | ✅ | ~24.5 GB | 5.52 | ✅ diverse | Best accuracy + speculative decoding (needs recent llama.cpp w/ bundled-MTP support) |
Huihui-Nex-N2-mini-APEX-Quality-MTP.gguf |
APEX Quality | ✅ | ~24.5 GB | 5.52 | ❌ | Lowest perplexity + speculative decoding |
Huihui-Nex-N2-mini-APEX-I-Quality-noMTP.gguf |
APEX I-Quality | ❌ | ~22.8 GB | 5.52 | ✅ diverse | Best accuracy, loads on any llama.cpp build |
Huihui-Nex-N2-mini-APEX-Quality-noMTP.gguf |
APEX Quality | ❌ | ~22.8 GB | 5.52 | ❌ | Lowest perplexity, max compatibility |
Which to pick?
- MTP variants → faster tok/s via speculative decoding, but need a recent
llama.cpp(bundled-MTP loader support, PR #22673 or newer;c1304d7confirmed working). - non-MTP variants → load on any
llama.cppbuild (incl. older), slightly smaller. Drop-in if your tool reportsmissing tensor 'blk.40.ssm_conv1d.weight'on the MTP file.
The MTP variants include the full draft block as blk.40.* + blk.40.nextn.{eh_proj,enorm,hnorm,shared_head_norm} tensors, with the MTP projection (eh_proj) kept at F16 and norms at F32 (per the APEX-MTP methodology: the draft head must stay high-precision or speculative-decoding acceptance drops). Non-MTP variants have block_count=40 (no blk.40.* / no nextn_predict_layers metadata).
APEX tier details (Quality profile)
APEX assigns precision per tensor role + per layer:
- Edge layers (L0-4, L35-39): routed experts
Q6_K - Near-edge (L5-9, L30-34): routed experts
Q5_K - Middle (L10-29): routed experts
IQ4_XS - Shared expert:
Q8_0(heavy-tailed, kurtosis ~13 → needs high precision) - Attention:
Q6_K - MTP draft head (blk.40.*):
F16(override)
I-variant (imatrix)
APEX-I-Quality is calibrated with a diverse importance matrix (chat + code + reasoning + tool-calling, no Wikipedia) rather than encyclopedic text. This trades a tiny wikitext perplexity increase for better real-world (assistant/coding/tool-use) accuracy and lower KL divergence. See the APEX technical report.
Note: the imatrix here is a locally-built diverse calibration set (~560 documents sampled from UltraChat, Alpaca, GSM8K chain-of-thought, Hermes function-calling, HumanEval code), not the private APEX
calibration_v1.2.txt. Methodology matches; results are not byte-identical to upstream APEX I-variants.
Loading / Inference
- Built-in MTP is load-bearing: use a recent
llama.cppbuild that supports bundled MTP / nextn (llama + spec: MTP SupportPR #22673 or newer —c1304d7confirmed working). Older builds fail withmissing tensor 'blk.40.ssm_conv1d.weight'. - MTP speculative decoding activates automatically on load → faster tok/s, no
-mddraft flag needed.
# MTP variant (needs recent llama.cpp with bundled-MTP support)
llama-server -m Huihui-Nex-N2-mini-APEX-I-Quality-MTP.gguf -ngl 12 -c 8192 --port 8080
# non-MTP variant (any llama.cpp build)
llama-server -m Huihui-Nex-N2-mini-APEX-I-Quality-noMTP.gguf -ngl 12 -c 8192 --port 8080
On a 12 GB VRAM GPU (e.g. RTX 3060), -ngl around 8-12 fits the ~23 GB model with partial CPU offload.
⚠️ Usage warnings (abliterated / uncensored)
This model is an abliterated (uncensored) derivative — its refusal direction has been removed via direction-ablation. The original huihui-ai warnings apply:
- Risk of sensitive/controversial outputs — safety filtering is significantly reduced.
- Not suitable for all audiences — outputs may be inappropriate for public settings, underage users, or high-security applications.
- Legal & ethical responsibility — ensure your usage complies with local laws. You are solely responsible for any consequences.
- Research / experimental use recommended — avoid unmonitored production or public-facing deployment.
- No default safety guarantees — this model has not undergone rigorous safety optimization. The uploader bears no responsibility for any consequences arising from its use.
Credits
- Original model:
nex-agi/Nex-N2-mini - Abliteration:
huihui-ai - Quantization method: APEX — Adaptive Precision for EXpert Models by the LocalAI team (Ettore Di Giacinto & Richard Palethorpe)
- Engine: llama.cpp
Donation
If this is useful, donations are appreciated:
BTC: bc1q6xxf0j3e7zn52cqrprc6gplql225wj8mnq75yw