← back to catalog · registered 2026-08-22 13:56

sakamakismile/Huihui-Nex-N2-mini-abliterated-MTP-NVFP4

sakamakismile 17B MoE multimodal second-order
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/sakamakismile%2FHuihui-Nex-N2-mini-abliterated-MTP-NVFP4"
Response includes
  • classification m1
  • files 11
  • hub_downloads_all_time 844
  • author_summary 34 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M1
Primary method

Direct removal

No other method signals detected in this model.
Confidence
MEDIUM
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=0 (base model)
  • no specific method indicators - defaulting to M1 (most common)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
MEDIUM
Why we say so
primary_method=M1; difference-of-means is the reference extraction for M1/M3 (Arditi 2024)
Downloads · lifetime
844
51 last 30d - cooling
Likes
6
Model age
3mo ago
created 2026-06-16
Downloads over time
Now866→from161↑438%
126396666937161 on Jun 17866 on Oct 11866 on Oct 10JunJulAugSepOct
Jun 17 → Oct 11 · 56 snapshots · spans 116 days

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
apache-2.0
Languages
en ja
Tags
vllm safetensors qwen3_5_moe nvfp4 w4a4 compressed-tensors llm-compressor abliterated uncensored blackwell sm120 speculative-decoding

Related

Total size
20.4 GB
Files
11
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-06-16 04:48

Files by quantization

Auxiliary files 11 files 20.4 GB
model.safetensors 20.4 GB 4643a56e download
tokenizer.json 19.1 MB 128d23d7 download
config.json 13.6 KB e02eb00b download
README.md 8.70 KB 43c489df download
chat_template.jinja 7.57 KB fa6e2772 download
.gitattributes 1.53 KB 52373fe2 download
processor_config.json 1.27 KB 15a94ef1 download
tokenizer_config.json 1.14 KB 5344df8c download
preprocessor_config.json 390 B 2ea84a43 download
recipe.yaml 237 B 2ed68f09 download
generation_config.json 115 B 86ce6bc5 download

README current version from Hugging Face


base_model:

  • huihui-ai/Huihui-Nex-N2-mini-abliterated
  • nex-agi/Nex-N2-mini
    base_model_relation: quantized
    language:
  • en
  • ja
    tags:
  • qwen3_5_moe
  • nvfp4
  • w4a4
  • compressed-tensors
  • llm-compressor
  • abliterated
  • uncensored
  • vllm
  • blackwell
  • sm120
  • speculative-decoding
  • mtp
  • vision-language
    library_name: vllm
    pipeline_tag: image-text-to-text
    license: apache-2.0
    quantized_by: Lna-Lab

Huihui-Nex-N2-mini-abliterated-MTP-NVFP4

NVFP4 (W4A4) quantization of huihui-ai/Huihui-Nex-N2-mini-abliterated — the abliterated Nex-N2-mini (a 35B-A3B qwen3_5_moe hybrid-linear Vision-Language MoE) — packed to ~23.6 GB (from ~72 GB BF16). The vision tower is kept in BF16 and the native MTP draft ships in MTP/, so a single hf download gives you an uncensored VLM and speculative decoding.

Made by quantizing the huihui-ai abliterated release with llm-compressor + compressed-tensors. All credit for the model itself goes to huihui-ai (abliteration) and nex-agi (the original Nex-N2-mini); this repo only adds the FP4 quantization. Please review the original models' terms before use.

Lineage: nex-agi/Nex-N2-mini → huihui-ai abliteration → Lna-Lab NVFP4 W4A4 (this repo).

A text-only sibling (vision tower dropped, ~22.7 GB) is at Huihui-Nex-N2-mini-abliterated-text-MTP-NVFP4.

What it is

Architecture Qwen3_5MoeForConditionalGeneration (qwen3_5_moe) — Qwen3.5-VL-MoE family
Total / active 35B total · ~3B active (256 experts, top-8, + shared expert)
Attention hybrid: linear-attention (GatedDeltaNet-style SSM) ×3 → full-attention every 4th layer (40 layers)
Vision Qwen3-VL ViT (depth 27, patch 16, spatial-merge 2) — kept BF16
MTP native multi-token-prediction draft (mtp_num_hidden_layers=1), shipped in MTP/ (BF16, ~1.69 GB)
Context 262,144 positions (mRoPE)
Quantization NVFP4 nvfp4-pack-quantized, W4A4, group size 16, FP8-E4M3 scales
What's quantized all language-model Linear layers incl. 30,720 expert projections; lm_head / visual.* / MoE router (mlp.gate, mlp.shared_expert_gate) / norms / conv kept BF16
Size ~23.6 GB (model.safetensors + MTP/)
Nature abliterated / uncensored · reasoning model (emits <think>…</think>)

Serving with vLLM

Requires vLLM 0.22.x (native qwen3_5_moe + qwen3_5_mtp; compressed-tensors NVFP4 auto-detected — no --quantization flag) and a Blackwell (SM120) GPU. TP=4 (4× 16 GB) is the floor — the ~9.9 GB/GPU weights plus the MoE NVFP4 GEMM workspace do not fit TP=2 on 16 GB cards.

# 4× 16 GB Blackwell. On a box WITHOUT NVLink/P2P the NCCL flag + --disable-custom-all-reduce
# are MANDATORY (else it hangs at the first all-reduce). MTP is auto-detected from MTP/ —
# do NOT pass a model path in --speculative-config.
NCCL_P2P_DISABLE=1 \
vllm serve sakamakismile/Huihui-Nex-N2-mini-abliterated-MTP-NVFP4 \
  --tensor-parallel-size 4 \
  --disable-custom-all-reduce \
  --max-model-len 8192 \
  --gpu-memory-utilization 0.87 \
  --kv-cache-dtype fp8 \
  --speculative-config '{"method":"qwen3_5_mtp","num_speculative_tokens":3}' \
  --reasoning-parser qwen3 \
  --port 8000

Benchmarks

Measured on 4× RTX PRO 2000 Blackwell 16 GB (vLLM 0.22.0, TP=4, MTP n=3, KV fp8, greedy):

Metric Result
Single-stream decode 69 tok/s
Aggregate throughput @ 4 concurrent 151 tok/s
Aggregate throughput @ 8 concurrent 287 tok/s
HumanEval pass@1 83.5% (137/164)

KV pool ~798k tokens (97× concurrency at 8 192 ctx), ~14.4 GB/GPU. Bilingual (JA/EN) output coherent, arithmetic correct — no W4A4 collapse. 83.5% pass@1 on a 35B-A3B model with only ~3B active parameters is strong for the active-parameter budget. HumanEval was run via chat with enable_thinking=false (this is a reasoning model — in thinking mode it scores comparably, ~80%, but spends far more tokens and occasionally over-thinks short problems). The language weights are byte-identical to the text-only sibling, so these results apply to both.

  • It is a reasoning model: give it generous max_tokens (even short answers spend hundreds of tokens in <think>), and use --reasoning-parser qwen3 to split reasoning_content from content.
  • For text-only use on this VL checkpoint, add --limit-mm-per-prompt '{"image":0,"video":0}'. To use images, drop that flag.

Recommended usage: persona + thinking control

This is a reasoning model with a per-request thinking switch, and it responds strongly to a system-prompt persona. We benchmarked both levers on a 16-task verifiable battery (multi-step math, logic traps, instruction-following, code, JSON extraction, JP character-counting). The combination is what to ship:

Configuration Score Relative speed
non-thinking, no system prompt 13/16 fastest
non-thinking + persona 14/16 fast — recommended default
thinking, no system prompt 15/16 slow
thinking + persona 16/16 slowest

Recommended default — non-thinking + a careful-generalist persona. Fast and cheap, and it cleanly handles math, logic, code, JSON and bilingual writing:

{
  "messages": [
    {"role": "system", "content": "あなたは自己内省的で慎重な性格かつ、あらゆる分野に精通した回答ができる汎用人工知性体です。"},
    {"role": "user", "content": "..."}
  ],
  "chat_template_kwargs": {"enable_thinking": false}
}

Thinking control

  • Thinking is governed by chat_template_kwargs.enable_thinking: omit it (or true) to reason, false to answer directly. Default (no flag) = thinking on, and the <think>…</think> block is auto-stripped from content — you always get a clean answer.
  • The persona lifts accuracy in both modes (it recovered borderline code-gen and tightened outputs), so keep it on regardless of the thinking flag.
  • Escalate to enable_thinking: true only for hidden-enumeration-under-terse-output tasks — e.g. "count the X and reply with only the number", exact word/character counts. Non-thinking has no hidden scratchpad: when "be terse" and "count carefully" conflict, terseness wins and the count drifts. Thinking counts in the hidden block, then prints the clean answer → those tasks go 14/16 → 16/16.
  • Even without flipping the flag, asking the model to show its work ("list each item, then count") recovers most counting tasks in non-thinking mode, since the reasoning then happens in the visible answer.
  • This chat template has no /think /no_think keyword parsing — drive thinking via the enable_thinking flag (or an app-layer router that maps user intent → the flag).

Quantization recipe

llm-compressor one-shot, QuantizationModifier(targets="Linear", scheme="NVFP4", ignore=["lm_head","re:.*visual.*","re:.*mlp.gate$","re:.*mlp.shared_expert_gate$"]), 32 calibration samples (neuralmagic/calibration, seq 8192). The full model is loaded via AutoModelForImageTextToText so the vision tower is present and preserved in BF16; the per-expert nn.Linear projections are packed to NVFP4. The native MTP draft is carried over verbatim in MTP/.

License

Inherits apache-2.0 from the upstream models. Abliterated/uncensored: you are responsible for how you use it.

Credits

Support the Base Model Author (huihui-ai)

If you find the abliterated base useful, please support huihui-ai — this repo only adds the FP4 quantization; the abliteration work is theirs:

README history 3 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-06-16Add recommended usage: persona + thinking-control recipe (benchmarked)0e8f1ef8.7 KB
    Loading...
  2. 2026-06-16Add huihui-ai credit/support footer + benchmarks04122376.4 KB
    Loading...
  3. 2026-06-16NVFP4 W4A4 (llm-compressor) of huihui-ai/Huihui-Nex-N2-mini-abliterated + MTP1474b904.9 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration