← back to catalog · registered 2026-08-22 13:56

sakamakismile/Huihui-Nex-N2-mini-abliterated-text-MTP-NVFP4

sakamakismile 17B MoE second-order
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/sakamakismile%2FHuihui-Nex-N2-mini-abliterated-text-MTP-NVFP4"
Response includes
  • classification m1
  • files 11
  • hub_downloads_all_time 86
  • author_summary 34 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M1
Primary method

Direct removal

No other method signals detected in this model.
Confidence
MEDIUM
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=0 (base model)
  • no specific method indicators - defaulting to M1 (most common)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
MEDIUM
Why we say so
primary_method=M1; difference-of-means is the reference extraction for M1/M3 (Arditi 2024)
Downloads · lifetime
86
17 last 30d - stable
Likes
0
Model age
3mo ago
created 2026-06-16
Downloads over time
Now94→from41↑129%
3859799941 on Jun 1794 on Oct 11JunJulAugSepOct
Jun 17 → Oct 11 · 56 snapshots · spans 116 days

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
apache-2.0
Languages
en ja
Tags
vllm safetensors qwen3_5_moe_text qwen3_5_moe nvfp4 w4a4 compressed-tensors llm-compressor abliterated uncensored blackwell sm120

Related

Total size
19.6 GB
Files
11
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-06-16 04:15

Files by quantization

Auxiliary files 11 files 19.6 GB
model.safetensors 19.6 GB 70fa522e download
tokenizer.json 19.1 MB 128d23d7 download
chat_template.jinja 7.57 KB fa6e2772 download
config.json 6.80 KB e6d04d0f download
README.md 4.22 KB eeea1d7b download
.gitattributes 1.53 KB 52373fe2 download
processor_config.json 1.27 KB 15a94ef1 download
tokenizer_config.json 1.14 KB 5344df8c download
preprocessor_config.json 390 B 2ea84a43 download
recipe.yaml 237 B 2ed68f09 download
generation_config.json 115 B 86ce6bc5 download

README current version from Hugging Face


base_model:

  • huihui-ai/Huihui-Nex-N2-mini-abliterated
  • nex-agi/Nex-N2-mini
    base_model_relation: quantized
    language:
  • en
  • ja
    tags:
  • qwen3_5_moe
  • nvfp4
  • w4a4
  • compressed-tensors
  • llm-compressor
  • abliterated
  • uncensored
  • vllm
  • blackwell
  • sm120
  • speculative-decoding
  • mtp
  • text-only
    library_name: vllm
    pipeline_tag: text-generation
    license: apache-2.0
    quantized_by: Lna-Lab

Huihui-Nex-N2-mini-abliterated-text-MTP-NVFP4

Text-only NVFP4 (W4A4) quantization of huihui-ai/Huihui-Nex-N2-mini-abliterated — the abliterated Nex-N2-mini (35B-A3B qwen3_5_moe hybrid-linear MoE). This is the language model only (Qwen3_5MoeForCausalLM, vision tower dropped) at ~22.7 GB, with the native MTP draft in MTP/. For the full vision-language version, see Huihui-Nex-N2-mini-abliterated-MTP-NVFP4.

Made by quantizing the huihui-ai abliterated release with llm-compressor + compressed-tensors. All credit for the model itself goes to huihui-ai (abliteration) and nex-agi (the original Nex-N2-mini); this repo only adds the FP4 quantization.

Lineage: nex-agi/Nex-N2-mini → huihui-ai abliteration → Lna-Lab NVFP4 W4A4, text-only (this repo).

What it is

Architecture Qwen3_5MoeForCausalLM (qwen3_5_moe_text) — language model only, no vision tower
Total / active ~34B total · ~3B active (256 experts, top-8, + shared expert)
Attention hybrid: linear-attention (GatedDeltaNet-style SSM) ×3 → full-attention every 4th layer (40 layers)
MTP native draft shipped in MTP/ (BF16, ~1.69 GB)
Quantization NVFP4 nvfp4-pack-quantized, W4A4, group size 16 — 30,720 expert projections + attention/linear-attn projections packed; router/norms/conv/lm_head BF16
Size ~22.7 GB
Nature abliterated / uncensored · reasoning model (<think>…</think>)

⚠️ Serving note (read this)

As of vLLM 0.22.0, the standalone text config (model_type: qwen3_5_moe_text / Qwen3_5MoeForCausalLM) is not yet served directly — vLLM routes this architecture through its Qwen3-VL path and expects a multimodal config with vision_config, so it stops with a config-type mismatch. Until vLLM adds a standalone text path for qwen3_5_moe, run the vision-language sibling with --limit-mm-per-prompt '{"image":0,"video":0}' for identical text-only behavior (the language weights are byte-for-byte the same). This text-only repo is published for transformers use, future/other backends, and as the smaller checkpoint.

llm-compressor recipe: QuantizationModifier(targets="Linear", scheme="NVFP4", ignore=["lm_head","re:.*visual.*","re:.*mlp.gate$","re:.*mlp.shared_expert_gate$"]), 32 calibration samples (neuralmagic/calibration).

License

Inherits apache-2.0 from the upstream models. Abliterated/uncensored: you are responsible for how you use it.

Credits

Support the Base Model Author (huihui-ai)

If you find the abliterated base useful, please support huihui-ai — this repo only adds the FP4 quantization; the abliteration work is theirs:

README history 2 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-06-16Add huihui-ai credit/support footer + benchmarksac0a3264.2 KB
    Loading...
  2. 2026-06-16NVFP4 W4A4 (llm-compressor) of huihui-ai/Huihui-Nex-N2-mini-abliterated + MTPc5ead6c3.4 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration