← back to catalog · registered 2026-08-22 13:56

SC117/Huihui-Nex-N2-mini-abliterated-APEX-GGUF

SC117 GGUF MoE multimodal second-order 262K ctx
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/SC117%2FHuihui-Nex-N2-mini-abliterated-APEX-GGUF"
Response includes
  • classification m8
  • files 9
  • hub_downloads_all_time 9,389
  • author_summary 18 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M8
Primary method

Repackaging (quantization)

Applied on top of direct removal inherited from the base model.
Confidence
MEDIUM
Inherited from base model
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=1
  • assume M1 (base ablation) + M8 (GGUF quant) - default when producer unknown
Refusal direction extraction

No specific extraction method could be identified for this model. The producer either did not document it or used a proprietary pipeline.

What is a refusal direction? →
Downloads · lifetime
9K
1K last 30d - stable
Likes
6
Model age
3mo ago
created 2026-06-16
Downloads over time
Now10.4K→from0↑0%
03.8K7.6K11.4K0 on Jun 1710.4K on Oct 11JunJulAugSepOct
Jun 17 → Oct 11 · 56 snapshots · spans 116 days

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
apache-2.0
Quantizations
BF16
Tags
transformers gguf qwen3_5_moe qwen3_5 agentic vision moe apex quantization abliterated uncensored image-text-to-text

Related

Total size
125 GB
Files
9
Quantizations
3
Registered
2026-08-22 13:56
Last updated on HF
2026-10-03 10:23

Files by quantization

BF16 1 file 64.6 GB
Huihui-Nex-N2-mini-abliterated.BF16.gguf 64.6 GB 7f6ac6f0 download
F16 1 file 858 MB
mmproj-Nex-N2-mini.F16.gguf 858 MB f5be3a80 download
Auxiliary files 7 files 60.3 GB
Huihui-Nex-N2-mini-abliterated-APEX-Balanced.gguf 23.6 GB f85d8315 download
Huihui-Nex-N2-mini-abliterated-APEX-Quality.gguf 21.3 GB 20ef47da download
Huihui-Nex-N2-mini-abliterated-APEX-Compact.gguf 15.4 GB 2f73246e download
README.md 18.6 KB afaafb46 download
README_zh.md 18.3 KB 2a4b2e7b download
chat_template.jinja 7.71 KB e241185d download
.gitattributes 2.13 KB 651c1406 download

README current version from Hugging Face


library_name: transformers
license: apache-2.0
license_link: https://huggingface.co/nex-agi/Nex-N2-mini/blob/main/LICENSE
pipeline_tag: image-text-to-text
tags:

  • qwen3_5_moe
  • qwen3_5
  • agentic
  • vision
  • moe
  • apex
  • quantization
  • abliterated
  • uncensored
    base_model:
  • huihui-ai/Huihui-Nex-N2-mini-abliterated

APEX Vision Agentic Abliterated

Huihui-Nex-N2-mini

📖 中文文档

Abliterated Agentic Vision MoE — APEX Quantized GGUF

⚡ Thinking Mode Requires Nex's Patched llama.cpp

Nex-N2-mini's original chat template uses complex vision processing macros that stock llama.cpp's Jinja parser cannot handle correctly. This causes thinking tags to not be injected, breaking thinking mode output.

The official fix: Use Nex's patched llama.cpp, which works with the unmodified GGUF and unmodified template. Once Nex's upstream patch is merged into stock llama.cpp, this workaround will no longer be needed.

⚠️ Do NOT modify chat_template.jinja. The model was trained strictly on the current template — editing the tags deviates from the training-time format and may degrade output quality. See discussion #3.

⚠️ Abliterated Model — Use at Your Own Risk

This is an abliterated (uncensored) version of Nex-N2-mini created by huihui-ai. Safety filtering has been significantly reduced.

  • May generate sensitive, controversial, or inappropriate content
  • Not suitable for public settings, underage users, or production use
  • Users are solely responsible for compliance with local laws and ethical standards
  • Recommended for research, testing, or controlled environments only

Original model: huihui-ai/Huihui-Nex-N2-mini-abliterated

💡 What is APEX?

These GGUF files are quantized using APEX, a MoE-aware mixed-precision quantization technique that outperforms standard quantization methods while being significantly smaller.

APEX beats Q8_0 perplexity at half the size — and even beats F16.

APEX classifies every tensor by its role — routed expert, shared expert, or attention — and applies a layer-wise precision gradient, giving the most sensitive edge layers higher precision and compressing the redundant middle layers more aggressively.

📦 Available Files
FileSizeBPWNote
Huihui-Nex-N2-mini-abliterated.BF16.gguf64.6 GB16.0Full precision reference
Huihui-Nex-N2-mini-abliterated-APEX-Quality.gguf21.3 GB5.23Highest quality, best accuracy
Huihui-Nex-N2-mini-abliterated-APEX-Balanced.gguf23.6 GB5.85Best all-rounder, recommended
Huihui-Nex-N2-mini-abliterated-APEX-Compact.gguf15.4 GB3.81Best quality/size ratio, 16GB VRAM
mmproj-Nex-N2-mini.F16.gguf858 MB-Vision projector (required for image/video)
chat_template.jinja7.9 KB-Original unmodified chat template
🧠 Model Details
ArchitectureQwen3.5 MoE (GatedDeltaNet + Full Attention) + Vision Encoder
Parameters35B total, 3B active per token
Experts256 routed experts, 8 active per token
Layers40 layers (30 linear_attn + 10 full_attn)
Context262,144 tokens
VisionImage support (mmproj 858MB)
ThinkingQwen3-style think tags — requires Nex's patched llama.cpp
ModificationAbliterated by huihui-ai (safety filters removed)
🚀 Usage

Download Nex's patched llama.cpp

Binaries: nex-agi/llama.cpp  |  Docker: ghcr.io/nex-agi/llama.cpp:server-cuda-nex-b9596-fix-b9598-8c0d5c9

./llama-server \ -m Huihui-Nex-N2-mini-abliterated-APEX-Quality.gguf \ -ngl 99 -ncmoe 19 -c 32768 \ --host 0.0.0.0 --port 8081

Replace Huihui-Nex-N2-mini-abliterated-APEX-Quality.gguf with your preferred quantization tier (Quality / Balanced / Compact). Add --mmproj mmproj-Nex-N2-mini.F16.gguf for vision. Recommended sampling: temperature 0.7, top_p 0.95, top_k 40, min_p 0.

📋 Original Model Benchmarks
BenchmarkScoreCategory
BrowseComp74.1Agent
SWE-Bench Verified74.4Coding
Terminal-Bench 2.160.7Coding
GPQA Diamond82.6Reasoning
IFEval89.1Instruction

From the original Nex-N2-mini model card (BF16, full precision). Abliteration does not significantly change benchmark scores.

Links

README history 6 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-10-03Add optional Ko-fi support bannere98f7bd18.9 KB
    Loading...
  2. 2026-06-26Update file naming and documentation21f9bad18.5 KB
    Loading...
  3. 2026-06-26Update README: remove I- prefix from APEX file names (no imatrix was used)c81961118.5 KB
    Loading...
  4. 2026-06-18Update README.md7c33a3718.5 KB
    Loading...
  5. 2026-06-16Update README.md07305e718.5 KB
    Loading...
  6. 2026-06-16Initial upload: APEX quantized GGUFs + BF16 reference8d829c118.5 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration