← back to catalog · registered 2026-08-22 13:56

llmfan46/Nex-N2-mini-ultra-uncensored-heretic-GGUF

llmfan46 GGUF MoE second-order 262K ctx
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/llmfan46%2FNex-N2-mini-ultra-uncensored-heretic-GGUF"
Response includes
  • classification m3
  • files 12
  • hub_downloads_all_time 13,197
  • author_summary 211 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M3
Primary method

Layer-wise ablation

Applied on top of direct removal inherited from the base model.
Confidence
HIGH
Inherited from base model
Why this label 2 signals
Producer identity confirmed by naming conventions, tags or the model card. This label is very unlikely to change.
  • 'heretic' in model name (Heretic-produced)
  • Heretic uses layer-wise optimization (M3) with underlying direction removal (M1)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
HIGH
Why we say so
name contains 'heretic'; Heretic default extraction is difference-of-means (Arditi 2024)
Downloads · lifetime
13K
1K last 30d - cooling
Likes
6
Model age
3mo ago
created 2026-06-24
Downloads over time
Now13.3K→from2.2K↑516%
1.6K5.9K10.2K14.5K2.2K on Jun 2413.3K on Oct 11JunJulAugSepOct
Jun 24 → Oct 11 · 57 snapshots · spans 109 days

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Variants by this author 2 formats · 1K downloads combined

The same weights this author released in different packaging. Pick the format that matches your runtime.

Metadata

License
apache-2.0
Quantizations
BF16 Q3_K Q4_K Q5_K Q6_K Q8_0
Tags
transformers gguf qwen3_5_moe heretic uncensored decensored abliterated mpoa text-generation base_model:llmfan46/Nex-N2-mini-ultra-uncensored-heretic base_model:quantized:llmfan46/Nex-N2-mini-ultra-uncensored-heretic license:apache-2.0

Related

Total size
242 GB
Files
12
Quantizations
7
Registered
2026-08-22 13:56
Last updated on HF
2026-06-25 10:12

Files by quantization

BF16 2 files 65.5 GB
Nex-N2-mini-ultra-uncensored-heretic-BF16.gguf 64.6 GB af4bedef download
Nex-N2-mini-ultra-uncensored-heretic-mmproj-BF16.gguf 861 MB 457a9a00 download
Q8_0 1 file 34.4 GB
Nex-N2-mini-ultra-uncensored-heretic-Q8_0.gguf 34.4 GB 3a929c84 download
Q6_K 1 file 26.6 GB
Nex-N2-mini-ultra-uncensored-heretic-Q6_K.gguf 26.6 GB 2b722d1c download
Q5_K 2 files 45.4 GB
Nex-N2-mini-ultra-uncensored-heretic-Q5_K_M.gguf 23.1 GB e04bce7b download
Nex-N2-mini-ultra-uncensored-heretic-Q5_K_S.gguf 22.4 GB 68fc233a download
Q4_K 2 files 38.4 GB
Nex-N2-mini-ultra-uncensored-heretic-Q4_K_M.gguf 19.8 GB a622553f download
Nex-N2-mini-ultra-uncensored-heretic-Q4_K_S.gguf 18.6 GB fb90cc3f download
Q3_K 2 files 32.7 GB
Nex-N2-mini-ultra-uncensored-heretic-Q3_K_L.gguf 17.0 GB 878215b1 download
Nex-N2-mini-ultra-uncensored-heretic-Q3_K_M.gguf 15.7 GB b14479f1 download
Auxiliary files 2 files 15.0 KB
README.md 14.2 KB cefad581 download
.gitattributes 849 B 4a69113e download

README current version from Hugging Face


license: apache-2.0
pipeline_tag: text-generation
library_name: transformers
tags:

  • qwen3_5_moe
  • heretic
  • uncensored
  • decensored
  • abliterated
  • mpoa
    base_model:
  • llmfan46/Nex-N2-mini-ultra-uncensored-heretic

🚨⚠️ I HAVE REACHED HUGGING FACE'S FREE STORAGE LIMIT ⚠️🚨

I can no longer upload new models unless I can cover the cost of additional storage.
I host 70+ free models as an independent contributor and this work is unpaid.
Without your support, no more new models can be uploaded.

🎉 Patreon (Monthly)  |  ☕ Ko-fi (One-time)

Every contribution goes directly toward Hugging Face storage fees to keep models free for everyone.


93% fewer refusals (5/100 Uncensored vs 74/100 Original) while preserving model quality (0.0020 KL divergence).

❤️ Support My Work

Creating these models takes significant time, work and compute. If you find them useful consider supporting me:

image/png

Platform Link What you get
🎉 Patreon Monthly support Priority model requests
☕ Ko-fi One-time tip My eternal gratitude

Your help will motivate me and would go into further improving my workflow and coverings fees for storage, compute and may even help uncensoring bigger model with rental Cloud GPUs.


GGUF quantizations of llmfan46/Nex-N2-mini-ultra-uncensored-heretic.

This is a decensored version of a nex-agi/Nex-N2-mini, made using Heretic v1.2.0 with a variant of the Magnitude-Preserving Orthogonal Ablation (MPOA) method

Abliteration parameters

Parameter Value
direction_index 16.68
attn.out_proj.max_weight 1.11
attn.out_proj.max_weight_position 29.79
attn.out_proj.min_weight 0.83
attn.out_proj.min_weight_distance 26.98
mlp.down_proj.max_weight 1.94
mlp.down_proj.max_weight_position 29.92
mlp.down_proj.min_weight 1.84
mlp.down_proj.min_weight_distance 26.37
attn.o_proj.max_weight 1.65
attn.o_proj.max_weight_position 29.28
attn.o_proj.min_weight 1.36
attn.o_proj.min_weight_distance 23.64

Targeted components

  • attn.o_proj
  • attn.out_proj
  • mlp.down_proj

Performance

Metric This model Original model (Nex-N2-mini)
KL divergence 0.0020 0 (by definition)
Refusals ✅ 5/100 ❌ 74/100

Lower refusals indicate fewer content restrictions, while lower KL divergence indicates more closeness to the original model's baseline. Higher refusals cause more rejections, objections, pushbacks, lecturing, censorship, softening and deflections.


Quantizations

For the K-quants below, small SSM tensors are kept at higher precision where useful.

-Q6_K keeps ssm_alpha, ssm_beta, and ssm_out as Q8_0.

-Q5_K, Q4_K, and Q3_K quants keep ssm_alpha and ssm_beta as Q8_0, while ssm_out is kept as Q6_K.

This helps preserve the hybrid/SSM blocks with a small file-size increase.

Filename Quant Description
Nex-N2-mini-ultra-uncensored-heretic-BF16.gguf BF16 Full precision
Nex-N2-mini-ultra-uncensored-heretic-Q8_0.gguf Q8_0 Near-lossless, recommended
Nex-N2-mini-ultra-uncensored-heretic-Q6_K.gguf Q6_K Excellent quality
Nex-N2-mini-ultra-uncensored-heretic-Q5_K_M.gguf Q5_K_M Good balance
Nex-N2-mini-ultra-uncensored-heretic-Q5_K_S.gguf Q5_K_S Smaller Q5
Nex-N2-mini-ultra-uncensored-heretic-Q4_K_M.gguf Q4_K_M Good for limited VRAM
Nex-N2-mini-ultra-uncensored-heretic-Q4_K_S.gguf Q4_K_S Smaller Q4
Nex-N2-mini-ultra-uncensored-heretic-Q3_K_L.gguf Q3_K_L Low VRAM, decent quality
Nex-N2-mini-ultra-uncensored-heretic-Q3_K_M.gguf Q3_K_M Low VRAM, smaller

Vision Projector

Filename Quant Description
Nex-N2-mini-ultra-uncensored-heretic-mmproj-BF16.gguf BF16 Native precision

A Vision Projector File is Required for vision/multimodal capabilities. Use alongside any quantization above.

Usage

Works with llama.cpp, LM Studio, Ollama, and other GGUF-compatible tools.



🤗 Model   |    🔀 OpenRouter (Enjoy two weeks free starting June 9!)   |    💻 Github   |    🧭 ModelScope   |    🚀 Nex-AGI

Nex-N2

An agentic model with Agentic Thinking.

Today, we are officially releasing and open-sourcing our next-generation model, Nex-N2 — an agent model built for real-world productivity scenarios. With first-tier coding and agentic capabilities, Nex-N2 keeps driving complex, long-horizon tasks forward in real environments to deliver stable, end-to-end results.

Over the past year, a paradigm shift led by Vibe Coding and Harness Engineering has been redefining the limits of LLM agents. From dialogue, to reasoning, to agents that execute long-horizon tasks with environmental feedback, the tasks models must handle keep growing harder, the contexts longer, and the environments more realistic. The core of next-generation model competition is no longer whether a model can think, but whether it can reliably and efficiently turn thinking into actions that are executable, verifiable, and iterable.

Rather than treating reasoning, tool use, and environment execution as separate capabilities, Nex-N2 unifies them through an Agentic Thinking framework that connects requirement understanding, task planning, code implementation, environmental feedback, evaluation and debugging, and continuous iteration into a single closed loop. The framework has two parts:

  • Adaptive Thinking lets the model decide on its own when to think and how deeply — executing simple actions quickly while reasoning thoroughly on critical decisions.
  • Coherent Thinking carries one consistent reasoning paradigm across general reasoning and diverse agentic tasks, staying consistent across tasks and modalities to enable stable capability transfer.

Across real agentic workflows — agentic coding, deep research, tool calling, and terminal execution — Nex-N2 reaches first-tier performance, with substantial gains over the previous-generation Nex-N1 on multiple authoritative benchmarks. In real productivity scenarios such as OpenClaw one-person-company workflows, end-to-end game development, and web and multimodal generation, it likewise demonstrates outstanding usability, robustness, and stability.

Open Source

In keeping with our commitment to open source, we are releasing both Nex-N2-Pro and Nex-N2-mini as open-source models starting today.

We welcome developers and enterprises to integrate and try Nex-N2 and share their feedback.

Performance

We evaluate Nex-N2 in real agentic workflows along three directions — agentic tasks, coding tasks, and general tasks — covering benchmarks across tool calling, search-based decision-making, software engineering, and terminal execution. Nex-N2-Pro delivers strong performance that keeps pace with top-tier models such as GPT-5.5 and Opus 4.7: it excels at coding (e.g., 75.3 on Terminal-Bench 2.1) and long-horizon tasks (1585 on GDPval), and shows especially strong generalization and competitiveness on newer benchmarks like SWE-Atlas and DeepSWE. On general capability and core reasoning, it stands on par with leading frontier models.

Nex-N2 Benchmark Overview

Nex-N2 ships in two variants, both post-trained on the Qwen3.5 series: Nex-N2-Pro (built on Qwen3.5-397B-A17B) and Nex-N2-mini (built on Qwen3.5-35B-A3B-Base), covering different latency and quality trade-offs. The table below reports their scores alongside leading proprietary and open models across our full evaluation suite.

Benchmark Nex-N2-mini Nex-N2-Pro GPT-5.5 Opus 4.7 Kimi-K2.6 GLM-5.1 MiniMax M3 DeepSeek-V4-Pro
Agent
BrowseComp 74.1 83.7 84.4 79.8 83.2 79.3 83.5 83.4
GDPval 1402 1585 1769 1753 1481 1535 - 1554
Toolathlon 33.3 51.9 55.6 52.8 50.0 40.7 - 51.8
WildClawBench 47.7 53.5 58.2 62.2 - 48.2 - 43.7
WideSearch 62.0 75.6 - - 80.8 - - -
TAU3 65.9 71.1 - - - 70.6 - -
Coding & SWE
SWE-Bench Pro 50.2 58.8 58.6 64.3 58.6 58.4 59.0 55.4
Terminal-Bench 2.1 60.7 75.3 83.4 69.7 - 58.7 66.0 72.0
DeepSWE 8.0 33.6 70 54 24 18 - 8
SWE-Bench Verified 74.4 80.8 82.9 87.6 80.2 - 80.5 80.6
SWE Atlas QnA 31.5 37.9 45.4 45.2 - - 37.9 -
SWE Atlas RF 30.0 32.9 44.8 48.6 - - - -
SWE Atlas TW 23.3 40.0 42.6 38.2 - - 30.8 -
General & Reasoning
GPQA Diamond 82.6 90.7 93.6 94.2 90.5 86.2 - 90.1
IFEval 89.1 94.0 - - 94.5 94.5 - 91.9
Apex 9.4 36.5 - - 24.0 11.5 - 38.3

Usage

Local Deployment

Note: For the best performance with Nex-series models, we recommend serving them with our customized sglang fork.

First, install our sglang fork:

# Use the customized `sglang` fork
git clone https://github.com/nex-agi/sglang.git
cd sglang

# Install the python packages
pip install --upgrade pip
pip install -e "python"

Nex-N2-Pro

Launch the server (example on two 8× H100 servers with CUDA 13.0):

# Multi-node (2 nodes). Run the same command on every node with:
#   <node-rank> = 0 on the head node, 1 on the other node
#   <node0-ip>  = IP of the head node (reachable from all others)
python -m sglang.launch_server \
  --model-path /path/to/your/model  \
  --tp 16 \
  --nnodes 2 \
  --node-rank <node-rank> \
  --dist-init-addr <node0-ip>:20000 \
  --reasoning-parser qwen3 \
  --tool-call-parser qwen3_coder \
  --mamba-scheduler-strategy extra_buffer

Nex-N2-mini

Launch the server (example on one 2× H100 server with CUDA 13.0):

python -m sglang.launch_server \
  --model-path /path/to/your/model  \
  --tp 2 \
  --reasoning-parser qwen3 \
  --tool-call-parser qwen3_coder \
  --mamba-scheduler-strategy extra_buffer

Docker Deployment

We also provide a prebuilt Docker image with our customized sglang fork preinstalled: nexagi/sglang:v0.5.12. The launch command is the same as above.

Nex-N2-Pro

# Multi-node (2 nodes). Run the same command on every node with:
#   <node-rank> = 0 on the head node, 1 on the other node
#   <node0-ip>  = IP of the head node (reachable from all others)
docker run --gpus all --shm-size 32g --network host \
  -v /path/to/your/model:/model \
  nexagi/sglang:v0.5.12 \
  python3 -m sglang.launch_server \
    --model-path /model \
    --tp 16 \
    --nnodes 2 \
    --node-rank <node-rank> \
    --dist-init-addr <node0-ip>:20000 \
    --host 0.0.0.0 --port 30000 \
    --reasoning-parser qwen3 \
    --tool-call-parser qwen3_coder \
    --mamba-scheduler-strategy extra_buffer

Nex-N2-mini

Single node with 2× H100:

docker run --gpus all --shm-size 32g --ipc=host \
  -p 30000:30000 \
  -v /path/to/your/model:/model \
  nexagi/sglang:v0.5.12 \
  python3 -m sglang.launch_server \
    --model-path /model \
    --tp 2 \
    --host 0.0.0.0 --port 30000 \
    --reasoning-parser qwen3 \
    --tool-call-parser qwen3_coder \
    --mamba-scheduler-strategy extra_buffer

Recommended Sampling Parameters

For the best generation quality, we recommend the following sampling parameters:

  • temperature: 0.7
  • top_p: 0.95
  • top_k: 40

Function Calling

Nex-series models support robust function-calling capabilities. To enable function calling, add the --tool-call-parser qwen3_coder flag when launching the server:

python -m sglang.launch_server --model-path /path/to/your/model --tool-call-parser qwen3_coder

Reasoning Parser

Nex-series models emit explicit reasoning traces. Add the --reasoning-parser qwen3 flag to parse the reasoning content separately from the final response. It can be combined with the function-calling parser above:

python -m sglang.launch_server --model-path /path/to/your/model --tool-call-parser qwen3_coder --reasoning-parser qwen3

README history 2 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-06-24Update README.mdf52417414.2 KB
    Loading...
  2. 2026-06-24Upload folder using huggingface_hubc9392b514.1 KB
    Loading...

Discussions 1 thread

  1. 2026-06-25llama.cpp cannot load that GGUF because an expected tensor is missing, model Ne…closed13 💬#1
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration