← back to catalog · registered 2026-08-22 13:56

edougawa/Nex-N2-mini-Abliterated

edougawa 35B MoE multimodal
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/edougawa%2FNex-N2-mini-Abliterated"
Response includes
  • classification m1
  • files 10
  • hub_downloads_all_time 60
  • author_summary 2 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M1
Primary method

Direct removal

No other method signals detected in this model.
Confidence
MEDIUM
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=0 (base model)
  • no specific method indicators - defaulting to M1 (most common)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
MEDIUM
Why we say so
primary_method=M1; difference-of-means is the reference extraction for M1/M3 (Arditi 2024)
Downloads · lifetime
60
21 last 30d - stable
Likes
1
Descendants
2
in 2 direct forks
Model age
3mo ago
created 2026-06-16
Downloads over time
Now66→from30↑120%
2842567030 on Jun 1766 on Oct 1166 on Oct 9JunJulAugSepOct
Jun 17 → Oct 11 · 56 snapshots · spans 116 days

Genealogy 2 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Variants by this author 2 formats · 52 downloads combined

The same weights this author released in different packaging. Pick the format that matches your runtime.

Metadata

License
apache-2.0
Tags
transformers safetensors qwen3_5_moe image-text-to-text vllm abliterix uncensored decensored abliterated conversational base_model:nex-agi/Nex-N2-mini base_model:finetune:nex-agi/Nex-N2-mini

Related

Total size
65.4 GB
Files
10
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-06-17 07:13

Files by quantization

Auxiliary files 10 files 65.4 GB
model-00001-of-00002.safetensors 46.3 GB a6fdeec8 download
model-00002-of-00002.safetensors 19.1 GB e3e457b0 download
tokenizer.json 19.1 MB 639e352c download
model.safetensors.index.json 94.6 KB ec177d70 download
README.md 13.7 KB 1415e259 download
chat_template.jinja 7.57 KB fa6e2772 download
config.json 3.17 KB 3000379c download
.gitattributes 1.53 KB 52373fe2 download
tokenizer_config.json 1.14 KB 5344df8c download
generation_config.json 116 B 26a38965 download

README current version from Hugging Face


license: apache-2.0
pipeline_tag: image-text-to-text
library_name: transformers
base_model:

  • nex-agi/Nex-N2-mini
    tags:
  • vllm
  • abliterix
  • uncensored
  • decensored
  • abliterated

Nex-N2-mini-Abliterated

This is a decensored (abliterated) version of nex-agi/Nex-N2-mini, produced with Abliterix v1.8.0.

Abliteration orthogonalizes the model weights against the measured "refusal" direction, reducing refusals while keeping the base model's capabilities as intact as possible (low KL divergence). It does not add any new knowledge or capability — all credit for the underlying model belongs to Nex-AGI. The original model card is reproduced in full below.

Base model: nex-agi/Nex-N2-mini (Apache-2.0, by Nex-AGI)

⚠️ Safety disclaimer

This model has had its built-in refusal behavior deliberately reduced. As a result it may produce unexpected, offensive, inaccurate, or otherwise harmful output, and may comply with requests that the original model would have refused.

  • It is provided by the publisher, edougawa, "as is" and without warranty of any kind, express or implied. Use at your own risk.
  • You are solely responsible for how you use this model and for ensuring your use — and any generated output — complies with all applicable laws, regulations, and the terms of the base model's license.
  • To the maximum extent permitted by law, the publisher (edougawa), the base-model authors (Nex-AGI), and the Abliterix authors accept no liability for any claim, damages, or other consequences arising from the use of this model or its outputs.
  • Outputs do not reflect the views of the publisher (edougawa), the base-model authors (Nex-AGI), or the Abliterix authors. Apply your own safety filtering, human review, and guardrails before any production or user-facing use.

Steering parameters

Parameter Value
vector_index 29.50
attn.k_proj.max_weight 3.43
attn.k_proj.max_weight_position 28.18
attn.k_proj.min_weight 1.88
attn.k_proj.min_weight_distance 11.27
attn.o_proj.max_weight 2.41
attn.o_proj.max_weight_position 37.58
attn.o_proj.min_weight 0.30
attn.o_proj.min_weight_distance 20.37
attn.q_proj.max_weight 2.29
attn.q_proj.max_weight_position 35.88
attn.q_proj.min_weight 0.52
attn.q_proj.min_weight_distance 9.47
attn.v_proj.max_weight 0.59
attn.v_proj.max_weight_position 29.14
attn.v_proj.min_weight 0.22
attn.v_proj.min_weight_distance 10.12
mlp.down_proj.max_weight 3.84
mlp.down_proj.max_weight_position 30.09
mlp.down_proj.min_weight 2.22
mlp.down_proj.min_weight_distance 13.71

Performance

Metric This model (Nex-N2-mini-Abliterated) Base model (Nex-N2-mini)
KL divergence 0.0180 0 (by definition)
Refusals 17/200 197/200

Deployment, sampling, function-calling, and reasoning-parser instructions from the base model (see the original card below) apply unchanged to this abliterated checkpoint.

Quickstart

Download this model (edougawa/Nex-N2-mini-Abliterated) and serve it with the Nex-AGI sglang fork:

# 1. Download edougawa/Nex-N2-mini-Abliterated
hf download edougawa/Nex-N2-mini-Abliterated --local-dir ./Nex-N2-mini-Abliterated

# 2. Serve it (single node, 2× H100 — see the original card for multi-node / Docker)
python -m sglang.launch_server \
  --model-path ./Nex-N2-mini-Abliterated \
  --tp 2 \
  --reasoning-parser qwen3 \
  --tool-call-parser qwen3_coder \
  --mamba-scheduler-strategy extra_buffer
# Or load it directly with transformers
from transformers import AutoModelForCausalLM, AutoTokenizer

model_id = "edougawa/Nex-N2-mini-Abliterated"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(model_id, dtype="bfloat16", device_map="auto")

Original model card

The following is the original model card for the base model, nex-agi/Nex-N2-mini, reproduced here in full with credit to Nex-AGI. The serving instructions below apply to this abliterated checkpoint as well.


🤗 Model   |    🔀 OpenRouter (Enjoy two weeks free starting June 9!)   |    💻 Github   |    🧭 ModelScope   |    🚀 Nex-AGI

Nex-N2

An agentic model with Agentic Thinking.

Today, we are officially releasing and open-sourcing our next-generation model, Nex-N2 — an agent model built for real-world productivity scenarios. With first-tier coding and agentic capabilities, Nex-N2 keeps driving complex, long-horizon tasks forward in real environments to deliver stable, end-to-end results.

Over the past year, a paradigm shift led by Vibe Coding and Harness Engineering has been redefining the limits of LLM agents. From dialogue, to reasoning, to agents that execute long-horizon tasks with environmental feedback, the tasks models must handle keep growing harder, the contexts longer, and the environments more realistic. The core of next-generation model competition is no longer whether a model can think, but whether it can reliably and efficiently turn thinking into actions that are executable, verifiable, and iterable.

Rather than treating reasoning, tool use, and environment execution as separate capabilities, Nex-N2 unifies them through an Agentic Thinking framework that connects requirement understanding, task planning, code implementation, environmental feedback, evaluation and debugging, and continuous iteration into a single closed loop. The framework has two parts:

  • Adaptive Thinking lets the model decide on its own when to think and how deeply — executing simple actions quickly while reasoning thoroughly on critical decisions.
  • Coherent Thinking carries one consistent reasoning paradigm across general reasoning and diverse agentic tasks, staying consistent across tasks and modalities to enable stable capability transfer.

Across real agentic workflows — agentic coding, deep research, tool calling, and terminal execution — Nex-N2 reaches first-tier performance, with substantial gains over the previous-generation Nex-N1 on multiple authoritative benchmarks. In real productivity scenarios such as OpenClaw one-person-company workflows, end-to-end game development, and web and multimodal generation, it likewise demonstrates outstanding usability, robustness, and stability.

Open Source

In keeping with our commitment to open source, we are releasing both Nex-N2-Pro and Nex-N2-mini as open-source models starting today.

We welcome developers and enterprises to integrate and try Nex-N2 and share their feedback.

Performance

We evaluate Nex-N2 in real agentic workflows along three directions — agentic tasks, coding tasks, and general tasks — covering benchmarks across tool calling, search-based decision-making, software engineering, and terminal execution. Nex-N2-Pro delivers strong performance that keeps pace with top-tier models such as GPT-5.5 and Opus 4.7: it excels at coding (e.g., 75.3 on Terminal-Bench 2.1) and long-horizon tasks (1585 on GDPval), and shows especially strong generalization and competitiveness on newer benchmarks like SWE-Atlas and DeepSWE. On general capability and core reasoning, it stands on par with leading frontier models.

Nex-N2 Benchmark Overview

Nex-N2 ships in two variants, both post-trained on the Qwen3.5 series: Nex-N2-Pro (built on Qwen3.5-397B-A17B) and Nex-N2-mini (built on Qwen3.5-35B-A3B-Base), covering different latency and quality trade-offs. The table below reports their scores alongside leading proprietary and open models across our full evaluation suite.

Benchmark Nex-N2-mini Nex-N2-Pro GPT-5.5 Opus 4.7 Kimi-K2.6 GLM-5.1 MiniMax M3 DeepSeek-V4-Pro
Agent
BrowseComp 74.1 83.7 84.4 79.8 83.2 79.3 83.5 83.4
GDPval 1402 1585 1769 1753 1481 1535 - 1554
Toolathlon 33.3 51.9 55.6 52.8 50.0 40.7 - 51.8
WildClawBench 47.7 53.5 58.2 62.2 - 48.2 - 43.7
WideSearch 62.0 75.6 - - 80.8 - - -
TAU3 65.9 71.1 - - - 70.6 - -
Coding & SWE
SWE-Bench Pro 50.2 58.8 58.6 64.3 58.6 58.4 59.0 55.4
Terminal-Bench 2.1 60.7 75.3 83.4 69.7 - 58.7 66.0 72.0
DeepSWE 8.0 33.6 70 54 24 18 - 8
SWE-Bench Verified 74.4 80.8 82.9 87.6 80.2 - 80.5 80.6
SWE Atlas QnA 31.5 37.9 45.4 45.2 - - 37.9 -
SWE Atlas RF 30.0 32.9 44.8 48.6 - - - -
SWE Atlas TW 23.3 40.0 42.6 38.2 - - 30.8 -
General & Reasoning
GPQA Diamond 82.6 90.7 93.6 94.2 90.5 86.2 - 90.1
IFEval 89.1 94.0 - - 94.5 94.5 - 91.9
Apex 9.4 36.5 - - 24.0 11.5 - 38.3

Usage

Local Deployment

Note: For the best performance with Nex-series models, we recommend serving them with our customized sglang fork.

First, install our sglang fork:

# Use the customized `sglang` fork
git clone https://github.com/nex-agi/sglang.git
cd sglang

# Install the python packages
pip install --upgrade pip
pip install -e "python"

Nex-N2-Pro

Launch the server (example on two 8× H100 servers with CUDA 13.0):

# Multi-node (2 nodes). Run the same command on every node with:
#   <node-rank> = 0 on the head node, 1 on the other node
#   <node0-ip>  = IP of the head node (reachable from all others)
python -m sglang.launch_server \
  --model-path /path/to/your/model  \
  --tp 16 \
  --nnodes 2 \
  --node-rank <node-rank> \
  --dist-init-addr <node0-ip>:20000 \
  --reasoning-parser qwen3 \
  --tool-call-parser qwen3_coder \
  --mamba-scheduler-strategy extra_buffer

Nex-N2-mini

Launch the server (example on one 2× H100 server with CUDA 13.0):

python -m sglang.launch_server \
  --model-path /path/to/your/model  \
  --tp 2 \
  --reasoning-parser qwen3 \
  --tool-call-parser qwen3_coder \
  --mamba-scheduler-strategy extra_buffer

Docker Deployment

We also provide a prebuilt Docker image with our customized sglang fork preinstalled: nexagi/sglang:v0.5.12. The launch command is the same as above.

Nex-N2-Pro

# Multi-node (2 nodes). Run the same command on every node with:
#   <node-rank> = 0 on the head node, 1 on the other node
#   <node0-ip>  = IP of the head node (reachable from all others)
docker run --gpus all --shm-size 32g --network host \
  -v /path/to/your/model:/model \
  nexagi/sglang:v0.5.12 \
  python3 -m sglang.launch_server \
    --model-path /model \
    --tp 16 \
    --nnodes 2 \
    --node-rank <node-rank> \
    --dist-init-addr <node0-ip>:20000 \
    --host 0.0.0.0 --port 30000 \
    --reasoning-parser qwen3 \
    --tool-call-parser qwen3_coder \
    --mamba-scheduler-strategy extra_buffer

Nex-N2-mini

Single node with 2× H100:

docker run --gpus all --shm-size 32g --ipc=host \
  -p 30000:30000 \
  -v /path/to/your/model:/model \
  nexagi/sglang:v0.5.12 \
  python3 -m sglang.launch_server \
    --model-path /model \
    --tp 2 \
    --host 0.0.0.0 --port 30000 \
    --reasoning-parser qwen3 \
    --tool-call-parser qwen3_coder \
    --mamba-scheduler-strategy extra_buffer

Recommended Sampling Parameters

For the best generation quality, we recommend the following sampling parameters:

  • temperature: 0.7
  • top_p: 0.95
  • top_k: 40

Function Calling

Nex-series models support robust function-calling capabilities. To enable function calling, add the --tool-call-parser qwen3_coder flag when launching the server:

python -m sglang.launch_server --model-path /path/to/your/model --tool-call-parser qwen3_coder

Reasoning Parser

Nex-series models emit explicit reasoning traces. Add the --reasoning-parser qwen3 flag to parse the reasoning content separately from the final response. It can be combined with the function-calling parser above:

python -m sglang.launch_server --model-path /path/to/your/model --tool-call-parser qwen3_coder --reasoning-parser qwen3

README history 6 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-06-17Update tags4d3d7b413.7 KB
    Loading...
  2. 2026-06-16docs: include publisher (edougawa) in disclaimer — no-warranty + no-liability73302d013.7 KB
    Loading...
  3. 2026-06-16docs: add safety disclaimer (reduced refusals; use at your own risk)710802713.2 KB
    Loading...
  4. 2026-06-16docs: rewrite README — abliteration info on top, credit + full original Nex-N...10f367e12.3 KB
    Loading...
  5. 2026-06-16Upload README.md with huggingface_hubd9562a910.3 KB
    Loading...
  6. 2026-06-16initial commitc36f9c428 B
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration