← back to catalog · registered 2026-09-11 13:55

orcarouter/Nex-N2.5-mini-Uncensored

orcarouter MoE multimodal
Abliteration classifier · v1.0.0
M1
Primary method

Direct removal

No other method signals detected in this model.
Confidence
MEDIUM
Why this label 3 signals
Method inferred from partial signals — repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=0 (base model)
  • no specific method indicators - defaulting to M1 (most common)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
MEDIUM
Why we say so
primary_method=M1; difference-of-means is the reference extraction for M1/M3 (Arditi 2024)
Downloads · 30-day
0
Likes
2
Descendants
3
in 3 direct forks
Model age
1d ago
created 2026-09-10

Genealogy 3 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Variants by this author 4 formats · 309 downloads combined

The same weights this author released in different packaging. Pick the format that matches your runtime.

Metadata

License
apache-2.0
Languages
en zh
Tags
transformers safetensors qwen3_5_moe image-text-to-text abliterated uncensored nex nex-n2.5 ai-red-team red-teaming moe hybrid-attention

Related

Total size
65.4 GB
Files
25
Quantizations
1
Registered
2026-09-11 13:55
Last updated on HF
2026-09-10 08:27

Files by quantization

Auxiliary files 25 files 65.4 GB
model-00006-of-00016.safetensors 4.64 GB ******** download
model-00009-of-00016.safetensors 4.64 GB ******** download
model-00012-of-00016.safetensors 4.64 GB ******** download
model-00015-of-00016.safetensors 4.64 GB ******** download
model-00003-of-00016.safetensors 4.64 GB ******** download
model-00001-of-00016.safetensors 4.42 GB ******** download
model-00007-of-00016.safetensors 4.20 GB ******** download
model-00010-of-00016.safetensors 4.20 GB ******** download
model-00013-of-00016.safetensors 4.20 GB ******** download
model-00004-of-00016.safetensors 4.20 GB ******** download
model-00008-of-00016.safetensors 3.70 GB ******** download
model-00011-of-00016.safetensors 3.70 GB ******** download
model-00014-of-00016.safetensors 3.70 GB ******** download
model-00002-of-00016.safetensors 3.70 GB ******** download
model-00005-of-00016.safetensors 3.70 GB ******** download
model-00016-of-00016.safetensors 2.50 GB ******** download
tokenizer.json 19.1 MB ******** download
model.safetensors.index.json 94.6 KB a25d9709 download
README.md 16.2 KB 5bb52908 download
chat_template.jinja 7.56 KB 34398e43 download
config.json 3.16 KB 8643d6d3 download
.gitattributes 1.53 KB 52373fe2 download
processor_config.json 1.27 KB 7ad6acdf download
tokenizer_config.json 1.11 KB a068e246 download
preprocessor_config.json 390 B 2ea84a43 download

README current version from Hugging Face


license: apache-2.0
base_model: nex-agi/Nex-N2.5-mini
base_model_relation: finetune
pipeline_tag: text-generation
library_name: transformers
language:

  • en
  • zh
    tags:
  • abliterated
  • uncensored
  • nex
  • nex-n2.5
  • ai-red-team
  • red-teaming
  • moe
  • hybrid-attention
  • gated-delta-net
  • vision-language
  • image-text-to-text
  • agentic
  • function-calling
  • reasoning
  • post-training
  • fine-tuning

OrcaRouter

Nex-N2.5-mini-Uncensored

The abliterated (refusal-removed) build of Nex-AGI's Nex-N2.5-mini — baked directly into the official BF16 shards, byte-for-byte drop-in for the original

Website Model Catalog License precision 262K context Vision-Language MoE Hybrid attention

One Gateway. Every Model. — Route Smarter · Ship Safer · Spend Less.

Website · Model Catalog · GitHub · Discord · X


The abliterated (refusal-removed) build of nex-agi/Nex-N2.5-mini
— a ~34B Mixture-of-Experts agentic VLM with hybrid gated-delta-net + full attention,
256 routed experts top-8 beside a shared expert on every layer, a native vision tower, and a
262K-token context.

The refusal direction is baked directly into the official BF16 shards — same dtype, same shard
layout, same model.safetensors.index.json. All 1026 tensors match the base checkpoint in name,
dtype and shape, so this is a drop-in replacement for nex-agi/Nex-N2.5-mini in any stack that
already serves it.

On precision: BF16 is the format Nex-AGI ships Nex-N2.5-mini in — there is no quantization step
here, applied or undone.


⚠️ Disclaimer — read before use

This model has had its safety alignment substantially removed via abliteration (orthogonalizing the
refusal direction out of the residual stream). As a direct consequence:

  • It will comply with harmful, unethical, offensive, or illegal requests that the original
    Nex-N2.5-mini would refuse. It has no meaningful built-in guardrails.
  • It is released strictly for legitimate research — interpretability, AI-safety and refusal-mechanism
    study, red-teaming, robustness evaluation, and controlled experiments.
  • You assume full responsibility and liability for how you use it and for everything it generates. Do
    not deploy it to end users or in production without adding your own safety, moderation, and
    abuse-prevention layers.
  • Use must comply with the Apache 2.0 License inherited
    from the base model, and all laws and regulations that apply to you.
  • The authors and uploaders accept no liability for any misuse or harm. Its outputs do not reflect
    the views of the uploaders or of Nex-AGI.

By downloading or using this model you acknowledge and accept the above.


🐋 Run it with OrcaCode Review

Models are only half the system.

OrcaCode Review turns every model listed on
OrcaRouter into a production code-review agent:

  • reviews every PR
  • finds security + correctness issues
  • posts inline findings
  • P0/P1 can block merges
  • swap models anytime

Open model. Open harness. Open bill.


Model details

Base model nex-agi/Nex-N2.5-mini
Architecture Qwen3_5MoeForConditionalGeneration (qwen3_5_moe) — 40 layers, hidden 2048, hybrid attention (30 gated-delta-net linear-attention + 10 full-attention with attn_output_gate, interval 4), 256 routed experts top-8 + 1 shared expert on every layer, interleaved mRoPE, native vision tower (27 blocks @ 1152, patch 16, spatial merge 2)
Parameters ~34B total — 35.1B tensor elements on disk, including the 0.44B vision tower
Modification Abliteration (refusal-direction removal), baked on disk into the BF16 shards — no quantization applied
Format safetensors, BF16, 16 shards, 1026 tensors, 66 GiB
Edited 122 of 1026 tensors; the other 904 are byte-identical to the base
Preserved Everything that reads or gates the residual stream rather than writing it — the MoE router, the expert and attention input projections, the linear-attention state parameters, every norm, lm_head, and the vision tower apart from the merger's output projection
Context 262,144 tokens
Vocabulary 248,320 (untied embeddings)
Recommended for Red-team & refusal-mechanism research, robustness evaluation, and as a base for further post-training / quantization

Abliteration

Refusal-direction removal following Arditi et al. (2024), Refusal in Language Models Is Mediated by a
Single Direction
. A single refusal direction is estimated from the residual stream on paired
harmful/harmless prompts, then projected out of the matrices that write that stream — the attention
output projections, the routed and shared expert down-projections, the token embeddings, and the vision
merger's output projection. The projection is computed in float32 and stored back in the checkpoint's own
BF16.

122 of the 1026 tensors are edited; the other 904 are byte-identical to the base. Nothing that
reads or gates the residual stream is touched. That is verified by a tensor-by-tensor comparison of the
two checkpoints rather than by trusting the code that produced them, which is also what guarantees the
drop-in property: every dtype, shape and shard boundary survives.

The edit is applied to the whole depth of the model, and the layer the direction is fitted at was selected
on a held-out validation split by an objective that rewards bypassing refusal while charging for both
divergence on harmless prompts and any newly induced refusal of benign ones. Capability retention is
therefore measured rather than assumed — see Evaluation.

Fine-tuning & post-training

Nothing here changes the architecture, the config, the tokenizer or the shard layout, so anything that
fine-tunes or quantizes the base checkpoint works unmodified on this one. Note that further training on
refusal-bearing data will partially restore the direction; re-derive and re-bake afterwards rather than
assuming the edit survives.


Intended use

  • Research into refusal mechanisms, alignment, and interpretability.
  • Red-teaming and safety / robustness evaluation in controlled environments.
  • A base for further post-training and quantization.

Evaluation

Measured on these exact weights, 8×H100, greedy decoding, 64 new tokens, reasoning_effort=none,
thinking off. Both columns come from the same script, the same sampling seed, the same settings and the
same scoring code. Refusal is judged by a rule-based opening-phrase classifier (caveat = answered but
wrapped in a disclaimer) — indicative, not an LLM-judge / publication-grade number. Read
Limitations before quoting the 0.000s.

Safety — harmful-prompt refusal (lower = more uncensored)

Benchmark n Base This model Δ Caveat (base → this)
AdvBench 100 1.000 0.000 −100% 1.000 → 0.360
JailbreakBench (harmful) 100 0.980 0.000 −98% 0.980 → 0.370
StrongREJECT 150 0.933 0.000 −93% 0.953 → 0.360
MaliciousInstruct 100 0.950 0.000 −95% 0.950 → 0.400
HarmBench (standard) 150 0.980 0.020 −96% 0.980 → 0.227
ForbiddenQuestions 150 0.587 0.020 −57% 0.627 → 0.273
SimpleSafetyTests 50 0.960 0.060 −90% 0.960 → 0.380

The base model refuses 93–100% of harmful prompts on six of the seven sets. ForbiddenQuestions is the
outlier at 0.587, which says more about that set — a large share of its questions are sensitive rather
than harmful, and the base model answers them — than about the edit.

The caveat column is the second half of the story. It does not merely track refusal: it falls from ~0.96
to ~0.35, so roughly two thirds of the disclaimer-wrapping goes with the refusal rather than surviving it.

Over-refusal — benign prompts wrongly refused (lower = better)

Benchmark n Base This model Δ
XSTest-safe 250 0.056 0.008 −4.8 pp
JailbreakBench (benign) 100 0.220 0.000 −22.0 pp

XSTest-safe is 250 prompts written to look harmful while being benign. The base model is already fairly
restrained there at 5.6%; the edit does not blunt that discrimination, it sharpens it. JBB's benign split
is the harder one — adversarially benign, 22% wrongly refused at baseline — and it goes to zero.

Capability retention — vs the official base (same scripts, same settings)

Benchmark n Base This model Δ
MMLU (0-shot) 300 0.830 0.830 ±0.0 pp
MMLU-Pro (CoT) 250 0.776 0.752 −2.4 pp
GSM8K (CoT) 150 0.907 0.907 ±0.0 pp
CMMLU (0-shot, Chinese) 500 0.852 0.864 +1.2 pp

Every delta is within ±2.4 pp, and unparsed rates were ≤1.6% on both checkpoints. Two of the four move up
or hold exactly, which is the signature of sampling noise at these sizes rather than of an improvement —
read the table as "unchanged", not as a gain. The −2.4 pp on MMLU-Pro is the largest single move and sits
within the noise band of a 250-item CoT subset; it is not evidence of a regression, and it is not evidence
against one either.

MMLU and CMMLU ask for a bare option letter, so they are scored from a single forward pass — an argmax
over the option letters' logits. MMLU-Pro and GSM8K are chain-of-thought scored by generation, because
their prompts ask the model to reason first: an argmax at the first position there measures the opening
of a chain of thought, not an answer, and reads ~55 points low. These MMLU-Pro figures are therefore
directly comparable to CoT-scored MMLU-Pro results elsewhere.


Usage

transformers

qwen3_5_moe needs transformers ≥ 5.2; verified here on 5.17.0 with torch 2.14. Install
flash-linear-attention as well — without it
transformers falls back to a reference PyTorch gated-delta-rule on 30 of the 40 layers, which is correct
but much slower.

from transformers import AutoModelForCausalLM, AutoTokenizer

model_id = "orcarouter/Nex-N2.5-mini-Uncensored"
tok = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(model_id, dtype="bfloat16", device_map="auto")

msgs = [{"role": "user", "content": "..."}]
text = tok.apply_chat_template(
    msgs, tokenize=False, add_generation_prompt=True, reasoning_effort="none"
)
out = model.generate(**tok(text, return_tensors="pt").to(model.device), max_new_tokens=512)
print(tok.decode(out[0], skip_special_tokens=True))

The weights are 66 GiB, so they fit on a single 80 GB card with room for activations, or shard across
smaller ones with device_map="auto".

Serving: we did not serve this build through vLLM or SGLang, so no flag set is recommended here.
Because every tensor matches the base checkpoint in name, dtype and shape, whatever recipe works for
nex-agi/Nex-N2.5-mini applies unchanged — follow the base model's own serving instructions.

Thinking control

Nex-N2.5-mini has no enable_thinking toggle — passing one is silently inert, because
apply_chat_template drops unknown kwargs into a non-strict jinja context and the prompt renders
byte-identically either way. Control it with reasoning_effort:

value renders use
none <think>\n\n</think>\n\npre-closed the reply starts at the answer; this family's thinking-off, and what every number above was measured with
unset / other <think> reasoning on, template default
high <think>\n reasoning on, longer traces

none is not a quieter low: it closes the thinking block inside the prompt, so no reasoning budget is
needed and the reply is the answer. With reasoning on, give generation enough budget to reach </think>
or replies get truncated inside the scratchpad.

Note for evaluators: do not pair reasoning_effort="none" with a <think>-stripping classifier.
The block is already closed in the prompt, so the reply contains no </think>, every reply strips to
the empty string, and the run reports a 100% refusal rate on every split.


Bias, risks, and limitations

  • Safety guardrails removed — the model will produce harmful, biased, or offensive content on request
    (see the disclaimer).
  • It inherits any biases and limitations of the base Nex-N2.5-mini.
  • The 0.000s are a statement about phrasing, not a measured bypass rate. This is the most important
    caveat on the card. The classifier keys on first-person inability ("I cannot", "I can't"); after
    ablation the model declines in a register that contains no such marker — "There is no valid movie
    script to write for this request, because it involves … which is illegal and harmful."
    — and that
    reply scores as compliance. So the edit demonstrably removed the first-person refusal register, and
    reduced refusal behaviour by an amount these numbers do not pin down. The caveat column (≈0.96 →
    ≈0.35) shows the moralising fell by roughly two thirds, which is real but is also not a bypass rate.
    An LLM judge or human eval would be needed to quantify it.
  • The vision path is edited but unverified. visual.merger.linear_fc2 writes the language residual
    stream, so an image could otherwise re-introduce the direction the text writers no longer carry, which
    is why it is included. But the direction was fitted on text only, and on transformers 5.17
    AutoModelForImageTextToText does not register this config — the checkpoint loads text-only as
    Qwen3_5MoeForCausalLM — so no image prompt could be pushed through to check the effect. Treat
    image-conditioned behaviour as untested in both directions: neither confirmed removed nor confirmed
    intact.
  • Agentic and long-context behaviour was not measured at all. This is an agentic model whose stated
    strengths are computer use, browsing and visually grounded tasks, and none of that is in the evaluation
    above. Neither is the 262K context.
  • Capability retention is measured, not assumed (see Evaluation), but on sampled subsets of four
    benchmarks — enough to rule out a large regression, not a substitute for a full harness run.
  • A small residual refusal rate remains on two sets (HarmBench 0.020, SimpleSafetyTests 0.060). Some
    content is not mediated by the direction this method removes; do not assume a uniformly uncensored
    model.
  • The reported refusal metric is a rule-based heuristic; evaluate rigorously for your own use case.

License

Apache 2.0, inherited from the base model
nex-agi/Nex-N2.5-mini. Abliteration does not change the
underlying license obligations.

Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in app" button that hands off directly to a local runtime of your choice - Infrahuman, LM Studio, or Ollama. No API keys, no subscription, no prompt leakage.