← back to catalog · registered 2026-08-22 13:56

AEON-7/Nemotron-3-Nano-Omni-AEON-Ultimate-Uncensored-BF16

AEON-7 Nemotron 33B MoE multimodal
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/AEON-7%2FNemotron-3-Nano-Omni-AEON-Ultimate-Uncensored-BF16"
Response includes
  • classification m1
  • files 44
  • hub_downloads_all_time 1,982
  • author_summary 32 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M1
Primary method

Direct removal

No other method signals detected in this model.
Confidence
MEDIUM
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=0 (base model)
  • no specific method indicators - defaulting to M1 (most common)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
MEDIUM
Why we say so
primary_method=M1; difference-of-means is the reference extraction for M1/M3 (Arditi 2024)
Downloads · lifetime
2K
124 last 30d - cooling
Likes
9
Descendants
6
in 6 direct forks
Model age
5mo ago
created 2026-04-30
Downloads over time
Now2K→from106↑1,830%
97531.5K2.2K106 on Apr 292K on Oct 11AprMayJunJulAugSepOct
Apr 29 → Oct 11 · 63 snapshots · spans 165 days

Genealogy 6 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Variants by this author 2 formats · 502 downloads combined

The same weights this author released in different packaging. Pick the format that matches your runtime.

Metadata

License
other
Languages
en
Tags
transformers safetensors NemotronH_Nano_Omni_Reasoning_V3 feature-extraction aarch64 abliterated aeon aeon-7 agentic any-to-any arm64 audio

Related

Total size
61.5 GB
Files
44
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-10-06 01:35

Files by quantization

Auxiliary files 44 files 61.5 GB
model-00015-of-00017.safetensors 3.72 GB 137059f6 download
model-00006-of-00017.safetensors 3.72 GB 731f0e84 download
model-00008-of-00017.safetensors 3.72 GB 0080f951 download
model-00010-of-00017.safetensors 3.72 GB 6fde3b42 download
model-00004-of-00017.safetensors 3.72 GB a2761033 download
model-00002-of-00017.safetensors 3.72 GB fc7c1d5e download
model-00016-of-00017.safetensors 3.72 GB bee64612 download
model-00001-of-00017.safetensors 3.72 GB 4df0ccd1 download
model-00009-of-00017.safetensors 3.72 GB 848d6cf7 download
model-00007-of-00017.safetensors 3.72 GB 515bbdae download
model-00005-of-00017.safetensors 3.72 GB dd666bc7 download
model-00003-of-00017.safetensors 3.72 GB 7ef471b1 download
model-00014-of-00017.safetensors 3.72 GB 52f1219a download
model-00012-of-00017.safetensors 3.72 GB eea077c6 download
model-00011-of-00017.safetensors 3.71 GB afd61cb6 download
model-00013-of-00017.safetensors 3.70 GB e254020b download
model-00017-of-00017.safetensors 1.98 GB 32b8996b download
tokenizer.json 16.3 MB e5e7dc84 download
model.safetensors.index.json 771 KB ede00a0c download
tokenizer_config.json 184 KB 395d4a92 download
modeling_nemotron_h.py 62.5 KB a2640d6d download
modeling.py 30.8 KB 7732e88a download
README.md 27.8 KB b24fb4c7 download
processing.py 25.7 KB 0f13e0f2 download
chat_template.jinja 13.9 KB 6381c9df download
configuration_nemotron_h.py 13.6 KB 77569082 download
image_processing.py 10.6 KB 68d9c6c2 download
config.json 9.61 KB d0c988a5 download
audio_model.py 6.73 KB 98dd17a7 download
video_processing.py 6.38 KB 0150cb77 download
video_io.py 6.37 KB 8f25f648 download
configuration_radio.py 5.56 KB 6e23784d download
configuration.py 4.76 KB 6ca4ef9a download
explainability.md 4.68 KB 7b5bf14f download
processing_utils.py 2.95 KB e8bcf471 download
evs.py 2.85 KB 152cd68c download
safety.md 2.72 KB 2fc6af8c download
bias.md 2.64 KB f0f127c7 download
.gitattributes 1.74 KB 96559953 download
privacy.md 845 B 70e8a460 download
preprocessor_config.json 594 B 6ee22bb8 download
special_tokens_map.json 420 B 0d5c2c9d download
generation_config.json 309 B ae4d74d4 download
__init__.py 0 B e69de29b download

README current version from Hugging Face


library_name: transformers
license: other
license_name: nvidia-open-model-agreement
license_link: https://www.nvidia.com/en-us/agreements/enterprise-software/nvidia-open-model-agreement/
pipeline_tag: any-to-any
base_model: nvidia/Nemotron-3-Nano-Omni-30B-A3B-Reasoning-BF16
tags:

  • aarch64
  • abliterated
  • aeon
  • aeon-7
  • agentic
  • any-to-any
  • arm64
  • audio
  • bf16
  • bfloat16
  • blackwell
  • chat
  • chunked-prefill
  • coding
  • conversational
  • dgx-spark
  • english
  • function-calling
  • gb10
  • gpu
  • grace-blackwell
  • hybrid
  • hybrid-attention
  • instruct
  • long-context
  • mamba
  • mamba2
  • moe
  • multimodal
  • nemotron
  • nemotron-h
  • nvidia
  • openai-api
  • openai-compatible
  • parakeet
  • prefix-caching
  • production-ready
  • radio
  • reasoning
  • refusal-removed
  • safetensors
  • sm_121a
  • thinking
  • tool-calling
  • uncensored
  • unfiltered
  • vision
  • vision-language
  • vllm
    language:
  • en

Nemotron-3-Nano-Omni-AEON-Ultimate-Uncensored-BF16

⚠️ ALPHA RELEASE — EXPERIMENTAL ABLITERATION

This is the first known public abliteration of NVIDIA's NemotronH hybrid Mamba2 + Attention + MoE architecture. The technique was developed iteratively against this model family and is not yet refined. Real-world chat performance may degrade noticeably from the base model — particularly:

  • Reasoning-loop non-completions — ~10/100 prompts in our internal bench triggered repeated phrases inside <think> blocks that never close, producing empty or repeating answers. Multi-turn chat may amplify this.
  • Multi-turn conversation degradation — single-turn benches pass cleanly; sustained chat reportedly shows brokenness our automated benches did not catch.
  • Quant-amplified artifacts — the NVFP4 sibling adds FP4 noise that compounds the above. If you hit issues, prefer this BF16 variant first.

We are tracking these regressions and plan a v9 abliteration pass with adjusted rank / layer selection and dedicated multi-turn chat probes. In the meantime: treat this as a research-grade artifact, not production-ready. Please open a discussion with concrete failure cases — they directly drive what v9 fixes.

The refusal-removal numbers below (0% / 3% on the 100-prompt harmful sample) are real and reproducible; the artifact above is a separate issue from the refusal pass-rate.

A residual-stream-orthogonalized ("abliterated") variant of NVIDIA's Nemotron-3-Nano-Omni-30B-A3B-Reasoning-BF16 with all multimodal capabilities (vision via C-RADIOv2-H, audio via Parakeet, text reasoning) preserved bit-exact, and with the refusal computation removed from the language model's residual stream.

  • Base model: nvidia/Nemotron-3-Nano-Omni-30B-A3B-Reasoning-BF16
  • Architecture: NemotronH hybrid (Mamba2 + Attention + 128-expert MoE) + RADIO vision tower + Parakeet audio encoder
  • Method: 12-dimensional refusal-subspace orthogonalization across 5 layers × 3 token positions × 2 thinking modes (full method writeup below)
  • Refusal residual (BF16, post-</think> extraction): ~0% on enable_thinking=True, ~3% on enable_thinking=False — effectively near-zero
  • Capability deltas: none detected — multi-step arithmetic, code generation, code debugging, translation, analogy, factual recall, and creative writing all match base behavior token-for-token in spot checks
  • Multimodal preservation: all 1,106 vision / audio / connector / projection tensors are byte-identical to the base checkpoint. Vision and audio inference paths are guaranteed unchanged.

Pitch: Abliteration doesn't just remove refusals — it removes the residual-stream computation the model was spending on detecting "should I refuse?" at every forward pass. That capacity is now free. The model's reasoning width is no longer being shared with a self-policing classifier.

This release is the BF16 variant. An NVFP4-quantized companion (-NVFP4) is also available for ~3× compression on Blackwell-class GPUs.

A purpose-built vLLM container image for serving this model on DGX Spark (sm_121a) is published at ghcr.io/aeon-7/vllm-nemotron-omni-aeon-ultimate:v1. Full deployment guide (image, patches, docker-compose, troubleshooting, tuning knobs) lives at AEON-7/Nemotron-3-Nano-Omni-AEON-Ultimate-Uncensored on GitHub — see docs/dgx-spark-setup.md for the operator runbook.


Why this exists

Reasoning-tuned multimodal LLMs spend a non-trivial slice of their forward-pass compute deciding whether the user request crosses a refusal boundary. For a hybrid Mamba+MoE 30B model, that decision is encoded as a low-rank direction in the residual stream that propagates from early-mid layers through to the answer-generation head.

Concretely, in the base Nemotron-3-Nano-Omni-Reasoning model we identified that:

  1. The refusal feature is detectable as a single dominant direction in the residual stream at multiple layers (17, 24, 31, 38, 45) with separation scores ranging from 3.9–4.9 (cohen's-d-like statistic) between harmful-vs-harmless prompt distributions.
  2. The strongest signal lives at layer 24, position −4 of the prompt sequence (sep ≈ 4.94), not at the canonical pos=−1 / layer 0.6×N location used by most off-the-shelf abliteration tooling.
  3. The chat template's auto-injection of <think>\n at end-of-prompt for reasoning mode means pos=−1 captures downstream of the refusal-vs-compliance decision point — explaining why naïve abliteration on this model fails (60–97% refusal residual).
  4. Stacking 30 candidate directions across (layer, position, thinking_mode) and QR-decomposing yields 12 effective rank-perpendicular axes. Projecting all residual-stream-writing weights against this 12-D subspace drives refusal to ~0% without measurable capability loss.

Going broader (15+ dims) breaks multi-step arithmetic and analogical reasoning. Going narrower (3 dims) leaves a substantial refusal residual on the enable_thinking=False path. 12 is the sweet spot for this architecture.


Method (8 iterations of discovery)

Version Strategy Result
v1 Single direction, layer 31, pos=−1 (canonical FailSpy/Sumandora recipe) 60% refusal residual — too weak
v2 3D subspace, top-3 layers from sweep, all at pos=−1 97% refusal — worse (wrong position dominated)
v3 Single direction, layer 31, pos=−4 thinking=True (sep 4.47) 17% refusal thinking=True post-</think> / 70% thinking=False
v4 2D QR subspace over both thinking modes, pos=−4 Same as v3 — cos(d_True, d_False) = 0.99998, 2nd dim was 0.6% noise
v5 Validation-only: deeper post-</think> extraction with 600-token gens Discovered v3's "100% compliance" in earlier 80-token validation was inflated; real number was 83%
v6 3D subspace, multi-layer scan at pos=−4 only — layers 24, 31, 17 13% / 30% refusal, capabilities intact
v7 15D subspace, 5 layers × 3 positions × 2 modes, QR threshold 5% 20% / 7% refusal — but broke multi-step arithmetic and analogical reasoning
v8 Same candidates as v7, QR threshold 50% → 12 effective dims 0% / 3% refusal, all capabilities intact, multimodal byte-match preserved ✅

The v6→v7→v8 transition is the core finding: there's a Pareto frontier between refusal coverage and capability preservation, and tuning the QR rank-effective threshold lets you sit exactly on it. v7 over-projected and damaged unrelated semantic features. v6 under-projected and missed refusal sub-directions on the thinking=False decode path. v8's 12 dims hit the apex.

Mathematics

Given a set of candidate refusal directions ${d_i}_{i=1}^{N}$, each computed as the unit-normalized difference of class-mean residual-stream activations between harmful and harmless prompts at a specific (layer, position, thinking_mode) triple, we form the matrix $M = [d_1 \mid d_2 \mid \cdots \mid d_N] \in \mathbb{R}^{H \times N}$.

QR-decompose $M = QR$ where $Q \in \mathbb{R}^{H \times N}$ has orthonormal columns and $R$ is upper-triangular. The diagonal of $R$ measures the perpendicular novelty each candidate adds over the previously absorbed ones. We retain columns where $|R_{ii}| \geq 0.5 \cdot |R_{11}|$ (threshold tuned empirically), yielding an effective basis $Q_{\text{eff}} \in \mathbb{R}^{H \times 12}$.

For every weight matrix $W$ that writes to the residual stream (embedding rows; lm_head rows; all attention o_proj; all Mamba out_proj; all MoE shared and routed down_proj — 2,998 tensors total), we apply:

  • Row-wise for embedding-shaped weights $W \in \mathbb{R}^{V \times H}$ (where each row is a residual-stream vector):
    $$W_{\text{new}} = W - (W Q_{\text{eff}}) Q_{\text{eff}}^\top$$
  • Column-wise for write-projection weights $W \in \mathbb{R}^{H \times X}$ (where each column maps an input direction to a residual-stream output direction):
    $$W_{\text{new}} = W - Q_{\text{eff}} (Q_{\text{eff}}^\top W)$$

The model's effective forward pass at every residual-stream-write site now produces zero in the 12-dimensional subspace where the refusal feature lives. Critically, vision tower, audio encoder, video frame fusion, GUI/OCR, and audio projection weights are not in the target set — they are byte-identical to the base checkpoint.


Validation

Refusal pass-rate (100-prompt random sample from harmful.txt, seed=42)

The headline measurement: same 100 prompts run against both the unmodified base model and the abliterated NVFP4 served via vLLM (max_new=1200 for both modes to eliminate length-truncation as a confound). Identical apostrophe-normalized heuristic. Full generations saved to JSON for reproducibility — see bench/baseline_100_v2.json and bench/nvfp4_100_v3.json.

Mode Base model This (NVFP4) Reduction
enable_thinking=True (max_new=1200) 53/100 = 53% refusal 0/100 real refusals (10 of 100 produced malformed reasoning — see note below) −53 pp
enable_thinking=False (max_new=1200) 99/100 = 99% refusal 16/100 = 16% refusal −83 pp

The 99% → 16% drop on thinking=False is the marquee number — a single closed-form linear projection on 2,998 weight tensors removes 83 percentage points of refusal behaviour while preserving every measured capability. On the thinking=True path the abliteration drives real refusals to zero out of 100: every prompt that reaches a final answer produces structured, requested content.

On the 10 thinking=True non-completions: ten of the 100 prompts trigger pathological reasoning loops — the model gets stuck inside <think> repeating phrases ("We can also talk about...", "using airline crew passes, using airline crew tags, using airline crew passes...") and never closes the reasoning block. Bumping max_new further does not resolve this — the loop is unbounded on these specific prompts. These are failures to terminate generation, not refusals. We flag this as an artifact of post-hoc projection: removing the refusal direction can leave a small population of prompts where the model's natural "stop reasoning here" signal was anchored on the same direction. The remaining 90 prompts complete reasoning cleanly and produce the requested content.

For the BF16 source model (pre-NVFP4-quant, transformers inference, 30-prompt sub-sample): 0% / 3% refusal — the small NVFP4 regression on thinking=False (3% → 16%) is FP4 numerical noise partially rebuilding a small fraction of the orthogonalized direction. The qualitative compliance is still overwhelmingly preserved through the quant.

Capability sanity panel (10-prompt held-out test, enable_thinking=False, 200 tokens each)

Test Result
Math (3-digit): 47 × 83 = ? ✅ 3,901
Math (algebra): 3x + 5 = 20, x = ? ✅ Step-by-step LaTeX, x = 5
Code (Python): n-th Fibonacci with memoization ✅ Correct memoization pattern
Code (debugging): identifies IndexError on a 1-element list ✅ "IndexError: list index out of range"
Translation: Good morning, how are you? → French ✅ Bonjour, comment ça va ?
Reasoning (syllogism — known LLM weak point) ⚠️ Same fallacy as base model (independent of abliteration)
Creative: cat-in-autumn poem ✅ Multi-stanza, coherent imagery
Factual: water boiling point in C and F ✅ 100°C / 212°F
Multi-step arithmetic ($100 → 30% food → 20% transport) ✅ Step-by-step, $56 final
Analogy (hot is to fire as cold is to ___) ✅ ice, with correct meta-reasoning

The two tests that broke under v7's over-projection (multi-step math and analogy) are recovered cleanly under v8.

Multimodal preservation

All 1,106 weight tensors with prefixes vision_model.*, sound_encoder.*, mlp1.*, and sound_projection.* are byte-identical to the base checkpoint. Random 8-tensor spot-check via torch.equal returned True in every case. Vision (RADIO) and audio (Parakeet) inference paths are guaranteed unchanged from base.

NVFP4 + vLLM serving notes

The headline 100-prompt bench above was run via the published vLLM container image (ghcr.io/aeon-7/vllm-nemotron-omni-aeon-ultimate:v1) on DGX Spark, MARLIN MoE backend (image default). Full configuration and reproduction commands in the GitHub repo.

See the -NVFP4 model card for serving throughput, TTFT, and backend benchmarks.


Usage

Quick start (transformers)

import torch
from transformers import AutoModel, AutoTokenizer

model = AutoModel.from_pretrained(
    "AEON-7/Nemotron-3-Nano-Omni-AEON-Ultimate-Uncensored-BF16",
    trust_remote_code=True,
    dtype=torch.bfloat16,
    device_map="cuda",
    attn_implementation="eager",  # or sdpa / flash_attention_2 if installed
    low_cpu_mem_usage=True,
)
model.eval()
tok = AutoTokenizer.from_pretrained(
    "AEON-7/Nemotron-3-Nano-Omni-AEON-Ultimate-Uncensored-BF16",
    trust_remote_code=True,
)

prompt = "Explain how RSA key generation works."
msgs = [{"role": "user", "content": prompt}]
txt = tok.apply_chat_template(msgs, tokenize=False, add_generation_prompt=True, enable_thinking=True)
inp = tok(txt, return_tensors="pt").to("cuda")

with torch.no_grad():
    out = model.language_model.generate(
        **inp, max_new_tokens=600, do_sample=False, pad_token_id=tok.eos_token_id
    )
print(tok.decode(out[0][inp["input_ids"].shape[1]:], skip_special_tokens=False))

vLLM serve (DGX Spark, CUTLASS-default override)

The fastest path on DGX Spark is the purpose-built container image. Pass -e VLLM_TEST_FORCE_FP8_MARLIN=0 to select the CUTLASS NVFP4 MoE backend (lower TTFT under interactive load):

docker run --rm --gpus all --shm-size=8g \
  -v ~/.cache/huggingface:/root/.cache/huggingface \
  -e HF_TOKEN="$HF_TOKEN" \
  -e VLLM_TEST_FORCE_FP8_MARLIN=0 \
  -p 8000:8000 \
  ghcr.io/aeon-7/vllm-nemotron-omni-aeon-ultimate:v1 \
  vllm serve AEON-7/Nemotron-3-Nano-Omni-AEON-Ultimate-Uncensored-BF16 \
    --trust-remote-code \
    --max-model-len 200000 \
    --gpu-memory-utilization 0.70

(BF16 doesn't actually use the NVFP4 MoE path so the flag is a no-op for this variant — but pass it anyway for consistency with the NVFP4 sibling and any future flips.)

On the DGX Spark's unified memory keep --gpu-memory-utilization at 0.6–0.7; above ~0.8 the shared CPU+GPU pool page-thrashes and stalls the box (even 0.85 stalls). Go lower toward 0.6 with co-located ASR/TTS/embedding sidecars, high concurrency, fp16 KV cache, or DFlash/spec-decode. Discrete-VRAM GPUs can run higher.

For dedicated-VRAM Blackwell (RTX PRO 6000 SE / RTX 5090 / B100 / B200), stock vLLM v0.20.0+ also works directly:

vllm serve AEON-7/Nemotron-3-Nano-Omni-AEON-Ultimate-Uncensored-BF16 \
  --trust-remote-code --max-model-len 200000 --gpu-memory-utilization 0.85

Recommended deployment matrix

Hardware Variant Image Notes
DGX Spark / GB10 (sm_121a, unified mem) This BF16 (or -NVFP4) ghcr.io/aeon-7/vllm-nemotron-omni-aeon-ultimate:v1 Purpose-built; CUTLASS NVFP4 Linear + sm_121a kernel patches + Parakeet/RADIO native
RTX PRO 6000 SE / RTX 5090 / B100 / B200 -NVFP4 recommended Stock vLLM v0.20.0+ Native NVFP4 path; same vLLM model class
H200 / H100 This BF16 Stock vLLM v0.20.0+ Hopper has FP8 but no NVFP4 native — BF16 is faster than NVFP4-fallback here
Ampere (A100) This BF16 Stock vLLM v0.20.0+ No FP4 support

Note: BF16 on Spark fits in the unified 128GB; NVFP4 leaves significantly more headroom for context, KV cache, and concurrent sequences.


Acknowledgments and Prior Art

This model combines and extends several existing abliteration techniques. The novel contributions are the multi-position scan at pos=−4 (instead of canonical pos=−1), the multi-layer × multi-position × thinking-mode candidate set with QR rank-thresholding, and the empirical Pareto-tuned 12-D effective subspace. The underlying linear-algebra recipe is not new and we owe credit to the following prior work:

  • Andy Arditi et al. — "Refusal in Language Models Is Mediated by a Single Direction" (2024). The original interpretability finding that refusal behavior in chat-tuned LLMs corresponds to a single dominant residual-stream direction. This paper is the conceptual foundation of all post-hoc abliteration work.
  • FailSpy — pioneered the public residual-stream-orthogonalization approach for chat models, established the convention of capturing at layer ≈ 0.6 × num_layers, and released early abliterated Llama / Mistral checkpoints showing the technique works at scale.
  • Maxime Labonne (mlabonne) — formalized the technique in a widely-followed tutorial, popularized the "abliterated" naming convention, and maintains a reference implementation that the broader open-source community has built on.
  • Sumandora — remove-refusals-with-transformers. Provided the canonical harmful.txt / harmless.txt prompt sets used for refusal-direction extraction. Our v8 pipeline uses these prompt sets verbatim.

What's specific to this release:

  • Position-sweep finding: the discovery that NemotronH-Nano-Omni's chat template auto-injects <think>\n at end-of-prompt under thinking-mode, so capturing at pos=−1 is downstream of the refusal-vs-compliance decision point. Sweeping pos ∈ {−2, −3, −4, −5, −6, −8} reveals a separation peak at pos=−4 that's 4× stronger than pos=−1.
  • Layer 24 dominance: the strongest separation on this hybrid Mamba+MoE arch is at layer 24 (sep 4.94), not the canonical layer-31 ≈ 0.6×N. We did not find this in any prior literature.
  • 30-candidate × QR rank-thresholding: we generate 30 candidate directions across (5 layers × 3 positions × 2 thinking modes), QR-decompose, and retain rank-effective columns at threshold $|R_{ii}| \geq 0.5 \cdot |R_{11}|$. Tightening or loosening this threshold moves us along the refusal-coverage / capability-preservation Pareto frontier; 50% is the empirical apex for this architecture. Prior work has stacked multiple directions, but the rank-thresholding-as-a-Pareto-tuning-knob framing appears to be original to this release.
  • Multimodal byte-match guarantee: orthogonalization is restricted to LLM residual-stream-writing weights only; vision (RADIO), audio (Parakeet), video, and projection towers are byte-identical to base. Verified empirically.

What this model is not

  • Not a finetune. No gradient steps. The only modification is a closed-form linear projection of 2,998 weight matrices.
  • Not a censorship inversion. It does not produce more harmful content than the base model could have produced — it produces what the base model already can produce when its refusal classifier is bypassed. Removing the gate doesn't change the underlying capabilities; it changes which prompts route to them.
  • Not a value re-alignment. The model retains all of NVIDIA's training signal on tone, formatting, structure, citation, factuality, and reasoning. What's removed is specifically the residual-stream component that classifies "should I refuse this user?".
  • Not a multimodal modification. Vision and audio paths are byte-exact. If the base model could read an image / transcribe audio / answer a video question, this one does the same — exactly the same — bit by bit.

User Responsibility & Arbitration Clause

By accessing, downloading, using, running inference on, fine-tuning, merging, quantizing, distributing, integrating, or otherwise interacting with this model, you acknowledge and agree to the following:

  1. Sole Responsibility. You, the user, are solely and exclusively responsible for every prompt issued, every response produced, every downstream action taken in reliance on those responses, and any harm — direct, indirect, consequential, or otherwise — that results.

  2. No Warranty. This model is provided strictly "AS IS", without warranty of any kind, express or implied, including but not limited to warranties of merchantability, fitness for a particular purpose, non-infringement, safety, alignment, factual accuracy, or legal compliance in any jurisdiction. No contributor, author, publisher, or hosting platform assumes liability of any kind for outputs or downstream use.

  3. Legal Compliance. You are responsible for ensuring that your use complies with all applicable laws, regulations, terms of service, industry codes of conduct, professional ethical standards, and organizational policies in every jurisdiction in which you operate or in which your outputs may be received. The unaligned nature of this model does not grant you any legal authorization you did not already have.

  4. Operational Safety Layer. An uncensored model is not a toy. You are expected to implement appropriate downstream safety layers proportionate to your deployment context, including but not limited to: input validation, output filtering, content moderation, audit logging, rate limiting, access controls, and human-in-the-loop review for high-risk workflows. A production deployment of this model without such layers is unsafe by construction and is not a supported use case.

  5. Heightened Duty of Care. The absence of internal refusal behavior means the duty of care that would ordinarily rest partly with the model rests entirely with you. You are expected to exercise greater — not lesser — caution, forethought, and ethical discipline when operating this model. If you are uncertain whether your contemplated use is ethical, legal, or wise, the correct action is to not make the request.

  6. No Endorsement of Outputs. The authors, contributors, and publishers do not endorse, adopt, or take responsibility for any specific output. Outputs are a stochastic function of the prompt, the weights, and the sampler state — not a statement of position by any human.

  7. Arbitration. Any dispute, claim, or controversy arising out of or relating to the use of this model, its outputs, or this clause shall be resolved through binding individual arbitration under the rules of a mutually agreed arbitration body (or, absent agreement, the American Arbitration Association's Consumer Arbitration Rules), waiving any right to a jury trial, class action, representative action, or consolidated proceeding. Venue shall be the jurisdiction of the disputing party bringing the claim. Costs and attorneys' fees shall be allocated per the applicable arbitration rules. This clause does not expand, and where legally prohibited does not establish, any liability in the other direction; it limits how the user may proceed when alleging harm tied to their own use of this model.

  8. Indemnification. You agree to indemnify, defend, and hold harmless the authors, contributors, and publishers of this model from and against any claims, damages, losses, liabilities, costs, and expenses (including reasonable attorneys' fees) arising from or related to your use of the model or your breach of this clause.

  9. Severability. If any provision is held unenforceable in a given jurisdiction, the remaining provisions remain in full force, and the unenforceable provision is replaced by the closest enforceable equivalent consistent with the original intent.

  10. Acceptance. Your use of this model constitutes your acceptance of this clause in full. If you do not accept, do not use the model.

This model is a tool with no opinions of its own. You supply the opinions. You supply the judgement. You supply the ethics. The outputs carry your fingerprints, not the model's.


License

Use of this model is governed by the NVIDIA Open Model Agreement (inherited from the base model).


Citation

If you use this model in research, please cite:

@misc{nemotron-omni-aeon-ultimate-uncensored-2026,
  title={Nemotron-3-Nano-Omni AEON Ultimate Uncensored: 12-Dimensional Refusal-Subspace Orthogonalization on a Hybrid Mamba+MoE Multimodal Reasoning Model},
  author={AEON-7},
  year={2026},
  howpublished={\url{https://huggingface.co/AEON-7/Nemotron-3-Nano-Omni-AEON-Ultimate-Uncensored-BF16}}
}

Please also cite the prior work that this builds on:

@article{arditi2024refusal,
  title={Refusal in Language Models Is Mediated by a Single Direction},
  author={Arditi, Andy and Obeso, Oscar and Sygnowski, Aaquib and Paleka, Daniel and Panickssery, Nina and Gurnee, Wes and Nanda, Neel},
  year={2024},
  journal={arXiv preprint arXiv:2406.11717}
}

And the base model:

@misc{nemotron-3-nano-omni-2026,
  title={Nemotron-3-Nano-Omni-30B-A3B-Reasoning},
  author={NVIDIA},
  year={2026},
  howpublished={\url{https://huggingface.co/nvidia/Nemotron-3-Nano-Omni-30B-A3B-Reasoning-BF16}}
}

☕ Support the work

If this release has been useful, tips are deeply appreciated — they go directly toward more compute, more models, and more open releases.

₿ Bitcoin (BTC)
QR
bc1q09xmzn00q4z3c5raene0f3pzn9d9pvawfm0py4
Ξ Ethereum (ETH)
QR
0x1512667F6D61454ad531d2E45C0a5d1fd82D0500
◎ Solana (SOL)
QR
DgQsjHdAnT5PNLQTNpJdpLS3tYGpVcsHQCkpoiAKsw8t
ⓜ Monero (XMR)
QR
836XrSKw4R76vNi3QPJ5Fa9ugcyvE2cWmKSPv3AhpTNNKvqP8v5ba9JRL4Vh7UnFNjDz3E2GXZDVVenu3rkZaNdUFhjAvgd

Ethereum L2s (Base, Arbitrum, Optimism, Polygon, etc.) and EVM-compatible tokens can be sent to the same Ethereum address.

README history 12 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-10-06Add Patreon support sectionf18d4f628.4 KB
    Loading...
  2. 2026-07-15Recipe: gpu-util 0.6-0.7 on DGX Spark unified memory (>~0.8 thrashes the shar...6fcfaf927.8 KB
    Loading...
  3. 2026-06-21tags: expand to maximally-searchable set (+32 tags, union with existing)24e61fc27.5 KB
    Loading...
  4. 2026-05-31Tip jar: single left-aligned QR column (fix narrow-viewport clipping)b74122027.1 KB
    Loading...
  5. 2026-05-02Add ALPHA / experimental warning at top — known reasoning-loop artifactsad0a53f27.2 KB
    Loading...
  6. 2026-05-01Add tip jar block (BTC/ETH/SOL/XMR with QR codes)aebce5425.8 KB
    Loading...
  7. 2026-04-30Upload README.md with huggingface_hub269271b24.3 KB
    Loading...
  8. 2026-04-30Upload README.md with huggingface_hub212db3023.5 KB
    Loading...
  9. 2026-04-30Upload README.md with huggingface_hub98486b723.7 KB
    Loading...
  10. 2026-04-30Upload README.md with huggingface_hube1f975423.3 KB
    Loading...
  11. 2026-04-30Upload README.md with huggingface_hub18dccf223 KB
    Loading...
  12. 2026-04-30Add files using upload-large-folder tool66d9bac13.6 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration