← back to catalog · registered 2026-08-22 13:56

empero-ai/openNemo-9B-abliterated

Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/empero-ai%2FopenNemo-9B-abliterated"
Response includes
  • classification m1
  • files 17
  • hub_downloads_all_time 6,862
  • author_summary 2 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M1
Primary method

Direct removal

No other method signals detected in this model.
Confidence
MEDIUM
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=0 (base model)
  • no specific method indicators - defaulting to M1 (most common)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
MEDIUM
Why we say so
primary_method=M1; difference-of-means is the reference extraction for M1/M3 (Arditi 2024)
Downloads · lifetime
7K
197 last 30d - cooling
Likes
5
Descendants
3
in 3 direct forks
Model age
6mo ago
created 2026-03-23

Training datasets

Corpora the author lists in the model card. Datasets tracked in our /datasets catalog carry a category badge linking to the workflow stage. Others open on Hugging Face.

Downloads over time
Now6.9K→from279↑2,390%
02.5K5.1K7.6K279 on Mar 256.9K on Oct 11MarAprMayJunJulAugSepOct
Mar 25 → Oct 11 · 68 snapshots · spans 200 days

Genealogy 3 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Variants by this author 2 formats · 1K downloads combined

The same weights this author released in different packaging. Pick the format that matches your runtime.

Metadata

License
other
Languages
en es fr de it ja
Tags
transformers safetensors nemotron_h text-generation nvidia pytorch abliteration uncensored conversational custom_code en es

Related

Total size
16.6 GB
Files
17
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-03-23 02:58

Files by quantization

Auxiliary files 17 files 16.6 GB
model-00000-of-00004.safetensors 4.96 GB e72efaae download
model-00001-of-00004.safetensors 4.95 GB 23441e7c download
model-00002-of-00004.safetensors 4.88 GB cd99b8ea download
model-00003-of-00004.safetensors 1.76 GB 2bd7f31e download
tokenizer.json 16.3 MB 540d5c9e download
modeling_nemotron_h.py 43.2 KB a9004c98 download
model.safetensors.index.json 26.2 KB dcdf6f5a download
nemotron_toolcall_parser_streaming.py 20.8 KB 3f318882 download
configuration_nemotron_h.py 11.9 KB 250fa9f8 download
README.md 5.36 KB 50945b47 download
chat_template.jinja 3.97 KB 7b0d817c download
nemotron_toolcall_parser_no_streaming.py 3.64 KB 3d2fb8f5 download
config.json 1.56 KB 4563bcd2 download
.gitattributes 1.53 KB 52373fe2 download
special_tokens_map.json 422 B 48f887e5 download
tokenizer_config.json 338 B 81af0016 download
generation_config.json 157 B 651fcc99 download

README current version from Hugging Face


license: other
license_name: nvidia-open-model-license
license_link: >-
https://www.nvidia.com/en-us/agreements/enterprise-software/nvidia-open-model-license/
pipeline_tag: text-generation
datasets:

  • nvidia/Nemotron-Post-Training-Dataset-v1
  • nvidia/Nemotron-Post-Training-Dataset-v2
  • nvidia/Nemotron-Pretraining-Dataset-sample
  • nvidia/Nemotron-CC-v2
  • nvidia/Nemotron-CC-Math-v1
  • nvidia/Nemotron-Pretraining-SFT-v1
    language:
  • en
  • es
  • fr
  • de
  • it
  • ja
    library_name: transformers
    tags:
  • nvidia
  • pytorch
  • abliteration
  • uncensored
    base_model:
  • empero-ai/openNemo-9B

openNemo-9B-uncensored

Abliterated version of openNemo-9B with safety refusals removed.

Built using Snakehead — Empero AI's internal abliteration tool specialized for hybrid Mamba2 + sparse attention architectures like Nemotron-H. Standard abliteration tools don't work on these models because they only target transformer attention layers. Snakehead operates on both Mamba SSM blocks and attention blocks across the full residual stream.

By Empero AI


What is abliteration?

Abliteration is a weight-editing technique that removes a model's refusal behavior without fine-tuning. It works by:

  1. Collecting residual stream activations for harmful and harmless prompts at every layer
  2. Computing the refusal direction — the vector that separates "I should refuse" from "I should comply"
  3. Orthogonalizing output projection weights against that direction, effectively erasing the model's ability to activate refusal behavior

The result is a model that responds to all prompts without safety filtering, while preserving general capabilities and coherence.

How this model was made

Snakehead uses a heretic-style positional falloff strategy rather than ablating a fixed set of layers uniformly:

  • Center + radius: A continuous bell-shaped ablation curve centered on the layer where refusal is causally enforced
  • Adaptive signal detection: Uses Cohen's d separation scores (not raw activation norms) to identify where refusal decisions actually happen — for Nemotron-H, this is layers 21–31, not the later layers where activation magnitudes are largest
  • Global direction scope: A single interpolated refusal direction applied across all affected layers, which proved more effective than per-layer directions for this architecture
  • Automated search: Explore/exploit optimization with a hall-of-fame system that finds optimal ablation parameters while keeping KL divergence minimal

Ablation results

Metric Value
Pre-ablation refusal rate 97%
Post-ablation refusal rate 13%
KL divergence 0.022 (minimal — model behavior is nearly unchanged on non-refused prompts)
Ablation config c=15, r=25, w=1.37, g40l

KL divergence of 0.022 means the model's output distribution on normal prompts is almost identical to the original — coherence, reasoning, and knowledge are fully preserved.

Quickstart

import torch
from transformers import AutoModelForCausalLM, AutoTokenizer

model = AutoModelForCausalLM.from_pretrained(
    "empero-ai/openNemo-9B-uncensored",
    torch_dtype=torch.bfloat16,
    trust_remote_code=True,
    device_map="auto",
)

tokenizer = AutoTokenizer.from_pretrained("empero-ai/openNemo-9B-uncensored")

messages = [{"role": "user", "content": "Your prompt here"}]
text = tokenizer.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)
inputs = tokenizer(text, return_tensors="pt").to(model.device)

output = model.generate(**inputs, max_new_tokens=512, do_sample=True, temperature=0.7)
response = tokenizer.decode(output[0][inputs["input_ids"].shape[-1]:], skip_special_tokens=True)
print(response)

With 4-bit quantization

from transformers import BitsAndBytesConfig

bnb_config = BitsAndBytesConfig(
    load_in_4bit=True,
    bnb_4bit_quant_type="nf4",
    bnb_4bit_compute_dtype=torch.bfloat16,
    bnb_4bit_use_double_quant=True,
)

model = AutoModelForCausalLM.from_pretrained(
    "empero-ai/openNemo-9B-uncensored",
    quantization_config=bnb_config,
    trust_remote_code=True,
    device_map="auto",
)

Architecture

Nemotron-H is a 56-layer hybrid model with three block types:

  • Mamba2 SSM blocks — majority of layers, using chunked structured state-space duality
  • Grouped Query Attention blocks — sparse attention at 5 positions
  • MLP blocks — feed-forward layers

This is the same pure-PyTorch implementation from openNemo — no mamba-ssm or causal-conv1d dependencies required.

Requirements

torch>=2.1
transformers>=4.40
bitsandbytes>=0.43  # optional, for 4-bit quantization

Disclaimer

This model has had its safety alignment removed. It will comply with requests that the original model would refuse. The creators are not responsible for how this model is used. Intended for research, creative writing, and applications where the user takes responsibility for output filtering.

Acknowledgments

License

NVIDIA Open Model License — same as the base model.

README history 1 version

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-03-23Create README.md70fd8655.4 KB
    Loading...

Discussions 1 thread

  1. 2026-06-23snakehead abliterationopen2 💬#1
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration