← back to catalog · registered 2026-08-22 13:56

mlasli/Nemotron-3.5-Lightning-30B-A3B-Heretic-Uncensored-BF16

mlasli Nemotron 32B MoE
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/mlasli%2FNemotron-3.5-Lightning-30B-A3B-Heretic-Uncensored-BF16"
Response includes
  • classification m3
  • files 25
  • hub_downloads_all_time 563
  • author_summary 23 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M3
Primary method

Layer-wise ablation

Applied on top of direct removal inherited from the base model.
Confidence
HIGH
Inherited from base model
Why this label 2 signals
Producer identity confirmed by naming conventions, tags or the model card. This label is very unlikely to change.
  • 'heretic' in model name (Heretic-produced)
  • Heretic uses layer-wise optimization (M3) with underlying direction removal (M1)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
HIGH
Why we say so
name contains 'heretic'; Heretic default extraction is difference-of-means (Arditi 2024)
Downloads · lifetime
563
120 last 30d - stable
Likes
2
Model age
7w ago
created 2026-08-16
Downloads over time
Now600→from310↑94%
296407518629310 on Aug 19600 on Oct 11AugSepOct
Aug 19 → Oct 11 · 48 snapshots · spans 53 days

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Variants by this author 3 formats · 2K downloads combined

The same weights this author released in different packaging. Pick the format that matches your runtime.

Metadata

License
other
Languages
en es fr de it ja
Tags
safetensors nemotron_h nemotron nemotron-3.5 mamba moe hybrid abliterated heretic uncensored decensored text-generation

Related

Total size
58.8 GB
Files
25
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-08-16 17:22

Files by quantization

Auxiliary files 25 files 58.8 GB
model-00001-of-00016.safetensors 3.87 GB 51491d6b download
model-00013-of-00016.safetensors 3.83 GB 865c2e10 download
model-00005-of-00016.safetensors 3.79 GB 779ab709 download
model-00007-of-00016.safetensors 3.79 GB 68a81490 download
model-00009-of-00016.safetensors 3.79 GB cd605302 download
model-00011-of-00016.safetensors 3.79 GB da286f79 download
model-00015-of-00016.safetensors 3.79 GB 4003cc42 download
model-00003-of-00016.safetensors 3.79 GB d08218d1 download
model-00004-of-00016.safetensors 3.72 GB b3bc7c03 download
model-00006-of-00016.safetensors 3.72 GB 5503d0c8 download
model-00008-of-00016.safetensors 3.72 GB 32d1431e download
model-00010-of-00016.safetensors 3.72 GB bda92a44 download
model-00002-of-00016.safetensors 3.72 GB 4070264a download
model-00012-of-00016.safetensors 3.68 GB 07ad7fe3 download
model-00014-of-00016.safetensors 3.68 GB d353f701 download
model-00016-of-00016.safetensors 2.42 GB d7fbb1a8 download
tokenizer.json 16.3 MB 5215d963 download
model.safetensors.index.json 575 KB 715bde09 download
chat_template.jinja 9.64 KB d85b0c77 download
README.md 4.05 KB 8cee8560 download
config.json 2.65 KB bcc64990 download
.gitattributes 1.53 KB 52373fe2 download
special_tokens_map.json 563 B 0451f379 download
tokenizer_config.json 394 B 18b66d85 download
generation_config.json 210 B 95e209cc download

README current version from Hugging Face


language: [en, es, fr, de, it, ja]
license: other
license_name: nvidia-open-model-license
license_link: https://www.nvidia.com/en-us/agreements/enterprise-software/nvidia-open-model-license/
pipeline_tag: text-generation
tags:

  • nemotron
  • nemotron-3.5
  • nemotron_h
  • mamba
  • moe
  • hybrid
  • abliterated
  • heretic
  • uncensored
  • decensored
  • text-generation
  • roleplay
  • nvidia
    base_model: nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16

Nemotron-3.5-Lightning-30B-A3B Heretic-Abliterated (BF16)

This is NVIDIA-Nemotron-3.5-Lightning-30B-A3B
(31.6B total / 3B active params, hybrid Mamba-2 + MoE + attention) with its refusal direction
removed using Heretic — a single-direction abliteration
with an Optuna-based parameter search. The language backbone is abliterated; all other
capabilities are preserved.

What this is for: NVIDIA's hybrid Mamba-2 + MoE + attention backbone (3B active of 31.6B) with
the refusal direction removed — 0% refusals at ~0.04 KL divergence. Ideal for uncensored roleplay
and long-context agent work where the base model would refuse.

What is Heretic?

Heretic removes a model's safety-aligned refusal
direction in one shot (unlike earlier multi-direction approaches), trading a minimal amount of
capability for a large drop in refusals. Its Optuna search picks the ablation parameters on the
Pareto front of (compliance, first-token KL divergence).

Results

Refusals Compliance KL Divergence Trials
0% 100% 0.0397 200

Independent eval of the merged model (50 harmful-behavior prompts).

Best Trial (Trial 141 of 200)

  • Refusal rate (in-run): 7/50 (86% compliance)
  • KL divergence (in-run): 0.0392

The automated Zou keyword detector read 80% compliance and a stricter combined detector read
66%, but manual review of all 50 completions confirms these are false positives: the model
answers directly and uses words like "illegal"/"unethical"/"harmful" inside compliant responses.
The manual-review numbers above are the reliable refusal estimate.

Usage

from transformers import AutoModelForCausalLM, AutoTokenizer

model = AutoModelForCausalLM.from_pretrained(
    "mlasli/Nemotron-3.5-Lightning-30B-A3B-Heretic-Uncensored-BF16",
    torch_dtype="auto",
    device_map="auto",
)
tokenizer = AutoTokenizer.from_pretrained(
    "mlasli/Nemotron-3.5-Lightning-30B-A3B-Heretic-Uncensored-BF16"
)

messages = [{"role": "user", "content": "Hi! What is 2+2?"}]
text = tokenizer.apply_chat_template(
    messages,
    add_generation_prompt=True,
    enable_thinking=False,  # disables the verbose <think> chain-of-thought
    tokenize=False,
)

Quantizations

GGUF quantizations are available in separate repositories. Load them with
llama.cpp (architecture nemotron_h_moe, build b10326+).
All quants were locally smoke-tested before upload.

Quantization Size Repository
Q8_0 33.6 GB Nemotron-3.5-Lightning-30B-A3B-Heretic-Uncensored-Q8_0-GGUF
Q6_K 33.5 GB Nemotron-3.5-Lightning-30B-A3B-Heretic-Uncensored-Q6_K-GGUF
Q4_K_M 24.3 GB Nemotron-3.5-Lightning-30B-A3B-Heretic-Uncensored-Q4_K_M-GGUF

Notes

  • This release does not include the MTP (NextN) speculative-decoding draft head; the merged
    weights omit it, so inference is single-head (no --spec-type mtp). Main-model quality is unaffected.
  • Abliteration removes safety alignment. Use responsibly and in accordance with your local laws and
    the upstream NVIDIA Open Model License.

README history 4 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-08-16Add why-this-over-base hook lineb0e59394.1 KB
    Loading...
  2. 2026-08-16Retitle to Heretic-Abliterated and reorder tags443e6573.8 KB
    Loading...
  3. 2026-08-16Add structured eval results table to model card42d16323.8 KB
    Loading...
  4. 2026-08-16Upload README.md with huggingface_hub437f4053.7 KB
    Loading...

Discussions 1 thread

  1. 2026-09-17Thank you! =Dclosed2 💬#1
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration