language: [en, es, fr, de, it, ja]
license: other
license_name: nvidia-open-model-license
license_link: https://www.nvidia.com/en-us/agreements/enterprise-software/nvidia-open-model-license/
pipeline_tag: text-generation
tags:
- nemotron
- nemotron-3.5
- nemotron_h
- mamba
- moe
- hybrid
- abliterated
- heretic
- uncensored
- nvfp4
- modelopt
- vllm
base_model: nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16
Nemotron-3.5-Lightning-30B-A3B — Abliterated, NVFP4
This is NVIDIA-Nemotron-3.5-Lightning-30B-A3B
(31.6B total / 3B active params, hybrid Mamba-2 + MoE + attention) with its refusal direction
removed using Heretic — a single-direction abliteration
with an Optuna-based parameter search — then quantized to NVFP4 (weight+activation, ModelOpt).
Both the abliteration and the NVFP4 quantization were done by us. Refusal/compliance/KL-divergence
numbers for this specific run were not measured (not published here).
What is Heretic?
Heretic removes a model's safety-aligned refusal
direction in one shot, trading a minimal amount of capability for a large drop in refusals. Its
Optuna search picks the ablation parameters on the Pareto front of (compliance, first-token KL
divergence).
Usage
from transformers import AutoModelForCausalLM, AutoTokenizer
model = AutoModelForCausalLM.from_pretrained(
"qpqpqpqpqpqqqq/Nemotron-3.5-Lightning-30B-A3B-Abliterated-NVFP4",
torch_dtype="auto",
device_map="auto",
)
tokenizer = AutoTokenizer.from_pretrained(
"qpqpqpqpqpqqqq/Nemotron-3.5-Lightning-30B-A3B-Abliterated-NVFP4"
)
messages = [{"role": "user", "content": "Hi! What is 2+2?"}]
text = tokenizer.apply_chat_template(
messages,
add_generation_prompt=True,
enable_thinking=False, # disables the verbose <think> chain-of-thought
tokenize=False,
)
Serving via vLLM (0.30+): --moe-backend marlin for the NVFP4 MoE path. See the DSpark draft
below for speculative decoding.
Speculative decoding draft
Pairs with nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-NVFP4-DSpark
(unmodified, NVIDIA's own release) — not bundled in this repo.
Notes
- Abliteration removes safety alignment. Use responsibly and in accordance with your local laws
and the upstream NVIDIA Open Model License.