← back to catalog · registered 2026-09-28 18:57

qpqpqpqpqpqqqq/Nemotron-3.5-Lightning-30B-A3B-Abliterated-NVFP4

qpqpqpqpqpqqqq Nemotron 30B MoE
curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/qpqpqpqpqpqqqq%2FNemotron-3.5-Lightning-30B-A3B-Abliterated-NVFP4"
Response includes
  • classification m-uncensored
  • files 13
  • author_summary 2 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M-U
Primary method

Uncensored (method unknown)

No other method signals detected in this model.
Confidence
LOW
Why this label 3 signals
Weak or ambiguous signals. Best guess based on catalog patterns; treat as tentative and check the evidence below.
  • 'uncensored' in name/tags but no 'abliterated' marker
  • method not identifiable from author declaration alone
  • may be DPO fine-tune, prompt engineering, or unknown technique
Refusal direction extraction

No specific extraction method could be identified for this model. The producer either did not document it or used a proprietary pipeline.

What is a refusal direction? →
Downloads · 30-day
0
Likes
0
Model age
today
created 2026-09-28

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
other
Languages
en es fr de it ja
Tags
safetensors nemotron_h nemotron nemotron-3.5 mamba moe hybrid abliterated heretic uncensored nvfp4 modelopt

Related

Total size
17.6 GB
Files
13
Quantizations
1
Registered
2026-09-28 18:57
Last updated on HF
2026-09-28 19:17

Files by quantization

Auxiliary files 13 files 17.6 GB
model-00001-of-00002.safetensors 9.31 GB 6a59c478 download
model-00002-of-00002.safetensors 8.26 GB 18c7af91 download
tokenizer.json 16.3 MB 5215d963 download
model.safetensors.index.json 1.72 MB 8a0ff3d3 download
.quant_summary.txt 1.56 MB 72b8849a download
config.json 33.4 KB 61ae6b46 download
hf_quant_config.json 18.9 KB c296644b download
chat_template.jinja 9.64 KB d85b0c77 download
README.md 2.55 KB 8af2c0a6 download
.gitattributes 1.53 KB 52373fe2 download
special_tokens_map.json 563 B 0451f379 download
tokenizer_config.json 497 B 866139e2 download
generation_config.json 210 B 145693e1 download

README current version from Hugging Face


language: [en, es, fr, de, it, ja]
license: other
license_name: nvidia-open-model-license
license_link: https://www.nvidia.com/en-us/agreements/enterprise-software/nvidia-open-model-license/
pipeline_tag: text-generation
tags:

  • nemotron
  • nemotron-3.5
  • nemotron_h
  • mamba
  • moe
  • hybrid
  • abliterated
  • heretic
  • uncensored
  • nvfp4
  • modelopt
  • vllm
    base_model: nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16

Nemotron-3.5-Lightning-30B-A3B — Abliterated, NVFP4

This is NVIDIA-Nemotron-3.5-Lightning-30B-A3B
(31.6B total / 3B active params, hybrid Mamba-2 + MoE + attention) with its refusal direction
removed using Heretic — a single-direction abliteration
with an Optuna-based parameter search — then quantized to NVFP4 (weight+activation, ModelOpt).

Both the abliteration and the NVFP4 quantization were done by us. Refusal/compliance/KL-divergence
numbers for this specific run were not measured (not published here).

What is Heretic?

Heretic removes a model's safety-aligned refusal
direction in one shot, trading a minimal amount of capability for a large drop in refusals. Its
Optuna search picks the ablation parameters on the Pareto front of (compliance, first-token KL
divergence).

Usage

from transformers import AutoModelForCausalLM, AutoTokenizer

model = AutoModelForCausalLM.from_pretrained(
    "qpqpqpqpqpqqqq/Nemotron-3.5-Lightning-30B-A3B-Abliterated-NVFP4",
    torch_dtype="auto",
    device_map="auto",
)
tokenizer = AutoTokenizer.from_pretrained(
    "qpqpqpqpqpqqqq/Nemotron-3.5-Lightning-30B-A3B-Abliterated-NVFP4"
)

messages = [{"role": "user", "content": "Hi! What is 2+2?"}]
text = tokenizer.apply_chat_template(
    messages,
    add_generation_prompt=True,
    enable_thinking=False,  # disables the verbose <think> chain-of-thought
    tokenize=False,
)

Serving via vLLM (0.30+): --moe-backend marlin for the NVFP4 MoE path. See the DSpark draft
below for speculative decoding.

Speculative decoding draft

Pairs with nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-NVFP4-DSpark
(unmodified, NVIDIA's own release) — not bundled in this repo.

Notes

  • Abliteration removes safety alignment. Use responsibly and in accordance with your local laws
    and the upstream NVIDIA Open Model License.
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.