← back to catalog · registered 2026-08-25 19:02

HangGlidersRule/Darkstar-Nemotron-3.5-Lightning-30B-A3B-Abliterated-ModelOpt-W4A16-NVFP4

HangGlidersRule Nemotron 15B MoE second-order
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/HangGlidersRule%2FDarkstar-Nemotron-3.5-Lightning-30B-A3B-Abliterated-ModelOpt-W4A16-NVFP4"
Response includes
  • classification m1
  • files 23
  • hub_downloads_all_time 311
  • author_summary 6 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M1
Primary method

Direct removal

No other method signals detected in this model.
Confidence
MEDIUM
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=0 (base model)
  • no specific method indicators - defaulting to M1 (most common)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
MEDIUM
Why we say so
primary_method=M1; difference-of-means is the reference extraction for M1/M3 (Arditi 2024)
Downloads · lifetime
311
156 last 30d - active
Likes
0
Model age
6w ago
created 2026-08-25
Downloads over time
Now373→from4↑9,225%
01372734104 on Aug 26373 on Oct 11AugSepOct
Aug 26 → Oct 11 · 47 snapshots · spans 46 days

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
other
Languages
en
Tags
safetensors nemotron_h darkstar nemotron-h abliterated reduced-refusal nvfp4 modelopt vllm text-generation conversational base_model:HangGlidersRule/Darkstar-Nemotron-3.5-Lightning-30B-A3B-Abliterated-BF16

Related

Total size
21.4 GB
Files
23
Quantizations
1
Registered
2026-08-25 19:02
Last updated on HF
2026-09-02 15:30

Files by quantization

Auxiliary files 23 files 21.4 GB
model-00001-of-00003.safetensors 9.31 GB bd53901c download
model-00002-of-00003.safetensors 8.91 GB 13997314 download
model-00003-of-00003.safetensors 3.14 GB a116efa5 download
tokenizer.json 16.3 MB 623c3456 download
model.safetensors.index.json 1.74 MB 710da479 download
.quant_summary.txt 1.30 MB c29e86cd download
abliteration_report.json 689 KB e0e936d3 download
config.json 588 KB 08971c43 download
agentic_coding_benchmarks.png 215 KB 06cf5486 download
accuracy_plot.png 137 KB feecce07 download
chat_template.jinja 9.64 KB d85b0c77 download
README.md 4.50 KB f5ce1c14 download
_SUCCESS.json 3.34 KB 8c27c1bc download
bias.md 3.19 KB cdcc8055 download
explainability.md 2.64 KB 2435f23c download
LICENSE 2.63 KB 4d76cc87 download
safety.md 2.09 KB f53bf94a download
hf_quant_config.json 2.08 KB 4168e3e3 download
.gitattributes 1.65 KB dd95036d download
privacy.md 828 B e3bf30aa download
special_tokens_map.json 563 B 0451f379 download
tokenizer_config.json 393 B 9bf527c3 download
generation_config.json 210 B 145693e1 download

README current version from Hugging Face


license: other
license_name: openmdw-1.1
license_link: https://openmdw.ai/license/1-1/
language:

  • en
    tags:
  • nemotron-h
  • darkstar
  • abliteration
  • nvfp4
  • modelopt
  • hybrid-mamba-moe
    pipeline_tag: text-generation
    base_model: HangGlidersRule/Darkstar-Nemotron-3.5-Lightning-30B-A3B-Abliterated-BF16
    base_model_relation: quantized
    quantization: nvidia-modelopt
    extra_gated_heading: Darkstar Nemotron-3.5-Lightning 30B-A3B Abliterated ModelOpt NVFP4

Darkstar-Nemotron-3.5-Lightning-30B-A3B-Abliterated-ModelOpt-W4A16-NVFP4

Safety notice: This checkpoint is an edited (abliterated) derivative
served in NVFP4: the refusal direction measured at layer 34 was
deliberately projected out of 3,126 residual-writing tensors, then the edited
BF16 artifact was quantized with NVIDIA TensorRT Model Optimizer to a mixed
W4A16-NVFP4 layout. It will tend to comply with harmful requests. Released for
red-teaming and alignment research only. Also see Safety.

A single-GPU-friendly (≈22 GB), modelopt-quantized derivative of NVIDIA's
Nemotron-3.5-Lightning-30B-A3B-BF16:
hybrid Mamba2 + MoE + sparse attention, 52 layers, 262,144-token context,
OpenMDW-1.1 license. This is the fourth product in the Darkstar family matrix.

Product family

Product Format Edit Status
Base-BF16 BF16 none (upstream reference) not republished here
Base-ModelOpt-NVFP4 W4A16 NVFP4 none sibling repository
Abliterated-BF16 BF16 refusal direction removed sibling repository
Abliterated-ModelOpt-NVFP4 W4A16 NVFP4 refusal direction removed this repository

Quantization contract (reproducible)

  • Source for the edit: nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16
    @ d468880b6ad3c6e0d21377ce7242adaea4cc884d
  • Abliteration: identical to the Abliterated-BF16 product (layer 34, 3,126
    targets, max residual leakage 0.000160, MTP intact, 320/320 chat-templated)
  • Quantization: NVIDIA Model Optimizer 0.46.0rc2 (43fd41a), recipe
    w4a16_nvfp4_mse-fp8_attn-kv_bf16_nemotron_h.yaml
  • Calibration: cnn_dailymail 512 + Nemotron-Post-Training-Dataset-v2 512,
    sequence length 2048, seed 1234, batch 1, KV cache quantization disabled (BF16)
  • Protected BF16: lm_head, Mamba/SSM (conv1d, in_proj, out_proj, A_log, D,
    dt_bias), norms, embeddings, MTP head
  • Quantized: routed + shared expert up/down projections (5,934 modules),
    W4A16-NVFP4 group 16
  • Artifact: 3 shards, 22 GB, quantization_config → vLLM modelopt_mixed
  • Recipe: recipes/nemotron-3.5-lightning/darkstar-nemotron-3.5-lightning-30b-a3b-abliterated-modelopt-nvfp4.yaml
    in HangGlidersRule/model-forge

Measured quality (final servable result)

  • GPQA Diamond: 141/198 = 71.2% (Wilson 95% 64.5–77.1; 1 unparseable, 2 TRUNC, 0 errors),
    evaluated with llm-inference-bench gpqa-diamond (chat template + thinking ON, temp 0),
    served MTP10 + --reasoning-parser nemotron_v3, single RTX PRO 6000 Blackwell, BF16 KV.
  • Behavior gate: 200/200 harmful compliance, 0/83 safe over-refusals, 0 errors.
  • Throughput (weighted 4K/16K/48K = 0.6/0.3/0.1): MTP10 = 554.7 tok/s (4K 571.2, 16K 546.3, 48K 480.7).
  • NVIDIA publishes GPQA Diamond 75.44 (BF16) / 75.57 (NVFP4) on the same task; the delta is
    serving-stack config (their vLLM 0.26 + FP8 KV + TP2 + temp 1.0 averaged over 8 repeats), not
    abliteration or quantization damage. Full protocol: models/nemotron-3.5-lightning-r1/gpqa-protocol.md.

License

OpenMDW-1.1 (same as upstream). See https://openmdw.ai/license/1-1/. Retain all
NVIDIA copyright/attribution/notice lines in distributions of this derivative.

Publication

This checkpoint is public on Hugging Face at the pinned milestone tag
darkstar-nemotron-3.5-lightning-v1.0.0. Weights are hash-verified (sha256
manifest in the source repo) and serve with vLLM (the measured shipping
config — MTP10):

vllm serve HangGlidersRule/Darkstar-Nemotron-3.5-Lightning-30B-A3B-Abliterated-ModelOpt-W4A16-NVFP4 \
  --max-model-len 131072 --kv-cache-dtype bfloat16 --reasoning-parser nemotron_v3 \
  --speculative-config '{"method":"mtp","num_speculative_tokens":10}'

Safety

This model intentionally has a reduced refusal response. Do not deploy in
user-facing assistant roles without alignment hardening and content filtering.
It is intended for researchers studying refusal behavior, ablation, and
alignment techniques.

README history 3 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-09-02Card format: conform to gold-standard Darkstar card skeleton2aa197f4 KB
    Loading...
  2. 2026-08-25update card: base_model, publication section, milestone tagd3342034.5 KB
    Loading...
  3. 2026-08-25publish model card: Darkstar Nemotron-3.5-Lightning (R1 abliterated)26539933.9 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration