For privacy reasons a browser tells us at most
"≥ 8 GB RAM, 8 cores" - same reading whether you have 8 GB
or 128 GB. It has no idea how much RAM is free right now, which
apps are open, or whether you have a GPU.
The Abliteration app is integrated with your machine
It reads your exact RAM, GPU model and VRAM,
free memory right now, and picks the sharpest quant that
still fits. Every model page lights up precisely for your rig.
And you can chat with any model, right now
The app is a full local runtime - no API keys, no subscription,
everything runs on your machine. Click any model on this site and
start a conversation in seconds.
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
'abliterated' in name/tags
is_gguf=0 (base model)
no specific method indicators - defaulting to M1 (most common)
NVFP4 weights of Nemotron-3-Ultra-550B-A55B-abliterated-uncensored — an abliterated, uncensored variant of NVIDIA's Nemotron-3-Ultra-550B-A55B (550B total / 55B active). The model keeps Nemotron-3's hybrid Mamba-2 / Attention / Latent-MoE reasoning stack fully intact — including the MTP speculative-decoding head and the enable_thinking reasoning mode — so this checkpoint is a drop-in replacement for the original at the architecture level and serves out-of-the-box on vLLM.
The pipeline:
Refusal Ablation — A residual-stream refusal direction was extracted by diff-in-means on a labeled harmful/harmless prompt set, read at the end of the model's own reasoning trace (</think>), then baked into the weights as an offline orthogonal projection on the residual-write modules — using our own custom abliteration framework that operates directly on the packed NVFP4 tensors (dequantize → project → requantize), with no full-precision decompress and no training.
Sampling:temperature=1.0, top_p=0.95 (the values in generation_config.json). A mild repetition_penalty (~1.1) is recommended for long generations.
Thinking mode: set enable_thinking=True in chat_template_kwargs; reasoning streams inside <think>…</think> before the answer. Do not feed previous-turn reasoning back into multi-turn history.
Hardware
NVFP4 weights are ~329 GB. Single-node 4× B200 / 4× B300 (or 8× H100) for full context; expert-parallel recommended. Smaller deployments work at reduced context with tensor-parallel-size 2 on high-VRAM Blackwell cards.
Notes
License: OpenMDW-1.1 (inherits from the base model)
Use is the responsibility of the user. Ensure your usage complies with applicable laws, platform rules, the OpenMDW-1.1 license terms, and your deployment requirements.
README history
6 versions
The author's README evolved over time. Click a version to see its content at that point.
2026-06-16Move Support & Community section to top (below banner); unify across models5f814695.3 KB
Loading...
2026-06-08Add OYM banner to top of model cardbb96e635.3 KB
Loading...
2026-06-08Add highlighted Buy Me a Coffee support sectiona2569fe5.2 KB
Loading...
2026-06-05Update README.md667fa7d4.8 KB
Loading...
2026-06-05Update README.md65547115.5 KB
Loading...
2026-06-05Upload folder using huggingface_hub8eae6945.4 KB
Loading...
Catalog is the map. Apps are the tools.
Run models on your own machine, not in the cloud.
Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.