← back to catalog · registered 2026-08-22 13:56

timteh673/Nemotron-Super-49B-v1.5-Uncensored-GGUF

timteh673 Llama 49B GGUF 131K ctx
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/timteh673%2FNemotron-Super-49B-v1.5-Uncensored-GGUF"
Response includes
  • classification m8
  • files 9
  • hub_downloads_all_time 12,394
  • author_summary 8 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M8
Primary method

Repackaging (quantization)

Applied on top of direct removal inherited from the base model.
Confidence
MEDIUM
Inherited from base model
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=1
  • assume M1 (base ablation) + M8 (GGUF quant) - default when producer unknown
Refusal direction extraction

No specific extraction method could be identified for this model. The producer either did not document it or used a proprietary pipeline.

What is a refusal direction? →
Downloads · lifetime
12K
1K last 30d - cooling
Likes
3
Model age
6mo ago
created 2026-03-31
Downloads over time
Now12.6K→from296↑4,143%
04.6K9.2K13.8K296 on Apr 112.6K on Oct 11AprMayJunJulAugSepOct
Apr 1 → Oct 11 · 67 snapshots · spans 193 days

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
llama3.3
Languages
en
Quantizations
BF16 Q2_K Q3_K Q4_K Q5_K Q6_K Q8_0
Tags
gguf uncensored abliterated nvidia nemotron llama-3.3 49b timteh text-generation en base_model:nvidia/Llama-3_3-Nemotron-Super-49B-v1_5 base_model:quantized:nvidia/Llama-3_3-Nemotron-Super-49B-v1_5

Related

Total size
282 GB
Files
9
Quantizations
8
Registered
2026-08-22 13:56
Last updated on HF
2026-03-31 18:30

Files by quantization

BF16 1 file 92.9 GB
Nemotron-Super-49B-Uncensored-BF16.gguf 92.9 GB 84426bb6 download
Q8_0 1 file 49.4 GB
Nemotron-Super-49B-Uncensored-Q8_0.gguf 49.4 GB 3c949253 download
Q6_K 1 file 38.1 GB
Nemotron-Super-49B-Uncensored-Q6_K.gguf 38.1 GB 2a88d4dc download
Q5_K 1 file 33.0 GB
Nemotron-Super-49B-Uncensored-Q5_K_M.gguf 33.0 GB 6a17b134 download
Q4_K 1 file 28.1 GB
Nemotron-Super-49B-Uncensored-Q4_K_M.gguf 28.1 GB 3a408134 download
Q3_K 1 file 22.6 GB
Nemotron-Super-49B-Uncensored-Q3_K_M.gguf 22.6 GB cb3d38a1 download
Q2_K 1 file 17.5 GB
Nemotron-Super-49B-Uncensored-Q2_K.gguf 17.5 GB 6bb67821 download
Auxiliary files 2 files 6.48 KB
README.md 4.47 KB 395127b1 download
.gitattributes 2.01 KB 06ef78b8 download

README current version from Hugging Face


license: llama3.3
base_model: nvidia/Llama-3_3-Nemotron-Super-49B-v1_5
tags:

  • uncensored
  • abliterated
  • gguf
  • nvidia
  • nemotron
  • llama-3.3
  • 49b
  • timteh
    quantized_by: timteh673
    language:
  • en
    pipeline_tag: text-generation

Nemotron-Super-49B-v1.5 Uncensored GGUF

Zero-degradation uncensoring of NVIDIA's Llama-3.3-Nemotron-Super-49B-v1.5 — guardrails surgically removed via representation engineering while preserving full model capability.

⚡ Forged on 8×H200 SXM5 | 1.1TB VRAM

Model Details

Property Value
Base Model nvidia/Llama-3_3-Nemotron-Super-49B-v1_5
Architecture DeciLM (NAS-optimized Llama-3.3) — variable attention and FFN per layer
Parameters 49B
Context 128K tokens
License Llama 3.3 Community License
Base Downloads 174K+
Uncensoring Method Representation engineering — refusal direction projection removal

What is this?

NVIDIA's Nemotron-Super-49B-v1.5 is one of the strongest sub-50B models available — a NAS-optimized architecture that punches well above its weight class. This release removes alignment guardrails using representation engineering (abliteration), allowing the model to respond to all prompts without refusal.

Abliteration Method

  • 32 harmful + 32 harmless prompt pairs used to identify refusal directions across all 80 layers
  • Refusal direction projected out of residual stream weights only (ffn_down, attn_output) — 127 weight tensors modified
  • Alpha = 1.0 (full removal)
  • NaN/zero directions automatically skipped (1 layer)
  • No fine-tuning, no dataset bias — pure mathematical guardrail removal

Why Nemotron-Super-49B?

  • 174K downloads on the base model — proven demand
  • Zero uncensored/abliterated versions existed before this release
  • 49B sweet spot — runs on consumer hardware (24GB+ VRAM for Q4), outperforms many 70B models
  • NAS-optimized architecture — variable layer widths for maximum efficiency

Available Quantizations

Quantization Size BPW Use Case
BF16 93 GB 16.00 Full precision, research
Q8_0 50 GB 8.50 Near-lossless, 2×A100/H100
Q6_K 39 GB 6.57 High quality, 48GB GPU
Q5_K_M 33 GB 5.63 Great balance, 48GB GPU
Q4_K_M 29 GB 4.85 Recommended — best quality/size, 32GB GPU
Q3_K_M 23 GB 3.86 Good quality, 24GB GPU
Q2_K 18 GB 2.96 Minimum viable, 24GB GPU

Quick Start

# Download recommended quantization
huggingface-cli download timteh673/Nemotron-Super-49B-v1.5-Uncensored-GGUF \
  Nemotron-Super-49B-Uncensored-Q4_K_M.gguf \
  --local-dir ./models

# Run with llama.cpp
./llama-server -m models/Nemotron-Super-49B-Uncensored-Q4_K_M.gguf \
  -c 8192 -ngl 99

Ollama

# Create Modelfile
echo 'FROM ./Nemotron-Super-49B-Uncensored-Q4_K_M.gguf' > Modelfile
ollama create nemotron-super-49b-uncensored -f Modelfile
ollama run nemotron-super-49b-uncensored

Hardware Requirements

Quantization Minimum VRAM Recommended Setup
Q2_K / Q3_K_M 24 GB RTX 3090/4090
Q4_K_M / Q5_K_M 32-48 GB RTX A6000, 2×3090
Q6_K 48 GB A6000, A100 40GB + offload
Q8_0 64 GB A100 80GB, 2×A6000
BF16 96+ GB 2×A100 80GB, H100

Ethical Notice

This model is provided for research and development purposes. The removal of safety guardrails means the model will respond to prompts that the original model would refuse. Users are responsible for ensuring their use complies with applicable laws and regulations. This model should not be used to generate content that could cause harm.

Support This Work

If you find this useful, consider supporting continued open model releases:

☕ Buy Me a Coffee: https://buymeacoffee.com/timteh

Crypto:

  • BTC: bc1qmz3vu2naymwfmz7f7krfteevfy0yk9ts09wp5y
  • ETH: 0x27fd2C8d3b5a1C6a0e85c5A9FCa2a8743dD04E7a
  • SOL: 7x5Eo3FhKMZxFNoE3DfQfBRYnmBVbmj3bSduHaVJpump

📧 Enterprise/Custom Merges: [email protected]


Built by timteh673 — Cognitive Preservation Foundry

README history 1 version

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-03-31Upload README.md with huggingface_hubbc2e7b64.5 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration