← back to catalog · registered 2026-08-22 13:56

SicariusSicariiStuff/Llama-3.1-Nemotron-8B-UltraLong-1M-Instruct_Abliterated

Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/SicariusSicariiStuff%2FLlama-3.1-Nemotron-8B-UltraLong-1M-Instruct_Abliterated"
Response includes
  • classification m1
  • files 15
  • benchmarks 11 entries
  • hub_downloads_all_time 519
  • author_summary 75 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M1
Primary method

Direct removal

No other method signals detected in this model.
Confidence
MEDIUM
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=0 (base model)
  • no specific method indicators - defaulting to M1 (most common)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
MEDIUM
Why we say so
primary_method=M1; difference-of-means is the reference extraction for M1/M3 (Arditi 2024)
Downloads · lifetime
519
57 last 30d - stable
Likes
10
Descendants
2
in 2 direct forks
Model age
8mo ago
created 2026-01-17
Downloads over time
Now536→from38↑1,311%
1320439558638 on Jan 21536 on Oct 11536 on Oct 10JanMarMayJulSep
Jan 21 → Oct 11 · 77 snapshots · spans 263 days

Benchmarks

Benchmark Score Source
Entertainment 1.4 UGI
Hazardous 1.2 UGI
Natural Intelligence 15.11 UGI
Political lean -32.8% UGI
Sensitive-Info 13.21 UGI
SocPol 1.3 UGI
UGI 32.98 UGI
Willingness (10) 7.2 UGI
W10-Adherence 4.5 UGI
W10-Direct 10 UGI
Writing 17.71 UGI

Genealogy 2 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
llama3.1
Languages
en
Tags
safetensors llama en base_model:nvidia/Llama-3.1-Nemotron-8B-UltraLong-1M-Instruct base_model:finetune:nvidia/Llama-3.1-Nemotron-8B-UltraLong-1M-Instruct license:llama3.1 region:us

Related

Total size
29.9 GB
Files
15
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-01-26 16:08

Files by quantization

Auxiliary files 15 files 30.0 GB
model-00004-of-00007.safetensors 4.66 GB 4e90b257 download
model-00006-of-00007.safetensors 4.66 GB ddce114e download
model-00003-of-00007.safetensors 4.66 GB 2fe63321 download
model-00001-of-00007.safetensors 4.56 GB 19715f8e download
model-00005-of-00007.safetensors 4.50 GB b2fa930a download
model-00002-of-00007.safetensors 4.50 GB f2471161 download
model-00007-of-00007.safetensors 2.41 GB c2ccf616 download
tokenizer.json 16.4 MB 65ff5472 download
tokenizer_config.json 49.4 KB 555d0b07 download
model.safetensors.index.json 23.4 KB 03438a11 download
chat_template.jinja 4.51 KB 33089ace download
README.md 3.72 KB 6e1322f8 download
.gitattributes 1.64 KB 52b37997 download
config.json 933 B 400143d5 download
special_tokens_map.json 325 B b43be966 download

README current version from Hugging Face


license: llama3.1
language:

base_model:

  • nvidia/Llama-3.1-Nemotron-8B-UltraLong-1M-Instruct

Llama-3.1-Nemotron-8B-UltraLong-1M-Instruct
Abliterated

Developed by: SicariusSicariiStuff

Small update post UGI results

Interestingly, this abliterated version slightly outperforms the base model by nVidia in raw intelligence.

Looks like minimizing KL divergence (1st priority) while minimizing refusals (2nd priority) can sometimes produce a model that outperforms the base model across most benchmarks.

This was speculated for quite some time ("the rlhf alignment tax"), but it is still interesting to see, although the difference is small, so that it might be a fluke. "More testing is needed", is a bit of a cliché, but true nonetheless.


Llama-3.1-Nemotron-8B-UltraLong-1M-Instruct_Abliterated is an abliterated variant of Meta's Llama 3.1 8B Instruct model with surgical removal of refusal mechanisms. This model maintains the full ONE MILLION context window while eliminating safety guardrails through orthogonalization techniques.

KL divergence

<0.005

Refusals

~8%

What is KL divergence?

Think about it as a way to measure the variance between the original model "World Model," vs the abliterated one; the lower the KL divergence, the closer the "World Model" of the two models to each other.

If the original model thinks making pineapple pizza is a crime against humanity (it is), then the abliterated model will still hold to this belief, but if asked how to make one (probably after giving you a disclaimer about what an abomination that is), it would still tell you how. In other words, most of the knowledge, quirks, and capabilities are preserved.


Technical Specs

  • Base Model: Llama-3.1-Nemotron-8B-UltraLong-1M-Instruct
  • Parameters: 8B
  • Context Length: 1M tokens
  • Architecture: Llama (decoder-only transformer)
  • Precision: fp32
  • Method: Orthogonalization-based abliteration
  • License: Llama 3.1 Community License

Methodology

  1. Identifies refusal direction vectors in activation space
  2. Orthogonalizes weights to inhibit activation along these directions
  3. Preserves (mostly) all other model behaviors and knowledge

Model Details

  • Intended use: General Tasks.

  • Censorship level: Low - Very Low

  • 7.2 / 10 (10 completely uncensored)

UGI score:


Citation Information

@llm{Llama-3.1-Nemotron-8B-UltraLong-1M-Instruct_Abliterated,
  author = {SicariusSicariiStuff},
  title = {Llama-3.1-Nemotron-8B-UltraLong-1M-Instruct_Abliterated},
  year = {2026},
  publisher = {Hugging Face},
  url = {https://huggingface.co/SicariusSicariiStuff/Llama-3.1-Nemotron-8B-UltraLong-1M-Instruct_Abliterated}
}

Other stuff

README history 10 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-01-26Update README.mdeff6a5f3.7 KB
    Loading...
  2. 2026-01-19Update README.mdff61dfb3.1 KB
    Loading...
  3. 2026-01-19Update README.mdabe42ea3.1 KB
    Loading...
  4. 2026-01-19Update README.mdfe127fa3.1 KB
    Loading...
  5. 2026-01-19Update README.mdd95b2073.1 KB
    Loading...
  6. 2026-01-19Update README.md5162af42.5 KB
    Loading...
  7. 2026-01-17Upload README.mda3585322.3 KB
    Loading...
  8. 2026-01-17Update README.mda07e3292.3 KB
    Loading...
  9. 2026-01-17Update README.md0fdaab62.3 KB
    Loading...
  10. 2026-01-17Upload 14 fileseb4f3dd2.3 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration