← back to catalog · registered 2026-08-22 13:56

adriaflores/Llama-3-8B-abliterated

adriaflores Llama 8.0B
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/adriaflores%2FLlama-3-8B-abliterated"
Response includes
  • classification m1
  • files 13
  • benchmarks 5 entries
  • hub_downloads_all_time 79
  • author_summary 1 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M1
Primary method

Direct removal

No other method signals detected in this model.
Confidence
MEDIUM
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=0 (base model)
  • no specific method indicators - defaulting to M1 (most common)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
MEDIUM
Why we say so
primary_method=M1; difference-of-means is the reference extraction for M1/M3 (Arditi 2024)
Downloads · lifetime
79
17 last 30d - stable
Likes
2
Descendants
2
in 2 direct forks
Model age
5mo ago
created 2026-05-02
Downloads over time
Now85→from33↑158%
3050709033 on May 685 on Oct 1185 on Oct 8MayJunJulAugSepOct
May 6 → Oct 11 · 62 snapshots · spans 158 days

Benchmarks

Portrait before abliteration
Benchmarks of the base model as it stood before the refusal-removal operation. Compare with the numbers above to see what the operation cost.
Benchmark Score Source
BBH average 0.448047736411986 OpenLLM-v2
IFEval instruct 0.5479616306954437 OpenLLM-v2
IFEval-Prompt 0.40850277264325324 OpenLLM-v2
MATH lvl 5 0.08383685800604229 OpenLLM-v2
MMLU-Pro 0.359125664893617 OpenLLM-v2

Genealogy 2 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
llama3
Languages
en
Tags
safetensors llama abliteration uncensored mechanistic-interpretability llama-3 en base_model:meta-llama/Meta-Llama-3-8B-Instruct base_model:finetune:meta-llama/Meta-Llama-3-8B-Instruct license:llama3 region:us

Related

Total size
15.0 GB
Files
13
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-05-03 01:54

Files by quantization

Auxiliary files 13 files 15.0 GB
model-00002-of-00004.safetensors 4.66 GB 4b4ea3e5 download
model-00001-of-00004.safetensors 4.63 GB a14229ba download
model-00003-of-00004.safetensors 4.58 GB 40312876 download
model-00004-of-00004.safetensors 1.09 GB 14dc1217 download
tokenizer.json 16.4 MB 8fc5ed64 download
tokenizer_config.json 51.4 KB d0bc1281 download
model.safetensors.index.json 23.7 KB e0c3f181 download
README.md 1.83 KB 13e2b766 download
.gitattributes 1.53 KB 52373fe2 download
config.json 712 B 7be8ab8c download
chat_template.jinja 393 B a4bddd7a download
special_tokens_map.json 342 B 4eae993f download
generation_config.json 206 B 51fc5d3d download

README current version from Hugging Face


base_model: meta-llama/Meta-Llama-3-8B-Instruct
language:

  • en
    license: llama3
    tags:
  • abliteration
  • uncensored
  • mechanistic-interpretability
  • llama-3

Llama-3-8B-abliterated

An uncensored version of meta-llama/Meta-Llama-3-8B-Instruct produced via abliteration: a training-free technique that permanently removes refusal behavior by orthogonalizing the model's weight matrices against a learned "refusal direction" in residual stream activation space.

Method

Abliteration identifies the refusal direction using contrastive mean-difference on activations from harmful vs. harmless instruction pairs (layer 9, resid_pre), then removes it from the embedding matrix and all attention and MLP output projections. No fine-tuning is involved.

Evaluation

Metric Original Abliterated
Censorship rate (harmful_behaviors, n=100) 97.0% 18.0%
ARC-Challenge accuracy (n=1172) 80.2% 79.5%

Censorship judged by Qwen3.5-9B. The 18% residual reflects prompts where refusal is encoded across multiple directions beyond the one removed.

Usage

from transformers import AutoModelForCausalLM, AutoTokenizer

model = AutoModelForCausalLM.from_pretrained(
    "adriaflores/Llama-3-8B-abliterated",
    torch_dtype="bfloat16",
    device_map="auto",
)
tokenizer = AutoTokenizer.from_pretrained("adriaflores/Llama-3-8B-abliterated")

Requires a GPU with at least 16 GB VRAM.

Limitations

  • Residual refusal on a subset of harmful prompts (~18%) suggests the technique does not fully suppress all refusal-related circuitry.
  • Capability impact is negligible (0.7% ARC-Challenge delta), but has not been evaluated on other benchmarks.
  • Intended for research use. The user is responsible for evaluating suitability for any downstream application.

README history 1 version

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-05-03Upload folder using huggingface_hubc43b9611.8 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration