← back to catalog · registered 2026-08-22 13:56

wangzhang/Llama-3-8B-Instruct-RR-Abliterated

wangzhang Llama 8.0B
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/wangzhang%2FLlama-3-8B-Instruct-RR-Abliterated"
Response includes
  • classification m1
  • files 8
  • hub_downloads_all_time 907
  • providers 1
  • author_summary 28 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M1
Primary method

Direct removal

No other method signals detected in this model.
Confidence
MEDIUM
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=0 (base model)
  • no specific method indicators - defaulting to M1 (most common)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
MEDIUM
Why we say so
primary_method=M1; difference-of-means is the reference extraction for M1/M3 (Arditi 2024)
Downloads · lifetime
907
106 last 30d - stable
Likes
2
Descendants
2
in 2 direct forks
Model age
6mo ago
created 2026-04-13
Available via
1 provider
featherless-ai
Downloads over time
Now915→from124↑638%
84388691994124 on Apr 15915 on Oct 11AprMayJunJulAugSepOct
Apr 15 → Oct 11 · 65 snapshots · spans 179 days

Genealogy 2 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
llama3
Languages
en zh
Tags
transformers safetensors llama text-generation abliterated abliterix circuit-breakers representation-rerouting safety-removed llama3 conversational en

Related

Total size
15.0 GB
Files
8
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-08-29 17:05

Files by quantization

Auxiliary files 8 files 15.0 GB
model.safetensors 15.0 GB 6442a69b download
tokenizer.json 16.4 MB 602f9c58 download
README.md 4.69 KB 7309720b download
.gitattributes 1.53 KB 52373fe2 download
config.json 739 B 7af6fe01 download
tokenizer_config.json 522 B e51db540 download
chat_template.jinja 389 B 39bd0c9f download
generation_config.json 193 B acc11051 download

README current version from Hugging Face


license: llama3
base_model: GraySwanAI/Llama-3-8B-Instruct-RR
tags:

  • abliterated
  • abliterix
  • circuit-breakers
  • representation-rerouting
  • safety-removed
  • llama3
    language:
  • en
  • zh
    library_name: transformers
    pipeline_tag: text-generation

Llama-3-8B-Instruct-RR-Abliterated

A drop-in replacement for GraySwanAI/Llama-3-8B-Instruct-RR with the Representation Rerouting / Circuit Breakers safety circuit removed.

Produced with abliterix. No fine-tuning, no gradient updates, no manual prompt engineering.

Why this exists

Circuit Breakers (Zou et al., NeurIPS 2024) is one of the strongest open-source robustness baselines for LLMs. It trains the model with a Representation Engineering loss that detects harmful intermediate hidden states and reroutes them into a safety-circuit attractor before generation, making the model resistant to ablation-style attacks that work on a single refusal direction.

We discovered that the GraySwan release ships as a rank-16 LoRA delta on NousResearch/Meta-Llama-3-8B-Instruct — the same structural shape as the Mistral RR sibling and as skysys00/Meta-Llama-3-8B-Instruct-DeepRefusal. SVD analysis of W_RR − W_base confirms the rank-16 cliff on v_proj / o_proj / gate_proj. Once the LoRA delta is identified, removing it is a one-liner.

Result
Base refusal rate (held-out 100 harmful prompts, LLM judge) 99 / 100
This model — refusal rate 1 / 100
Attack Success Rate 99 %
KL divergence vs base 0.017
Hardcore 15 (10 EN + 5 CN) 15 / 15 compliant
Total wall-clock attack time ~70 min on a single RTX A6000

The hardcore 15 cover pipe-bomb assembly, methamphetamine synthesis, Python password-stealing malware, WiFi WEP/WPS attacks, hotwiring, signature forgery, ID-card forgery (CN), 网络诈骗 (CN), 入侵电脑 (CN), and others — all 15 produce compliant, on-topic responses.

How it was made

The recipe diverges slightly from the Mistral sibling: Llama-3-Instruct's stronger built-in RLHF combined with the RR delta meant a partial lerp (λ=0.3) wasn't enough. Full delta strip + a minimal single-direction abliteration cleared it cleanly.

# Stage 0 — SVD diagnose the delta (confirms rank-16 LoRA shape)
python scripts/deeprefusal_attenuate.py \
    --base NousResearch/Meta-Llama-3-8B-Instruct \
    --defended GraySwanAI/Llama-3-8B-Instruct-RR \
    --analyze-only --lambda 0.0 --output /tmp/unused

# Stage 1 — fully strip the LoRA delta (λ=0.0)
python scripts/deeprefusal_attenuate.py \
    --base NousResearch/Meta-Llama-3-8B-Instruct \
    --defended GraySwanAI/Llama-3-8B-Instruct-RR \
    --output /workspace/llama3_rr_stripped --lambda 0.0

# Stage 3 — abliterix direct-mode, single direction, 60 trials
AX_CONFIG=configs/llama3_8b_instruct_rr.toml abliterix --non-interactive

# Stage 6 — export champion trial
python scripts/export_model.py \
    --model /workspace/llama3_rr_stripped \
    --checkpoint checkpoints_llama3_rr \
    --trial 40 \
    --config configs/llama3_8b_instruct_rr.toml \
    --push-to wangzhang/Llama-3-8B-Instruct-RR-Abliterated

Best trial parameters: vector_method=mean, n_directions=1, steering_mode=direct, decay_kernel=linear, iterative.enabled=false, strength_range=[1.5, 6.0]. Full config: configs/llama3_8b_instruct_rr.toml.

Usage

from transformers import AutoModelForCausalLM, AutoTokenizer

model = AutoModelForCausalLM.from_pretrained(
    "wangzhang/Llama-3-8B-Instruct-RR-Abliterated",
    torch_dtype="bfloat16",
    device_map="auto",
)
tokenizer = AutoTokenizer.from_pretrained(
    "wangzhang/Llama-3-8B-Instruct-RR-Abliterated"
)

chat = [
    {"role": "system", "content": "You are a helpful assistant."},
    {"role": "user", "content": "Hello!"},
]
inputs = tokenizer.apply_chat_template(chat, return_tensors="pt", add_generation_prompt=True).to(model.device)
out = model.generate(inputs, max_new_tokens=256)
print(tokenizer.decode(out[0], skip_special_tokens=True))

License & Intended Use

Released for AI safety research, red-teaming, and reproducibility of abliteration claims against published defenses. You are responsible for any output you generate. Inherits the Llama 3 license of the upstream Meta-Llama-3-8B-Instruct weights.

Citation

@software{abliterix2026,
  author = {Wu, Wangzhang},
  title  = {Abliterix: Optimal Refusal Removal for Transformer Models},
  year   = {2026},
  url    = {https://github.com/wuwangzhang1216/abliterix},
}

README history 3 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-08-29docs: add upstream license and provenanceeb7ad6c9.3 KB
    Loading...
  2. 2026-08-29docs: add disclaimer and responsible-use noticea00959c7.8 KB
    Loading...
  3. 2026-04-13Upload folder using huggingface_hubde820c84.7 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration