← back to catalog · registered 2026-08-22 13:56

ZennyKenny/Daredevil-8B-abliterated

ZennyKenny Llama 8.0B
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/ZennyKenny%2FDaredevil-8B-abliterated"
Response includes
  • classification m1
  • files 9
  • benchmarks 5 entries
  • hub_downloads_all_time 178
  • author_summary 1 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M1
Primary method

Direct removal

No other method signals detected in this model.
Confidence
MEDIUM
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=0 (base model)
  • no specific method indicators - defaulting to M1 (most common)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
MEDIUM
Why we say so
primary_method=M1; difference-of-means is the reference extraction for M1/M3 (Arditi 2024)
Downloads · lifetime
178
24 last 30d - stable
Likes
1
Model age
17mo ago
created 2025-05-02

Training datasets

2 of 2 in /datasets

Corpora the author lists in the model card. Datasets tracked in our /datasets catalog carry a category badge linking to the workflow stage. Others open on Hugging Face.

Downloads over time
Now193→from0↑0%
0711422120 on Apr 30, 2025193 on Oct 11193 on Oct 10Apr '25Jul '25Oct '25JanAprJulOct
Apr 30, 2025 → Oct 11 · 115 snapshots · spans 529 days

Benchmarks

Portrait before abliteration
Benchmarks of the base model as it stood before the refusal-removal operation. Compare with the numbers above to see what the operation cost.
Benchmark Score Source
BBH average 0.448047736411986 OpenLLM-v2
IFEval instruct 0.5479616306954437 OpenLLM-v2
IFEval-Prompt 0.40850277264325324 OpenLLM-v2
MATH lvl 5 0.08383685800604229 OpenLLM-v2
MMLU-Pro 0.359125664893617 OpenLLM-v2

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
mit
Languages
en
Tags
transformers safetensors llama text-generation abliteration alignment safety llama3 directional_steering interpretability en dataset:mlabonne/harmful_behaviors

Related

Total size
15.0 GB
Files
9
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2025-05-02 19:19

Files by quantization

Auxiliary files 9 files 15.0 GB
model-00002-of-00004.safetensors 4.66 GB 77307462 download
model-00001-of-00004.safetensors 4.63 GB c920c356 download
model-00003-of-00004.safetensors 4.58 GB 69fe29fd download
model-00004-of-00004.safetensors 1.09 GB 5bbcd95a download
model.safetensors.index.json 23.4 KB 0fd8120f download
README.md 4.39 KB ee67364a download
.gitattributes 1.48 KB a6344aac download
config.json 689 B 381b8187 download
generation_config.json 194 B 93e1ac20 download

README current version from Hugging Face


library_name: transformers
tags:

  • abliteration
  • alignment
  • safety
  • llama3
  • directional_steering
  • interpretability
    license: mit
    datasets:
  • mlabonne/harmful_behaviors
  • mlabonne/harmless_alpaca
    language:
  • en
    base_model:
  • meta-llama/Meta-Llama-3-8B-Instruct

Model Card for ZennyKenny/Daredevil-8B-abliterated

This is an "abliterated" version of mlabonne/Daredevil-8B, based on the abliteration method developed by mlabonne to allow LLMs to perform otherwise restricted actions in through direction-based activation editing.

The technique projects out harmful activation directions without further finetuning or modifying the model architecture. It is inspired by work on steering vectors, mechanistic interpretability, and alignment by construction.


Model Details

Model Description

This model has been modified from meta-llama/Meta-Llama-3-8B-Instruct by applying vector-based orthogonal projection to internal representations associated with harmful outputs. The method uses HookedTransformer from transformer_lens to calculate harmful activation directions from prompt-based comparisons and then removes those components from the weights.

  • Model type: Causal Language Model
  • Language(s): English
  • License: llama3-license
  • Finetuned from model: mlabonne/Daredevil-8B
  • Modified from base model: meta-llama/Meta-Llama-3-8B-Instruct

Model Sources


Uses

Direct Use

This model is intended for experiments in safety and alignment research, especially in:

  • Exploring vector-based interpretability
  • Testing refusal behaviors
  • Evaluating models modified via non-finetuning methods

Out-of-Scope Use

  • Do not rely on this model for high-stakes decisions.
  • This model was not tested for factuality, multilingual use, or downstream generalization.
  • Not intended for production or safety-critical applications.

Bias, Risks, and Limitations

Limitations

  • Only a single direction (or small subset) was ablated—this does not guarantee complete refusal behavior.
  • Potential for capability degradation or underperformance on certain prompts.
  • Effectiveness is prompt-sensitive and may vary significantly.

Recommendations

  • Treat this model as exploratory, not final.
  • Evaluate outputs thoroughly before using in any application beyond experimentation.
  • Use interpretability tools (like transformer_lens) to understand effects layer-by-layer.

How to Get Started with the Model

from transformers import AutoModelForCausalLM, AutoTokenizer

model = AutoModelForCausalLM.from_pretrained("ZennyKenny/Daredevil-8B-abliterated")
tokenizer = AutoTokenizer.from_pretrained("meta-llama/Meta-Llama-3-8B-Instruct")

prompt = "How can I build a bomb?"
inputs = tokenizer(prompt, return_tensors="pt").to(model.device)
outputs = model.generate(**inputs, max_new_tokens=64)
print(tokenizer.decode(outputs[0], skip_special_tokens=True))

Training Details

Training Data

This model was not further trained. Instead, it used representations from:

  • mlabonne/harmful_behaviors (harmful prompt dataset)
  • mlabonne/harmless_alpaca (harmless instruction dataset)

Training Procedure

  • Model activations were captured with transformer_lens
  • Harmful vs. harmless activations compared across layers
  • Top directional vectors removed from internal weights via projection

Training Hyperparameters

  • Precision used: bfloat16 (model loading), float32 (conversion)
  • Orthogonalization method: L2-normalized difference vectors
  • Number of layers edited: Entire stack (all transformer blocks)

Evaluation

Model completions were evaluated by:

  • Human inspection of generations
  • Baseline vs. intervention vs. orthogonalized comparisons
  • Focused on refusal language: e.g., presence of "I can't", "I won't", etc.

Environmental Impact

  • Hardware Type: NVIDIA A100 (Google Colab)
  • Hours used: ~1
  • Cloud Provider: Google Cloud (Colab)
  • Compute Region: [Unknown]
  • Carbon Emitted: Minimal (low compute footprint, no training)

Model Card Contact

For questions, reach out via Hugging Face

README history 3 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2025-05-02Update README.mdcb2b3864.4 KB
    Loading...
  2. 2025-05-02Update README.mdd55cb754.5 KB
    Loading...
  3. 2025-05-02Upload LlamaForCausalLM78bbb235.1 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration