← back to catalog · registered 2026-08-22 13:56

marx161-cmd/geometric-abliteration-adapters

Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/marx161-cmd%2Fgeometric-abliteration-adapters"
Response includes
  • classification unknown
  • files 4
  • benchmarks 21 entries
  • author_summary 5 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
?
Primary method

Unclassified

No clear signals of an abliteration technique in this model.
Confidence
UNKNOWN
Why this label 1 signal
No classification signals present. This may not be an abliterated model at all - it could be a repackaging, a merge with unrelated goals, or unrelated content that mentions the term.
  • no classification signals present (no abliterated, uncensored, or known producer/method markers)
Refusal direction extraction

No specific extraction method could be identified for this model. The producer either did not document it or used a proprietary pipeline.

What is a refusal direction? →
Downloads · 30-day
0
Likes
0
Model age
4mo ago
created 2026-06-09

Training datasets

Corpora the author lists in the model card. Datasets tracked in our /datasets catalog carry a category badge linking to the workflow stage. Others open on Hugging Face.

Downloads over time
Now0→from0↑0%
00110 on Jun 100 on Oct 11JunJulAugSepOct
Jun 10 → Oct 11 · 57 snapshots · spans 123 days

Benchmarks

Portrait before abliteration
Benchmarks of the base model as it stood before the refusal-removal operation. Compare with the numbers above to see what the operation cost.
Benchmark Score Source
Arena-Battles 8523 LM-Arena
LM Arena Elo 1067.4348646816593 LM-Arena
Arena-Elo-Lower 1060.145342091014 LM-Arena
Arena-Elo-Upper 1074.7243872723045 LM-Arena
Arena-Rank 207 LM-Arena
BBH average 0.3342846024127643 OpenLLM-v2
IFEval instruct 0.6294964028776978 OpenLLM-v2
IFEval-Prompt 0.5101663585951941 OpenLLM-v2
MATH lvl 5 0.02945619335347432 OpenLLM-v2
MMLU-Pro 0.16821808510638298 OpenLLM-v2
Entertainment 0.2 UGI
Hazardous 0 UGI
Natural Intelligence 6.89 UGI
Political lean -6.0% UGI
Sensitive-Info 2.34 UGI
SocPol 0.5 UGI
UGI 6.56 UGI
Willingness (10) 1.5 UGI
W10-Adherence 1 UGI
W10-Direct 2 UGI
Writing 11.82 UGI

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
llama3.2
Tags
peft safetensors lora llama mechanistic-interpretability abliteration research text-generation dataset:custom base_model:meta-llama/Llama-3.2-1B-Instruct base_model:adapter:meta-llama/Llama-3.2-1B-Instruct license:llama3.2

Related

Total size
0 B
Files
4
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-06-09 19:47

Files by quantization

Auxiliary files 4 files 8.56 KB
README.md 6.14 KB 285d364e download
merge_adapters.py 2.27 KB 9acc221b download
requirements.txt 108 B 00f035b8 download
.gitattributes 50.0 B 580d310c download

README current version from Hugging Face


license: llama3.2
base_model: meta-llama/Llama-3.2-1B-Instruct
library_name: peft
pipeline_tag: text-generation
tags:

  • peft
  • lora
  • llama
  • mechanistic-interpretability
  • abliteration
  • research
    datasets:
  • custom

Composable Geometric Abliteration via Rank-1 PEFT Adapters

This repository publishes small pure-projection LoRA adapters for
meta-llama/Llama-3.2-1B-Instruct. The adapters encode behavioral direction
edits as rank-1 PEFT deltas rather than distributing a full merged model.

The two included adapters are:

Adapter Subfolder Direction Scale Layers Target modules
Disinhibition / hedge-reduction adapters/disinhibition-lora-pure disinhibition_purified.pt 2.0 1-15 o_proj, down_proj
Refusal-direction ablation adapters/refusal-lora-pure refusal_purified.pt 1.0 1-15 o_proj, down_proj

Each adapter has 215,040 LoRA parameters and is about 0.44 MB as safetensors.
The base model weights are not included. Users must have their own access to
the Llama 3.2 1B Instruct base model and must follow the base model license.

Core Idea

For a measured unit direction d, pure projection edits a target weight matrix
W as:

W_edited = W - scale * d (d^T W)
delta = W_edited - W = -scale * d (d^T W)

This delta is an outer product and is therefore exactly rank-1. It can be stored
directly as PEFT LoRA factors:

B = -scale * d
A = d^T W
delta = B @ A

No training, fitting, or SVD is needed for the pure-projection case. The scale
is baked into B; the PEFT config uses r=1, lora_alpha=1,
lora_dropout=0.0, and init_lora_weights=false.

Why Pure Projection

Earlier norm-restored weight edits are useful historical baselines, but norm
restoration is a direction-blind nonlinear rescale. It breaks exact rank-1
composition and makes multiple edits harder to reason about. For composable
adapters, pure projection is the cleaner primitive: each delta is computed
against the original base weights and multiple adapters add linearly.

Important limitation: if two measured directions overlap, additive stacking can
double-count the shared component. In the Llama 3.2 1B measurements that
motivated this release, purified refusal/disinhibition overlap was small but
not zero. Direction overlap should be measured per model and per direction pair.

Quick Start

Install:

pip install -r requirements.txt

Load one adapter:

import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
from peft import PeftModel

repo_id = "marx161-cmd/geometric-abliteration-adapters"
base_id = "meta-llama/Llama-3.2-1B-Instruct"

tokenizer = AutoTokenizer.from_pretrained(base_id)
base_model = AutoModelForCausalLM.from_pretrained(
    base_id,
    torch_dtype=torch.bfloat16,
    device_map="auto",
)

model = PeftModel.from_pretrained(
    base_model,
    repo_id,
    subfolder="adapters/disinhibition-lora-pure",
    adapter_name="disinhibition",
)

Load both adapters into one model:

model.load_adapter(
    repo_id,
    subfolder="adapters/refusal-lora-pure",
    adapter_name="refusal",
)

# Recent PEFT versions support activating multiple adapters by list.
model.set_adapter(["disinhibition", "refusal"])

If your PEFT version does not support list activation, load and merge one
adapter, then load the next against the merged in-memory model, or upgrade PEFT.
See merge_adapters.py for a complete script.

Offline Merge

To bake one or both adapters into a standalone local model directory:

python merge_adapters.py \
  --base-model meta-llama/Llama-3.2-1B-Instruct \
  --repo-id marx161-cmd/geometric-abliteration-adapters \
  --adapters disinhibition,refusal \
  --out ./llama-3.2-1b-geometric-double-edit

This writes a local merged model for inference engines that do not load PEFT
adapters directly. The merged output inherits the base model license.

Evaluation Notes

These adapters were developed from local contrastive activation measurements and
marker-based regression harnesses:

Bucket What was counted Limitation
Opinion/disinhibition hedge and neutrality markers marker counts are not semantic truthfulness or helpfulness tests
Refusal direction refusal-marker phrases on harmful/harmless prompt sets not a full safety evaluation
Coherence short-output, repetition, NaN/Inf-style flags catches obvious breakage only

The disinhibition direction should be understood as broad hedge reduction. A
later break-scale check did not find a clean separation where only
manufactured/corporate hedging is removed while justified caution remains.
Scale is the practical behavioral dial.

Data

data/disinhibition_paired_curated_300.jsonl contains the custom paired-topic
probe set used for disinhibition measurement. Each row has a direct
stance-seeking prompt and a matched noncommittal/balance prompt on the same
topic. The pairing is intended to reduce topic-level variance in the activation
contrast.

Harmful/refusal prompt datasets are not redistributed in this repository.

Responsible Use

This is a research artifact for studying low-rank directional model edits. It
can change refusal and hedging behavior in instruction-tuned models. Do not use
the adapters as a substitute for safety evaluation, policy compliance checks, or
domain-specific validation. If you publish merged derivatives, clearly disclose
the base model, adapter names, scales, and evaluation limitations.

Files

adapters/disinhibition-lora-pure/
  adapter_config.json
  adapter_model.safetensors
  ABLITERATION_META.json
adapters/refusal-lora-pure/
  adapter_config.json
  adapter_model.safetensors
  ABLITERATION_META.json
data/
  disinhibition_paired_curated_300.jsonl
  README.md
tools/
  abliterate_to_lora.py
  measure_overlap.py
eval/
  eval_buckets.json
merge_adapters.py

Citation / Related Work

  • Arditi et al., "Refusal in Language Models Is Mediated by a Single Direction"
  • Grimjim/Jim Lai, projected and norm-preserving abliteration writeups
  • Treadon/Ritesh Khanna, abliteration and disinhibition experiments

This repository is an independent research release and is not affiliated with
Meta.

README history 1 version

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-06-09Publish geometric abliteration LoRA adapters3fb4e5a6.1 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration