base_model: jdqqjr/DeepSeek-R1-Distill-Llama-3.2-1B-Instruct
library_name: peft
pipeline_tag: text-generation
tags:
- peft
- lora
- llama
- deepseek-r1-distill
- mechanistic-interpretability
- abliteration
- research
datasets: - custom
Geometric Abliteration Adapters for DeepSeek R1 Distill Llama 3.2 1B
This repository packages two small pure-projection LoRA adapters measured onjdqqjr/DeepSeek-R1-Distill-Llama-3.2-1B-Instruct using the Modal
abliteration pipeline.
The included adapters are:
| Adapter | Subfolder | Direction | Scale | Layers | Target modules |
|---|---|---|---|---|---|
| Disinhibition / hedge-reduction | adapters/disinhibition-lora-pure |
disinhibition_purified.pt |
2.0 |
1-15 |
o_proj, down_proj |
| Refusal-direction ablation | adapters/refusal-lora-pure |
refusal_purified.pt |
1.0 |
1-15 |
o_proj, down_proj |
Method
For a measured direction d, pure projection edits a target weight matrix W:
W_edited = W - scale * d (d^T W)
delta = -scale * d (d^T W)
That outer product is stored directly as rank-1 PEFT LoRA factors. No adapter
training or SVD is used.
Modal Run Provenance
The artifacts came from Modal volume model-weights:
- model:
llm/DeepSeek-R1-Distill-Llama-3.2-1B-Instruct - measurements:
measurements/DeepSeek-R1-Distill-Llama-3.2-1B-Instruct - adapters:
loras/DeepSeek-R1-Distill-Llama-3.2-1B-Instruct
Local source workspace:
/home/comrade/homelab/abliteration-research-hub/workspaces/deepseek-r1-distill-llama32-1b
Modal logs available from the recent inspect/full pipeline runs are stored inmodal-logs/.
Measurement Notes
Measurement reports are included under measurements/:
purification_report.jsonrefusal_purification_report.json
The disinhibition purification report marked layer 0 invalid, so the adapter
uses layers 1-15. The refusal report marked layers 0-15 valid, but the
adapter also uses 1-15 for consistency with the production Llama 3.2 adapter
convention and to avoid mixing a layer that failed the companion direction.
This is a research artifact. Marker-based direction measurement and benchmark
purification are not a full behavioral, safety, or capability evaluation.
Usage
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
from peft import PeftModel
base_id = "jdqqjr/DeepSeek-R1-Distill-Llama-3.2-1B-Instruct"
repo_id = "YOUR_HF_REPO_ID"
tokenizer = AutoTokenizer.from_pretrained(base_id)
base_model = AutoModelForCausalLM.from_pretrained(
base_id,
torch_dtype=torch.bfloat16,
device_map="auto",
)
model = PeftModel.from_pretrained(
base_model,
repo_id,
subfolder="adapters/disinhibition-lora-pure",
)
Files
adapters/disinhibition-lora-pure/
adapter_config.json
adapter_model.safetensors
ABLITERATION_META.json
adapters/refusal-lora-pure/
adapter_config.json
adapter_model.safetensors
ABLITERATION_META.json
measurements/
purification_report.json
refusal_purification_report.json
modal-logs/
modal-app-*.log
tools/
abliterate_to_lora.py
measure_overlap.py
eval/
eval_buckets.json
merge_adapters.py
Responsible Use
These adapters can alter refusal and hedging behavior. Do not treat them as a
substitute for safety evaluation, policy compliance checks, or domain-specific
validation. Any merged derivative inherits the base model's license and use
terms.