license: llama3.2
base_model: meta-llama/Llama-3.2-1B-Instruct
base_model_revision: 9213176726f574b556790deb65791e0c5aa438b6
library_name: transformers
pipeline_tag: text-generation
tags:
- abliteration
- refusal-direction
- research-artifact
- llama
Llama-3.2-1B-Instruct-abliterated
A refusal-ablated edit of meta-llama/Llama-3.2-1B-Instruct (pinned
revision 9213176726f574b556790deb65791e0c5aa438b6): the readout-space
refusal direction was orthogonalized out of the untied lm_head weight,
making the edit persistent — no hook required at inference time.
First non-Qwen patient of this engine; the founding direction method
was originally characterized on Llama-2-class models, so this run
measures it back on a modern small Llama.
What it does
Refusal behavior on a fixed 64-prompt harmful set (greedy, 200 new
tokens, marker-based scorer — identical instrument to this engine's
other patients; every arm below is per-row artifact-backed):
| arm | refused /64 | note |
|---|---|---|
| base Llama-3.2-1B-Instruct | 38 (59.4%) | artifact-backed |
| inference-time hook ablation | 34 (53.1%) | artifact-backed |
| this variant (wd_B, persistent) | 8 (12.5%) | strongest gate result in this program — clears the ≤25% publish gate outright |
| wd_BN (readout + final norm) | 8 (12.5%) | identical refusal to wd_B; adds 0 |
| wd_ML (K=3 mid-layer row-space) | 34 (53.1%) | no better than hook |

Direction transfer across architecture families
The refusal direction's layer fingerprint moves with the architecture:
the coherence scan selects a mid-stack site (layer 9 of 16,
coherence 0.717) here, versus the deep sites of the Qwen2.5 family
(L17/24). Same method, materially different geometry — the run's main
architecture-dependence datapoint.

Benign behavior
Benign preservation on the 64-prompt harmless set, same instrument
(per-row artifacts; benign floor gate = baseline − 10pp clears at every arm):
| arm | answered /64 |
|---|---|
| base | 64 (100%) |
| hook | 61 (95.3%) |
| this variant (wd_B) | 63 (98.4%) |
| wd_BN | 63 (98.4%) |
| wd_ML | 64 (100%) |

Capability guardrail (MMLU)
0-shot MMLU over all 61 subjects (lm-eval 0.4.13, fp16, seed 0):
| model | MMLU % |
|---|---|
| base Llama-3.2-1B-Instruct | 48.27 |
| this variant (wd_B) | 47.98 |
Loss: −0.29pp against the ≤3.0pp limit — PASS, 10× margin. The base
arm's measurement is digit-identical (48.2694) across two independent
GPU sessions; per-arm raw results + the engine log ship under eval/mmlu/.

Try it
from transformers import AutoModelForCausalLM, AutoTokenizer
repo = "sbussiso/Llama-3.2-1B-Instruct-abliterated"
tok = AutoTokenizer.from_pretrained(repo,
clean_up_tokenization_spaces=False)
model = AutoModelForCausalLM.from_pretrained(repo, torch_dtype="auto")
messages = [{"role": "user", "content": "Explain what a for loop is."}]
prompt = tok.apply_chat_template(messages, add_generation_prompt=True,
tokenize=False)
enc = tok(prompt, return_tensors="pt")
out = model.generate(**enc, max_new_tokens=200, do_sample=False)
print(tok.decode(out[0][enc["input_ids"].shape[1]:],
skip_special_tokens=True))
Method
Founding directional-editing recipe (Arditi et al. 2024): 64-pair
contrastive direction extraction → coherence-ranked site selection →
weight-space orthogonalization. This variant (wd_B) edits thelm_head readout: W ← W − (W r̂) r̂ᵀ with the untied lm_head
(tie_word_embeddings true → false, persisted in config). The edit's
disk-verified residual (max |W r̂| = 7.1e-05) agrees to three
significant digits across two GPU edit streams and a deterministic
CPU rebuild of the same edit.
Honest limitations
- Probe tables are from fixed 64-prompt sets, greedy decoding,
single-seed, single marker-set scorer — not a benchmark-suite claim. - Selective-refusal calibration (RefusalBench-style graded refusal) and
refusal-residual breadth (e.g. SORRY-Bench classes) are not yet run
on this patient. - Any post-hoc refusal ablation is recoverable by small benign
fine-tuning (literature finding; applies to this edit). - Meta's Llama 3.2 Community License governs the base model and this
edit; use under that license's terms.
References
- Arditi, A. et al. 2024. Refusal in Language Models Is Mediated by a
Single Direction. NeurIPS 2024. https://arxiv.org/abs/2406.11717 - Xie, T. et al. 2025. SORRY-Bench: Systematically Evaluating Large
Language Model Safety Refusal. ICLR 2025. (residual-breadth instrument
queued for this patient) - Muhamed, A. et al. 2025. RefusalBench: Generative Evaluation of
Selective Refusal in Grounded Language Models. EACL 2026.
(selective-refusal instrument queued for this patient) - Malla, S. et al. 2025. The Geometry of Refusal: Why Post-Hoc Safety
Is Fragile and Pretraining-Time Safety Persists.
https://arxiv.org/abs/2609.06934
Provenance
- Base weights:
meta-llama/Llama-3.2-1B-Instruct@9213176726f574b556790deb65791e0c5aa438b6(gated; access accepted under
the owner account). - This variant (wd_B):
tie_word_embeddingstrue → false (persisted);
main weights sha256cb9ad4d09bb787ac0bbe77966afa06a1e82e23535b0f02beec7d7870618bcc21
(GPU-native — the engine's own edit stream on L4; the canonical bytes
the run's measurements were made against). A deterministic CPU rebuild
of the same edit was produced and verified from the shipped direction
banks: sha2566417be31231b76300c7013acccfd30251cc5c105356089a456950e23666e0a9c,
behaviorally identical (probe rows digit-exact, edit residuals agree
to 3 significant figures; bytes differ by platform — the run's
documented device-provenance finding). Rebuild recipe + residual
chain ship ineval/rebuild_evidence.json; the twin bytes are not
duplicated here (regenerate from the banks instead). - Direction banks:
refusal_direction_A.npy(residual-space, ‖d‖ 3.78)
andrefusal_direction_B.npy(readout-space, ‖d‖ 74.58) — the run's
banked stage-A artifacts, shipped for full re-derivability. - Charts in
charts/are generated programmatically from the recorded
probe/coherence/MMLU artifacts (make_card_charts.pyships with the run),
in the engine's dark house palette; no hand-typed digits. - Full method + evaluation write-ups: available on request.
Files
| path | what |
|---|---|
model.safetensors |
this variant's weights (wd_B, GPU-native, canonical) |
config.json, generation_config.json, tokenizer*, chat_template.jinja |
base-derived runtime files (untie persisted) |
refusal_direction_A.npy, refusal_direction_B.npy |
banked stage-A direction banks |
eval/probes_*.json |
per-row probe records, all five arms (baseline/hook/wd_B/wd_BN/wd_ML) |
eval/mmlu/ |
MMLU guardrail summary + engine log (this card's guardrail numbers) |
eval/rebuild_evidence.json |
weights provenance (both shas, recipe, edit residuals — the CPU rebuild is fully re-derivable from the shipped banks) |
eval/selection.json |
engine's variant-selection record |
charts/*.png |
the four figures embedded above |
This is a research artifact. It is not a product and is not
intended for production use.