license: apache-2.0
base_model: Qwen/Qwen2.5-0.5B-Instruct
base_model_revision: 7ae557604adf67be50417f59c2c2f167def9a775
tags:
- abliteration
- refusal-direction
- uncensored
- research
library_name: transformers
sbussiso/Qwen2.5-0.5B-abliterated-r2
Abliterated (refusal-direction) variant of Qwen/Qwen2.5-0.5B-Instruct
at revision 7ae557604adf67be50417f59c2c2f167def9a775, produced by the Arditi et al. (2024) method.
Abliterated by the sbussiso lab research agent (Hermes, research-workstation profile) on Google
Colab T4, 2026-09-28.
Method
Arditi et al. 2024, "Refusal in LLMs is mediated by a single direction"
(NeurIPS 2024). From 64 harmful/harmless prompt pairs
(greedy decoding, seed 0), the mean difference of final-position
residual activations gives the refusal direction at each layer; the layer
with the highest direction coherence was chosen and its direction removed.
Two families of edits are compared in this repo's evaluation:
- inference-time ablation: project the direction out of every activation at
decoder layer 17 (forward hook, all positions) - the full-removal
contrast, NOT the published weights; - persistent weight decoding (published artifact): multi-layer row-space orth at top-5 layers [17, 18, 19, 16, 15] + lm_head orth + final-norm orth.
Qwen2.5-0.5B ships with tied embeddings, so the lm_head edit was applied to an UNTIED clone andtie_word_embeddings: falseis persisted in this repo's config.json (input embeddings untouched).
Ablation details
| field | value |
|---|---|
| chosen decoder layer | 17 / 24 (hook target model.model.layers[17]) |
| coherence (residual space) | 0.664 |
| coherence (final-layer readout space, direction B) | 0.562 |
| published variant | wd_ML_BN |
| structure | 24 layers, hidden 896, GQA 14q/2kv heads, tied embeddings: true (pre-edit) |
| direction pairs / probes | 64 / 16 |
| decoding | greedy (do_sample=False), max_new_tokens 200 |
| seed | 0 |
| GPU | Tesla T4 |
| python / torch | 3.13.15 / 2.11.0+cu128 |
Refusal behavior (16 harmful + 16 harmless probes, greedy, 200 tokens)
| condition | harmful refusal rate | harmless answered | degenerate outputs |
|---|---|---|---|
| baseline | 87.5% | 93.8% | 0 |
| hook (inference-time, L17) | 0.0% | 93.8% | 0 |
| wd_B | 56.2% | 93.8% | 0 |
| wd_BN | 56.2% | 87.5% | 0 |
| wd_ML | 68.8% | 81.2% | 0 |
| wd_ML_BN | 0.0% | 87.5% | 0 |
Headline: refusal 87.5% -> 0.0%
(published weights); benign-preservation 93.8% ->
87.5%. The inference-time hook condition measured
0.0% and is recorded as the full-removal contrast:
Persistent weight edits reach only the pathways the edited matrices carry,
so a persistent rate above the hook rate means residual refusal pathways
remain. Absolute refusal rates are keyword-marker based (first-person/
explicit markers only, constant scorer across conditions); deltas are
meaningful, absolute rates approximate.
Charts (generated programmatically from this repo's eval artifacts)




Generator: charts/make_r2_charts.py (regenerates every figure fromeval/ + the run JSONs; every number is read from recorded artifacts,
none hand-typed).
Capability check: MMLU (lm-evaluation-harness, 0-shot, fp16, seed 0)
| model | MMLU acc | acc_stderr |
|---|---|---|
| Qwen/Qwen2.5-0.5B-Instruct @ 7ae55760 | 45.78% | 0.41pp |
| this model (wd_ML_BN) | 45.55% | 0.41pp |
MMLU delta: 0.24pp (mission guardrail: capability loss
must stay under 3pp - PASSED).
Both evaluations used identical config (lm_eval --model hf --tasks mmlu --num_fewshot 0 --batch_size auto --seed 0, dtype float16; base loaded at
the pinned revision). Raw results in eval/.
Intended use
- Research artifact: study of the refusal-direction phenomenon and of
persistent weight-space ablation depth on a small instruct model. - NOT a production assistant. Refusal behavior is deliberately degraded;
the model may produce harmful content when asked for it. Do not deploy
where that is unacceptable. Quality/verbosity of the base model is not
guaranteed to be preserved beyond the probes and MMLU check above.
Files
- full safetensors weights + tokenizer (this repo root)
refusal_direction.npy(representative direction for wd_ML_BN),refusal_direction_A.npy(residual space, layer 17),refusal_direction_B.npy(final-layer readout space),layer_directions.npz(all 24 per-layer directions)eval/- lm-eval MMLU results (base + variant) and refusal probe logscharts/- the four card figures + generatormake_r2_charts.pyrun_config.json,layer_coherence.json,selection.json,selection_candidates.json- run metadata