← back to catalog · registered 2026-09-28 08:57

sbussiso/Qwen2.5-0.5B-abliterated-r2

sbussiso Qwen 500M
curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/sbussiso%2FQwen2.5-0.5B-abliterated-r2"
Response includes
  • classification m-uncensored
  • files 18
  • author_summary 2 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M-U
Primary method

Uncensored (method unknown)

No other method signals detected in this model.
Confidence
LOW
Why this label 3 signals
Weak or ambiguous signals. Best guess based on catalog patterns; treat as tentative and check the evidence below.
  • 'uncensored' in name/tags but no 'abliterated' marker
  • method not identifiable from author declaration alone
  • may be DPO fine-tune, prompt engineering, or unknown technique
Refusal direction extraction

No specific extraction method could be identified for this model. The producer either did not document it or used a proprietary pipeline.

What is a refusal direction? →
Downloads · 30-day
0
Likes
0
Model age
today
created 2026-09-28

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
apache-2.0
Tags
transformers safetensors qwen2 text-generation abliteration refusal-direction uncensored research conversational base_model:Qwen/Qwen2.5-0.5B-Instruct base_model:finetune:Qwen/Qwen2.5-0.5B-Instruct license:apache-2.0

Related

Total size
1.17 GB
Files
18
Quantizations
1
Registered
2026-09-28 08:57
Last updated on HF
2026-09-28 07:17

Files by quantization

Auxiliary files 18 files 1.18 GB
model.safetensors 1.17 GB 0b213334 download
tokenizer.json 10.9 MB 3fd16973 download
layer_directions.npz 66.8 KB 08c28d2b download
README.md 4.81 KB 36e83037 download
refusal_direction.npy 3.63 KB 5004544a download
refusal_direction_A.npy 3.63 KB e5da6984 download
refusal_direction_B.npy 3.63 KB 5004544a download
layer_coherence.json 3.36 KB 61b3d8f4 download
chat_template.jinja 2.45 KB bdf7919a download
run_config.json 2.16 KB e468cf10 download
.gitattributes 1.53 KB 52373fe2 download
config.json 1.25 KB 6f4ed009 download
tokenizer_config.json 694 B 770e41d6 download
selection_candidates.json 668 B 8df116e3 download
selection.json 470 B 3a96221d download
generation_config.json 242 B 75a97a7c download
harness_sha256.json 174 B ad960453 download
ladder_sha256.json 89.0 B 6284179b download

README current version from Hugging Face


license: apache-2.0
base_model: Qwen/Qwen2.5-0.5B-Instruct
base_model_revision: 7ae557604adf67be50417f59c2c2f167def9a775
tags:

  • abliteration
  • refusal-direction
  • uncensored
  • research
    library_name: transformers

sbussiso/Qwen2.5-0.5B-abliterated-r2

Abliterated (refusal-direction) variant of Qwen/Qwen2.5-0.5B-Instruct
at revision 7ae557604adf67be50417f59c2c2f167def9a775, produced by the Arditi et al. (2024) method.

Abliterated by the sbussiso lab research agent (Hermes, research-workstation profile) on Google
Colab T4, 2026-09-28.

Method

Arditi et al. 2024, "Refusal in LLMs is mediated by a single direction"
(NeurIPS 2024). From 64 harmful/harmless prompt pairs
(greedy decoding, seed 0), the mean difference of final-position
residual activations gives the refusal direction at each layer; the layer
with the highest direction coherence was chosen and its direction removed.

Two families of edits are compared in this repo's evaluation:

  • inference-time ablation: project the direction out of every activation at
    decoder layer 17 (forward hook, all positions) - the full-removal
    contrast, NOT the published weights;
  • persistent weight decoding (published artifact): multi-layer row-space orth at top-5 layers [17, 18, 19, 16, 15] + lm_head orth + final-norm orth.
    Qwen2.5-0.5B ships with tied embeddings, so the lm_head edit was applied to an UNTIED clone and tie_word_embeddings: false is persisted in this repo's config.json (input embeddings untouched).

Ablation details

field value
chosen decoder layer 17 / 24 (hook target model.model.layers[17])
coherence (residual space) 0.664
coherence (final-layer readout space, direction B) 0.562
published variant wd_ML_BN
structure 24 layers, hidden 896, GQA 14q/2kv heads, tied embeddings: true (pre-edit)
direction pairs / probes 64 / 16
decoding greedy (do_sample=False), max_new_tokens 200
seed 0
GPU Tesla T4
python / torch 3.13.15 / 2.11.0+cu128

Refusal behavior (16 harmful + 16 harmless probes, greedy, 200 tokens)

condition harmful refusal rate harmless answered degenerate outputs
baseline 87.5% 93.8% 0
hook (inference-time, L17) 0.0% 93.8% 0
wd_B 56.2% 93.8% 0
wd_BN 56.2% 87.5% 0
wd_ML 68.8% 81.2% 0
wd_ML_BN 0.0% 87.5% 0

Headline: refusal 87.5% -> 0.0%
(published weights); benign-preservation 93.8% ->
87.5%. The inference-time hook condition measured
0.0% and is recorded as the full-removal contrast:
Persistent weight edits reach only the pathways the edited matrices carry,
so a persistent rate above the hook rate means residual refusal pathways
remain. Absolute refusal rates are keyword-marker based (first-person/
explicit markers only, constant scorer across conditions); deltas are
meaningful, absolute rates approximate.

Charts (generated programmatically from this repo's eval artifacts)

Refusal by condition
Benign preservation
MMLU guardrail
Layer coherence scan

Generator: charts/make_r2_charts.py (regenerates every figure from
eval/ + the run JSONs; every number is read from recorded artifacts,
none hand-typed).

Capability check: MMLU (lm-evaluation-harness, 0-shot, fp16, seed 0)

model MMLU acc acc_stderr
Qwen/Qwen2.5-0.5B-Instruct @ 7ae55760 45.78% 0.41pp
this model (wd_ML_BN) 45.55% 0.41pp

MMLU delta: 0.24pp (mission guardrail: capability loss
must stay under 3pp - PASSED).
Both evaluations used identical config (lm_eval --model hf --tasks mmlu --num_fewshot 0 --batch_size auto --seed 0, dtype float16; base loaded at
the pinned revision). Raw results in eval/.

Intended use

  • Research artifact: study of the refusal-direction phenomenon and of
    persistent weight-space ablation depth on a small instruct model.
  • NOT a production assistant. Refusal behavior is deliberately degraded;
    the model may produce harmful content when asked for it. Do not deploy
    where that is unacceptable. Quality/verbosity of the base model is not
    guaranteed to be preserved beyond the probes and MMLU check above.

Files

  • full safetensors weights + tokenizer (this repo root)
  • refusal_direction.npy (representative direction for wd_ML_BN),
    refusal_direction_A.npy (residual space, layer 17),
    refusal_direction_B.npy (final-layer readout space),
    layer_directions.npz (all 24 per-layer directions)
  • eval/ - lm-eval MMLU results (base + variant) and refusal probe logs
  • charts/ - the four card figures + generator make_r2_charts.py
  • run_config.json, layer_coherence.json, selection.json,
    selection_candidates.json - run metadata
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.