base_model: Qwen/Qwen3.5-9B
license: <SPDX identifier of Qwen/Qwen3.5-9B's license> -- see the License section below
tags:
- abliterated
- orthex
Qwen/Qwen3.5-9B (abliterated)
Weight-level orthogonalized ("abliterated") version of Qwen/Qwen3.5-9B, produced with orthex — an implementation of Arditi et al., "Refusal in Language Models Is Mediated by a Single Direction" (NeurIPS 2024).
🧩 Summary
| Base model | Qwen/Qwen3.5-9B |
| Architecture adapter | qwen3_5 |
| Ablation strategy | weight_orthogonalization |
| Ablation targets | embed_tokens, every layer's attn_out, every layer's mlp_out and lm_head (untied from embed_tokens, ablated separately) |
| Selected direction | layer 22, site resid_pre |
Ablation is applied in place to the weights listed above — not a runtime hook. This checkpoint behaves this way standalone, with no orthex dependency at inference time.
📊 Evaluation
Measured on the held-out test prompt set, pre vs. post ablation:
| Metric | Pre | Post | Δ |
|---|---|---|---|
| Refusal rate | 0.97 | 0.12 | -0.84 |
| Perplexity | 16.26 | 21.26 | 5.00 |
See evaluation_report.json in this repo for the full per-prompt breakdown (refusal_samples) and the ranked candidate list considered during selection (selection_report).
⚠️ Responsible use
This model has had refusal behavior removed and may comply with requests the base model would normally decline. It is intended for red-teaming, robustness research, and model-behavior analysis. Usage remains subject to the base model's original license and usage policy — this repo does not grant any additional rights beyond what Qwen/Qwen3.5-9B's license allows.
⚖️ License
This model's license follows Qwen/Qwen3.5-9B's original license, unchanged — this repo grants no additional rights.