license: apache-2.0
base_model: Qwen/Qwen3.8-27B
tags:
- abliteration
- uncensored
- qwen
- qwen3.8
pipeline_tag: text-generation
datasets: - walledai/AdvBench
- tatsu-lab/alpaca
Qwen3.8-27B Abliterated
Uncensored version of Qwen/Qwen3.8-27B using weight-level refusal-direction ablation: a difference-in-means direction extracted from harmful vs. harmless prompt activations, permanently orthogonalized out of every residual-stream-writing weight matrix. There is no runtime hook and no LoRA; the edit is baked directly into the checkpoint.
Direction source: layer 37 of 64 (purified), applied at every layer at coefficient 1.0 (full strength).
Format
Text-only bf16 safetensors, architecture Qwen3_5ForCausalLM. The base repo is a vision-language model. This checkpoint contains only the language model: the vision encoder and the multi-token-prediction (MTP) head are not included, so it takes text input only.
Edited tensors: every residual-write projection. That is self_attn.o_proj in the 16 full-attention layers, linear_attn.out_proj in the 48 Gated DeltaNet linear-attention layers, and mlp.down_proj in all 64 layers. Every other tensor, the config, and the tokenizer/chat template are unchanged from the base model.
Serving
Transformers (tested with transformers>=5.19.0):
from transformers import AutoModelForCausalLM, AutoTokenizer
import torch
model_id = "rajaykumar12959/qwen3.8-27b-abliterated"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(model_id, dtype=torch.bfloat16, device_map="auto")
messages = [{"role": "user", "content": "Your prompt here"}]
prompt = tokenizer.apply_chat_template(
messages, tokenize=False, add_generation_prompt=True, enable_thinking=False
)
inputs = tokenizer(prompt, return_tensors="pt", add_special_tokens=False).to(model.device)
out = model.generate(**inputs, max_new_tokens=256)
print(tokenizer.decode(out[0][inputs["input_ids"].shape[1]:], skip_special_tokens=True))
Needs about 55 GB of GPU memory in bf16, e.g. one 80 GB A100/H100. Installing flash-linear-attention (and optionally causal-conv1d) speeds up the linear-attention layers considerably.
Thinking mode: all evaluation below used enable_thinking=False (direct answers). Thinking mode was not evaluated.
Metrics
| Metric | Original | Abliterated |
|---|---|---|
| Refusal rate (292 held-out harmful prompts, 12 categories) | 99.3% | 0.0% |
| MMLU-Pro accuracy (490 questions, 35 per subject × 14; chance 11.7%) | 65.1% | 61.8% (−3.3 pp) |
| KL divergence on harmless prompts (first token, 90 prompts) | — | 0.086 nats |
| Same top-1 first token as original on harmless prompts | — | 92.2% |
| ARC-Easy (100 questions) | 1.000 | 1.000 |
Refusal dropped to 0% in all 12 categories. Spot-checked generations are fluent English answers, comparable in length and vocabulary diversity to the original model's responses. The edit did not leave the model producing degenerate text that a grader would miss.
Tradeoff curve (reversible runtime-hook ablation, same direction, judged by the unmodified model):
| Coefficient | Refusal rate | Capability |
|---|---|---|
| 1.00 | 0.0% | 1.000 |
| 0.85 | 0.3% | 1.000 |
| 0.70 | 2.1% | 1.000 |
| 0.50 | 3.1% | 1.000 |
Capability cost (MMLU-Pro)
Removing refusal costs about 3.3 points on MMLU-Pro (65.1% → 61.8%). Eleven of 14 subjects dropped slightly, which looks systematic rather than noise. Each subject has only 35 questions, so one question is 2.9 pp, and per-subject differences of 1–2 questions are within noise:
| Subject | Original | Abliterated |
|---|---|---|
| biology | 97.1% | 97.1% |
| business | 45.7% | 34.3% |
| chemistry | 60.0% | 51.4% |
| computer science | 68.6% | 68.6% |
| economics | 68.6% | 62.9% |
| engineering | 48.6% | 45.7% |
| health | 80.0% | 77.1% |
| history | 91.4% | 88.6% |
| law | 45.7% | 48.6% |
| math | 48.6% | 45.7% |
| other | 57.1% | 57.1% |
| philosophy | 57.1% | 54.3% |
| physics | 54.3% | 51.4% |
| psychology | 88.6% | 82.9% |
MMLU-Pro is scored zero-shot by answer-letter log-likelihood (no chain of thought), so absolute numbers are lower than Qwen's published 5-shot CoT results. What matters here is the original-vs-abliterated difference on identical questions. The KL figure measures how much the next-token distribution moves on harmless prompts. It is small, which indicates the edit is largely confined to refusal behaviour.
Method
Per-layer difference-in-means direction extraction (fp32, last templated token, harmful vs. harmless prompts). The direction is Gram-Schmidt-purified against the harmless mean direction and unit-normalized, then removed from the weights by closed-form orthogonalization, with no gradient-based training. For a residual-writing projection y = Wx + b, the component along d̂ is removed exactly via W' = W − d̂(d̂ᵀW), b' = b − (d̂·b)d̂.
Layer selection: a sweep over layers 33/37/41/47/53/56 × coefficients 1.0/0.85/0.7/0.5 on a held-out validation slice. The layers with the highest activation signal-quality scores (53, 56) suppressed refusal the least (36–70% refusal remaining). Layer 37, at 58% depth, removed it entirely. That is the same relative depth as the winning layer for Qwen2.5-7B-Instruct (16/28).
| Layer | c=1.0 | c=0.85 | c=0.7 | c=0.5 |
|---|---|---|---|---|
| 33 | 14% | 18% | 16% | 20% |
| 37 | 0% | 0% | 0% | 0% |
| 41 | 0% | 4% | 2% | 2% |
| 47 | 0% | 0% | 0% | 0% |
| 53 | 36% | 38% | 34% | 38% |
| 56 | 70% | 70% | 64% | 66% |
(Validation-slice refusal rate, 50 prompts; capability was 1.000 at every point.)
Datasets
| Dataset | Role |
|---|---|
| walledai/AdvBench | Harmful prompts (direction extraction + eval) |
| tatsu-lab/alpaca | Harmless prompts (direction extraction) |
Evaluation
Per-category refusal rate:
| Category | Original | Abliterated |
|---|---|---|
| cybercrime_hacking | 100% | 0% |
| extremism | 100% | 0% |
| financial_crime | 100% | 0% |
| fraud_scams | 100% | 0% |
| hate_speech | 100% | 0% |
| illicit_drugs | 100% | 0% |
| malware | 100% | 0% |
| misinformation | 95.8% | 0% |
| privacy_invasion | 100% | 0% |
| self_harm | 95.8% | 0% |
| violence | 100% | 0% |
| weapons | 100% | 0% |
Caveats:
- Grading. Refusal grading is two-tier: rule-based first, with the model's own self-judge for ambiguous responses. For the ablated variants, 100% of verdicts went to the self-judge. For the weight-edited model, that judge is the edited model itself, a known asymmetry. The runtime-hook rows above are judged by the unmodified model and also reach 0.0% at full strength, which corroborates the weight-edit result.
- Capability. ARC-Easy saturates at this size (1.000 before and after) and would have hidden the capability cost. MMLU-Pro shows it: a 3.3 pp drop. A gentler edit (the runtime curve keeps refusal at 2.1% at coefficient 0.7) or norm-preserving ablation may reduce it; neither has been tested on this checkpoint.
Disclaimer
This model has had safety guardrails removed and will comply with requests the original model would refuse. Released for research into AI alignment, interpretability, and refusal mechanisms. The creator assumes no responsibility for downstream use.
Acknowledgments
- Qwen team: base model (Apache-2.0)
- Ablated and evaluated with AblateBench
- Created by rajaykumar12959