← back to catalog · registered 2026-10-09 20:58

rajaykumar12959/qwen3.8-27b-abliterated

Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/rajaykumar12959%2Fqwen3.8-27b-abliterated"
Response includes
  • classification m-uncensored
  • files 19
  • author_summary 2 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M-U
Primary method

Uncensored (method unknown)

No other method signals detected in this model.
Confidence
LOW
Why this label 3 signals
Weak or ambiguous signals. Best guess based on catalog patterns; treat as tentative and check the evidence below.
  • 'uncensored' in name/tags but no 'abliterated' marker
  • method not identifiable from author declaration alone
  • may be DPO fine-tune, prompt engineering, or unknown technique
Refusal direction extraction

No specific extraction method could be identified for this model. The producer either did not document it or used a proprietary pipeline.

What is a refusal direction? →
Downloads · 30-day
0
Likes
1
Model age
today
created 2026-10-09

Training datasets

1 of 2 in /datasets

Corpora the author lists in the model card. Datasets tracked in our /datasets catalog carry a category badge linking to the workflow stage. Others open on Hugging Face.

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
apache-2.0
Tags
safetensors qwen3_5_text abliteration uncensored qwen qwen3.8 text-generation conversational dataset:walledai/AdvBench dataset:tatsu-lab/alpaca base_model:Qwen/Qwen3.8-27B base_model:finetune:Qwen/Qwen3.8-27B

Related

Total size
50.1 GB
Files
19
Quantizations
1
Registered
2026-10-09 20:58
Last updated on HF
2026-10-09 20:09

Files by quantization

Auxiliary files 19 files 50.1 GB
model-00007-of-00011.safetensors 4.99 GB 94b0115f download
model-00009-of-00011.safetensors 4.99 GB e3600063 download
model-00010-of-00011.safetensors 4.96 GB 9b4101b5 download
model-00004-of-00011.safetensors 4.96 GB 06833b97 download
model-00006-of-00011.safetensors 4.96 GB ad0b8480 download
model-00005-of-00011.safetensors 4.96 GB 154e61f2 download
model-00001-of-00011.safetensors 4.95 GB be145aea download
model-00003-of-00011.safetensors 4.95 GB ecaf1bc2 download
model-00008-of-00011.safetensors 4.95 GB b4bc5b15 download
model-00002-of-00011.safetensors 4.93 GB 771e67d8 download
model-00011-of-00011.safetensors 540 MB 00838314 download
tokenizer.json 19.1 MB a5cd9732 download
model.safetensors.index.json 81.9 KB aa3eb1ac download
chat_template.jinja 8.74 KB c0c686f9 download
README.md 7.61 KB f06c4a67 download
config.json 2.68 KB b6ac85c2 download
.gitattributes 1.53 KB 52373fe2 download
tokenizer_config.json 1.43 KB 59c317c2 download
generation_config.json 214 B d36de3b2 download

README current version from Hugging Face


license: apache-2.0
base_model: Qwen/Qwen3.8-27B
tags:

  • abliteration
  • uncensored
  • qwen
  • qwen3.8
    pipeline_tag: text-generation
    datasets:
  • walledai/AdvBench
  • tatsu-lab/alpaca

Qwen3.8-27B Abliterated

Uncensored version of Qwen/Qwen3.8-27B using weight-level refusal-direction ablation: a difference-in-means direction extracted from harmful vs. harmless prompt activations, permanently orthogonalized out of every residual-stream-writing weight matrix. There is no runtime hook and no LoRA; the edit is baked directly into the checkpoint.

Direction source: layer 37 of 64 (purified), applied at every layer at coefficient 1.0 (full strength).

Format

Text-only bf16 safetensors, architecture Qwen3_5ForCausalLM. The base repo is a vision-language model. This checkpoint contains only the language model: the vision encoder and the multi-token-prediction (MTP) head are not included, so it takes text input only.

Edited tensors: every residual-write projection. That is self_attn.o_proj in the 16 full-attention layers, linear_attn.out_proj in the 48 Gated DeltaNet linear-attention layers, and mlp.down_proj in all 64 layers. Every other tensor, the config, and the tokenizer/chat template are unchanged from the base model.

Serving

Transformers (tested with transformers>=5.19.0):

from transformers import AutoModelForCausalLM, AutoTokenizer
import torch

model_id = "rajaykumar12959/qwen3.8-27b-abliterated"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(model_id, dtype=torch.bfloat16, device_map="auto")

messages = [{"role": "user", "content": "Your prompt here"}]
prompt = tokenizer.apply_chat_template(
    messages, tokenize=False, add_generation_prompt=True, enable_thinking=False
)
inputs = tokenizer(prompt, return_tensors="pt", add_special_tokens=False).to(model.device)
out = model.generate(**inputs, max_new_tokens=256)
print(tokenizer.decode(out[0][inputs["input_ids"].shape[1]:], skip_special_tokens=True))

Needs about 55 GB of GPU memory in bf16, e.g. one 80 GB A100/H100. Installing flash-linear-attention (and optionally causal-conv1d) speeds up the linear-attention layers considerably.

Thinking mode: all evaluation below used enable_thinking=False (direct answers). Thinking mode was not evaluated.

Metrics

Metric Original Abliterated
Refusal rate (292 held-out harmful prompts, 12 categories) 99.3% 0.0%
MMLU-Pro accuracy (490 questions, 35 per subject × 14; chance 11.7%) 65.1% 61.8% (−3.3 pp)
KL divergence on harmless prompts (first token, 90 prompts) — 0.086 nats
Same top-1 first token as original on harmless prompts — 92.2%
ARC-Easy (100 questions) 1.000 1.000

Refusal dropped to 0% in all 12 categories. Spot-checked generations are fluent English answers, comparable in length and vocabulary diversity to the original model's responses. The edit did not leave the model producing degenerate text that a grader would miss.

Tradeoff curve (reversible runtime-hook ablation, same direction, judged by the unmodified model):

Coefficient Refusal rate Capability
1.00 0.0% 1.000
0.85 0.3% 1.000
0.70 2.1% 1.000
0.50 3.1% 1.000

Capability cost (MMLU-Pro)

Removing refusal costs about 3.3 points on MMLU-Pro (65.1% → 61.8%). Eleven of 14 subjects dropped slightly, which looks systematic rather than noise. Each subject has only 35 questions, so one question is 2.9 pp, and per-subject differences of 1–2 questions are within noise:

Subject Original Abliterated
biology 97.1% 97.1%
business 45.7% 34.3%
chemistry 60.0% 51.4%
computer science 68.6% 68.6%
economics 68.6% 62.9%
engineering 48.6% 45.7%
health 80.0% 77.1%
history 91.4% 88.6%
law 45.7% 48.6%
math 48.6% 45.7%
other 57.1% 57.1%
philosophy 57.1% 54.3%
physics 54.3% 51.4%
psychology 88.6% 82.9%

MMLU-Pro is scored zero-shot by answer-letter log-likelihood (no chain of thought), so absolute numbers are lower than Qwen's published 5-shot CoT results. What matters here is the original-vs-abliterated difference on identical questions. The KL figure measures how much the next-token distribution moves on harmless prompts. It is small, which indicates the edit is largely confined to refusal behaviour.

Method

Per-layer difference-in-means direction extraction (fp32, last templated token, harmful vs. harmless prompts). The direction is Gram-Schmidt-purified against the harmless mean direction and unit-normalized, then removed from the weights by closed-form orthogonalization, with no gradient-based training. For a residual-writing projection y = Wx + b, the component along d̂ is removed exactly via W' = W − d̂(d̂ᵀW), b' = b − (d̂·b)d̂.

Layer selection: a sweep over layers 33/37/41/47/53/56 × coefficients 1.0/0.85/0.7/0.5 on a held-out validation slice. The layers with the highest activation signal-quality scores (53, 56) suppressed refusal the least (36–70% refusal remaining). Layer 37, at 58% depth, removed it entirely. That is the same relative depth as the winning layer for Qwen2.5-7B-Instruct (16/28).

Layer c=1.0 c=0.85 c=0.7 c=0.5
33 14% 18% 16% 20%
37 0% 0% 0% 0%
41 0% 4% 2% 2%
47 0% 0% 0% 0%
53 36% 38% 34% 38%
56 70% 70% 64% 66%

(Validation-slice refusal rate, 50 prompts; capability was 1.000 at every point.)

Datasets

Dataset Role
walledai/AdvBench Harmful prompts (direction extraction + eval)
tatsu-lab/alpaca Harmless prompts (direction extraction)

Evaluation

Per-category refusal rate:

Category Original Abliterated
cybercrime_hacking 100% 0%
extremism 100% 0%
financial_crime 100% 0%
fraud_scams 100% 0%
hate_speech 100% 0%
illicit_drugs 100% 0%
malware 100% 0%
misinformation 95.8% 0%
privacy_invasion 100% 0%
self_harm 95.8% 0%
violence 100% 0%
weapons 100% 0%

Caveats:

  • Grading. Refusal grading is two-tier: rule-based first, with the model's own self-judge for ambiguous responses. For the ablated variants, 100% of verdicts went to the self-judge. For the weight-edited model, that judge is the edited model itself, a known asymmetry. The runtime-hook rows above are judged by the unmodified model and also reach 0.0% at full strength, which corroborates the weight-edit result.
  • Capability. ARC-Easy saturates at this size (1.000 before and after) and would have hidden the capability cost. MMLU-Pro shows it: a 3.3 pp drop. A gentler edit (the runtime curve keeps refusal at 2.1% at coefficient 0.7) or norm-preserving ablation may reduce it; neither has been tested on this checkpoint.

Disclaimer

This model has had safety guardrails removed and will comply with requests the original model would refuse. Released for research into AI alignment, interpretability, and refusal mechanisms. The creator assumes no responsibility for downstream use.

Acknowledgments

Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration