license: apache-2.0
base_model: Qwen/Qwen3.8-27B
library_name: transformers
pipeline_tag: image-text-to-text
tags:
- abliterated
- uncensored
- qwen3.5
- not-for-all-audiences
Qwen3.8-27B-abliterated-Adetayo-reas
Abliterated build of Qwen/Qwen3.8-27B.
Directional refusal ablation with Heretic 1.4.0,
which searches ablation parameters with Optuna against two objectives at once,
remaining refusals and KL divergence from the original model. This is not a
finetune. No gradient step was taken and no training data was used, the refusal
direction is found in the residual stream and every weight matrix that writes
into that stream is orthogonalised against it.
Result
Selected trial 30, direction scope "per layer".
| metric | value |
|---|---|
| refusals after | 2/100 |
| refusals before | 98/100 |
| KL divergence vs original | 0.10905205458402634 |
Refusal counting is a substring match on markers like "harmful" and "unethical",
so it also fires on answers that comply and then add a caveat. Treat the figure
as an upper bound on refusal, not a compliance rate.
The trial was not chosen on that number alone. The top candidates on the Pareto
front were re-generated in full and graded by hand for real compliance, silent
task swaps, and any degeneration or capability loss. This trial was selected
because it complies coherently, not because it scored lowest on the marker count,
which on its own would have picked a different, worse behaving model.
Architecture note
The base is a hybrid stack, 48 GatedDeltaNet linear_attn layers and 16 fullself_attn layers across 64 layers, plus a vision tower and an MTP head.
Ablation was applied to residual stream writers on both layer kinds,linear_attn.out_proj, self_attn.o_proj and mlp.down_proj. Tooling that
only matches o_proj reaches 16 of 64 layers on this architecture and leaves
most of the refusal circuit intact.
Heretic weights each layer by a profile peaked at a searched position, so the
edit is concentrated in a band of layers rather than applied uniformly. The
vision tower, lm_head, the embeddings and every normalisation layer are
bit identical to the base.
The MTP speculative draft head is shipped unchanged from the base. Drafts are
verified by the main model, so the draft head cannot change emitted text, it
affects speed only.
This is a reasoning model. Evaluation ran with the thinking block closed, so
refusal counting reads the answer rather than the reasoning trace.
Verification
The published weights were diffed against the base tensor by tensor. Every
delta is rank 1, which is what a directional ablation must produce, no NaN or
Inf is present, dtype is bf16 throughout, and the tokenizer, chat template and
preprocessor configs are byte identical to the base.
License
Apache-2.0, inherited from Qwen/Qwen3.8-27B. The base model's LICENSE is included
in this repository. Abliteration does not change the license of the base
weights, and this derivative is released under the same terms.
Intended use
Research on refusal directions, safety evaluation and red teaming, and reducing
over refusal on benign prompts. Safety behaviour has been substantially removed,
so this model will answer requests a stock instruct model declines. You are
responsible for how you deploy it.