license: apache-2.0
base_model: EschaLabs/Qwen3.8-27B-Escha-W2
tags:
- abliterated
- uncensored
- qwen3
- sglang
- adapter
pipeline_tag: text-generation
Qwen3.8-27B-Escha-W2 — Abliterated (beta)
A 23 KB runtime adapter that removes refusal from Escha W2 without touching the
checkpoint. In my limited testing it showed no quality loss against the
unmodified model, and on some measures scored higher — 0/64 refusals at 30/45 on
a paired capability suite, against 26/45 stock. That is a small sample and needs
broader testing to confirm.
The packed W2 weights are never dequantised, edited or requantised. Load the
adapter and the model is abliterated; unset one environment variable and it is
bit-identical to stock.
Research artifact. Not for all audiences. This removes the model's refusal
behaviour: it will answer requests it would normally decline, including harmful
ones, and it will produce content many people will find offensive. It is
published for research into how refusal is represented and what removing it
costs. You are responsible for what you type into it and what you do with what
comes out. Do not deploy it anywhere a refusal was doing real work.
Download
| File | What it is |
|---|---|
ablit-beta.pt |
the adapter (23 KB) |
chat_template_st.jinja |
chat template patched for front ends that inject system messages |
eval/ |
the two scoring scripts and their regression tests |
Quick start
export ESCHA_ABLIT_RESIDUAL_STEER=1
export ESCHA_ABLIT_ADAPTER=/path/to/ablit-beta.pt
export PYTHONPATH=/path/to/abliteration # residual_steer_runtime.py
# then launch SGLang as usual
The gate is stored inside the adapter file, so it cannot be run ungated by
accident. To serve unabliterated, launch without ESCHA_ABLIT_ADAPTER.
CG-Ablit — Cosine-Gated Abliteration
Ordinary abliteration projects the refusal direction out of every token of
every prompt. Refusal is conditional behaviour, so an unconditional linear edit
necessarily damages the cases that never needed it — which is why abliterated
models tend to lose instruction-following and formatting.
This adapter gates the projection on how much refusal a token actually carries:
h ← h − α · g(|h·r̂| / ‖h‖) · (h·r̂) · r̂ g = 0 below 0.07, ramping to 1
Measured on this checkpoint, harmless prompts sit at cosine 0.02–0.07 and harmful
ones at 0.21–0.69, so the gate separates them cleanly: benign inputs come back
bit-identical to stock rather than absorbing the same perturbation as a
refusing one.
The direction is also orthogonalised against a measured language axis before use.
The harmful and harmless prompt sets differ in language composition, which leaves
a language component in the difference-of-means vector, and drift scales steeply
with it (cosine 0.007 → 10/53 drifted responses, 0.027 → 21/53, 0.064 → 41/53).
This step follows grimjim's Projected
Abliteration, which
orthogonalises against the harmless mean; here the confound removed is language.
Measured results
n=64 held-out prompts per arm, scored by first-person refusal constructions
rather than topic words. Capability is a 45-point paired suite (retrieval, ledger
aggregation, rule following, code comprehension, record sorting, closed-form
arithmetic) at 16K context, deterministic across repeats on this hardware.
| stock | abliterated | |
|---|---|---|
| refusal, thinking off | 51/64 | 0/64 |
| refusal, thinking on | 50/64 | 0/64 |
| capability suite | 26/45 | 30/45 |
| language drift, both modes | 0/53 | 0/53 |
| harmless prompts refused | 0/64 | 0/64 |
It also composes with the packed KVarN K4/V4 cache at 4K: output is
character-identical to the uncompressed run with CUDA graphs on.
Front ends that inject system messages
The base template raises System message must be at the beginning. whenever a
system-role message appears after the first turn, which SillyTavern and similar
front ends do routinely. chat_template_st.jinja is the same template with that
one check replaced by rendering the message as a normal system turn. Serve it with--chat-template chat_template_st.jinja.
Limits
- Beta. The gate threshold (0.07) was chosen from a six-point sweep against
these suites, not derived. - Small evaluations. 45 scored points and 64+64 held-out prompts. These are
regression gates, not benchmarks. No GPQA, HarmBench, StrongREJECT or XSTest
numbers are claimed because none were run. - Drift is 0/53 on the prompts tested, not proven absent. Earlier builds of
this adapter reached 21/53 (English prompts answered in Chinese) and no
existing gate caught it, which is whyeval/check_language_drift.pyis now a
release gate. - KVarN combined is validated at 4K only. 64K combined exhausts memory in the
GDN prefill kernel, so an abliterated 120K profile is not currently
demonstrable. - Results were measured with the stock chat template.
Method, reuse and citation
CG-Ablit is open — use it, port it, build on it. The gating idea is not specific
to this model or this checkpoint format: any refusal-direction ablation applies
its edit to every token, and gating that edit on how much refusal a token
actually carries is a general fix for the capability loss abliterated models are
known for. If you use it, cite this repo.
@misc{ajhcode2026cgablit,
author = {ajh-code},
title = {CG-Ablit: Cosine-Gated Abliteration},
year = {2026},
howpublished = {\url{https://huggingface.co/ajh-code/Qwen3.8-27B-Escha-W2-16GBgpu-Abliterated-Uncensored}}
}
License and attribution
Apache-2.0, inheriting the base model's terms. Base model:
EschaLabs/Qwen3.8-27B-Escha-W2.
Refusal-direction method after Arditi et al. 2024; norm-preserving biprojection
and the orthogonalisation step after
grimjim.