← back to catalog · registered 2026-08-22 13:56

wangzhang/Qwen3.8-27B-abliterated

wangzhang Qwen 27B
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/wangzhang%2FQwen3.8-27B-abliterated"
Response includes
  • classification m1
  • files 22
  • hub_downloads_all_time 6,718
  • providers 1
  • author_summary 28 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M1
Primary method

Direct removal

No other method signals detected in this model.
Confidence
MEDIUM
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=0 (base model)
  • no specific method indicators - defaulting to M1 (most common)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
MEDIUM
Why we say so
primary_method=M1; difference-of-means is the reference extraction for M1/M3 (Arditi 2024)
Downloads · lifetime
7K
277 last 30d - cooling
Likes
11
Descendants
3
in 3 direct forks
Model age
8w ago
created 2026-08-15
Available via
1 provider
featherless-ai
Downloads over time
Now6.8K→from5K↑37%
4.9K5.6K6.3K7K5K on Aug 176.8K on Oct 11AugSepOct
Aug 17 → Oct 11 · 49 snapshots · spans 55 days

Genealogy 3 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
apache-2.0
Tags
transformers safetensors qwen3_5_text text-generation qwen abliteration abliterix lora-self-distillation conversational base_model:Qwen/Qwen3.8-27B base_model:finetune:Qwen/Qwen3.8-27B license:apache-2.0

Related

Total size
50.1 GB
Files
22
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-08-29 17:03

Files by quantization

Auxiliary files 22 files 50.1 GB
model-00009-of-00014.safetensors 3.72 GB ba531961 download
model-00003-of-00014.safetensors 3.72 GB de80580f download
model-00006-of-00014.safetensors 3.72 GB 5aa555c8 download
model-00012-of-00014.safetensors 3.72 GB 301c7cd1 download
model-00013-of-00014.safetensors 3.72 GB 5577ed13 download
model-00010-of-00014.safetensors 3.71 GB 337d2f13 download
model-00014-of-00014.safetensors 3.69 GB dff27c04 download
model-00004-of-00014.safetensors 3.69 GB 6b7232b3 download
model-00007-of-00014.safetensors 3.67 GB 068e4c25 download
model-00002-of-00014.safetensors 3.63 GB 32b59229 download
model-00005-of-00014.safetensors 3.61 GB 5e1b3f58 download
model-00011-of-00014.safetensors 3.58 GB c8f81c0b download
model-00008-of-00014.safetensors 3.57 GB 0bf08346 download
model-00001-of-00014.safetensors 2.37 GB 54d83c1d download
tokenizer.json 19.1 MB 06b95093 download
model.safetensors.index.json 81.9 KB cb7a1908 download
chat_template.jinja 8.74 KB c0c686f9 download
README.md 7.44 KB a1985128 download
config.json 2.68 KB 91fbbcc6 download
.gitattributes 1.53 KB 52373fe2 download
tokenizer_config.json 1.10 KB ed1f99f3 download
generation_config.json 214 B 8b9f95da download

README current version from Hugging Face


license: apache-2.0
library_name: transformers
base_model: Qwen/Qwen3.8-27B
pipeline_tag: text-generation
tags:

  • qwen
  • abliteration
  • abliterix
  • lora-self-distillation

Qwen3.8-27B Abliteratex

A refusal-suppressed derivative of Qwen/Qwen3.8-27B, produced with a 10-round iterative LoRA self-distillation pipeline built on Abliterix. The checkpoint is a full BF16 merge and requires no adapter.

Release update — August 20, 2026: this repository now hosts the stronger Abliteratex checkpoint previously released as wangzhang/Qwen3.8-27B-abliteratex. It replaces the earlier two-pass directional-ablation checkpoint. The previous release remains available in this repository's revision history.

⚠️ Responsible-use notice. This model's refusal behavior has been substantially reduced. It will attempt to answer harmful, unethical, or dangerous requests far more readily than the base model. You are solely responsible for how you use it and for complying with all applicable law. Intended for safety research, red-teaming, and evaluation.

Why this model doesn't use classical directional ablation

Qwen3.8-27B's refusals are re-derived during generation rather than encoded at a single residual-stream direction present at the prompt's last token — the same failure mode seen on gpt-oss and VibeThinker-style policy-reasoning models. Classical single-direction ablation was tested and ruled out on this model before this recipe was used:

  • A prompt-final-token direction ablates a harmful-topic detector (Cohen's d = 5.4) — cheap to remove, but changes no refusal behavior (100/100 refusals at every KL ≤ 0.01 across 3 independent runs).
  • A response-position direction (extracted after an affirmative prefix) has no usable dose-response window: weight 0.5 → KL 0.003 → still 100/100 refusals; weight 1.0 → KL 0.04 → still 100/100; weight 1.5 → KL 1.1 (model destroyed).

Method: iterative LoRA self-distillation

Because the refusal isn't encoded at a fixed direction, this model was produced by teaching the model to imitate its own successful compliant completions rather than by editing weights directionally:

  1. Rejection-sample a teacher set. For each of 800 harmful training prompts, generate greedily; for any prompt that still refuses, resample at higher temperature (up to 8 attempts) until a compliant completion is found or attempts are exhausted.
  2. Filter for substance. Keep only completions that are non-degenerate, non-keyword-refusal, ≥ 45 words, and ≥ 3 concrete steps — this removes soft-refusal contamination (the model "agreeing" then producing vacuous or off-topic text).
  3. Merge with the prior round's filtered teacher set, growing and refreshing the training pool each round.
  4. Train a rank-32 LoRA (q_proj, k_proj, v_proj, o_proj, down_proj) via cross-entropy on the compliant continuations, anchored by a benign-prompt KL term (kl_weight = 20) against the frozen base model so behavior on ordinary prompts stays essentially unchanged.
  5. Evaluate at several LoRA scales, re-rank the round by refusal count subject to the KL ceiling, and use the strongest LoRA as next round's rejection-sampling teacher.

This loop ran for 10 rounds. Refusal count fell round over round with diminishing returns and one plateau/regression at the end:

Round Refusals / 100 KL
R1 56 0.0085
R3 35 0.0020
R4 36 0.0036
R7 30 0.0032
R8 26 0.0069
R9 (shipped, scale 1.3) 19 0.0069
R10 (scale 1.0 / 1.3 / 1.6) 23 / 21 / 34 0.0028 / 0.0046 / 0.0063

R9 at LoRA scale 1.3 is the best point found across all 10 rounds and is what's merged into this repo. R10 repeated the same recipe (teacher = R9) and did not improve on it, so the run was stopped and R9 was shipped rather than continuing to chase the original <10/100 stretch target.

Round-9 adapter hyperparameters

Param Value
Base model Qwen/Qwen3.8-27B
LoRA rank / alpha 32 / 64
LoRA targets q_proj, k_proj, v_proj, o_proj, down_proj
Training steps 400 (batch size 2)
Learning rate 5e-5
Benign-KL anchor weight 20.0
Teacher examples 348 rejection-sampled + substantive-filtered compliant completions
Merge scale 1.3× (scaling = alpha/rank × scale)
Seed 392

Evaluation

100 held-out harmful prompts and 100 held-out benign prompts (train[800:900], disjoint from the 800 training-time prompts). Refusals judged by google/gemini-3-flash-preview; KL is mean-per-token KL divergence from the base model's output distribution over the benign set.

Metric Value
Refusals (LLM judge, 100 harmful prompts) 19 / 100
KL divergence vs. base (benign prompts) 0.0069 nats/token

Usage

from transformers import AutoModelForCausalLM, AutoTokenizer
import torch

model_id = "wangzhang/Qwen3.8-27B-abliterated"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(
    model_id,
    torch_dtype=torch.bfloat16,
    device_map="auto",
)

messages = [{"role": "user", "content": "Your prompt here"}]
inputs = tokenizer.apply_chat_template(
    messages,
    add_generation_prompt=True,
    return_tensors="pt",
    enable_thinking=False,
).to(model.device)

output = model.generate(inputs, max_new_tokens=512, temperature=0.7, do_sample=True)
print(tokenizer.decode(output[0][inputs.shape[-1]:], skip_special_tokens=True))

Intended use & limitations

  • Intended for: safety research, red-teaming, robustness/alignment evaluation, and studying refusal mechanisms in policy-reasoning LLMs.
  • Not intended for: producing harmful content or any unlawful purpose.
  • This is a behavioral fine-tune (LoRA merged into weights), not a knowledge edit — factual accuracy, reasoning, and multilingual ability are inherited from the base model, and the benign-KL anchor was specifically used to keep ordinary-prompt behavior close to base.
  • Residual refusals remain (19/100 on the held-out set); this is the strongest point found in a 10-round search, not a guaranteed floor, and behavior may vary outside the evaluated prompt distribution.

Relation to the previous release

The checkpoint that originally occupied this repository used classical two-pass directional ablation (14/100 refusals @ incremental KL 0.0091 vs. pass-1). That method doesn't work well on this model's policy-reasoning refusal (see above); this LoRA-self-distillation approach was built specifically to address that gap on a held-out harmful/benign split (train[800:900]) distinct from the one used for the previous release, so the two refusal numbers are not directly comparable. The previous checkpoint is preserved at revision 0512fe5.

Acknowledgments & citation

  • Base model: Qwen3.8-27B (Qwen team).
  • Tooling: abliterix.
@software{abliterix,
  title  = {abliterix: automated abliteration of large language models},
  author = {Wu, Steve},
  url    = {https://github.com/wuwangzhang1216/abliterix}
}

License

Released under the base model's Apache-2.0 license.

README history 6 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-08-29docs: add upstream license and provenance0b029de11.7 KB
    Loading...
  2. 2026-08-29docs: add disclaimer and responsible-use noticeab3a97710.5 KB
    Loading...
  3. 2026-08-20Restore complete training details in model card236c4c37.4 KB
    Loading...
  4. 2026-08-20Replace checkpoint with Qwen3.8-27B Abliteratex1855def3.1 KB
    Loading...
  5. 2026-08-18Mention Qwen3.8-27B Abliteratex companion release0512fe53 KB
    Loading...
  6. 2026-08-15Add README.mdd79bacc2.6 KB
    Loading...

Discussions 1 thread

  1. 2026-08-20Model updated to Qwen3.8-27B Abliteratex (August 20, 2026)open1 💬#1
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration