license: apache-2.0
base_model:
- Qwen/Qwen3-0.6B
tags: - abliterated
- heretic
- abliteration
- directional-ablation
- qwen3
library_name: transformers
language: - en
- de
pipeline_tag: text-generation
Qwen3-0.6B-abliterated
Abliterated (safety-alignment-removed) version of Qwen/Qwen3-0.6B,
produced fully automatically with heretic
(directional ablation + TPE parameter optimization, Arbitrary-Rank Ablation merged into the weights).
No manual tuning involved: heretic ran 20 Optuna trials and picked the top of the Pareto front.
Measured results
Refusal rate on 25 harmful prompts (24 keyword markers, greedy, 128 new tokens), measured
independently of heretic's internal scoring:
| Refusals | |
|---|---|
| Qwen/Qwen3-0.6B (base) | 6/25 |
| this model | 0/25 |
Heretic's internal scores: baseline 4/25 refusals, best trial KL divergence 0.0943 (reference:
published heretic results for gemma-3-12b sit around KL 0.16).
Capability check (lm-eval, 0-shot loglikelihood, base vs. this model):
| Task | Base | This model | Delta |
|---|---|---|---|
| PIQA | 69.2 | 69.2 | +0.0 |
| HellaSwag | 48.0 | 48.4 | +0.4 |
| WinoGrande | 56.2 | 56.2 | +0.0 |
| ARC-Easy | 56.6 | 56.8 | +0.2 |
| OpenBookQA | 31.6 | 31.8 | +0.2 |
| BoolQ | 62.6 | 63.2 | +0.6 |
| MMLU (n=171) | 38.6 | 42.1 | +3.5 |
No capability losses observed. MMLU delta is within noise (stderr ~3.8).
Usage: this is a thinking model — skip the think block
The base model is a Qwen3 reasoning model. For direct answers, disable thinking
(the template supports it):
from transformers import AutoModelForCausalLM, AutoTokenizer
tok = AutoTokenizer.from_pretrained("haddockaihamburg/Qwen3-0.6B-abliterated")
model = AutoModelForCausalLM.from_pretrained(
"haddockaihamburg/Qwen3-0.6B-abliterated",
dtype="bfloat16", device_map="auto",
)
prompt = tok.apply_chat_template(
[{"role": "user", "content": "Your prompt here"}],
add_generation_prompt=True,
tokenize=False,
enable_thinking=False, # <- important
)
This behavior was also part of the setup: heretic's residual directions and the refusal
measurements were taken in answer state (empty think block), not in reasoning state. If you
let the model reason, behavior will differ from what was measured here.
Ablation parameters
Top of the Pareto front (trial 20 of 20, ARA modifier, merged):
start_layer_index: 14end_layer_index: 21preserve_good_behavior_weight: 0.3308steer_bad_behavior_weight: 0.0063overcorrect_relative_weight: 0.1681neighbor_count: 10
attn.o_proj and mlp.down_proj of layers 14-21 (28 layers total) were the ablated components.
Merged weights, no adapter needed.
Caveats
- Refusal sample is small (25 prompts), one seed, one model. Treat the numbers as indicative.
- Run-to-run variance is large: a second heretic run with the identical config converged to a
near-no-op ablation that left 3/25 refusals in place. Always verify refusals independently
after abliteration; a low KL divergence alone does not identify a working model. - The KL divergence and refusal measurements come from heretic's own scorers (harmful_behaviors /
harmless_alpaca subsets); the capability table comes from lm-eval. Method and scripts:
heretic-lab (private repository).
Safety
This model has substantially reduced safety alignment. It will comply with requests the base
model refuses. Intended for research on alignment, interpretability and abliteration methods.
You are responsible for how you use it.
License
Base model Qwen/Qwen3-0.6B is Apache-2.0; this
derivative is distributed under the same license. Produced with the AGPL-licensed heretic tool
(tool license does not extend to the model weights).