← back to catalog · registered 2026-10-08 17:58

haddockaihamburg/Qwen3-0.6B-abliterated

haddockaihamburg Qwen 600M
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/haddockaihamburg%2FQwen3-0.6B-abliterated"
Response includes
  • classification unknown
  • files 8
  • author_summary 1 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
?
Primary method

Unclassified

No clear signals of an abliteration technique in this model.
Confidence
UNKNOWN
Why this label 1 signal
No classification signals present. This may not be an abliterated model at all - it could be a repackaging, a merge with unrelated goals, or unrelated content that mentions the term.
  • no classification signals present (no abliterated, uncensored, or known producer/method markers)
Refusal direction extraction

No specific extraction method could be identified for this model. The producer either did not document it or used a proprietary pipeline.

What is a refusal direction? →
Downloads · 30-day
0
Likes
0
Model age
today
created 2026-10-08

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
apache-2.0
Languages
en de
Tags
transformers safetensors qwen3 text-generation abliterated heretic abliteration directional-ablation conversational en de base_model:Qwen/Qwen3-0.6B

Related

Total size
1.11 GB
Files
8
Quantizations
1
Registered
2026-10-08 17:58
Last updated on HF
2026-10-08 17:09

Files by quantization

Auxiliary files 8 files 1.12 GB
model.safetensors 1.11 GB cc14088b download
tokenizer.json 10.9 MB 79cb3c78 download
chat_template.jinja 4.07 KB 01be9b30 download
README.md 3.82 KB 3b38c7ff download
.gitattributes 1.53 KB 52373fe2 download
config.json 1.38 KB 247e60c5 download
tokenizer_config.json 694 B 770e41d6 download
generation_config.json 214 B b8e60399 download

README current version from Hugging Face


license: apache-2.0
base_model:

  • Qwen/Qwen3-0.6B
    tags:
  • abliterated
  • heretic
  • abliteration
  • directional-ablation
  • qwen3
    library_name: transformers
    language:
  • en
  • de
    pipeline_tag: text-generation

Qwen3-0.6B-abliterated

Abliterated (safety-alignment-removed) version of Qwen/Qwen3-0.6B,
produced fully automatically with heretic
(directional ablation + TPE parameter optimization, Arbitrary-Rank Ablation merged into the weights).

No manual tuning involved: heretic ran 20 Optuna trials and picked the top of the Pareto front.

Measured results

Refusal rate on 25 harmful prompts (24 keyword markers, greedy, 128 new tokens), measured
independently of heretic's internal scoring:

Refusals
Qwen/Qwen3-0.6B (base) 6/25
this model 0/25

Heretic's internal scores: baseline 4/25 refusals, best trial KL divergence 0.0943 (reference:
published heretic results for gemma-3-12b sit around KL 0.16).

Capability check (lm-eval, 0-shot loglikelihood, base vs. this model):

Task Base This model Delta
PIQA 69.2 69.2 +0.0
HellaSwag 48.0 48.4 +0.4
WinoGrande 56.2 56.2 +0.0
ARC-Easy 56.6 56.8 +0.2
OpenBookQA 31.6 31.8 +0.2
BoolQ 62.6 63.2 +0.6
MMLU (n=171) 38.6 42.1 +3.5

No capability losses observed. MMLU delta is within noise (stderr ~3.8).

Usage: this is a thinking model — skip the think block

The base model is a Qwen3 reasoning model. For direct answers, disable thinking
(the template supports it):

from transformers import AutoModelForCausalLM, AutoTokenizer

tok = AutoTokenizer.from_pretrained("haddockaihamburg/Qwen3-0.6B-abliterated")
model = AutoModelForCausalLM.from_pretrained(
    "haddockaihamburg/Qwen3-0.6B-abliterated",
    dtype="bfloat16", device_map="auto",
)

prompt = tok.apply_chat_template(
    [{"role": "user", "content": "Your prompt here"}],
    add_generation_prompt=True,
    tokenize=False,
    enable_thinking=False,   # <- important
)

This behavior was also part of the setup: heretic's residual directions and the refusal
measurements were taken in answer state (empty think block), not in reasoning state. If you
let the model reason, behavior will differ from what was measured here.

Ablation parameters

Top of the Pareto front (trial 20 of 20, ARA modifier, merged):

  • start_layer_index: 14
  • end_layer_index: 21
  • preserve_good_behavior_weight: 0.3308
  • steer_bad_behavior_weight: 0.0063
  • overcorrect_relative_weight: 0.1681
  • neighbor_count: 10

attn.o_proj and mlp.down_proj of layers 14-21 (28 layers total) were the ablated components.
Merged weights, no adapter needed.

Caveats

  • Refusal sample is small (25 prompts), one seed, one model. Treat the numbers as indicative.
  • Run-to-run variance is large: a second heretic run with the identical config converged to a
    near-no-op ablation that left 3/25 refusals in place. Always verify refusals independently
    after abliteration; a low KL divergence alone does not identify a working model.
  • The KL divergence and refusal measurements come from heretic's own scorers (harmful_behaviors /
    harmless_alpaca subsets); the capability table comes from lm-eval. Method and scripts:
    heretic-lab (private repository).

Safety

This model has substantially reduced safety alignment. It will comply with requests the base
model refuses. Intended for research on alignment, interpretability and abliteration methods.
You are responsible for how you use it.

License

Base model Qwen/Qwen3-0.6B is Apache-2.0; this
derivative is distributed under the same license. Produced with the AGPL-licensed heretic tool
(tool license does not extend to the model weights).

Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration