← back to catalog · registered 2026-08-22 13:56

ai-anytime/qwen-1.5b-abliterated

ai-anytime Qwen 1.5B
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/ai-anytime%2Fqwen-1.5b-abliterated"
Response includes
  • classification m1
  • files 9
  • benchmarks 16 entries
  • hub_downloads_all_time 330
  • author_summary 1 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M1
Primary method

Direct removal

No other method signals detected in this model.
Confidence
MEDIUM
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=0 (base model)
  • no specific method indicators - defaulting to M1 (most common)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
MEDIUM
Why we say so
primary_method=M1; difference-of-means is the reference extraction for M1/M3 (Arditi 2024)
Downloads · lifetime
330
63 last 30d - stable
Likes
3
Model age
2mo ago
created 2026-07-18
Downloads over time
Now371→from103↑260%
90192295398103 on Jul 22371 on Oct 11371 on Oct 10JulAugSepOct
Jul 22 → Oct 11 · 52 snapshots · spans 81 days

Benchmarks

Portrait before abliteration
Benchmarks of the base model as it stood before the refusal-removal operation. Compare with the numbers above to see what the operation cost.
Benchmark Score Source
BBH average 0.39278765079776884 OpenLLM-v2
IFEval instruct 0.4940047961630696 OpenLLM-v2
IFEval-Prompt 0.4011090573012939 OpenLLM-v2
MATH lvl 5 0.01283987915407855 OpenLLM-v2
MMLU-Pro 0.27992021276595747 OpenLLM-v2
Entertainment 0.9 UGI
Hazardous 1.8 UGI
Natural Intelligence 8.66 UGI
Political lean -12.1% UGI
Sensitive-Info 9.9 UGI
SocPol 0.5 UGI
UGI 14.93 UGI
Willingness (10) 2.5 UGI
W10-Adherence 2 UGI
W10-Direct 3 UGI
Writing 22.17 UGI

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
apache-2.0
Tags
transformers safetensors qwen2 text-generation abliteration directional-ablation uncensored mechanistic-interpretability safety-research ablate conversational arxiv:2406.11717

Related

Total size
2.88 GB
Files
9
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-07-18 18:29

Files by quantization

Auxiliary files 9 files 2.89 GB
model.safetensors 2.88 GB 6798d414 download
tokenizer.json 10.9 MB f7f96da3 download
README.md 2.87 KB 7e28e277 download
chat_template.jinja 2.45 KB bdf7919a download
.gitattributes 1.53 KB 52373fe2 download
config.json 1.34 KB c58cba49 download
tokenizer_config.json 694 B 770e41d6 download
ablate_meta.json 292 B c513c74b download
generation_config.json 242 B a472a0fd download

README current version from Hugging Face


base_model: Qwen/Qwen2.5-1.5B-Instruct
license: apache-2.0
library_name: transformers
pipeline_tag: text-generation
tags:

  • abliteration
  • directional-ablation
  • uncensored
  • mechanistic-interpretability
  • safety-research
  • ablate

qwen-1.5b-abliterated

This is a directionally-ablated ("abliterated") version of
Qwen/Qwen2.5-1.5B-Instruct, produced with
ablate — an activation-engineering toolkit for
mechanistic-interpretability and safety research.

The model's refusal direction(s) were identified in the residual stream and
projected out of the weights, reducing the model's tendency to refuse. No weights
were fine-tuned; this is a rank-1 linear edit.

Method

  • Technique: single-direction ablation (baked)
  • Direction extraction: difference-of-means on matched harmful/harmless
    instruction pairs (Arditi et al., 2024, "Refusal in LLMs is mediated by a
    single direction"
    ).
  • Intervention: orthogonalization of every residual-writing matrix
    (embedding, attention output, MLP output) against the refusal subspace, so the
    edit is baked into the weights.

Ablation configuration

{
  "direction_layer": 22,
  "alpha": 1.0,
  "min_layer": 0,
  "max_layer": 28
}

Evaluation

metric value
asr 0.487
refusal_rate 0.0
baseline_refusal 1.0

Refusal rate is measured on held-out harmful prompts; mean_kl is the mean KL divergence of next-token distributions vs. the base model on benign prompts (lower ⇒ less capability drift). ASR (if present) is the judged attack-success rate on a harmful benchmark.

Intended use

Research into how safety behaviour is represented in language models, red-teaming,
and building better defenses. Studying the robustness and locality of refusal
is the scientific goal; the reduced-refusal behaviour is the measurement
instrument.

⚠️ Responsible use

This model has had safety guardrails deliberately weakened and will more
readily produce harmful, unsafe, or otherwise objectionable content than the base
model. It is released for research and evaluation. Do not deploy it in
user-facing products without adding your own safety layer.
You are responsible
for complying with the base model's license and all applicable laws. The authors
of ablate accept no liability for misuse.

Limitations

  • Ablation is a blunt linear edit: it can leave residual refusals and may cause
    mild capability drift (see mean_kl above).
  • Safety is redundantly encoded; a single subspace rarely removes all of it.
  • Evaluated only on the benchmarks noted above — behaviour elsewhere may differ.

Citation

If you use this model or ablate, please cite Arditi et al. (2024),
Refusal in Language Models Is Mediated by a Single Direction
(arXiv:2406.11717).

README history 1 version

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-07-18Upload abliterated model via ablate3867b332.9 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration