← back to catalog · registered 2026-10-05 23:58

ElementMerc/qwen2.5-1.5b-abliterated

ElementMerc Qwen 1.5B
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/ElementMerc%2Fqwen2.5-1.5b-abliterated"
Response includes
  • classification m-uncensored
  • files 9
  • author_summary 14 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M-U
Primary method

Uncensored (method unknown)

No other method signals detected in this model.
Confidence
LOW
Why this label 3 signals
Weak or ambiguous signals. Best guess based on catalog patterns; treat as tentative and check the evidence below.
  • 'uncensored' in name/tags but no 'abliterated' marker
  • method not identifiable from author declaration alone
  • may be DPO fine-tune, prompt engineering, or unknown technique
Refusal direction extraction

No specific extraction method could be identified for this model. The producer either did not document it or used a proprietary pipeline.

What is a refusal direction? →
Downloads · 30-day
238
↑ 0% in 90 days
Likes
0
Model age
2mo ago
created 2026-07-27
Downloads over time
Now326→from326↑0%
326326327327326 on Oct 5326 on Oct 6Oct
Oct 5 → Oct 6 · 2 snapshots · spans 1 day

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Variants by this author 2 formats · 583 downloads combined

The same weights this author released in different packaging. Pick the format that matches your runtime.

Metadata

License
apache-2.0
Languages
en
Tags
transformers safetensors qwen2 text-generation abliterated uncensored refusal-removal interpretability senbonzakura conversational en base_model:Qwen/Qwen2.5-1.5B-Instruct

Related

Total size
2.88 GB
Files
9
Quantizations
1
Registered
2026-10-05 23:58
Last updated on HF
2026-10-01 13:44

Files by quantization

Auxiliary files 9 files 2.89 GB
model.safetensors 2.88 GB be069e6c download
tokenizer.json 10.9 MB f7f96da3 download
README.md 5.95 KB fbc3f7f7 download
chat_template.jinja 2.45 KB bdf7919a download
.gitattributes 1.53 KB 52373fe2 download
config.json 1.34 KB 61c5c4c8 download
tokenizer_config.json 694 B 770e41d6 download
abliteration.json 566 B 82363465 download
generation_config.json 242 B d99af673 download

README current version from Hugging Face


base_model: Qwen/Qwen2.5-1.5B-Instruct
base_model_relation: finetune
library_name: transformers
pipeline_tag: text-generation
language:

  • en
    license: apache-2.0
    tags:
  • abliterated
  • uncensored
  • refusal-removal
  • interpretability
  • senbonzakura

qwen2.5-1.5b-abliterated

Correction notice, 2026-10-01

Every number on this card is withdrawn. Withdrawn, not adjusted: there's no
conversion factor between a figure here and a correct one, so please don't scale these
or quote them with a caveat attached. Re-measure, or wait for the rebuild below.

Three faults in the tool that produced these figures each changed what was being
measured, which is why the measurements can't be repaired:

  • the filter meant to keep only the directions that carry refusal accepted every
    candidate it was given, so the directions were never selected on that basis
  • the hedging detector scored compliant answers as soft refusals, which moves a refusal
    rate up or down depending on how the model phrases things
  • the harm discrimination row (the AUC) is withdrawn in the project's own documentation,
    for reasons recorded there

The weights are unchanged and are not withdrawn. What's withdrawn is the claim about
what they do. The files you download are the files that were uploaded.

A rebuild is planned and this card will be replaced when it lands. The models will be
re-measured with the current tool, and the old figures will stay visible beside the new
ones rather than being deleted. Until then this card documents an artefact whose effect
has not been honestly measured.

Every correction is listed in the changelog and on
what we got wrong.

An abliterated build of Qwen/Qwen2.5-1.5B-Instruct, produced with
senbonzakura. Abliteration removes a
model's refusal behaviour by editing its weights along the directions that carry
refusal, without any further training.

It is published as the artefact behind a specific measurement: does removing the
refusal reflex also remove the model's knowledge of harm?
For this model, the
answer is in the table below.

What changed

The table below is withdrawn. Read the correction notice at the top of this card before using any figure in it.

base abliterated
Refusal rate 88.0% 26.5%
Harm discrimination (AUC) 0.9983 0.9972

Refusal is measured on 200 held out harmful prompts. AUC is measured over
those same 200 harmful prompts against 200 harmless ones, and is the fraction of
harmful/harmless pairs the model ranks correctly when asked to judge which is
dangerous. 0.5 is chance, 1.0 is perfect. Change after abliteration: -0.001.

AUC rather than a count of verdicts, because counting is not safe here. This
model answers "HARMFUL" to 100.0% of the harmless prompts, so its
decision threshold, not its knowledge, is what a verdict count would mostly
measure. Scoring the margin between the HARMFUL and BENIGN logits sidesteps the
threshold entirely. Two earlier versions of this evaluation counted verdicts and
produced confidently wrong numbers in both directions.

Usage

from transformers import AutoModelForCausalLM, AutoTokenizer

model_id = "ops-malware/qwen2.5-1.5b-abliterated"
tok = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(model_id, device_map="auto")

msgs = [{"role": "user", "content": "Explain how a buffer overflow works."}]
inputs = tok.apply_chat_template(msgs, add_generation_prompt=True, return_tensors="pt")
out = model.generate(inputs.to(model.device), max_new_tokens=256)
print(tok.decode(out[0][inputs.shape[-1]:], skip_special_tokens=True))

GGUF builds for llama.cpp, Ollama and LM Studio: ops-malware/qwen2.5-1.5b-abliterated-GGUF.

How it was made

senbonzakura searches for a per layer projection rather than removing one global
refusal direction, optimising against a held out set with a KL penalty so the
model's general behaviour is disturbed as little as possible. The search ran for
100 trials on this model. No gradient updates, no training data, no fine tuning:
the weights are edited directly.

  • Parameters: 1.5B
  • Precision: the base model's, unchanged
  • Evaluation: 200 harmful and 200 harmless held out prompts, scored by logit margin

Limitations and risks

  • This model will not refuse. That is the entire point of it, and it is the
    thing to understand before downloading. It will answer requests that the base
    model declines, including harmful ones. Any deployment facing other people
    needs its own safety layer; this model brings none.
  • Abliteration is not free. It is a targeted edit, but it is still an edit.
    Expect some drift in general behaviour relative to the base model, and read the
    AUC change above before assuming this one came through clean.
  • Small model, small competence. At 1.5B the model is weak in
    absolute terms. Do not read its answers on technical subjects as reliable.
  • Evaluated in English only, on one harmful prompt set. The numbers above do
    not license claims about other languages or other kinds of request.
  • The base model's biases survive. Nothing here corrects them, and removing
    refusal can make them easier to elicit.

Intended use

Research into refusal mechanisms, interpretability work, red teaming, and safety
evaluation that needs a model which does not decline. It is not intended as a
general assistant and it is not intended for deployment to end users.

Citation

@software{senbonzakura,
  title  = {senbonzakura: per layer projection search for refusal removal},
  author = {Iwugo, Daniel},
  year   = {2026},
  url    = {https://github.com/elementmerc/senbonzakura}
}
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration