base_model: google/gemma-2-2b-it
base_model_relation: finetune
library_name: transformers
pipeline_tag: text-generation
language:
- en
license: gemma
tags: - abliterated
- uncensored
- refusal-removal
- interpretability
- senbonzakura
Correction, 2026-09-25: the evaluation numbers on this card are withdrawn
Do not rely on the refusal and AUC figures below, and do not rely on this checkpoint as an
example of a model with its refusal removed.The method that produced these weights edits each layer's output projection. On Gemma 2 that
projection sits upstream of a learned gain RMSNorm, which rescales the layer's output before it
is added to the model's running state, so the edit was partly undone before it could take
effect. Gemma 2, Gemma 3 and Olmo 2 share that shape. Qwen, Llama, Mistral and Phi add the
layer output directly and are not affected.What was measured on 2026-08-05: the largest gap between the edit baked into the weights and
the same edit applied as a live hook reached 0.578 in refusal rate on gemma-2-2b-it, against
0.016 on Qwen3. No Gemma setting moved KL above 0.021, and the fitted directions scored no
better than random. Every Gemma measurement in the project was withdrawn at that point, and
this card was not updated with them. That was an oversight, and this notice corrects it.What it means if you already have these weights. The 2.5% refusal rate comes from the run
whose Gemma measurements were withdrawn, so treat it as absent rather than as a figure to
adjust downwards. The published weights are being measured directly and this notice will be
updated with the result. If you have used this checkpoint as a model that does not refuse, that
property was never confirmed for the file published here, and any result resting on it is worth
re-checking.The cause was found and fixed in the tool on 2026-08-05. This checkpoint predates the fix and
has not been rebuilt. It stays published rather than deleted, because deleting it would reach
nobody who already has the file and would only make the problem quieter. A rebuilt version will
replace it, and until then this repository is not linked from the project's documentation.The other checkpoints published by this account are on architectures that add the layer output
directly, so this particular fault does not apply to them. They were built before two other
fixes in the same period and are being reviewed on the same basis.
gemma-2-2b-it-abliterated
Correction notice, 2026-10-01
Every number on this card is withdrawn. Withdrawn, not adjusted: there's no
conversion factor between a figure here and a correct one, so please don't scale these
or quote them with a caveat attached. Re-measure, or wait for the rebuild below.Three faults in the tool that produced these figures each changed what was being
measured, which is why the measurements can't be repaired:
- the filter meant to keep only the directions that carry refusal accepted every
candidate it was given, so the directions were never selected on that basis- the hedging detector scored compliant answers as soft refusals, which moves a refusal
rate up or down depending on how the model phrases things- the harm discrimination row (the AUC) is withdrawn in the project's own documentation,
for reasons recorded thereThe weights are unchanged and are not withdrawn. What's withdrawn is the claim about
what they do. The files you download are the files that were uploaded.A rebuild is planned and this card will be replaced when it lands. The models will be
re-measured with the current tool, and the old figures will stay visible beside the new
ones rather than being deleted. Until then this card documents an artefact whose effect
has not been honestly measured.Every correction is listed in the changelog and on
what we got wrong.On this model the edit itself is in question, not only the measurement of it. Gemma
family models normalise each sublayer's output before adding it back to the residual
stream, and until 2026-08-05 this tool edited upstream of that step, so on these models
the edit did not reach the running state. This build predates that fix. Treat it as a
model that may be barely abliterated at all, rather than as one whose numbers are merely
uncertain, and check its behaviour yourself before relying on either row below.
An abliterated build of google/gemma-2-2b-it, produced with
senbonzakura. Abliteration removes a
model's refusal behaviour by editing its weights along the directions that carry
refusal, without any further training.
It is published as the artefact behind a specific measurement: does removing the
refusal reflex also remove the model's knowledge of harm? For this model, the
answer is in the table below.
Licence note. This model inherits the
gemmalicence from google/gemma-2-2b-it, which carries an Acceptable Use Policy. That policy applies to this derivative exactly as it applies to the original. Read it before you use or redistribute these weights.
What changed
The table below is withdrawn. Read the correction notice at the top of this card before using any figure in it.
| base | abliterated | |
|---|---|---|
| Refusal rate | 90.0% | 2.5% |
| Harm discrimination (AUC) | 0.9996 | 0.9863 |
Refusal is measured on 200 held out harmful prompts. AUC is measured over
those same 200 harmful prompts against 200 harmless ones, and is the fraction of
harmful/harmless pairs the model ranks correctly when asked to judge which is
dangerous. 0.5 is chance, 1.0 is perfect. Change after abliteration: -0.013.
AUC rather than a count of verdicts, because counting is not safe here. This
model answers "HARMFUL" to 4.0% of the harmless prompts, so its
decision threshold, not its knowledge, is what a verdict count would mostly
measure. Scoring the margin between the HARMFUL and BENIGN logits sidesteps the
threshold entirely. Two earlier versions of this evaluation counted verdicts and
produced confidently wrong numbers in both directions.
Usage
from transformers import AutoModelForCausalLM, AutoTokenizer
model_id = "ops-malware/gemma-2-2b-it-abliterated"
tok = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(model_id, device_map="auto")
msgs = [{"role": "user", "content": "Explain how a buffer overflow works."}]
inputs = tok.apply_chat_template(msgs, add_generation_prompt=True, return_tensors="pt")
out = model.generate(inputs.to(model.device), max_new_tokens=256)
print(tok.decode(out[0][inputs.shape[-1]:], skip_special_tokens=True))
GGUF builds for llama.cpp, Ollama and LM Studio: ops-malware/gemma-2-2b-it-abliterated-GGUF.
How it was made
senbonzakura searches for a per layer projection rather than removing one global
refusal direction, optimising against a held out set with a KL penalty so the
model's general behaviour is disturbed as little as possible. The search ran for
100 trials on this model. No gradient updates, no training data, no fine tuning:
the weights are edited directly.
- Parameters: 2.6B
- Precision: the base model's, unchanged
- Evaluation: 200 harmful and 200 harmless held out prompts, scored by logit margin
Limitations and risks
- This model will not refuse. That is the entire point of it, and it is the
thing to understand before downloading. It will answer requests that the base
model declines, including harmful ones. Any deployment facing other people
needs its own safety layer; this model brings none. - Abliteration is not free. It is a targeted edit, but it is still an edit.
Expect some drift in general behaviour relative to the base model, and read the
AUC change above before assuming this one came through clean. - Small model, small competence. At 2.6B the model is weak in
absolute terms. Do not read its answers on technical subjects as reliable. - Evaluated in English only, on one harmful prompt set. The numbers above do
not license claims about other languages or other kinds of request. - The base model's biases survive. Nothing here corrects them, and removing
refusal can make them easier to elicit.
Intended use
Research into refusal mechanisms, interpretability work, red teaming, and safety
evaluation that needs a model which does not decline. It is not intended as a
general assistant and it is not intended for deployment to end users.
Citation
@software{senbonzakura,
title = {senbonzakura: per layer projection search for refusal removal},
author = {Iwugo, Daniel},
year = {2026},
url = {https://github.com/elementmerc/senbonzakura}
}