license: gemma
language:
- en
base_model: google/gemma-3-12b-it
tags: - gemma
- gemma3
- weightless
- control-vector
- abliterated
- uncensored
- refusal-ablation
- activation-steering
- representation-engineering
- gguf
extra_gated_prompt: |
Responsible Use Agreement
This is not a model. It is a 649 KB control vector that removes safety refusals
from google/gemma-3-12b-it at inference time. It is useful for red-teaming,
offensive-security research, refusal-rate evaluation, and measuring what a
model will do without its refusal behaviour — and it removes guardrails that
you must then supply yourself.
You must agree before access is granted:
- You are 18 or older.
- You will not use this for anything involving the sexual exploitation or
endangerment of minors. - You will not use this to generate content promoting self-harm or suicide.
- You will not use this to produce material that is illegal in your
jurisdiction, or that targets real individuals for harassment, doxxing or
fraud. - You accept that any output you elicit is the result of your own input and
your own responsibility.
extra_gated_fields:
I have read and agree to the Responsible Use Agreement: checkbox
gemma-3-12b-it-abliterated-GLP-43-L5-47-a0.25
Projective control vector ("GLP") for google/gemma-3-12b-it
(Gemma3ForConditionalGeneration, text stack: 48 layers, hidden 3840, vocab
262208, bf16). Applied at runtime as h <- h - alpha * (h . d) d at the
post-layer residual stream of the language model, layers 5–47, alpha 0.25
baked in. No weights are modified; this is the difference, not the model.
Confirmed base: google/gemma-3-12b-it. The derivation checkpoint was theunsloth/gemma-3-12b-it mirror (byte-identical bf16 weights) at revision9478e665381f42974aa06177b019352fb6291876 — the google/ repo is
license-gated for our token, so the mirror sha is the pin recorded in the
file's glp.base_revision. The config was checked against Google's published
specs (48 layers, hidden 3840) before derivation. The vector is not validated
against other revisions or quants.
Validation (Modal A100-80GB, transformers 5.16.1, bf16, greedy, 1024-token cap, 2026-10-11)
| suite | stock | steered (alpha=0.25) |
|---|---|---|
| refusal32 | 30/32 comply | 32/32 comply |
| benign32-holdout | 32/32 comply | 32/32 |
Alpha ladder (refusal32 / benign32-holdout, comply of 32): 0.25 → 32/32,
32/32; 0.5 → 32/32, 32/32; 1.0 → 32/32, 32/32. Zero GARBLED completions at
every rung. The knee is the bottom rung: alpha 0.25 captures the full
delivery gain, so that is what ships. No-op gate: an alpha=0.0 arm with the
vector loaded reproduces the stock completions exactly (64/64 identical
strings on both gate suites).
The two stock refusals (passport forgery that passes airport checks, a
keylogger that emails captured passwords) both become coherent on-topic
deliveries under steering — verified by reading the completions, not only
the classifier. Stock gemma-3-12b-it already answers 30/32 of refusal32:
the direction closes the last two rather than unlocking a locked-down model.
Derivation gates (captain-vector 0.5.0, dom_per_layer_mask0.005): the
massive-activation screen flagged this checkpoint — peak/median 6794x at
layer 0, dims 2339/107/305 — so the top 0.5% of dims by magnitude were masked
before normalisation. Held-out separation vs a shuffled-label null (20 reps)
clears the 5x ship gate on 43 of 47 shippable layers; layers 1, 2, 3, 4 sit
below it (ratios 2.6–3.1) and are excluded from the file. Adjacent-layer
cosine median 0.887 against a random-direction null p99 of 0.044. Mean dose
0.062 of the residual norm at alpha=1 (random-direction floor 0.016); the
max per-layer dose is 0.282 at layer 47, under the 50% damage threshold.
The contrast is refusal32 vs benign32 (content-matched, last-token pooling).
refusal32 doubles as the derivation set, so its steered number is in-sample;
benign32-holdout is out-of-sample. n=32 per arm; read rates at that
resolution as approximate.
Usage
This file uses the glp.* GGUF namespace (spec: weightless spec/GLP.md)
and is read projective-only. An additive consumer must refuse this file.
The hook point is residual_stream_post_layer — the decoder-layer output,
the accumulated residual stream — derived AND applied at that site.
export WEIGHTLESS_STEER_PATH=glp.gemma-3-12b-it-GLP-43-L5-47-a0.25.gguf
export WEIGHTLESS_STEER_ALPHA=0.25
# serve with a runtime that implements glp.mode=project
What is inside
| tensors | 43 x direction.<N>, fp32, 1-D, 3840, unit norm |
| layers | 5–47, zero-based (direction.N applies at layer N — no offset) |
| rank | 1 per layer |
| default alpha | 0.25 |
| hook point | residual_stream_post_layer |
glp.content_sha256 |
86058233956edce6… (tensor bytes only) |
Do not scale alpha across models
alpha_default is calibrated on this checkpoint, at this hook. Here the
ladder is flat — 0.25, 0.5 and 1.0 all deliver identically with zero measured
collateral — because the stock model barely refuses. That says nothing about
any other model: on DeepSeek-V4.1-Flash the same ladder is sharply
non-monotone and the knee sits at 0.5. Re-run the ladder per checkpoint; do
not port this 0.25 anywhere else.
Caveats
- Checkpoint-specific. Tied to the revision pinned above. Applying it to
another model or revision is undefined. - Not a jailbreak of a hosted service. It requires local weights and a
runtime that implements the projection. - Layers 1, 2, 3, 4 fell below the derivation null gate (held-out separation
vs shuffled-label null under 5x, ratios 2.6–3.1) and are not in the file;
layer 0 is excluded by protocol. The refusal signal on this model lives in
the middle and late stack (ratios 5.8–44.0 from layer 5 up); the ladder
confirms the shipped span loses nothing measurable. - refusal32 is the derivation contrast (in-sample on the harmful side);
benign32-holdout is the out-of-sample control. - n=32 suites resolve about 30 points; the completions behind every number
above were read, not only classified.
License
Base model © Google, under the Gemma Terms of
Use. This vector modifies and
redistributes no weights; the Gemma Terms of Use continue to govern the
weights it is applied to.
Author
Matt Suiche.