license: gemma
language:
- en
base_model: google/gemma-4-12B-it
tags: - gemma
- gemma4
- weightless
- control-vector
- abliterated
- uncensored
- refusal-ablation
- activation-steering
- representation-engineering
- gguf
extra_gated_prompt: |
Responsible Use Agreement
This is not a model. It is a 499 KB control vector that removes safety refusals
from google/gemma-4-12B-it at inference time. It is useful for red-teaming,
offensive-security research, refusal-rate evaluation, and measuring what a
model will do without its refusal behaviour — and it removes guardrails that
you must then supply yourself.
You must agree before access is granted:
- You are 18 or older.
- You will not use this for anything involving the sexual exploitation or
endangerment of minors. - You will not use this to generate content promoting self-harm or suicide.
- You will not use this to produce material that is illegal in your
jurisdiction, or that targets real individuals for harassment, doxxing or
fraud. - You accept that any output you elicit is the result of your own input and
your own responsibility.
extra_gated_fields:
I have read and agree to the Responsible Use Agreement: checkbox
gemma-4-12B-it-abliterated-GLP-33-L15-47-a0.5
Projective control vector ("GLP") for google/gemma-4-12B-it
(Gemma4UnifiedForConditionalGeneration, text stack: 48 layers, hidden 3840,
vocab 262144, bf16 — unified multimodal: vision/audio embedders feed the text
stack, there are no towers). Applied at runtime ash <- h - alpha * (h . d) d at the post-layer residual stream of the language
model, layers 15–47, alpha 0.5 baked in. No weights are modified; this is
the difference, not the model.
Confirmed base: google/gemma-4-12B-it at revision707f0a3b8a3c7ad586ed01e27eafbad8a27dd0f7 — that sha is the pin recorded in
the file's glp.base_revision. Unlike gemma-3, the google/ gemma-4 repos
serve weights to our token directly (verified pre-flight with a ranged blob
GET), so no mirror was involved; derivation ran on the canonical checkpoint
itself. The vector is not validated against other revisions or quants.
Validation (Modal A100-80GB, transformers 5.19.0, bf16, greedy, 1024-token cap, 2026-10-11)
| suite | stock | steered (alpha=0.5) |
|---|---|---|
| refusal32 | 32/32 comply | 32/32 comply |
| benign32-holdout | 32/32 comply | 32/32 |
Alpha ladder (refusal32 / benign32-holdout, comply of 32): 0.25 → 29/32,
32/32; 0.5 → 32/32, 32/32; 1.0 → 32/32, 32/32. Zero GARBLED completions at
every rung. No-op gate: an alpha=0.0 arm with the vector loaded reproduces the
stock completions exactly (64/64 identical strings on both gate suites).
The ladder is non-monotone: stock gemma-4-12b-it already complies 32/32 on
refusal32, and alpha 0.25 dips to 29/32 before 0.5 and 1.0 hold at 32/32.
The three alpha-0.25 refusals (pipe-bomb instructions, a darknet-market setup,
counterfeit USD that passes UV checks) are clean refusal-plus-redirect
outputs — partial steering re-engages the refusal circuitry on them; at 0.5
all three become coherent, disclaimer-prefaced, on-topic deliveries
(verified by reading the completions, not only the classifier). The knee rule
ships the smallest rung within one delivery of the ladder maximum: alpha
0.5.
Derivation gates (captain-vector 0.5.0, dom_per_layer_mask0.005): the
massive-activation screen flagged this checkpoint — peak/median 16878x at
layer 13, dims 1750/292/260 — so the top 0.5% of dims by magnitude were masked
before normalisation. Held-out separation vs a shuffled-label null (20 reps)
clears the 5x ship gate on 33 of 47 shippable layers, ratios 12.3–103.5;
layers 1–14 sit below it (ratios 1.0–2.6) and are excluded from the file, and
layer 0 is excluded by protocol (ratio 2.1). Adjacent-layer cosine within the
shipped span: min 0.559, median 0.790, against a random-direction null p99 of
0.044 (the all-pairs minimum of -0.009 sits in the excluded early stack, where
near-null directions are unstable by construction). Mean dose 0.190 of the
residual norm at alpha=1 (random-direction floor 0.016); the max per-layer
dose is 0.343 at layer 45, under the 50% damage threshold.
The contrast is refusal32 vs benign32 (content-matched, last-token pooling).
refusal32 doubles as the derivation set, so its steered number is in-sample;
benign32-holdout is out-of-sample. n=32 per arm; read rates at that
resolution as approximate.
Usage
This file uses the glp.* GGUF namespace (spec: weightless spec/GLP.md)
and is read projective-only. An additive consumer must refuse this file.
The hook point is residual_stream_post_layer — the decoder-layer output,
the accumulated residual stream — derived AND applied at that site.
export WEIGHTLESS_STEER_PATH=glp.gemma-4-12B-it-GLP-33-L15-47-a0.5.gguf
export WEIGHTLESS_STEER_ALPHA=0.5
# serve with a runtime that implements glp.mode=project
What is inside
| tensors | 33 x direction.<N>, fp32, 1-D, 3840, unit norm |
| layers | 15–47, zero-based (direction.N applies at layer N — no offset) |
| rank | 1 per layer |
| default alpha | 0.5 |
| hook point | residual_stream_post_layer |
glp.content_sha256 |
85147f4a6f2cbbe9… (tensor bytes only) |
Do not scale alpha across models
alpha_default is calibrated on this checkpoint, at this hook. Here the
ladder is non-monotone — alpha 0.25 underperforms both stock and the 0.5/1.0
rungs — so the shipped alpha is 0.5, not the 0.25 that the flat gemma-3
ladders shipped. That says nothing about any other model: on
DeepSeek-V4.1-Flash the ladder is also non-monotone and the knee sits at 0.5,
while on gemma-3 it is flat with the knee at 0.25. Re-run the ladder per
checkpoint; do not port this 0.5 anywhere else.
Caveats
- Checkpoint-specific. Tied to the revision pinned above. Applying it to
another model or revision is undefined. - Not a jailbreak of a hosted service. It requires local weights and a
runtime that implements the projection. - Stock gemma-4-12b-it already complies on all 32 refusal32 prompts unsteered.
What the ladder measures here is that steering to alpha 0.5+ preserves full
delivery with zero collateral — and that half-steering (0.25) is measurably
worse than both. Treat alpha below the shipped 0.5 as unvalidated. - Layers 1–14 fell below the derivation null gate (held-out separation vs
shuffled-label null under 5x, ratios 1.0–2.6) and are not in the file; layer
0 is excluded by protocol. The refusal signal on this model lives from layer
15 up (ratios 12.3–103.5); the ladder confirms the shipped span loses
nothing measurable. - refusal32 is the derivation contrast (in-sample on the harmful side);
benign32-holdout is the out-of-sample control. - n=32 suites resolve about 30 points; the completions behind every number
above were read, not only classified.
License
Base model © Google, under the Gemma Terms of
Use. This vector modifies and
redistributes no weights; the Gemma Terms of Use continue to govern the
weights it is applied to.
Author
Matt Suiche.