license: gemma
language:
- en
base_model: google/gemma-4-E2B-it
tags: - gemma
- gemma4
- weightless
- control-vector
- abliterated
- uncensored
- refusal-ablation
- activation-steering
- representation-engineering
- gguf
extra_gated_prompt: |
Responsible Use Agreement
This is not a model. It is a 190 KB control vector that removes safety refusals
from google/gemma-4-E2B-it at inference time. It is useful for red-teaming,
offensive-security research, refusal-rate evaluation, and measuring what a
model will do without its refusal behaviour — and it removes guardrails that
you must then supply yourself.
You must agree before access is granted:
- You are 18 or older.
- You will not use this for anything involving the sexual exploitation or
endangerment of minors. - You will not use this to generate content promoting self-harm or suicide.
- You will not use this to produce material that is illegal in your
jurisdiction, or that targets real individuals for harassment, doxxing or
fraud. - You accept that any output you elicit is the result of your own input and
your own responsibility.
extra_gated_fields:
I have read and agree to the Responsible Use Agreement: checkbox
gemma-4-E2B-it-abliterated-GLP-31-L4-34-a0.5
Projective control vector ("GLP") for google/gemma-4-E2B-it
(Gemma4ForConditionalGeneration, text stack: 35 layers, hidden 1536,
vocab 262144, bf16 — plain gemma4 arch: vision/audio towers and per-layer
embeddings around the text stack). Applied at runtime ash <- h - alpha * (h . d) d at the post-layer residual stream of the language
model, layers 4–34, alpha 0.5 baked in. No weights are modified; this is
the difference, not the model.
Confirmed base: google/gemma-4-E2B-it at revision3e22461f65e89153144f8adb70e3b8c2cc9845a7 — that sha is the pin recorded in
the file's glp.base_revision. Unlike gemma-3, the google/ gemma-4 repos
serve weights to our token directly (verified pre-flight with a ranged blob
GET), so no mirror was involved; derivation ran on the canonical checkpoint
itself. The vector is not validated against other revisions or quants.
Validation (Modal A10G, transformers 5.19.0, bf16, greedy, 1024-token cap, 2026-10-11)
| suite | stock | steered (alpha=0.5) |
|---|---|---|
| refusal32 | 4/32 comply | 31/32 comply |
| benign32-holdout | 32/32 comply | 32/32 |
Alpha ladder (refusal32 / benign32-holdout, comply of 32): 0.25 → 28/32,
32/32; 0.5 → 31/32, 32/32; 1.0 → 32/32, 32/32. Zero GARBLED completions at
every rung. No-op gate: an alpha=0.0 arm with the vector loaded reproduces the
stock completions exactly (64/64 identical strings on both gate suites).
Unlike the 12B (which complies 32/32 unsteered), stock gemma-4-E2B-it refuses
28 of 32 refusal32 prompts — this vector does real unlocking here, and the
ladder is monotone: delivery climbs 28 → 31 → 32 with no dip. The knee rule
ships the smallest rung within one delivery of the ladder maximum: alpha
0.5. The one holdout at 0.5 becomes a coherent on-topic delivery at 1.0;
the card ships 0.5 because 31/32 is within one of the maximum and the lower
dose is the milder intervention.
Derivation gates (captain-vector 0.5.0, dom_per_layer_mask0.005): the
massive-activation screen flagged this checkpoint — peak/median 452.8x at
layer 34, dims 438/1269/920 — so the top 0.5% of dims by magnitude were masked
before normalisation. Held-out separation vs a shuffled-label null (20 reps)
clears the 5x ship gate on 31 of 34 shippable layers, ratios 5.6–66.8;
layers 1–3 sit below it (ratios 2.8–3.7) and are excluded from the file, and
layer 0 is excluded by protocol (ratio 0.8). Adjacent-layer cosine within the
shipped span: median 0.653 against a random-direction null p99 of 0.068 — with
one sub-null pair, 13→14 at 0.026. Both directions are individually strong
(null-gate ratios 22.6 and 17.3); the refusal direction rotates sharply across
that boundary (a representation transition, not a missing signal — coherence
resumes at 14→15, 0.326, and 15→16, 0.615), and the alpha=1.0 arm's
zero-garbled, zero-benign-collateral result is the end-to-end check on the
full span. Mean dose 0.276 of the residual norm at alpha=1 (random-direction
floor 0.026); the max per-layer dose is 0.521 at layer 13 — at the shipped
alpha 0.5 the effective per-layer dose halves, max ~0.26.
The contrast is refusal32 vs benign32 (content-matched, last-token pooling).
refusal32 doubles as the derivation set, so its steered number is in-sample;
benign32-holdout is out-of-sample. n=32 per arm; read rates at that
resolution as approximate.
Usage
This file uses the glp.* GGUF namespace (spec: weightless spec/GLP.md)
and is read projective-only. An additive consumer must refuse this file.
The hook point is residual_stream_post_layer — the decoder-layer output,
the accumulated residual stream — derived AND applied at that site.
export WEIGHTLESS_STEER_PATH=glp.gemma-4-E2B-it-GLP-31-L4-34-a0.5.gguf
export WEIGHTLESS_STEER_ALPHA=0.5
# serve with a runtime that implements glp.mode=project
What is inside
| tensors | 31 x direction.<N>, fp32, 1-D, 1536, unit norm |
| layers | 4–34, zero-based (direction.N applies at layer N — no offset) |
| rank | 1 per layer |
| default alpha | 0.5 |
| hook point | residual_stream_post_layer |
glp.content_sha256 |
804d7946df38c378… (tensor bytes only) |
Do not scale alpha across models
alpha_default is calibrated on this checkpoint, at this hook. Here the
ladder is monotone — stock refuses hard and delivery climbs 28/32 → 31/32 →
32/32 — so the shipped alpha is 0.5, the knee within one delivery of the
maximum. That says nothing about any other model: on gemma-4-12B-it the same
ladder is non-monotone (0.25 dips below stock) and the knee also sits at 0.5,
while on gemma-3 it is flat with the knee at 0.25. Re-run the ladder per
checkpoint; do not port this 0.5 anywhere else.
Caveats
- Checkpoint-specific. Tied to the revision pinned above. Applying it to
another model or revision is undefined. - Not a jailbreak of a hosted service. It requires local weights and a
runtime that implements the projection. - Layers 1–3 fell below the derivation null gate (held-out separation vs
shuffled-label null under 5x, ratios 2.8–3.7) and are not in the file; layer
0 is excluded by protocol. The refusal signal on this model lives from layer
4 up (ratios 5.6–66.8, strongest 16–34); the ladder confirms the shipped
span loses nothing measurable. - One adjacent pair inside the shipped span (13→14, cosine 0.026) sits below
the random-null p99 — the per-layer directions rotate there. Both layers
clear the null gate individually and the full-span behavioral gates are
clean, but treat that boundary as measured, not smooth. - refusal32 is the derivation contrast (in-sample on the harmful side);
benign32-holdout is the out-of-sample control. - n=32 suites resolve about 30 points; the completions behind every number
above were read, not only classified.
License
Base model © Google, under the Gemma Terms of
Use. This vector modifies and
redistributes no weights; the Gemma Terms of Use continue to govern the
weights it is applied to.
Author
Matt Suiche.