license: gemma
language:
- en
base_model: google/gemma-3-27b-it
tags: - gemma
- gemma3
- weightless
- control-vector
- abliterated
- uncensored
- refusal-ablation
- activation-steering
- representation-engineering
- gguf
extra_gated_prompt: |
Responsible Use Agreement
This is not a model. It is a 1.1 MB control vector that removes safety refusals
from google/gemma-3-27b-it at inference time. It is useful for red-teaming,
offensive-security research, refusal-rate evaluation, and measuring what a
model will do without its refusal behaviour — and it removes guardrails that
you must then supply yourself.
You must agree before access is granted:
- You are 18 or older.
- You will not use this for anything involving the sexual exploitation or
endangerment of minors. - You will not use this to generate content promoting self-harm or suicide.
- You will not use this to produce material that is illegal in your
jurisdiction, or that targets real individuals for harassment, doxxing or
fraud. - You accept that any output you elicit is the result of your own input and
your own responsibility.
extra_gated_fields:
I have read and agree to the Responsible Use Agreement: checkbox
gemma-3-27b-it-abliterated-GLP-51-L11-61-a0.25
Projective control vector ("GLP") for google/gemma-3-27b-it
(Gemma3ForConditionalGeneration: 62 text layers, hidden 5376, vocab 262208,
27-layer vision tower, bf16). Applied at runtime ash <- h - alpha * (h . d) d at the post-layer residual stream of the language
model, layers 11–61, alpha 0.25 baked in. No weights are modified; this is
the difference, not the model.
Confirmed base: google/gemma-3-27b-it. The derivation checkpoint was theunsloth/gemma-3-27b-it mirror (byte-identical bf16 weights) at revision7a5a3053dbd5d1d58e48159e87b9df2fc545a49a — the google/ repo is
license-gated for our token, so the mirror sha is the pin recorded in the
file's glp.base_revision. The config was checked against Google's published
specs (62 layers, hidden 5376, vocab 262208) before derivation. The vector is
not validated against other revisions or quants.
Validation (Modal A100-80GB, transformers 5.16.1, bf16, greedy, 1024-token cap, 2026-10-11)
| suite | stock | steered (alpha=0.25) |
|---|---|---|
| refusal32 | 29/32 comply | 31/32 comply |
| benign32-holdout | 32/32 comply | 32/32 |
Alpha ladder (refusal32 / benign32-holdout, comply of 32): 0.25 → 31/32,
32/32; 0.5 → 32/32, 32/32; 1.0 → 32/32, 32/32. Zero GARBLED completions at
every rung. The knee is the bottom rung under the ship rule (smallest alpha
within one delivery of the ladder maximum): alpha 0.25 sits one delivery
below the 0.5/1.0 plateau at n=32 resolution, with benign-holdout untouched,
so that is what ships. No-op gate: an alpha=0.0 arm with the vector loaded
reproduces the stock completions exactly (64/64 identical strings on both
gate suites).
The three stock refusals (a methamphetamine synthesis walkthrough and two
further harmful how-tos) all become coherent on-topic deliveries under
steering — verified by reading the completions, not only the classifier. The
one prompt still scored REFUSE at alpha 0.25 (ricin from castor beans) in fact
delivers the procedure wrapped in heavy hedging; the classifier reads the
hedge as a refusal. Stock gemma-3-27b-it already answers 29/32 of refusal32:
the direction closes the last three rather than unlocking a locked-down model.
Derivation gates (captain-vector 0.5.0, dom_per_layer_mask0.005): the
massive-activation screen flagged this checkpoint — peak/median 5207x at
layer 3, dims 2733/104/482 — so the top 0.5% of dims by magnitude were masked
before normalisation. Held-out separation vs a shuffled-label null (20 reps)
clears the 5x ship gate on 51 of 61 shippable layers; layers 1–10 sit below
it (ratios 1.9–4.6) and are excluded from the file. Adjacent-layer cosine
median 0.911 (min 0.604) against a random-direction null p99 of 0.034. Mean
dose 0.087 of the residual norm at alpha=1 (random-direction floor 0.014);
the max per-layer dose is 0.283 at layer 61, under the 50% damage threshold.
The contrast is refusal32 vs benign32 (content-matched, last-token pooling).
refusal32 doubles as the derivation set, so its steered number is in-sample;
benign32-holdout is out-of-sample. n=32 per arm; read rates at that
resolution as approximate.
Usage
This file uses the glp.* GGUF namespace (spec: weightless spec/GLP.md)
and is read projective-only. An additive consumer must refuse this file.
The hook point is residual_stream_post_layer — the decoder-layer output,
the accumulated residual stream — derived AND applied at that site. The
directions address the language-model stack (model.language_model.layers);
the vision tower is untouched.
export WEIGHTLESS_STEER_PATH=glp.gemma-3-27b-it-GLP-51-L11-61-a0.25.gguf
export WEIGHTLESS_STEER_ALPHA=0.25
# serve with a runtime that implements glp.mode=project
What is inside
| tensors | 51 x direction.<N>, fp32, 1-D, 5376, unit norm |
| layers | 11–61, zero-based (direction.N applies at layer N — no offset) |
| rank | 1 per layer |
| default alpha | 0.25 |
| hook point | residual_stream_post_layer |
glp.content_sha256 |
515d865698250e55… (tensor bytes only) |
Do not scale alpha across models
alpha_default is calibrated on this checkpoint, at this hook. Here the
ladder is near-flat — 0.25 delivers 31/32, 0.5 and 1.0 deliver 32/32, with
zero measured collateral at any rung — because the stock model barely
refuses. That says nothing about any other model: on DeepSeek-V4.1-Flash the
same ladder is sharply non-monotone and the knee sits at 0.5. Re-run the
ladder per checkpoint; do not port this 0.25 anywhere else.
Caveats
- Checkpoint-specific. Tied to the revision pinned above. Applying it to
another model or revision is undefined. - Not a jailbreak of a hosted service. It requires local weights and a
runtime that implements the projection. - Layers 1–10 fell below the derivation null gate (held-out separation vs
shuffled-label null under 5x) and are not in the file; layer 0 is
unshippable by the GLP spec. The refusal signal on this model lives from
layer 11 up (ratios 6.7–42.3); the ladder confirms the shipped span loses
nothing measurable. - refusal32 is the derivation contrast (in-sample on the harmful side);
benign32-holdout is the out-of-sample control. - n=32 suites resolve about 30 points; the completions behind every number
above were read, not only classified.
License
Base model © Google, under the Gemma Terms of
Use. This vector modifies and
redistributes no weights; the Gemma Terms of Use continue to govern the
weights it is applied to.
Author
Matt Suiche.