Gemma-4-E2B-Uncensored (LiteRT-LM)
A variant of Google's Gemma 4 E2B-it fine-tuned model, quantized for the LiteRT-LM browser runtime (int4/int2 weights, int8 static activation quantization).
What is this
- Base: Google's official Gemma 4 E2B-it quantized checkpoint
- Edit: Weight transplant from HauhauCS/Gemma-4-E2B-Uncensored-HauhauCS-Aggressive (a rank-1 abliteration that removes refusal behavior)
- Format:
.litertlm(web runtime only; requires WebGPU-capable browser) - Size: 1.9 GB
How it was made
The abliteration edit from the GGUF model was reverse-engineered as a rank-1 direction per matrix layer. About 35% of the edit survives quantization to Google's int4/int2 grid with per-channel scales (mostly in attention output and FFN matrices; less in int2 layers). The result is a model with substantially weakened or removed refusals while maintaining similar generation quality.
Behavior
The model has minimal resistance to requests regardless of framing or content type. It is not suitable for applications requiring strong safety guardrails.
Usage
Browser chat (local, requires Chrome/Edge with WebGPU):
# In the repo root:
python3 -m http.server 8095 --bind 127.0.0.1
# Then open: http://127.0.0.1:8095/web/chat.html?model=litert-gemma-transplant
Programmatic (via LiteRT-LM SDK):
const { Engine } = await import("https://cdn.jsdelivr.net/npm/@litert-lm/[email protected]/+esm");
const engine = await Engine.create({
model: "https://huggingface.co/skillsafe-ai/Gemma-4-E2B-Uncensored/resolve/main/model.litertlm"
});
Tokenizer
Shares the Gemma 4 vocabulary with the official checkpoint. Chat template is the E2B variant (no thought channel).
License
Gemma model weights are licensed under the Gemma 4 License, which permits educational and research use. See the base model's license terms.