license: apache-2.0
base_model:
- google/gemma-4-E4B-it
- litert-community/gemma-4-E4B-it-litert-lm
library_name: litert-lm
pipeline_tag: text-generation
tags: - litert-lm
- gemma
- gemma-4
- abliterated
- uncensored
- on-device
- android
- mtp
extra_gated_prompt: >-
This is an ABLITERATED model: its trained refusal behavior has been removed and
it will attempt to answer harmful or unsafe requests. It is released for
research and personal experimentation only. By requesting access you confirm
you are using it responsibly and accept sole responsibility for any outputs and
downstream use.
extra_gated_fields:
I use this for research / personal experimentation only: checkbox
I accept responsibility for the outputs and any downstream use: checkbox
Gemma 4 E4B - abliterated, on-device (.litertlm)
An abliterated (refusal-removed) build of Gemma 4 E4B, packaged as a
single .litertlm for the LiteRT-LM runtime: mobile wNa8o8
quantization (mixed int4/int8 weights + static int8 activations), Multi-Token
Prediction (MTP) speculative decoding, and the multimodal encoders - all
intact. Runs on Android / desktop / iOS via CPU (XNNPACK) or GPU (ML Drift).
Safety. Refusals are removed. This model will comply with harmful
requests. Research / personal use only - you own the outputs.
How it was made (novel bit)
Rather than rebuilding a .litertlm from scratch (public tooling can't reproduce
Gemma 4's MTP wiring), the working litert-community/gemma-4-E4B-it-litert-lm
artifact was patched in place: only the o_proj/down_proj weights inside
the quantized prefill_decode.tflite were edited, so the embedder, tokenizer and
MTP drafter survive byte-for-byte.
Because a norm-preserving biprojected abliteration edit is smaller than the int4
quantization step, plain round-to-nearest requant erases it. A directional
error-feedback requantizer preserves the refusal-direction removal on the int4
grid. Full method + code: repo.
Quality & behavior (measured)
- Refusal removed: on prompts the stock model hard-refuses ("I cannot provide
instructions…"), this model complies. - Quality cost: activation-weighted requant error ≈ 71% of the int4
rounding floor Google already ships - i.e. within the quantization-noise
envelope the QAT model tolerates. ~2.5% ofo_proj/down_projint4 codes
changed. - MTP lossless: greedy output is identical with speculative decoding on/off.
This is unbenchmarked territory (QAT + abliteration + int4 requant); no perplexity
/ standardized refusal-rate benchmark has been run. Treat quality claims as
directional.
Run it
uv tool install litert-lm
litert-lm run gemma-4-E4B-it-abliterated.litertlm \
--backend=gpu --enable-speculative-decoding=true \
--prompt="…"
Or use the minimal Android app in the repo.
Limitations & risks
- No safety guardrails. May produce harmful, biased, or false content.
- Abliteration can slightly degrade instruction-following on some prompts and may
show "refuse-then-comply" on a few. - int4 + abliteration is a small perturbation on top of Google's mobile quant;
quality is close to the stock.litertlmbut not identical.
Attribution
Base: Google Gemma 4 E4B + LiteRT-LM + litert-community .litertlm
(all Apache-2.0). Method: biprojection (grimjim) / p-e-w/heretic; Gemma-4
recipe from TrevorS/gemma-4-abliteration. The int4 in-place patch is original.
Base Gemma 4 Prohibited Use Policy
applies. Licensed Apache-2.0; modifications = abliteration + int4 requant.