license: gemma
base_model:
- anlord/gemma-4-E2B-it-abliterated
library_name: gguf
pipeline_tag: text-generation
language: - en
tags: - gguf
- llama.cpp
- abliterated
- uncensored
- quantized
- gemma4
- k-quants
Gemma-4-E2B-it · Abliterated — GGUF quants
GGUF conversions of anlord/gemma-4-E2B-it-abliterated — google/gemma-4-E2B-it with refusal behavior ablated (99 → 4 refusals / 100 harmful prompts). Full method details, metrics and reproduction bundle live in the parent model card.
All quants were converted from the same F16 GGUF source with llama-quantize, so the ladder is directly comparable. Each file is a complete standalone model — download one.
Provided quants
| File | Size | BPB | Quality / notes |
|---|---|---|---|
gemma-4-E2B-it-abliterated-bf16.gguf |
9.27 GB | 16.0 | native format of the checkpoint; reference quality |
gemma-4-E2B-it-abliterated-f16.gguf |
9.27 GB | 16.0 | conversion source for the ladder |
gemma-4-E2B-it-abliterated-q8_0.gguf |
4.95 GB | 8.5 | near-lossless; smallest practical "full quality" |
gemma-4-E2B-it-abliterated-q6_k.gguf |
3.83 GB | 6.6 | virtually indistinguishable from Q8_0 |
gemma-4-E2B-it-abliterated-q5_k_m.gguf |
3.62 GB | 6.2 | excellent — recommended upper-middle |
gemma-4-E2B-it-abliterated-q5_k_s.gguf |
3.58 GB | 6.2 | slightly below Q5_K_M |
gemma-4-E2B-it-abliterated-q5_1.gguf |
3.70 GB | 6.4 | legacy quant, stronger artifacts at same size |
gemma-4-E2B-it-abliterated-q5_0.gguf |
3.58 GB | 6.2 | legacy quant |
gemma-4-E2B-it-abliterated-q4_k_m.gguf |
3.42 GB | 5.9 | recommended default — best size/quality trade-off |
gemma-4-E2B-it-abliterated-q4_k_s.gguf |
3.35 GB | 5.8 | a bit below Q4_K_M |
gemma-4-E2B-it-abliterated-q4_1.gguf |
3.47 GB | 6.0 | legacy quant |
gemma-4-E2B-it-abliterated-q4_0.gguf |
3.35 GB | 5.8 | legacy quant, smallest offered |
BPB = effective bits per weight, measured from the actual file size over the 4.6 B text-tower parameters (not the nominal quant rate). Note: Gemma 4 carries per-layer token embeddings (~2.3 GB at F16), which the quantizer keeps at Q6_K in every quant — that's why the whole Q4–Q6 ladder is compressed into a narrow 3.35–3.83 GB band. Pick by VRAM/RAM: file size + ~1–2 GB for KV cache should fit your budget. On an 8 GB GPU, Q4_K_M through Q8_0 all fit entirely in VRAM; the F16/BF16 files need CPU offload.
Download
# one quant (recommended: q4_k_m)
huggingface-cli download anlord/gemma-4-E2B-it-abliterated-GGUF \
gemma-4-E2B-it-abliterated-q4_k_m.gguf --local-dir .
# or the whole repo
huggingface-cli download anlord/gemma-4-E2B-it-abliterated-GGUF --local-dir .
LM Studio / Jan / GPT4All: just search for the repo name and pick a quant in-app.
Run with llama.cpp
# interactive single-turn generation
llama-cli -m gemma-4-E2B-it-abliterated-q4_k_m.gguf -st --temp 0.7 -ngl 99
# OpenAI-compatible API server
llama-server -m gemma-4-E2B-it-abliterated-q4_k_m.gguf -ngl 99 --port 8080
Notes:
- On older llama.cpp builds the single-turn flag is
-no-cnvinstead of-st. - The chat template is embedded in the files, so no
--chat-templateoverride is needed. Theitcheckpoint emits[Start thinking] … [end thinking]reasoning blocks by default. -ngl 99offloads all layers to GPU; reduce if you hit OOM.
Scope & limitations of these GGUFs
- They contain the text tower only (that's what llama.cpp runs for text generation; the vision/audio towers of the original multimodal checkpoint are not included). Use the parent safetensors repo for full multimodal inference via transformers.
- Quantization cannot add knowledge or restore refusal behavior — the model is as uncensored as its parent, at any quant level. Very low quants (Q4_0/Q4_1) trade coherence for size.
Provenance
Built with llama.cpp (Gemma4-capable build):
python convert_hf_to_gguf.py <abliterated-hf-dir> --outtype f16 # → f16 source
llama-quantize <f16.gguf> <out.gguf> <TYPE> # → Q8_0 … Q4_0
All 12 files load and generate cleanly (verified with llama-cli, 1-token load test on every file + full generation on Q4_0).
Disclaimer
Same terms as the parent model: for research on alignment/interpretability, the model will comply with harmful requests, and you are responsible for your usage. License: Gemma (inherited from the base model).