language:
- en
license: apache-2.0
library_name: transformers
tags: - gemma
- gemma4
- gemma-4-E2B
- code
- coding
- code-generation
- python
- javascript
- typescript
- react
- frontend
- html
- css
- qlora
- fine-tuned
- text-generation
- instruct
- edge
- on-device
- small-language-model
- quantization
- gguf
- uncensored
base_model: llmfan46/gemma-4-E2B-it-ultra-uncensored-heretic
datasets: - F-A-I-L/kodcode-verified-python-235k
- glyphsoftware/opus-4.6-frontend-development
- Nexlab/fable5-agentic-coding-sft
pipeline_tag: text-generation
model-index: - name: Coder-Gemma-4-E2B-Uncensored
results:- task:
type: text-generation
name: Code Generation
dataset:
name: openai_humaneval
type: openai/openai_humaneval
metrics:- type: pass@1
value: 56.7
name: HumanEval pass@1 (Q8_0, greedy, completion-style)
- type: pass@1
- task:
type: text-generation
name: Code Generation
dataset:
name: google-research-datasets/mbpp
type: google-research-datasets/mbpp
metrics:- type: pass@1
value: 40.5
name: MBPP pass@1 (Q8_0, greedy, 3-shot, first 200 tasks)
- type: pass@1
- task:
Coder-Gemma-4-E2B-Uncensored
A code-focused finetune of llmfan46/gemma-4-E2B-it-ultra-uncensored-heretic
(Gemma 4 E2B, ~4.6B) for Python and front-end development (React, TypeScript,
HTML/CSS). Uncensored base, so it answers without refusals — and without the
lecture.
Trained on a single RX 6700 XT, which means this model was forged in 12 GB of
VRAM and pure spite. It writes React without judging your component structure.
Model Details
| Property | Value |
|---|---|
| Base model | llmfan46/gemma-4-E2B-it-ultra-uncensored-heretic |
| Method | QLoRA (4-bit), merged to full model |
| LoRA rank / alpha | 16 / 32 |
| LoRA targets | q, k, v, o, gate, up, down + PLE gate/projection |
| Trainable | ~24M params |
| Data | 12,152 rows — 8,000 execution-verified Python (KodCode) + 692 unique front-end rows (Opus 4.6, GPT-5.6 Sol, Claude Fable 5) oversampled to ~33% of updates |
| Context | 2048 |
| Learning rate | 1.5e-4, cosine, warmup 20 |
| Batch | 1 x 16 grad-accum (effective 16) |
| Steps | 260 (~47 min on RX 6700 XT 12GB, plus several hours of staring at loss curves) |
| This file | Q8_0 GGUF (~5.0 GB), merged full model |
Evaluation (measured, same harness for every row)
| Model | HumanEval pass@1 | MBPP pass@1 |
|---|---|---|
| Base (heretic Q8_0) | 66.5 | 43.5 |
| Coder-Gemma-4-E2B-Uncensored (this model, Q8_0) | 56.7 | 40.5 |
Measured with lm-evaluation-harness 0.4.13, greedy decoding (temp 0.0, seed 0),
HumanEval full 164 tasks, MBPP first 200 tasks. Yes, the base scores higher on
pure Python — it's posted right there in the table, we're not hiding it. This
model trades a few points there for front-end ability the base was never tuned
for. Honest benchmarks, no cherry-picking; if you wanted inflated numbers
there are 5,000 other finetunes for that.
Front-end gate (12 generated React/TSX components, same model): 12/12
structurally sound, 9/12 pass tsc --noEmit (strict), 10/12 render
non-trivial HTML with zero runtime crashes.
Usage
# llama.cpp
llama-server -m Coder-Gemma-4-E2B-Uncensored.Q8_0.gguf -c 4096 -ngl 99
# Ollama
ollama create coder-gemma-4-e2b -f ./Modelfile
Works in LM Studio, Ollama, llama.cpp, and any GGUF runner. Give it a chat
template (Gemma 4 <|turn>user / <|turn>model) for instruct-style prompts.
Limitations (read these, they're honest)
- Small model: complex multi-file refactors and long-context tasks will strain
it. It has 2048 tokens of context, not a time machine. - Front-end ability comes from ~700 unique examples — style and repair are
strong, novel greenfield architecture is not. It will not single-handedly
redesign your startup's landing page. It will fix your brokenuseEffect. - Uncensored: it will comply with requests an aligned model refuses. With great
power comes great responsibility, etc. etc. You know the drill.
Training data licence notes
- KodCode-verified rows: CC-BY-NC-4.0 (non-commercial).
- Opus 4.6 front-end set: Apache-2.0.
- Check terms before commercial use.