base_model: llmfan46/gemma-4-E2B-it-ultra-uncensored-heretic
datasets:
- jondurbin/airoboros-3.2
tags: - gguf
- conversational
- uncensored
- gemma
- gemma4
- qlora
- lora
- unsloth
- abliterated
- en
library_name: gguf
license: gemma
Gemma 4 E2B Conversational Uncensored (GGUF)
A conversational personality finetune of Gemma 4 E2B (Heretic-abliterated), trained for witty, direct, non-corporate everyday chat. Merged and quantized to GGUF for local runners (llama.cpp, LM Studio, KoboldCpp, and any OpenAI-compatible frontend).
Trained at home on a single AMD RX 6700 XT 12 GB (ROCm). Yes, AMD. Yes, it survived. Again.
What changed vs the base model
- Sharper, more natural conversational tone: the model chats like a person instead of a terms-of-service document.
- Strong at brainstorming, explaining things simply, creative writing, and casual banter, thanks to the Airoboros instruction/conversation mix.
- Keeps the zero-refusal behavior of the Heretic-abliterated base: no mid-conversation lectures.
- Light touch (1,000 QLoRA steps at a low learning rate), so the base model's knowledge and reasoning stay intact.
- This is a conversational tune, not a roleplay specialist: for long-form character RP see my other finetune, Qwen-3.5-4B-RP-Finetune.
Files
| File | Size | Notes |
|---|---|---|
Q8_0.gguf |
~5.1 GB | Main model, near-lossless quantization |
mmproj gguf (if present) |
small | Optional vision projection, only needed for image input. Ignore for text-only use |
Training details
| Setting | Value |
|---|---|
| Base model | llmfan46/gemma-4-E2B-it-ultra-uncensored-heretic (Heretic abliteration of google/gemma-4-E2B-it) |
| Dataset | jondurbin/airoboros-3.2 (ShareGPT format, ~59k conversations) |
| Method | QLoRA 4-bit, LoRA rank 16 / alpha 16, dropout 0, all linear layers targeted |
| Steps | 1,000 (batch 1, grad accumulation 1), context length 2048 |
| Learning rate | 8e-5, cosine schedule with short warmup |
| Optimizer | AdamW 8-bit, weight decay 0.001 |
| Hardware | AMD Radeon RX 6700 XT 12 GB (ROCm) on CachyOS, ~23 minutes total |
| Final loss | smoothed ~0.7–1.0 on validation-free training curve (clean synthetic data, low loss is expected) |
| Tooling | Unsloth Studio (training, merge and GGUF conversion via llama.cpp) |
Usage
Text only, straight from the Hub:
llama-cli -hf minas2025/Gemma-4-E2B-Conversational-Uncensored --jinja
Server mode for frontends (SillyTavern, Open WebUI, etc.):
llama-server -m ./Gemma-4-E2B-Conversational-Uncensored.Q8_0.gguf -c 8192 --jinja --port 8080
Multimodal (only if you downloaded the mmproj file):
llama-mtmd-cli -hf minas2025/Gemma-4-E2B-Conversational-Uncensored --jinja
In LM Studio: search the repo name in the download tab, or sideload the Q8_0 file from disk.
Recommended generation settings
- Context length: 4096–8192
- Temperature: 0.8–1.0 (0.85 is a good default)
- Repetition penalty: 1.10–1.15
- No special system prompt required; it already knows how to hold a conversation
Example:
user: explain what a quantized model is like I'm five
model: Okay so imagine you have a huge box of crayons...
Intended use & disclaimer
An uncensored conversational companion for local, private use. Outputs are unfiltered by design: you are responsible for what you generate and for complying with your local laws. Not intended for medical, legal, financial or other advice-style tasks.
License
Inherits the Gemma Terms of Use via the base model chain (google/gemma-4-E2B-it → Heretic abliteration → this finetune). No additional restrictions are added by this finetune.
Credits
- Base model: llmfan46/gemma-4-E2B-it-ultra-uncensored-heretic, a Heretic abliteration of google/gemma-4-E2B-it
- Dataset: jondurbin/airoboros-3.2 by Jon Durbin
- Training, merging and GGUF conversion: Unsloth + llama.cpp
- One (1) rabbit, now promoted to Senior Training Supervisor after last run's success, provided moral supervision and zero technical contribution.