base_model: huihui-ai/Huihui-gemma-4-12B-coder-fable5-composer2.5-v1-abliterated
base_model_relation: quantized
license: gemma
library_name: gguf
pipeline_tag: text-generation
tags:
- gguf
- llama.cpp
- ollama
- gemma4
- code
- coder
- abliterated
- uncensored
language: - en
Huihui‑gemma‑4‑12B‑coder‑fable5‑composer2.5‑v1‑abliterated · GGUF
GGUF quantization for fast local inference of an abliterated (uncensored) coding model built on Gemma 4 12B.
🚀 Quick start · 📦 Files · 🔧 Conversion · ⚠️ Notice · 📜 Credits
✨ Overview
This repo provides a GGUF build ofhuihui-ai/Huihui-gemma-4-12B-coder-fable5-composer2.5-v1-abliterated
so you can run it locally with Ollama and llama.cpp.
The upstream model is a coding-focused fine-tune of Gemma 4 12B, trained on verifiable Python data with reasoning traces, then abliterated to remove refusals. Only the text tower is converted here (the multimodal vision/audio projector is omitted) — exactly what you want for code generation.
💡 The model emits a short reasoning pass (
Thinking…) before the final answer, then the code.
📦 Files
| File | Quant | Size | Bits/weight | Recommended for |
|---|---|---|---|---|
gemma4-coder-abliterated-Q4_K_M.gguf |
Q4_K_M | ~6.9 GB | 4.95 | Best quality/size balance · runs on ~8–12 GB VRAM |
🚀 Quick start (Ollama)
Requires Ollama ≥ 0.30 (built-in gemma4 renderer/parser).
1. Download the GGUF and create a Modelfile:
FROM ./gemma4-coder-abliterated-Q4_K_M.gguf
TEMPLATE {{ .Prompt }}
RENDERER gemma4
PARSER gemma4
PARAMETER temperature 1
PARAMETER top_k 64
PARAMETER top_p 0.95
2. Build and run:
ollama create huihui-gemma4-coder-abliterated -f Modelfile
ollama run huihui-gemma4-coder-abliterated "Write a Python function for binary search."
🦙 Quick start (llama.cpp)
llama-cli -m gemma4-coder-abliterated-Q4_K_M.gguf \
-p "Write a Python quicksort." -ngl 99 -c 8192
🔧 Conversion details
- Converted from the original fp16
safetensorswith llama.cppconvert_hf_to_gguf.py(build b9775), then quantized withllama-quantize→ Q4_K_M. - Architecture:
Gemma4UnifiedForConditionalGeneration— text tower only. - Tokenizer conversion requires transformers ≥ 5.10 (older 4.x breaks on the Gemma 4 tokenizer config).
- ⚠️ The upstream model uses a
proportionalRoPE type on full-attention layers that currentllama.cppdoes not yet implement; it falls back to standard RoPE. Fine for typical use, but very long contexts (>32k) may differ from the originaltransformersbehavior.
⚠️ Uncensored model
This is an abliterated model — its safety filtering has been significantly reduced and it may generate sensitive, controversial, or otherwise inappropriate content. Use responsibly, review outputs, and you are solely responsible for your use. This is provided as-is for research and local experimentation.
📜 License & credits
Derivative of a Gemma model, distributed under the Gemma Terms of Use (including the Prohibited Use Policy). All credit for the model itself goes to:
| Layer | Author |
|---|---|
| Abliterated version | huihui-ai |
| Base fine-tune | yuxinlu1/gemma-4-12B-coder-fable5-composer2.5-v1 |
| Foundation model | Google — Gemma 4 |
This repository contributes only the GGUF quantization for local inference.