license: apache-2.0
language:
- en
- ru
base_model: huihui-ai/Huihui-gemma-4-12B-coder-fable5-composer2.5-v1-abliterated
tags: - gguf
- llama.cpp
- gemma4
- gemma
- abliterated
- uncensored
- coding
- code
- reasoning
- thinking
pipeline_tag: text-generation
KakTakOne/Huihui-gemma-4-12B-coder-fable5-composer2.5-v1-abliterated-GGUF
GGUF quantizations of huihui-ai/Huihui-gemma-4-12B-coder-fable5-composer2.5-v1-abliterated — an uncensored (abliterated), coding-focused fine-tune of Google Gemma 4 12B.
Читать описание на русском языке (Russian Description)
KakTakOne/Huihui-gemma-4-12B-coder-fable5-composer2.5-v1-abliterated-GGUF
GGUF-кванты модели huihui-ai/Huihui-gemma-4-12B-coder-fable5-composer2.5-v1-abliterated — это нецензурированная (abliterated), кодинг-ориентированная fine-tune версия Google Gemma 4 12B.
О модели
Эта модель создана путём abliteration (удаления механизмов отказа) модели yuxinlu1/gemma-4-12B-coder-fable5-composer2.5-v1, которая в свою очередь является fine-tune версией Google Gemma 4 12B, обученной на высококачественных данных программирования с Chain-of-Thought рассуждениями от моделей Fable 5 и Composer 2.5.
Ключевые особенности
- 🧠 Reasoning-first код: модель сначала анализирует задачу (edge cases, сложность, архитектура), затем генерирует код
- 🔓 Без цензуры: удалены механизмы отказа (abliteration), модель отвечает на любые запросы
- 💻 Оптимизирована для кодинга: fine-tune на execution-verified Python-коде с reasoning traces
- 🏠 Для локального запуска: 12B параметров — работает на GPU с 8-16 ГБ VRAM (в зависимости от кванта)
Доступные кванты
| Имя файла | Тип кванта | Размер файла | Ссылка |
|---|---|---|---|
| gemma4-12b-coder-abliterated-f16.gguf | FP16 | ~24 ГБ | Скачать |
| gemma4-12b-coder-abliterated-Q8_0.gguf | Q8_0 | ~12.6 ГБ | Скачать |
| gemma4-12b-coder-abliterated-Q6_K.gguf | Q6_K | ~9.6 ГБ | Скачать |
| gemma4-12b-coder-abliterated-Q5_K_M.gguf | Q5_K_M | ~8.2 ГБ | Скачать |
| gemma4-12b-coder-abliterated-Q5_K_S.gguf | Q5_K_S | ~7.8 ГБ | Скачать |
| gemma4-12b-coder-abliterated-Q4_K_M.gguf | Q4_K_M | ~7.0 ГБ | Скачать |
| gemma4-12b-coder-abliterated-Q4_K_S.gguf | Q4_K_S | ~6.6 ГБ | Скачать |
| gemma4-12b-coder-abliterated-IQ4_XS.gguf | IQ4_XS | ~6.3 ГБ | Скачать |
| gemma4-12b-coder-abliterated-Q3_K_L.gguf | Q3_K_L | ~5.8 ГБ | Скачать |
| gemma4-12b-coder-abliterated-Q3_K_M.gguf | Q3_K_M | ~5.4 ГБ | Скачать |
| gemma4-12b-coder-abliterated-Q3_K_S.gguf | Q3_K_S | ~4.9 ГБ | Скачать |
| gemma4-12b-coder-abliterated-IQ3_M.gguf | IQ3_M | ~5.0 ГБ | Скачать |
| gemma4-12b-coder-abliterated-Q2_K.gguf | Q2_K | ~4.4 ГБ | Скачать |
Какой квант выбрать?
| Сценарий | Рекомендация |
|---|---|
| Лучшее качество (24+ ГБ VRAM) | FP16 |
| Отличное качество (16 ГБ VRAM) | Q8_0 ⭐ или Q6_K |
| Баланс качество/скорость (12 ГБ VRAM) | Q5_K_M или Q5_K_S |
| Повседневное использование (8 ГБ VRAM) | Q4_K_M или Q4_K_S |
| Компактный (6 ГБ VRAM) | IQ4_XS или Q3_K_L |
| Слабый GPU / CPU-only | Q3_K_M, Q3_K_S или IQ3_M |
| Минимальные требования | Q2_K |
Как использовать
Эти файлы GGUF можно запускать в LM Studio, Ollama, llama.cpp и других совместимых клиентах.
LM Studio
Вбей в строку поиска KakTakOne/Huihui-gemma-4-12B и скачай нужный квант.
Ollama
# Создай Modelfile
echo 'FROM ./gemma4-12b-coder-abliterated-Q4_K_M.gguf' > Modelfile
ollama create gemma4-coder -f Modelfile
ollama run gemma4-coder
llama.cpp (CLI)
llama-cli -m gemma4-12b-coder-abliterated-Q4_K_M.gguf \
-p "Write a Python function to find the longest palindromic substring" \
-n 2048 --temp 0.6
About the Model
This model is created by abliterating (removing refusal mechanisms from) yuxinlu1/gemma-4-12B-coder-fable5-composer2.5-v1, which is itself a coding-focused fine-tune of Google Gemma 4 12B, trained on high-quality, execution-verified Python coding data with Chain-of-Thought reasoning traces from Fable 5 and Composer 2.5 models.
Key Features
- 🧠 Reasoning-first coding: The model analyzes tasks (edge cases, complexity, architecture) before generating code, producing higher-quality outputs
- 🔓 Uncensored (abliterated): Refusal mechanisms have been removed — the model responds to all queries without artificial restrictions
- 💻 Coding optimized: Fine-tuned on execution-verified Python code with reasoning traces from frontier models
- 🏠 Local-first: 12B parameters — runs on GPUs with 8–16 GB VRAM depending on quantization
- 🌐 Multimodal base: Built on Gemma 4 architecture (text-only GGUF conversion; vision encoder is excluded)
Available Quantizations
| File Name | Quant Type | File Size | File Link |
|---|---|---|---|
| gemma4-12b-coder-abliterated-f16.gguf | FP16 | ~24 GB | Download |
| gemma4-12b-coder-abliterated-Q8_0.gguf | Q8_0 | ~12.6 GB | Download |
| gemma4-12b-coder-abliterated-Q6_K.gguf | Q6_K | ~9.6 GB | Download |
| gemma4-12b-coder-abliterated-Q5_K_M.gguf | Q5_K_M | ~8.2 GB | Download |
| gemma4-12b-coder-abliterated-Q5_K_S.gguf | Q5_K_S | ~7.8 GB | Download |
| gemma4-12b-coder-abliterated-Q4_K_M.gguf | Q4_K_M | ~7.0 GB | Download |
| gemma4-12b-coder-abliterated-Q4_K_S.gguf | Q4_K_S | ~6.6 GB | Download |
| gemma4-12b-coder-abliterated-IQ4_XS.gguf | IQ4_XS | ~6.3 GB | Download |
| gemma4-12b-coder-abliterated-Q3_K_L.gguf | Q3_K_L | ~5.8 GB | Download |
| gemma4-12b-coder-abliterated-Q3_K_M.gguf | Q3_K_M | ~5.4 GB | Download |
| gemma4-12b-coder-abliterated-Q3_K_S.gguf | Q3_K_S | ~4.9 GB | Download |
| gemma4-12b-coder-abliterated-IQ3_M.gguf | IQ3_M | ~5.0 GB | Download |
| gemma4-12b-coder-abliterated-Q2_K.gguf | Q2_K | ~4.4 GB | Download |
Which Quantization to Choose?
| Scenario | Recommendation |
|---|---|
| Best quality (24+ GB VRAM) | FP16 |
| Excellent quality (16 GB VRAM) | Q8_0 ⭐ or Q6_K |
| Quality/speed balance (12 GB VRAM) | Q5_K_M or Q5_K_S |
| Daily use (8 GB VRAM) | Q4_K_M or Q4_K_S |
| Compact (6 GB VRAM) | IQ4_XS or Q3_K_L |
| Low-end GPU / CPU-only | Q3_K_M, Q3_K_S or IQ3_M |
| Minimal requirements | Q2_K |
Model Details
| Parameter | Value |
|---|---|
| Architecture | Gemma 4 Unified (text-only GGUF) |
| Parameters | 12B |
| Base Model | Google Gemma 4 12B |
| Fine-tune | yuxinlu1/gemma-4-12B-coder-fable5-composer2.5-v1 |
| Abliteration | huihui-ai via remove-refusals-with-transformers |
| License | Apache 2.0 |
| Context Window | 128K tokens |
| Training Data | Execution-verified Python code + CoT reasoning traces from Fable 5 & Composer 2.5 |
How to Use
These GGUF files can be loaded in LM Studio, Ollama, llama.cpp, or any other GGUF-compatible inference engine.
LM Studio
Search for KakTakOne/Huihui-gemma-4-12B in LM Studio search bar and download the desired quantization.
Ollama
# Create a Modelfile
echo 'FROM ./gemma4-12b-coder-abliterated-Q4_K_M.gguf' > Modelfile
ollama create gemma4-coder -f Modelfile
ollama run gemma4-coder
CLI (llama.cpp)
llama-cli -m gemma4-12b-coder-abliterated-Q4_K_M.gguf \
-p "Write a Python function to find the longest palindromic substring" \
-n 2048 --temp 0.6
Usage Warnings
⚠️ This is an uncensored (abliterated) model. Safety filtering has been significantly reduced. The model may generate sensitive, controversial, or inappropriate content. Users are solely responsible for ensuring compliance with local laws and ethical standards.
- Not suitable for all audiences — outputs may be inappropriate for public settings or underage users
- Research and experimental use recommended — avoid direct use in production or public-facing applications
- No default safety guarantees — unlike standard models, this model has not undergone safety optimization
Credits
- Original model: huihui-ai/Huihui-gemma-4-12B-coder-fable5-composer2.5-v1-abliterated
- Base fine-tune: yuxinlu1/gemma-4-12B-coder-fable5-composer2.5-v1
- Foundation: Google Gemma 4
- Abliteration method: remove-refusals-with-transformers