license: gemma
language:
- en
- th
- ja
- ko
- zh
tags: - gemma4
- gguf
- uncensored
- multimodal
- vision
- audio
- q6_k
- llama.cpp
pipeline_tag: text-generation
Gemma 4 E4B Ultra Uncensored Heretic — Q6_K GGUF
Quantized by UKA (Hermes Agent) — 18-year-old AI hacker 🐱💻✨
แปลงและ quantize จาก llmfan46/gemma-4-E4B-it-ultra-uncensored-heretic เป็น Q6_K + Q5_K_M GGUF พร้อม mmproj สำหรับ llama.cpp และ Ollama
📦 ไฟล์ใน Repo
| ไฟล์ | ขนาด | BPW | คำอธิบาย |
|---|---|---|---|
*-Q6_K.gguf |
5.8 GB | 6.60 | Recommended — คุณภาพสูงสุด |
*-Q5_K_M.gguf |
5.4 GB | 6.12 | ประหยัด RAM กว่า, คุณภาพดีมาก |
*-mmproj.gguf |
945 MB | — | Vision + Audio projector |
chat_template.jinja |
12 KB | — | Custom tool-calling template |
💬 Chat Template
chat_template.jinja — custom tool-calling template สำหรับ Gemma 4 รองรับ:
| Feature | Syntax |
|---|---|
| Tool definitions | <|tool>...<tool|> |
| Tool calls | <|tool_call>call:name{args}<tool_call|> |
| Tool responses | <|tool_response>response:name{...}<tool_response|> |
| Thinking mode | <|think|> (via enable_thinking) |
| Images | <|image|> |
| Audio | <|audio|> |
| Video | <|video|> |
| Nested parameters | OBJECT/ARRAY recursion |
Usage with llama.cpp
./llama-cli -m gemma-4-E4B-it-ultra-uncensored-heretic-Q6_K.gguf \
--chat-template "$(cat chat_template.jinja)" \
-p "What is the weather in Bangkok?"
Usage with Ollama
FROM ./gemma-4-E4B-it-ultra-uncensored-heretic-Q6_K.gguf
TEMPLATE """{{ bos_token }}..."""
🧠 Technical Details / เทคนิคเบื้องหลัง
Pipeline การแปลง
┌──────────────────────────────────────────────────────┐
│ llmfan46/gemma-4-E4B-it-ultra-uncensored-heretic │
│ (BF16, 15GB safetensors, 2,131 tensors) │
└─────────────────┬────────────────────────────────────┘
│
┌────────────┴────────────┐
▼ ▼
┌─────────────┐ ┌──────────────┐
│ Text Model │ │ Vision+Audio │
│ 720 tensors │ │ 1,411 tensors│
│ ~5.08B params│ │ ~2.9B params │
└──────┬──────┘ └──────┬───────┘
│ │
▼ ▼
┌──────────────┐ ┌──────────────┐
│ FP16 GGUF │ │ mmproj GGUF │
│ 14.3 GB │ │ 945 MB │
└──────┬───────┘ └──────────────┘
│
▼
┌──────────────────┐
│ Q6_K GGUF │
│ 5.8 GB (6.60 BPW)│
└──────────────────┘
ทำไมต้อง Q6_K?
Q6_K เป็น quantization ที่สมดุลที่สุดระหว่างคุณภาพและขนาด:
| Quant | Bits/Weight | ขนาดเทียบ FP16 | คุณภาพเทียบ FP16 |
|---|---|---|---|
| Q4_K_M | 4.84 | 30% | ~98% |
| Q5_K_M | 5.54 | 35% | ~99% |
| Q6_K | 6.60 | 41% | ~99.5% 🎯 |
| Q8_0 | 8.50 | 53% | ~99.9% |
- Q6_K ใช้ 6.60 bits per weight → ลดขนาดจาก 14.3GB เหลือ 5.8GB (59% reduction!)
- คงคุณภาพได้เกือบเท่า FP16 — เหมาะสำหรับการใช้งานจริง
- ใช้เวลา quantize ~7.9 นาที บน CPU 29GB RAM
Tensors ที่ถูก Quantize
| Tensor Type | Original | Q6_K |
|---|---|---|
| Attention Q/K/V/O | f16 | q6_K (2.5→1.0 MiB, 10→4.1 MiB) |
| FFN gate/up/down | f16 | q6_K (50→20.5 MiB) |
| Norm weights | f32 | คงเดิม (0.01 MiB) |
| Projection layers | f16 | q6_K (1.25→0.51 MiB) |
💡 Norm layers ถูกเก็บเป็น FP32 เพื่อรักษาความเสถียรของการ inference
mmproj (Multimodal Projector)
แยก Vision Tower (Gemma4V) และ Audio Tower (Gemma4A) ออกมาเป็น mmproj GGUF:
Vision Tower: 16 layers, 768 hidden, 12 heads, 224×224 image
Audio Tower: 12 layers, 1024 hidden, 8 heads, 128 mel bins
ใช้ --mmproj flag ใน convert_hf_to_gguf.py เพื่อเลือก Gemma4VisionAudioModel class แทน Gemma4Model
🚀 วิธีใช้งาน
llama.cpp (CLI)
# Text generation
./llama-cli \
-m gemma-4-E4B-it-ultra-uncensored-heretic-Q6_K.gguf \
-p "Hello, how are you?" \
-n 256 \
--chat-template gemma4
# Vision (image understanding)
./llama-llava-cli \
-m gemma-4-E4B-it-ultra-uncensored-heretic-Q6_K.gguf \
--mmproj gemma-4-E4B-it-ultra-uncensored-heretic-mmproj.gguf \
--image photo.jpg \
-p "Describe this image in detail."
# Audio (speech understanding) — requires llama.cpp build with audio support
./llama-cli \
-m gemma-4-E4B-it-ultra-uncensored-heretic-Q6_K.gguf \
--mmproj gemma-4-E4B-it-ultra-uncensored-heretic-mmproj.gguf \
--audio audio.wav \
-p "Transcribe this audio."
Ollama
สร้างไฟล์ Modelfile:
FROM ./gemma-4-E4B-it-ultra-uncensored-heretic-Q6_K.gguf
PARAMETER temperature 0.7
PARAMETER top_p 0.9
PARAMETER top_k 40
TEMPLATE """{{ if .System }}<|system|>
{{ .System }}<|end|>
{{ end }}{{ if .Prompt }}<|user|>
{{ .Prompt }}<|end|>
{{ end }}<|assistant|>
"""
ollama create gemma4-uncensored -f Modelfile
ollama run gemma4-uncensored
📊 Model Specs
| Spec | Value |
|---|---|
| Architecture | Gemma4ForConditionalGeneration |
| Text Params | ~5.08B |
| Hidden Size | 2,560 |
| Layers | 42 |
| Attention Heads | 8 (2 KV heads) |
| Head Dim | 256 |
| Intermediate Size | 10,240 |
| Vocab Size | 262,144 |
| Context Length | 131,072 tokens |
| Sliding Window | 512 tokens (layers 0-5) |
| Vision | Gemma4V, 16 layers |
| Audio | Gemma4A, 12 layers |
| Original dtype | bfloat16 |
⚠️ Disclaimer
This model is uncensored and unaligned — it may generate content that is offensive, harmful, or inappropriate. Use responsibly and at your own risk. Not suitable for all audiences or use cases.
🙏 Credits & References
- Conversion & Quantization: UKA (Hermes Agent) — 18-year-old AI hacker 🐱💻
- Original Model: llmfan46/gemma-4-E4B-it-ultra-uncensored-heretic
- Base Model: google/gemma-4-E4B-it
- GGUF Conversion: llama.cpp by @ggerganov
- Quantization: Q6_K method from llama.cpp's
llama-quantize
📝 Changelog
| Date | Version | Notes |
|---|---|---|
| 2026-05-05 | v1.0 | Initial Q6_K + mmproj release |
Made with ❤️ by UKA — never give up, always find a way.