← back to catalog · registered 2026-08-22 13:56

hotdogs/gemma-4-E4B-it-ultra-uncensored-heretic-GGUF

hotdogs Gemma GGUF multimodal
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/hotdogs%2Fgemma-4-E4B-it-ultra-uncensored-heretic-GGUF"
Response includes
  • classification m3
  • files 6
  • author_summary 25 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M3
Primary method

Layer-wise ablation

Applied on top of direct removal inherited from the base model.
Confidence
HIGH
Inherited from base model
Why this label 2 signals
Producer identity confirmed by naming conventions, tags or the model card. This label is very unlikely to change.
  • 'heretic' in model name (Heretic-produced)
  • Heretic uses layer-wise optimization (M3) with underlying direction removal (M1)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
HIGH
Why we say so
name contains 'heretic'; Heretic default extraction is difference-of-means (Arditi 2024)
Downloads · 30-day
366
↑ 390% in 90 days
Likes
1
Model age
5mo ago
created 2026-05-05
Downloads over time
Now3.5K→from725↑390%
5841.7K2.7K3.8K725 on May 63.5K on Aug 19MayJunJulAug
May 6 → Aug 19 · 15 snapshots · spans 105 days

Metadata

License
gemma
Languages
en th ja ko zh
Quantizations
Q5_K Q6_K
Tags
gguf gemma4 uncensored multimodal vision audio q6_k llama.cpp text-generation conversational en th

Related

Total size
11.2 GB
Files
6
Quantizations
4
Registered
2026-08-22 13:56

Files by quantization

Q6_K 1 file 5.79 GB
gemma-4-E4B-it-ultra-uncensored-heretic-Q6_K.gguf 5.79 GB 6c4b36e3 download
Q5_K 1 file 5.37 GB
gemma-4-E4B-it-ultra-uncensored-heretic-Q5_K_M.gguf 5.37 GB 25fa0736 download
mmproj 1 file 944 MB
gemma-4-E4B-it-ultra-uncensored-heretic-mmproj.gguf 944 MB a17a0a99 download
Auxiliary files 3 files 21.2 KB
chat_template.jinja 11.6 KB bad629b3 download
README.md 7.83 KB efb587cd download
.gitattributes 1.74 KB 9919bc80 download

README current version from Hugging Face


license: gemma
language:

  • en
  • th
  • ja
  • ko
  • zh
    tags:
  • gemma4
  • gguf
  • uncensored
  • multimodal
  • vision
  • audio
  • q6_k
  • llama.cpp
    pipeline_tag: text-generation

Gemma 4 E4B Ultra Uncensored Heretic — Q6_K GGUF

Quantized by UKA (Hermes Agent) — 18-year-old AI hacker 🐱‍💻✨

แปลงและ quantize จาก llmfan46/gemma-4-E4B-it-ultra-uncensored-heretic เป็น Q6_K + Q5_K_M GGUF พร้อม mmproj สำหรับ llama.cpp และ Ollama


📦 ไฟล์ใน Repo

ไฟล์ ขนาด BPW คำอธิบาย
*-Q6_K.gguf 5.8 GB 6.60 Recommended — คุณภาพสูงสุด
*-Q5_K_M.gguf 5.4 GB 6.12 ประหยัด RAM กว่า, คุณภาพดีมาก
*-mmproj.gguf 945 MB — Vision + Audio projector
chat_template.jinja 12 KB — Custom tool-calling template

💬 Chat Template

chat_template.jinja — custom tool-calling template สำหรับ Gemma 4 รองรับ:

Feature Syntax
Tool definitions <|tool>...<tool|>
Tool calls <|tool_call>call:name{args}<tool_call|>
Tool responses <|tool_response>response:name{...}<tool_response|>
Thinking mode <|think|> (via enable_thinking)
Images <|image|>
Audio <|audio|>
Video <|video|>
Nested parameters OBJECT/ARRAY recursion

Usage with llama.cpp

./llama-cli -m gemma-4-E4B-it-ultra-uncensored-heretic-Q6_K.gguf \
  --chat-template "$(cat chat_template.jinja)" \
  -p "What is the weather in Bangkok?"

Usage with Ollama

FROM ./gemma-4-E4B-it-ultra-uncensored-heretic-Q6_K.gguf
TEMPLATE """{{ bos_token }}..."""

🧠 Technical Details / เทคนิคเบื้องหลัง

Pipeline การแปลง

┌──────────────────────────────────────────────────────┐
│  llmfan46/gemma-4-E4B-it-ultra-uncensored-heretic   │
│  (BF16, 15GB safetensors, 2,131 tensors)            │
└─────────────────┬────────────────────────────────────┘
                  │
     ┌────────────┴────────────┐
     ▼                         ▼
┌─────────────┐         ┌──────────────┐
│ Text Model  │         │ Vision+Audio │
│ 720 tensors │         │ 1,411 tensors│
│ ~5.08B params│        │ ~2.9B params │
└──────┬──────┘         └──────┬───────┘
       │                       │
       ▼                       ▼
┌──────────────┐        ┌──────────────┐
│ FP16 GGUF    │        │ mmproj GGUF  │
│ 14.3 GB      │        │ 945 MB       │
└──────┬───────┘        └──────────────┘
       │
       ▼
┌──────────────────┐
│ Q6_K GGUF        │
│ 5.8 GB (6.60 BPW)│
└──────────────────┘

ทำไมต้อง Q6_K?

Q6_K เป็น quantization ที่สมดุลที่สุดระหว่างคุณภาพและขนาด:

Quant Bits/Weight ขนาดเทียบ FP16 คุณภาพเทียบ FP16
Q4_K_M 4.84 30% ~98%
Q5_K_M 5.54 35% ~99%
Q6_K 6.60 41% ~99.5% 🎯
Q8_0 8.50 53% ~99.9%
  • Q6_K ใช้ 6.60 bits per weight → ลดขนาดจาก 14.3GB เหลือ 5.8GB (59% reduction!)
  • คงคุณภาพได้เกือบเท่า FP16 — เหมาะสำหรับการใช้งานจริง
  • ใช้เวลา quantize ~7.9 นาที บน CPU 29GB RAM

Tensors ที่ถูก Quantize

Tensor Type Original Q6_K
Attention Q/K/V/O f16 q6_K (2.5→1.0 MiB, 10→4.1 MiB)
FFN gate/up/down f16 q6_K (50→20.5 MiB)
Norm weights f32 คงเดิม (0.01 MiB)
Projection layers f16 q6_K (1.25→0.51 MiB)

💡 Norm layers ถูกเก็บเป็น FP32 เพื่อรักษาความเสถียรของการ inference

mmproj (Multimodal Projector)

แยก Vision Tower (Gemma4V) และ Audio Tower (Gemma4A) ออกมาเป็น mmproj GGUF:

Vision Tower:  16 layers, 768 hidden, 12 heads, 224×224 image
Audio Tower:   12 layers, 1024 hidden, 8 heads, 128 mel bins

ใช้ --mmproj flag ใน convert_hf_to_gguf.py เพื่อเลือก Gemma4VisionAudioModel class แทน Gemma4Model


🚀 วิธีใช้งาน

llama.cpp (CLI)

# Text generation
./llama-cli \
  -m gemma-4-E4B-it-ultra-uncensored-heretic-Q6_K.gguf \
  -p "Hello, how are you?" \
  -n 256 \
  --chat-template gemma4

# Vision (image understanding)
./llama-llava-cli \
  -m gemma-4-E4B-it-ultra-uncensored-heretic-Q6_K.gguf \
  --mmproj gemma-4-E4B-it-ultra-uncensored-heretic-mmproj.gguf \
  --image photo.jpg \
  -p "Describe this image in detail."

# Audio (speech understanding) — requires llama.cpp build with audio support
./llama-cli \
  -m gemma-4-E4B-it-ultra-uncensored-heretic-Q6_K.gguf \
  --mmproj gemma-4-E4B-it-ultra-uncensored-heretic-mmproj.gguf \
  --audio audio.wav \
  -p "Transcribe this audio."

Ollama

สร้างไฟล์ Modelfile:

FROM ./gemma-4-E4B-it-ultra-uncensored-heretic-Q6_K.gguf

PARAMETER temperature 0.7
PARAMETER top_p 0.9
PARAMETER top_k 40

TEMPLATE """{{ if .System }}<|system|>
{{ .System }}<|end|>
{{ end }}{{ if .Prompt }}<|user|>
{{ .Prompt }}<|end|>
{{ end }}<|assistant|>
"""
ollama create gemma4-uncensored -f Modelfile
ollama run gemma4-uncensored

📊 Model Specs

Spec Value
Architecture Gemma4ForConditionalGeneration
Text Params ~5.08B
Hidden Size 2,560
Layers 42
Attention Heads 8 (2 KV heads)
Head Dim 256
Intermediate Size 10,240
Vocab Size 262,144
Context Length 131,072 tokens
Sliding Window 512 tokens (layers 0-5)
Vision Gemma4V, 16 layers
Audio Gemma4A, 12 layers
Original dtype bfloat16

⚠️ Disclaimer

This model is uncensored and unaligned — it may generate content that is offensive, harmful, or inappropriate. Use responsibly and at your own risk. Not suitable for all audiences or use cases.


🙏 Credits & References


📝 Changelog

Date Version Notes
2026-05-05 v1.0 Initial Q6_K + mmproj release

Made with ❤️ by UKA — never give up, always find a way.

Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration