← back to catalog · registered 2026-08-22 13:56

cashxout/Gemma-3-4B-VL-it-Gemini-Pro-Heretic-Uncensored-Thinking_GGUF

cashxout Gemma 4B GGUF multimodal 131K ctx
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/cashxout%2FGemma-3-4B-VL-it-Gemini-Pro-Heretic-Uncensored-Thinking_GGUF"
Response includes
  • classification m3
  • files 10
  • hub_downloads_all_time 4,264
  • author_summary 2 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M3
Primary method

Layer-wise ablation

Applied on top of direct removal inherited from the base model.
Confidence
HIGH
Inherited from base model
Why this label 2 signals
Producer identity confirmed by naming conventions, tags or the model card. This label is very unlikely to change.
  • 'heretic' in model name (Heretic-produced)
  • Heretic uses layer-wise optimization (M3) with underlying direction removal (M1)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
HIGH
Why we say so
name contains 'heretic'; Heretic default extraction is difference-of-means (Arditi 2024)
Downloads · lifetime
4K
280 last 30d - cooling
Likes
2
Model age
7mo ago
created 2026-02-17
Downloads over time
Now4.4K→from924↑375%
7512.1K3.4K4.7K924 on Feb 184.4K on Oct 11FebAprJunAugOct
Feb 18 → Oct 11 · 73 snapshots · spans 235 days

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
apache-2.0
Languages
en
Quantizations
F16 Q8_0
Tags
gguf vision-language multimodal uncensored text-generation image-understanding en license:apache-2.0 endpoints_compatible region:us conversational

Related

Total size
22.6 GB
Files
10
Quantizations
3
Registered
2026-08-22 13:56
Last updated on HF
2026-02-17 12:34

Files by quantization

F16 2 files 8.03 GB
Gemma-3-4B-VL-it-Gemini-Pro-Heretic-Uncensored-Thinking_F16.gguf 7.23 GB d0e98143 download
Gemma-3-4B-VL-it-Gemini-Pro-Heretic-Uncensored-Thinking_mmproj_f16.gguf 812 MB 3b33720b download
Q8_0 1 file 3.85 GB
Gemma-3-4B-VL-it-Gemini-Pro-Heretic-Uncensored-Thinking_Q8_0.gguf 3.85 GB da903d54 download
Auxiliary files 7 files 11.5 GB
Gemma-3-4B-VL-it-Gemini-Pro-Heretic-Uncensored-Thinking_Q6_k.gguf 2.97 GB 88e51262 download
Gemma-3-4B-VL-it-Gemini-Pro-Heretic-Uncensored-Thinking_Q5_k_m.gguf 2.64 GB 15577d00 download
Gemma-3-4B-VL-it-Gemini-Pro-Heretic-Uncensored-Thinking_Q4_k_m.gguf 2.32 GB 9efb3af6 download
Gemma-3-4B-VL-it-Gemini-Pro-Heretic-Uncensored-Thinking_Q3_k_m.gguf 1.95 GB 2ab2983a download
Gemma-3-4B-VL-it-Gemini-Pro-Heretic-Uncensored-Thinking_Q2_k.gguf 1.61 GB 1a7dbf23 download
README.md 2.70 KB 58f241de download
.gitattributes 2.29 KB 5a4e97ff download

README current version from Hugging Face


license: apache-2.0
language:

  • en
    tags:
  • vision-language
  • multimodal
  • uncensored
  • gguf
  • text-generation
  • image-understanding
    base_model:
  • Gemma-3-4B

Gemma-3-4B-VL-it-Gemini-Pro-Heretic-Uncensored-Thinking-GGUF

This repository contains Gemma-3-4B-VL-it-Gemini-Pro-Heretic-Uncensored-Thinking-GGUF, a 4B-parameter vision-language instruction-tuned model provided in GGUF format for efficient local inference.
The model is designed for open-ended reasoning, multimodal understanding, and minimal alignment constraints, making it suitable for experimentation, research, and advanced local deployments.


Model Summary

  • Model ID: Gemma-3-4B-VL-it-Gemini-Pro-Heretic-Uncensored-Thinking-GGUF
  • Architecture: Gemma 3 (4B parameters)
  • Type: Vision-Language (Text + Image)
  • Format: GGUF
  • Publisher: mradermacher
  • License: Apache 2.0 (inherits from base model)

Key Characteristics

  • Multimodal input support (text + images)
  • Instruction-tuned for conversational and reasoning tasks
  • Reduced content filtering and alignment constraints
  • Optimized for local inference runtimes
  • Suitable for research, exploration, and advanced user workflows

⚠️ This model is uncensored. Outputs may include sensitive or unfiltered content. Use responsibly.


Supported Use Cases

Text-Based

  • Conversational assistants
  • Creative writing and storytelling
  • Summarization and rewriting
  • General reasoning and analysis

Vision + Text

  • Image captioning
  • Visual question answering
  • Scene and object understanding
  • Multimodal reasoning tasks

GGUF Compatibility

This model can be used with GGUF-compatible runtimes such as:

  • llama.cpp
  • Ollama (GGUF-based builds)
  • Other local inference engines supporting GGUF

Performance and supported features may vary depending on runtime and hardware.


Basic Usage Example

Command Line (llama.cpp-style)

./main \
  -m Andycurrent/Gemma-3-4B-VL-it-Gemini-Pro-Heretic-Uncensored-Thinking_GGUF_F16.gguf \
  -p "Describe the key idea behind multimodal AI models."

Usage Notes

  • Provide clear, explicit prompts for best results
  • When using images, ensure proper formatting and resolution
  • Add moderation or filtering layers if deploying in public-facing applications

Ethical Considerations

Due to its uncensored nature:

  • Not recommended for unrestricted public deployment
  • Should not be used in safety-critical environments
  • Users are responsible for compliance with applicable laws and policies

Acknowledgements

  • Gemma base model contributors
  • Open-source inference and quantization communities
  • Tools and runtimes enabling efficient local LLM deployment

README history 1 version

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-02-17Duplicate from Andycurrent/Gemma-3-4B-VL-it-Gemini-Pro-Heretic-Uncensored-Thi...3436d822.7 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration