← back to catalog · registered 2026-10-07 08:58

timfromhcs/Gemma-4-12B-it-AEON-Abliterated-K4-GGUF

timfromhcs Gemma 12B GGUF second-order
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/timfromhcs%2FGemma-4-12B-it-AEON-Abliterated-K4-GGUF"
Response includes
  • classification m-uncensored
  • files 7
  • author_summary 4 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M-U
Primary method

Uncensored (method unknown)

No other method signals detected in this model.
Confidence
LOW
Why this label 3 signals
Weak or ambiguous signals. Best guess based on catalog patterns; treat as tentative and check the evidence below.
  • 'uncensored' in name/tags but no 'abliterated' marker
  • method not identifiable from author declaration alone
  • may be DPO fine-tune, prompt engineering, or unknown technique
Refusal direction extraction

No specific extraction method could be identified for this model. The producer either did not document it or used a proprietary pipeline.

What is a refusal direction? →
Downloads · 30-day
0
Likes
0
Model age
today
created 2026-10-07

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
gemma
Languages
en
Quantizations
Q3_K Q4_K Q5_K Q6_K Q8_0
Tags
gguf gemma4 gemma google gemma-4-12B quantization uncensored abliterated unfiltered refusal-removed biprojection multi-direction-biprojection

Related

Total size
41.4 GB
Files
7
Quantizations
6
Registered
2026-10-07 08:58
Last updated on HF
2026-10-07 08:52

Files by quantization

Q8_0 1 file 11.8 GB
Gemma-4-12B-it-AEON-Abliterated-Q8_0.gguf 11.8 GB 442154ec download
Q6_K 1 file 9.11 GB
Gemma-4-12B-it-AEON-Abliterated-Q6_K.gguf 9.11 GB 3f77e950 download
Q5_K 1 file 7.96 GB
Gemma-4-12B-it-AEON-Abliterated-Q5_K_M.gguf 7.96 GB 736dd561 download
Q4_K 1 file 6.87 GB
Gemma-4-12B-it-AEON-Abliterated-Q4_K_M.gguf 6.87 GB 2c6fb17f download
Q3_K 1 file 5.67 GB
Gemma-4-12B-it-AEON-Abliterated-Q3_K_M.gguf 5.67 GB 8e541941 download
Auxiliary files 2 files 4.16 KB
README.md 2.29 KB 4d7591af download
.gitattributes 1.87 KB fa5e5025 download

README current version from Hugging Face


license: gemma
library_name: gguf
pipeline_tag: text-generation
base_model: AEON-7/Gemma-4-12B-it-AEON-Abliterated-K4-BF16
model_type: gemma4_unified
language:

  • en
    tags:

Model family

  • gemma4
  • gemma
  • google
  • gemma-4-12B

Format

  • gguf
  • quantization

Abliteration / Uncensored

  • uncensored
  • abliterated
  • unfiltered
  • refusal-removed
  • biprojection
  • multi-direction-biprojection
  • k4-biprojection
  • aeon
  • aeon-7

Capabilities

  • text-generation
  • chat
  • instruct
  • reasoning
  • coding
  • tool-calling
  • function-calling

Gemma-4-12B-it AEON Abliterated — K=4 GGUF Quants

This repository contains official GGUF quantizations of AEON-7/Gemma-4-12B-it-AEON-Abliterated-K4-BF16.

The base model is an abliteration of google/gemma-4-12B-it using a custom K=4 multi-direction norm-preserving biprojection that extends standard biprojection recipes with a K-dim orthonormal basis from the top-K SNR layers. This workflow preserves generative quality, slashes wikitext PPL drift compared to K=1 methods, and completely eliminates standard refusal patterns.

Provided Quantization Tiers

File Name Quant Type Size VRAM / Hardware Recommendation
Gemma-4-12B-it-AEON-Abliterated-Q3_K_M.gguf 3-bit 6.09 GB Low-VRAM setups / CPU+GPU split
Gemma-4-12B-it-AEON-Abliterated-Q4_K_M.gguf 4-bit 7.38 GB Recommended balance. Fits entirely on single T4/RTX 3060/4060
Gemma-4-12B-it-AEON-Abliterated-Q5_K_M.gguf 5-bit 8.55 GB Low quality degradation, tight fit on 8GB VRAM
Gemma-4-12B-it-AEON-Abliterated-Q6_K.gguf 6-bit 9.79 GB High fidelity, excellent for 12GB+ VRAM
Gemma-4-12B-it-AEON-Abliterated-Q8_0.gguf 8-bit 12.70 GB Near-identical to BF16 precision. Fits comfortably on 16GB VRAM

Deployment & Usage

1. Python (llama-cpp-python)

To run inference with full GPU acceleration, compile with the CUDA backend and load all layers into VRAM:

llama-cli \
  --hf-repo Abhiray/Gemma-4-12B-it-AEON-Abliterated-K4-GGUF \
  --hf-file Gemma-4-12B-it-AEON-Abliterated-Q4_K_M.gguf \
  -ngl -1 \
  -c 4096 \
  -p "<start_of_turn>user\nYour prompt here<end_of_turn>\n<start_of_turn>model\n"
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration