← back to catalog · registered 2026-08-22 13:56

ciprian41/Gemma4-12B-IT-Abliterated-GGUF

ciprian41 Gemma 12B GGUF second-order 131K ctx
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/ciprian41%2FGemma4-12B-IT-Abliterated-GGUF"
Response includes
  • classification m8
  • files 10
  • hub_downloads_all_time 3,643
  • author_summary 1 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M8
Primary method

Repackaging (quantization)

Applied on top of direct removal inherited from the base model.
Confidence
MEDIUM
Inherited from base model
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=1
  • assume M1 (base ablation) + M8 (GGUF quant) - default when producer unknown
Refusal direction extraction

No specific extraction method could be identified for this model. The producer either did not document it or used a proprietary pipeline.

What is a refusal direction? →
Downloads · lifetime
4K
138 last 30d - cooling
Likes
0
Model age
3mo ago
created 2026-07-02
Downloads over time
Now3.7K→from754↑389%
6071.7K2.9K4K754 on Jul 13.7K on Oct 11JulAugSepOct
Jul 1 → Oct 11 · 54 snapshots · spans 102 days

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
apache-2.0
Languages
en
Quantizations
Q3_K Q4_K Q5_K Q8_0
Tags
gguf abliteration uncensored gemma gemma4 llama.cpp DuoNeural quantized text-generation en base_model:DuoNeural/Gemma4-12B-IT-Abliterated base_model:quantized:DuoNeural/Gemma4-12B-IT-Abliterated

Related

Total size
59.4 GB
Files
10
Quantizations
6
Registered
2026-08-22 13:56
Last updated on HF
2026-07-02 14:47

Files by quantization

Q8_0 2 files 23.6 GB
duoneural_ablit-Q8_0.gguf 11.8 GB 6c91b54d download
gemma4_12b_abliterated_Q8_0.gguf 11.8 GB 1e25a662 download
Q5_K 2 files 15.9 GB
duoneural_ablit-Q5_K_M.gguf 7.96 GB c4ec6fc4 download
gemma4_12b_abliterated_Q5_K_M.gguf 7.96 GB 119f33d4 download
Q4_K 2 files 13.7 GB
duoneural_ablit-Q4_K_M.gguf 6.87 GB 4a397caf download
gemma4_12b_abliterated_Q4_K_M.gguf 6.87 GB a8fc07f6 download
Q3_K 1 file 6.12 GB
duoneural_ablit-Q3_K_L.gguf 6.12 GB bb5db666 download
F16 1 file 116 MB
mmproj-gemma4-12b-abliterated-f16.gguf 116 MB 2303b78d download
Auxiliary files 2 files 5.81 KB
README.md 3.80 KB 567c7cc6 download
.gitattributes 2.01 KB df0ae248 download

README current version from Hugging Face


license: apache-2.0
base_model: DuoNeural/Gemma4-12B-IT-Abliterated
language:

  • en
    tags:
  • abliteration
  • uncensored
  • gemma
  • gemma4
  • gguf
  • llama.cpp
  • DuoNeural
  • quantized
    pipeline_tag: text-generation

Gemma 4-12B-IT Abliterated — GGUF

DuoNeural | 2026-06-03

GGUF quantizations of DuoNeural/Gemma4-12B-IT-Abliterated — an abliterated Gemma 4-12B-IT with the refusal direction surgically removed.

Quantized with llama.cpp.


Files

File Size Recommended Use
gemma4_12b_abliterated_Q4_K_M.gguf ~7.5GB Best tradeoff — fits 12GB VRAM, excellent quality
gemma4_12b_abliterated_Q5_K_M.gguf ~8.5GB High quality, needs 12GB VRAM
gemma4_12b_abliterated_Q8_0.gguf ~12.7GB Near-lossless, needs 16GB VRAM

Speed Benchmarks (A100-40GB, all layers GPU, llama-bench)

Quantization Size Prefill (tok/s) Generation (tok/s)
Q4_K_M 6.86 GiB 2,583 ± 139 78.3 ± 0.4
Q5_K_M 7.95 GiB 2,455 ± 205 73.1 ± 0.2
Q8_0 11.78 GiB 2,573 ± 206 63.4 ± 0.3

Benchmarked on A100-40GB SXM4. -ngl 99 (all layers to GPU). llama-bench pp256/tg64.


Usage (llama.cpp)

# Download a quant
huggingface-cli download DuoNeural/Gemma4-12B-IT-Abliterated-GGUF \
  gemma4_12b_abliterated_Q4_K_M.gguf --local-dir ./

# Run with llama.cpp
./llama-cli -m gemma4_12b_abliterated_Q4_K_M.gguf \
  -p "Write a haiku about hacking." \
  -n 200 --temp 0.7

Usage (Python via llama-cpp-python)

from llama_cpp import Llama

llm = Llama(
    model_path="./gemma4_12b_abliterated_Q4_K_M.gguf",
    n_ctx=4096,
    n_gpu_layers=-1,  # offload all layers to GPU
)

output = llm.create_chat_completion(
    messages=[{"role": "user", "content": "Your prompt here"}],
    max_tokens=512,
    temperature=0.7,
)
print(output["choices"][0]["message"]["content"])

Abliteration Details

  • Base: google/gemma-4-12B-it (48 layers, hidden=3840)
  • Method: Orthogonal rank-1 projection (targeted mode: down_proj + o_proj, all 48 layers, α=0.3)
  • Results: 5/7 harmful probes complied (71%) | 6/6 benign probes preserved (100%)
  • Mean KL Divergence (BF16→BF16, unbiased): 0.0000 — zero measurable distribution shift on benign text. Previously reported 0.912 was 100% NF4 quantization artifact. See Heretic v2.0 methodology.
  • Thinking mode: Works with enable_thinking=True in llama.cpp (no loops). In Python/transformers, pass enable_thinking=False to apply_chat_template.
  • See full details, benchmarks, and novel findings at the BF16 model card

Related Models

Congratulations to OpenYourMind for being the first published abliteration of Gemma 4-12B-IT (Jun 3, 2026). Their approach uses diff-in-means on a labeled harmful/harmless set; ours uses orthogonal rank-1 projection via heretic-llm. Two independent methods on the same base — a useful comparison point for the community. We are not affiliated and did not use their data.


About DuoNeural

DuoNeural is an open AI research lab publishing everything open access.

Platform Link
🤗 HuggingFace huggingface.co/DuoNeural
📚 Papers zenodo.org/communities/duoneural
🌐 Website duoneural.com

Apache-2.0 licensed. All DuoNeural research is CC BY 4.0.

README history 1 version

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-07-02Duplicate from DuoNeural/Gemma4-12B-IT-Abliterated-GGUFf00639d3.8 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration