← back to catalog · registered 2026-08-22 13:56

hfacesp/gemma-4-12b-it-uncensored-W4A16-AutoRound

hfacesp Gemma 1.4B second-order
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/hfacesp%2Fgemma-4-12b-it-uncensored-W4A16-AutoRound"
Response includes
  • classification m-uncensored
  • files 14
  • hub_downloads_all_time 1,584
  • author_summary 3 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M-U
Primary method

Uncensored (method unknown)

No other method signals detected in this model.
Confidence
LOW
Why this label 3 signals
Weak or ambiguous signals. Best guess based on catalog patterns; treat as tentative and check the evidence below.
  • 'uncensored' in name/tags but no 'abliterated' marker
  • method not identifiable from author declaration alone
  • may be DPO fine-tune, prompt engineering, or unknown technique
Refusal direction extraction

No specific extraction method could be identified for this model. The producer either did not document it or used a proprietary pipeline.

What is a refusal direction? →
Downloads · lifetime
2K
91 last 30d - cooling
Likes
1
Model age
4mo ago
created 2026-06-06
Downloads over time
Now1.6K→from331↑382%
2687531.2K1.7K331 on Jun 101.6K on Oct 111.6K on Oct 9JunJulAugSepOct
Jun 10 → Oct 11 · 57 snapshots · spans 123 days

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
gemma
Tags
safetensors gemma4_unified quantization autoround w4a16 vllm base_model:zaakirio/gemma-4-12b-it-uncensored base_model:quantized:zaakirio/gemma-4-12b-it-uncensored license:gemma 4-bit auto-round region:us

Related

Total size
7.25 GB
Files
14
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-06-06 19:03

Files by quantization

Auxiliary files 14 files 7.28 GB
model-00001-of-00004.safetensors 1.99 GB edf0df9f download
model-00002-of-00004.safetensors 1.99 GB f5dec295 download
model-00004-of-00004.safetensors 1.97 GB 123c9e1a download
model-00003-of-00004.safetensors 1.30 GB aa93f7dc download
tokenizer.json 30.7 MB a2619fe1 download
model.safetensors.index.json 127 KB b9e83e6b download
chat_template.jinja 17.1 KB e61bbfe9 download
config.json 4.57 KB 07d5841c download
tokenizer_config.json 2.76 KB 5941f228 download
.gitattributes 1.53 KB 52373fe2 download
README.md 1.50 KB 9b624bec download
processor_config.json 1.35 KB b889adcd download
quantization_config.json 293 B 6c4ceb27 download
generation_config.json 255 B 683ff358 download

README current version from Hugging Face


license: gemma
base_model: zaakirio/gemma-4-12b-it-uncensored
base_model_relation: quantized
tags:

  • gemma4_unified
  • quantization
  • autoround
  • w4a16
  • vllm

gemma-4-12b-it-uncensored-W4A16-AutoRound

4-bit weight / 16-bit activation (W4A16) quantization of
zaakirio/gemma-4-12b-it-uncensored.

Attribution

Quantization

  • Tool: Intel AutoRound (v0.14.0)
  • Scheme: W4A16 — 4-bit weights, group size 128, sym=True
  • Mode: RTN (iters=0) — required for Gemma 4 numerical stability
  • Format: native auto_round (quant_method=auto-round)
  • Multimodal projections (vision/audio embedders) kept unquantized
    (quant_nontext_module=False) to preserve the encoder-free multimodal path.

Serving (vLLM)

Encoder-free Gemma 4 Unified; serve with a vLLM build that supports
Gemma4UnifiedForConditionalGeneration. On Turing (Tesla T4) the Triton
attention TILE_SIZE patch (vllm-project/vllm#39018) is required.

README history 1 version

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-06-06AutoRound W4A16 RTN quant (Gemma 4 12B Unified abliterated)f7298651.5 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration