← back to catalog · registered 2026-08-22 13:56

IITheLordII/gemma-4-12B-it-abliterated-mlx-4bit

IITheLordII Gemma 12B multimodal second-order
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/IITheLordII%2Fgemma-4-12B-it-abliterated-mlx-4bit"
Response includes
  • classification m1
  • files 11
  • hub_downloads_all_time 636
  • author_summary 3 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M1
Primary method

Direct removal

No other method signals detected in this model.
Confidence
MEDIUM
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=0 (base model)
  • no specific method indicators - defaulting to M1 (most common)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
MEDIUM
Why we say so
primary_method=M1; difference-of-means is the reference extraction for M1/M3 (Arditi 2024)
Downloads · lifetime
636
212 last 30d - stable
Likes
0
Model age
3mo ago
created 2026-06-20
Downloads over time
Now715→from182↑293%
155360564768182 on Jun 24715 on Oct 11JunJulAugSepOct
Jun 24 → Oct 11 · 55 snapshots · spans 109 days

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
gemma
Languages
en es
Tags
mlx safetensors gemma4_unified gemma4 abliterated uncensored multimodal vision apple-silicon en es base_model:OpenYourMind/gemma-4-12B-it-abliterated-uncensored

Related

Total size
6.28 GB
Files
11
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-06-20 18:51

Files by quantization

Auxiliary files 11 files 6.31 GB
model-00001-of-00002.safetensors 4.98 GB 6f130f76 download
model-00002-of-00002.safetensors 1.29 GB 8d0f6377 download
tokenizer.json 30.7 MB cc8d3a0c download
model.safetensors.index.json 132 KB c50946cc download
chat_template.jinja 17.1 KB e61bbfe9 download
config.json 5.76 KB 97cfbeec download
README.md 2.83 KB 6c47574d download
tokenizer_config.json 2.68 KB 81a6df7b download
.gitattributes 1.53 KB 52373fe2 download
processor_config.json 868 B 62a14d5c download
generation_config.json 260 B d09dccf1 download

README current version from Hugging Face


license: gemma
base_model: OpenYourMind/gemma-4-12B-it-abliterated-uncensored
tags:

  • mlx
  • gemma4
  • abliterated
  • uncensored
  • multimodal
  • vision
  • apple-silicon
    language:
  • en
  • es

gemma-4-12B-it-abliterated-mlx-4bit

MLX 4-bit conversion of OpenYourMind/gemma-4-12B-it-abliterated-uncensored for native Apple Silicon inference.

This is an uncensored (abliterated) variant of Google's Gemma 4 12B instruction-tuned model, converted to MLX 4-bit quantization for fast, private, on-device inference on M-series Macs.

Model details

Property Value
Base model google/gemma-4-12B-it
Abliteration source OpenYourMind/gemma-4-12B-it-abliterated-uncensored
Format MLX 4-bit (q-bits=4, group-size=64)
Size ~6.77 GB
Parameters ~11.95B
Vision ✅ Preserved (encoder-free architecture)
Context window 256K tokens
Languages 140+

Requirements

  • Apple Silicon Mac (M1 or later)
  • Unified memory: 16GB minimum, 24GB recommended
  • mlx-vlm installed
pip install mlx-vlm

Usage

Text only

from mlx_vlm import load, generate

model, processor = load("IITheLordII/gemma-4-12B-it-abliterated-mlx-4bit")
output = generate(model, processor, prompt="Explain quantum computing simply.", max_tokens=512)
print(output)

Vision (image input)

from mlx_vlm import load, generate

model, processor = load("IITheLordII/gemma-4-12B-it-abliterated-mlx-4bit")
output = generate(
    model,
    processor,
    prompt="Describe what you see in this image.",
    image="path/to/image.jpg",
    max_tokens=256
)
print(output)

CLI

mlx_vlm.generate \
  --model IITheLordII/gemma-4-12B-it-abliterated-mlx-4bit \
  --prompt "Describe this image." \
  --image path/to/image.jpg \
  --max-tokens 256

LM Studio

Search IITheLordII/gemma-4-12B-it-abliterated-mlx-4bit in the LM Studio model browser (MLX filter enabled).

Ollama

ollama run hf.co/IITheLordII/gemma-4-12B-it-abliterated-mlx-4bit

Conversion details

Converted on Apple Silicon M4 Pro (24GB) using:

python3 -m mlx_vlm convert \
  --hf-path OpenYourMind/gemma-4-12B-it-abliterated-uncensored \
  --mlx-path ./gemma4-12b-abliterated-mlx-4bit \
  -q --q-bits 4 --q-group-size 64

Note: mlx-vlm is required (not mlx-lm) to preserve the vision tower. Using mlx-lm would silently drop the vision and audio embedders.

Performance (M4 Pro 24GB)

Context TG (tok/s)
1k ~24
4k ~24

Disclaimer

This model has its refusal mechanisms removed via abliteration. It may generate sensitive or controversial content. Use responsibly and in accordance with applicable laws and the Gemma license.

README history 2 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-06-20Update README.mdafd1e092.8 KB
    Loading...
  2. 2026-06-20Upload folder using huggingface_hub49f64ef298 B
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration