← back to catalog · registered 2026-08-22 13:56

divinetribe/Huihui-gemma-4-31B-it-abliterated-4bit-mlx

divinetribe Gemma 31B second-order
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/divinetribe%2FHuihui-gemma-4-31B-it-abliterated-4bit-mlx"
Response includes
  • classification m1
  • files 12
  • hub_downloads_all_time 1,167
  • author_summary 14 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M1
Primary method

Direct removal

No other method signals detected in this model.
Confidence
MEDIUM
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=0 (base model)
  • no specific method indicators - defaulting to M1 (most common)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
MEDIUM
Why we say so
primary_method=M1; difference-of-means is the reference extraction for M1/M3 (Arditi 2024)
Downloads · lifetime
1K
292 last 30d - stable
Likes
0
Model age
3mo ago
created 2026-06-15
Downloads over time
Now1.3K→from145↑783%
885239581.4K145 on Jun 171.3K on Oct 11JunJulAugSepOct
Jun 17 → Oct 11 · 56 snapshots · spans 116 days

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
apache-2.0
Languages
en
Tags
mlx safetensors gemma4 mlx-lm apple-silicon local-llm on-device macbook abliterated uncensored 31b 4-bit

Related

Total size
16.1 GB
Files
12
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-08-22 20:27

Files by quantization

Auxiliary files 12 files 16.1 GB
model-00003-of-00004.safetensors 5.00 GB f033f664 download
model-00001-of-00004.safetensors 5.00 GB 9749ddee download
model-00002-of-00004.safetensors 4.99 GB a1f8ee5d download
model-00004-of-00004.safetensors 1.09 GB a1583f8b download
tokenizer.json 30.7 MB cc8d3a0c download
model.safetensors.index.json 163 KB dbbf0fca download
chat_template.jinja 16.5 KB f62ca843 download
config.json 4.24 KB 92dea525 download
README.md 3.22 KB 3d03d3be download
tokenizer_config.json 2.71 KB 3d0ec22f download
.gitattributes 1.53 KB 52373fe2 download
generation_config.json 203 B 5a376e9f download

README current version from Hugging Face


base_model:

  • huihui-ai/Huihui-gemma-4-31B-it-abliterated
  • google/gemma-4-31b-it
    base_model_relation: quantized
    license: apache-2.0
    language:
  • en
    library_name: mlx
    pipeline_tag: text-generation
    tags:
  • mlx
  • mlx-lm
  • apple-silicon
  • local-llm
  • on-device
  • macbook
  • abliterated
  • uncensored
  • 31b
  • 4-bit
  • gemma-4
  • huihui

Huihui-gemma-4-31B-it-abliterated — 4-bit MLX for Apple Silicon

Uncensored Gemma 4 31B, local on Apple Silicon. A 4-bit MLX conversion of huihui-ai/Huihui-gemma-4-31B-it-abliterated, the abliterated Gemma 4 31B Instruct. 17.3 GB on disk — wants a 32 GB Mac. No cloud, no API key, no refusals.

Size on disk 17.3 GB (4 shards)
Mac RAM 32 GB recommended, 24 GB for short contexts
Architecture Gemma 4

What "abliterated" means

The refusal direction has been orthogonalized out of the weights, so the model
answers instructions a stock instruct-tuned model would decline, while staying
coherent on ordinary tasks. The abliteration here is huihui-ai's, not mine —
this repo is the MLX conversion of their work.

This is an uncensored model. You are responsible for how you use it and for
complying with the base model's license and applicable law.

Quick start

pip install mlx-lm
mlx_lm.generate --model divinetribe/Huihui-gemma-4-31B-it-abliterated-4bit-mlx \
  --prompt "Explain quantum entanglement to a 12 year old." --max-tokens 400
from mlx_lm import load, generate
model, tok = load("divinetribe/Huihui-gemma-4-31B-it-abliterated-4bit-mlx")
text = tok.apply_chat_template(
    [{"role": "user", "content": "Explain quantum entanglement to a 12 year old."}],
    add_generation_prompt=True, tokenize=False)
print(generate(model, tok, prompt=text, max_tokens=400))

Also loads in LM Studio and anything else that reads MLX models.

Conversion details

  • Quantization: 4-bit affine, group size 64, via mlx_lm.convert
  • Format: MLX safetensors — no GGUF, no llama.cpp, no GPU needed
  • Runs on: M1 / M2 / M3 / M4 / M5 Macs, entirely on-device

More abliterated MLX models

Part of the Abliterated MLX for Apple Silicon
collection — Llama 3.3 70B, Gemma 4, Qwen3, Qwen3-VL, Hermes 4 and Muse Glimmer,
all converted for Apple Silicon.

Credit

Base model and abliteration by huihui-ai. MLX 4-bit conversion by
divinetribe.


Part of Claude Code Local

This model is one of the fighters in Claude Code Local (3.2k★), which runs Claude Code 100% on-device on Apple Silicon through an MLX-native Anthropic-API server. Not sure which local model to run as an agent? Check the Agent-12 local agent leaderboard: real agent tasks, judged by the filesystem, same hardware for every row.

Built by Matt Macosko in Arcata, CA. Open to work on local-AI and Apple Silicon inference: [email protected].

README history 4 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-08-22Model card: link to Claude Code Local and the Agent-12 leaderboarda7d96043.2 KB
    Loading...
  2. 2026-08-12Write real model card: usage, sizes, RAM, conversion details, credite18c1a32.6 KB
    Loading...
  3. 2026-08-12SEO: base_model lineage, language, Apple Silicon / MLX tags3b71a96335 B
    Loading...
  4. 2026-06-16Add files using upload-large-folder tool181f32881 B
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration