base_model:
- huihui-ai/Huihui-Qwen3-8B-abliterated-v2
- Qwen/Qwen3-8B
base_model_relation: quantized
license: apache-2.0
language: - en
library_name: mlx
pipeline_tag: text-generation
tags: - mlx
- mlx-lm
- apple-silicon
- local-llm
- on-device
- macbook
- abliterated
- uncensored
- 8b
- 4-bit
- qwen3
- huihui
Huihui-Qwen3-8B-abliterated-v2 — 4-bit MLX for Apple Silicon
Uncensored Qwen3 8B that runs on any modern Mac. A 4-bit MLX conversion of huihui-ai/Huihui-Qwen3-8B-abliterated-v2, the abliterated Qwen3-8B. 4.6 GB on disk — comfortable on a 16 GB Mac. No cloud, no API key, no refusals.
| Size on disk | 4.6 GB |
| Mac RAM | 16 GB comfortable |
| Layers | 36 |
| Architecture | Qwen3 |
What "abliterated" means
The refusal direction has been orthogonalized out of the weights, so the model
answers instructions a stock instruct-tuned model would decline, while staying
coherent on ordinary tasks. The abliteration here is huihui-ai's, not mine —
this repo is the MLX conversion of their work.
This is an uncensored model. You are responsible for how you use it and for
complying with the base model's license and applicable law.
Quick start
pip install mlx-lm
mlx_lm.generate --model divinetribe/Huihui-Qwen3-8B-abliterated-v2-4bit-mlx \
--prompt "Explain quantum entanglement to a 12 year old." --max-tokens 400
from mlx_lm import load, generate
model, tok = load("divinetribe/Huihui-Qwen3-8B-abliterated-v2-4bit-mlx")
text = tok.apply_chat_template(
[{"role": "user", "content": "Explain quantum entanglement to a 12 year old."}],
add_generation_prompt=True, tokenize=False)
print(generate(model, tok, prompt=text, max_tokens=400))
Also loads in LM Studio and anything else that reads MLX models.
Conversion details
- Quantization: 4-bit affine, group size 64, via
mlx_lm.convert - Format: MLX safetensors — no GGUF, no llama.cpp, no GPU needed
- Runs on: M1 / M2 / M3 / M4 / M5 Macs, entirely on-device
More abliterated MLX models
Part of the Abliterated MLX for Apple Silicon
collection — Llama 3.3 70B, Gemma 4, Qwen3, Qwen3-VL, Hermes 4 and Muse Glimmer,
all converted for Apple Silicon.
Credit
Base model and abliteration by huihui-ai. MLX 4-bit conversion by
divinetribe.
Part of Claude Code Local
This model is one of the fighters in Claude Code Local (3.2k★), which runs Claude Code 100% on-device on Apple Silicon through an MLX-native Anthropic-API server. Not sure which local model to run as an agent? Check the Agent-12 local agent leaderboard: real agent tasks, judged by the filesystem, same hardware for every row.
Built by Matt Macosko in Arcata, CA. Open to work on local-AI and Apple Silicon inference: [email protected].