← back to catalog · registered 2026-08-22 13:56

divinetribe/Huihui-Qwen3-8B-abliterated-v2-4bit-mlx

divinetribe Qwen 8.2B second-order
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/divinetribe%2FHuihui-Qwen3-8B-abliterated-v2-4bit-mlx"
Response includes
  • classification m1
  • files 9
  • hub_downloads_all_time 604
  • author_summary 14 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M1
Primary method

Direct removal

No other method signals detected in this model.
Confidence
MEDIUM
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=0 (base model)
  • no specific method indicators - defaulting to M1 (most common)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
MEDIUM
Why we say so
primary_method=M1; difference-of-means is the reference extraction for M1/M3 (Arditi 2024)
Downloads · lifetime
604
254 last 30d - stable
Likes
0
Model age
3mo ago
created 2026-06-15
Downloads over time
Now699→from50↑1,298%
1826651576450 on Jun 17699 on Oct 11JunJulAugSepOct
Jun 17 → Oct 11 · 56 snapshots · spans 116 days

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
apache-2.0
Languages
en
Tags
mlx safetensors qwen3 mlx-lm apple-silicon local-llm on-device macbook abliterated uncensored 8b 4-bit

Related

Total size
4.29 GB
Files
9
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-08-22 20:27

Files by quantization

Auxiliary files 9 files 4.30 GB
model.safetensors 4.29 GB 427dd678 download
tokenizer.json 10.9 MB be756060 download
model.safetensors.index.json 62.6 KB d5bad744 download
chat_template.jinja 4.02 KB 699ff8df download
README.md 3.16 KB 7082d88f download
.gitattributes 1.53 KB 52373fe2 download
config.json 1023 B 216a8c62 download
tokenizer_config.json 384 B 397f359a download
generation_config.json 227 B a38edb1b download

README current version from Hugging Face


base_model:

  • huihui-ai/Huihui-Qwen3-8B-abliterated-v2
  • Qwen/Qwen3-8B
    base_model_relation: quantized
    license: apache-2.0
    language:
  • en
    library_name: mlx
    pipeline_tag: text-generation
    tags:
  • mlx
  • mlx-lm
  • apple-silicon
  • local-llm
  • on-device
  • macbook
  • abliterated
  • uncensored
  • 8b
  • 4-bit
  • qwen3
  • huihui

Huihui-Qwen3-8B-abliterated-v2 — 4-bit MLX for Apple Silicon

Uncensored Qwen3 8B that runs on any modern Mac. A 4-bit MLX conversion of huihui-ai/Huihui-Qwen3-8B-abliterated-v2, the abliterated Qwen3-8B. 4.6 GB on disk — comfortable on a 16 GB Mac. No cloud, no API key, no refusals.

Size on disk 4.6 GB
Mac RAM 16 GB comfortable
Layers 36
Architecture Qwen3

What "abliterated" means

The refusal direction has been orthogonalized out of the weights, so the model
answers instructions a stock instruct-tuned model would decline, while staying
coherent on ordinary tasks. The abliteration here is huihui-ai's, not mine —
this repo is the MLX conversion of their work.

This is an uncensored model. You are responsible for how you use it and for
complying with the base model's license and applicable law.

Quick start

pip install mlx-lm
mlx_lm.generate --model divinetribe/Huihui-Qwen3-8B-abliterated-v2-4bit-mlx \
  --prompt "Explain quantum entanglement to a 12 year old." --max-tokens 400
from mlx_lm import load, generate
model, tok = load("divinetribe/Huihui-Qwen3-8B-abliterated-v2-4bit-mlx")
text = tok.apply_chat_template(
    [{"role": "user", "content": "Explain quantum entanglement to a 12 year old."}],
    add_generation_prompt=True, tokenize=False)
print(generate(model, tok, prompt=text, max_tokens=400))

Also loads in LM Studio and anything else that reads MLX models.

Conversion details

  • Quantization: 4-bit affine, group size 64, via mlx_lm.convert
  • Format: MLX safetensors — no GGUF, no llama.cpp, no GPU needed
  • Runs on: M1 / M2 / M3 / M4 / M5 Macs, entirely on-device

More abliterated MLX models

Part of the Abliterated MLX for Apple Silicon
collection — Llama 3.3 70B, Gemma 4, Qwen3, Qwen3-VL, Hermes 4 and Muse Glimmer,
all converted for Apple Silicon.

Credit

Base model and abliteration by huihui-ai. MLX 4-bit conversion by
divinetribe.


Part of Claude Code Local

This model is one of the fighters in Claude Code Local (3.2k★), which runs Claude Code 100% on-device on Apple Silicon through an MLX-native Anthropic-API server. Not sure which local model to run as an agent? Check the Agent-12 local agent leaderboard: real agent tasks, judged by the filesystem, same hardware for every row.

Built by Matt Macosko in Arcata, CA. Open to work on local-AI and Apple Silicon inference: [email protected].

README history 4 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-08-22Model card: link to Claude Code Local and the Agent-12 leaderboard2cafc1c3.2 KB
    Loading...
  2. 2026-08-12Write real model card: usage, sizes, RAM, conversion details, credit7ba02eb2.5 KB
    Loading...
  3. 2026-08-12SEO: base_model lineage, language, Apple Silicon / MLX tags4f809ad321 B
    Loading...
  4. 2026-06-15Add files using upload-large-folder tool608464b81 B
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration