← back to catalog · registered 2026-08-22 13:56

divinetribe/Huihui-Qwen3-Coder-Next-Opus-4.6-Reasoning-Distilled-abliterated-4bit-mlx

divinetribe Qwen 10.0B
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/divinetribe%2FHuihui-Qwen3-Coder-Next-Opus-4.6-Reasoning-Distilled-abliterated-4bit-mlx"
Response includes
  • classification m1
  • files 17
  • hub_downloads_all_time 1,028
  • author_summary 14 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M1
Primary method

Direct removal

No other method signals detected in this model.
Confidence
MEDIUM
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=0 (base model)
  • no specific method indicators - defaulting to M1 (most common)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
MEDIUM
Why we say so
primary_method=M1; difference-of-means is the reference extraction for M1/M3 (Arditi 2024)
Downloads · lifetime
1K
16 last 30d - cooling
Likes
1
Model age
3mo ago
created 2026-06-16
Downloads over time
Now1K→from388↑166%
3566028491.1K388 on Jun 171K on Oct 111K on Oct 8JunJulAugSepOct
Jun 17 → Oct 11 · 56 snapshots · spans 116 days

Metadata

License
apache-2.0
Tags
safetensors qwen3_next deprecated do-not-use cautionary-tale license:apache-2.0 4-bit region:us

Related

Total size
41.8 GB
Files
17
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-08-22 20:27

Files by quantization

Auxiliary files 17 files 41.8 GB
model-00008-of-00009.safetensors 4.90 GB 140d16ed download
model-00004-of-00009.safetensors 4.90 GB a9e77b9d download
model-00007-of-00009.safetensors 4.90 GB 081ff554 download
model-00005-of-00009.safetensors 4.90 GB 81d2cd93 download
model-00002-of-00009.safetensors 4.90 GB d4c86ba2 download
model-00003-of-00009.safetensors 4.88 GB 1b036796 download
model-00006-of-00009.safetensors 4.88 GB 7bc69c40 download
model-00001-of-00009.safetensors 4.78 GB aa503b2f download
model-00009-of-00009.safetensors 2.73 GB c4cee043 download
tokenizer.json 10.9 MB be756060 download
model.safetensors.index.json 169 KB 4ab1dfa6 download
config.json 23.1 KB 7c0b0a94 download
chat_template.jinja 5.93 KB 1ac848e3 download
README.md 3.08 KB 3efd43c8 download
.gitattributes 1.53 KB 52373fe2 download
tokenizer_config.json 702 B b1ec3422 download
generation_config.json 187 B dc6c662e download

README current version from Hugging Face


license: apache-2.0
tags:

  • deprecated
  • do-not-use
  • cautionary-tale

⚠️ DEPRECATED — DO NOT USE THIS MODEL

This model does not work for real coding tasks. I benchmarked it in June 2026 and it failed essentially every test — one-shot and agentic. I'm leaving it up only as a documented cautionary example so nobody wastes a 42 GB download. Use a normal coder model instead (e.g. Qwen3-Coder-30B-A3B-Instruct).

What this was

An MLX 4-bit conversion of huihui-ai/Huihui-Qwen3-Coder-Next-Opus-4.6-Reasoning-Distilled-abliterated — a Qwen3-Coder-Next 80B (A3B MoE) with Claude Opus 4.6 reasoning distilled in, then abliterated. On paper it sounds great: "a coder that reasons like Opus." In practice it cannot stop reasoning long enough to actually finish anything.

The benchmark — it lost everything

Head-to-head against the plain Qwen3-Coder-30B on the same machine (Apple Silicon, MLX), same prompts:

Task Qwen3-Coder-30B This model (80B reasoning)
Asteroids game (one-shot HTML) ✅ working game ❌ wrote a planning monologue, never produced the game
Snake game ✅ 14s ✅ but 8× slower, ~7× the tokens for the same result
Calculator ✅ clean, complete ❌ bloated to 56 KB and hit the token cap unfinished
Analog clock ✅ working ❌ bloated to 56 KB and hit the token cap unfinished
Hard expression parser (18 hidden tests, no eval()) ✅ 16/18 ❌ 0/18 — produced no code at all, just an unclosed <think> block
Agentic task: write code, run it, fix until tests pass (with tools) ✅ completed ❌ called zero tools, wrote zero files, just said "DONE"

Why it fails

A reasoning model distilled onto a coder is the worst of both worlds: it can't stop reasoning to converge. On one-shot tasks it over-generates and exhausts its token budget before finishing the answer. In agentic mode — its supposedly superior use case ("tool calling + bug detection") — it won't even pick up a tool; it narrates what it would do and declares itself done. The smaller, plain 30B coder beat it on speed, on completion, and even on the hard reasoning task.

Use instead

Any standard coder, e.g. Qwen3-Coder-30B-A3B-Instruct.


I converted this on reputation — "Opus reasoning distilled into a coder!" — before actually testing it. Lesson learned: benchmark before you adopt. Sharing the result so you don't repeat my mistake. — divinetribe


Part of Claude Code Local

This model is one of the fighters in Claude Code Local (3.2k★), which runs Claude Code 100% on-device on Apple Silicon through an MLX-native Anthropic-API server. Not sure which local model to run as an agent? Check the Agent-12 local agent leaderboard: real agent tasks, judged by the filesystem, same hardware for every row.

Built by Matt Macosko in Arcata, CA. Open to work on local-AI and Apple Silicon inference: [email protected].

README history 3 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-08-22Model card: link to Claude Code Local and the Agent-12 leaderboard4bdd9893.1 KB
    Loading...
  2. 2026-06-23DEPRECATED: does not work (one-shot or agentic) — benchmark + warning. Use a ...09a32182.5 KB
    Loading...
  3. 2026-06-16Add files using upload-large-folder tool881409881 B
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration