← back to catalog · registered 2026-08-22 13:56

OpenYourMind/Qwopus3.5-122B-A10B-Kimi-K2.6-destill-healed-abliterated-MTPLX-4bit

OpenYourMind Qwen 122B MoE multimodal second-order
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/OpenYourMind%2FQwopus3.5-122B-A10B-Kimi-K2.6-destill-healed-abliterated-MTPLX-4bit"
Response includes
  • classification m4
  • files 27
  • hub_downloads_all_time 12,053
  • author_summary 11 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M4
Primary method

Abliterate + heal

Applied on top of direct removal inherited from the base model.
Confidence
MEDIUM
Inherited from base model
Why this label 2 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'healed'/'orpo'/'dpo' in name suggests heal step after abliteration
  • M4 = abliterate + heal pipeline
Refusal direction extraction

No specific extraction method could be identified for this model. The producer either did not document it or used a proprietary pipeline.

What is a refusal direction? →
Downloads · lifetime
12K
673 last 30d - cooling
Likes
8
Model age
4mo ago
created 2026-05-25
Downloads over time
Now12.3K→from2.5K↑395%
2K5.8K9.5K13.3K2.5K on Jun 1012.3K on Oct 11JunJulAugSepOct
Jun 10 → Oct 11 · 58 snapshots · spans 123 days

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
other
Tags
mlx safetensors qwen3_5_moe mtplx mtp speculative-decoding qwen qwen3 qwen3.5 moe abliterated uncensored

Related

Total size
69.5 GB
Files
27
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-06-16 09:28

Files by quantization

Auxiliary files 27 files 69.6 GB
model-00011-of-00014.safetensors 4.84 GB fdec5259 download
model-00008-of-00014.safetensors 4.84 GB 11b1fbd7 download
model-00005-of-00014.safetensors 4.84 GB f527899e download
model-00012-of-00014.safetensors 4.84 GB 0adcf79a download
model-00006-of-00014.safetensors 4.84 GB 5e9222d3 download
model-00009-of-00014.safetensors 4.84 GB bad0a01f download
model-00003-of-00014.safetensors 4.84 GB 06f685c2 download
model-00002-of-00014.safetensors 4.84 GB 571a4927 download
model-00013-of-00014.safetensors 4.80 GB e92cb3a4 download
model-00004-of-00014.safetensors 4.79 GB 685e89af download
model-00007-of-00014.safetensors 4.79 GB 46aa2ff8 download
model-00010-of-00014.safetensors 4.79 GB 6f3b2563 download
model-00001-of-00014.safetensors 4.77 GB 8b85826a download
mtp.safetensors 4.70 GB 15c8ecba download
model-00014-of-00014.safetensors 2.14 GB e68bd36d download
tokenizer.json 19.1 MB 639e352c download
OYM_banner.png 1.58 MB a714b89b download
model.safetensors.index.json 247 KB 00324b36 download
config.json 22.9 KB e500d948 download
chat_template.jinja 7.57 KB a585dec8 download
README.md 5.65 KB 7c3b242d download
.gitattributes 1.58 KB b6da95e2 download
processor_config.json 1.27 KB 7ad6acdf download
tokenizer_config.json 1.24 KB 9de16b5b download
preprocessor_config.json 390 B 2ea84a43 download
video_preprocessor_config.json 385 B 3ba673a5 download
generation_config.json 213 B 318011ae download

README current version from Hugging Face


license: other
library_name: mlx
base_model: OpenYourMind/Qwopus3.5-122B-A10B-Kimi-K2.6-destill-healed-abliterated
tags:

  • mlx
  • mtplx
  • mtp
  • speculative-decoding
  • qwen
  • qwen3
  • qwen3.5
  • moe
  • abliterated
  • uncensored
  • dpo
  • opus
  • qwopus
  • kimi
  • kimi-k2
  • multimodal
  • vision
  • 4-bit
    pipeline_tag: image-text-to-text

OpenYourMind

Support & Community

☕ If these models are useful to you, consider supporting my work — it funds compute for more & larger abliterations.

Buy Me A Coffee

buymeacoffee.com/oym.kuato

💬 Discord: discord.gg/rhUZY5GEZr  ·  ₿ Bitcoin: bc1qsvfduzj9fjs9fugpc52yver3f2g8fp7xjxecdv


Qwopus3.5-122B-A10B-Kimi-K2.6-destill-healed-abliterated — MTPLX 4-bit (MoE MTP head)

Overview

This is the MLX 4-bit build of OpenYourMind/Qwopus3.5-122B-A10B-Kimi-K2.6-destill-healed-abliterated with the Qwen3.5 MoE Multi-Token-Prediction (MTP) head included, packaged for MTPLX native MTP speculative decoding on Apple Silicon.

The language/vision weights are byte-identical to the -MLX-4bit build. The only additions are the MTP head (mtp.safetensors, BF16, 4.7 GB) and a config.json pointer (mlx_lm_extra_tensors.mtp_file). That sidecar is ignored by plain mlx-lm/mlx-vlm, so this folder still loads as an ordinary MLX model — but with MTPLX it also drives speculative decoding.

  • Language: 4-bit, group size 64 (MoE routing gates kept at higher precision by the model's quant predicate), ≈ 4.5 bits/weight.
  • MTP head: 1 layer, MoE (router + 256 experts / 8 active + shared expert), full self-attention, BF16 (785 tensors). MTPLX stacks the experts into switch_mlp at load and verifies every drafted token against the target model.
  • Vision: the BF16 vision tower from the base build is still present; MTPLX runs the text path only. For image input, use the -MLX-4bit repo with mlx-vlm.

⚠️ Requires MTPLX with Qwen3.5-MoE MTP support

The Qwen3.5 MTP head is an MoE block. MTPLX ≤ 0.3.7 only supported a dense Qwen MTP head and will reject this model with invalid-mtp-tensor-layout. Support is added in MTPLX PR #84.

Until that lands in a release, install from the branch:

pip install "git+https://github.com/janfeddersen-wq/MTPLX.git@qwen3-5-moe-mtp"
# after the PR is merged & released:  pip install -U mtplx

Usage

MODEL=OpenYourMind/Qwopus3.5-122B-A10B-Kimi-K2.6-destill-healed-abliterated-MTPLX-4bit

# one-shot, with acceptance stats
mtplx ask --model "$MODEL" --prompt "Explain Rayleigh scattering simply." --mtp --stats --yes

# interactive terminal chat
mtplx start cli --model "$MODEL" --yes

# OpenAI-compatible server
mtplx quickstart --model "$MODEL" --port 8000 --yes

--yes accepts the "family-compatible-unverified" gate (no recorded exactness baseline is shipped). Add --no-mtp to compare against plain autoregressive decoding.

Measured (M5 Max, 128 GB)

120-token greedy run: depth-1 acceptance ≈ 70 %, accepted_by_depth = [40, 19, 3] of [57, 57, 56] drafted → 120 tokens in 57 target verify passes (≈ 2.1 tokens/verify), ~52 decode tok/s. (Contrary to the earlier note on the base card, the MoE MTP head does yield a real speedup once a runtime can consume it.)

Known limitation — MoE exactness

At temperature 0, MTP vs non-MTP greedy output is ~98 % identical and re-converges immediately, but occasionally flips a single token. This is the MoE router hitting a near-tie that resolves differently under batched verification vs single-token decode (an inherent MoE/FP effect), not a drafting error — the target model verifies every token. Strict bit-exactness for MoE heads is still being worked out (e.g. fp32 router logits during verify); see PR #84.

Files

File Description Size
model-*-of-00014.safetensors 4-bit language weights + BF16 vision tower ~65 GB
mtp.safetensors MoE MTP head (BF16) 4.7 GB
config.json Qwen3_5MoeForConditionalGeneration + quantization + mlx_lm_extra_tensors.mtp_file —
tokenizer*, chat_template.jinja, generation_config.json, processor configs Standard —

Total on disk: ~70 GB.

Hardware

Needs roughly ≥ 80 GB unified memory to load with usable context (65 GB base + ~5 GB BF16 MTP + KV cache). Runs comfortably on 96 GB+ M-series Macs.

Notes

Disclaimer

Use is the responsibility of the user. Ensure your usage complies with applicable laws, platform rules, and deployment requirements.

README history 4 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-06-16Move Support & Community section to top (below banner); unify across models7600ece5.7 KB
    Loading...
  2. 2026-06-08Add OYM banner to top of model card7c113405.6 KB
    Loading...
  3. 2026-06-08Add highlighted Buy Me a Coffee support sectiond13e2595.5 KB
    Loading...
  4. 2026-05-25Add files using upload-large-folder tool6e0876f5.1 KB
    Loading...

Discussions 1 thread

  1. 2026-07-23MPTLXopen1 💬#1
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration