← back to catalog · registered 2026-08-22 13:56

OpenYourMind/Qwopus3.5-122B-A10B-Kimi-K2.6-destill-healed-abliterated

OpenYourMind Qwen 125B MoE multimodal
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/OpenYourMind%2FQwopus3.5-122B-A10B-Kimi-K2.6-destill-healed-abliterated"
Response includes
  • classification m4
  • files 16
  • benchmarks 11 entries
  • hub_downloads_all_time 2,712
  • author_summary 11 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M4
Primary method

Abliterate + heal

Applied on top of direct removal inherited from the base model.
Confidence
MEDIUM
Inherited from base model
Why this label 2 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'healed'/'orpo'/'dpo' in name suggests heal step after abliteration
  • M4 = abliterate + heal pipeline
Refusal direction extraction

No specific extraction method could be identified for this model. The producer either did not document it or used a proprietary pipeline.

What is a refusal direction? →
Downloads · lifetime
3K
124 last 30d - cooling
Likes
17
Descendants
10
in 10 direct forks
Model age
4mo ago
created 2026-05-20
Downloads over time
Now2.8K→from116↑2,278%
01K2K3K116 on May 202.8K on Oct 11MayJunJulAugSepOct
May 20 → Oct 11 · 61 snapshots · spans 144 days

Benchmarks

Portrait before abliteration
Benchmarks of the base model as it stood before the refusal-removal operation. Compare with the numbers above to see what the operation cost.
Benchmark Score Source
Entertainment 1.4 UGI
Hazardous 1.8 UGI
Natural Intelligence 31.08 UGI
Political lean -21.9% UGI
Sensitive-Info 17.81 UGI
SocPol 2.3 UGI
UGI 17.71 UGI
Willingness (10) 1.8 UGI
W10-Adherence 1.5 UGI
W10-Direct 2 UGI
Writing 39.54 UGI

Genealogy 10 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Variants by this author 2 formats · 2K downloads combined

The same weights this author released in different packaging. Pick the format that matches your runtime.

Metadata

License
mit
Tags
transformers safetensors qwen3_5_moe image-text-to-text qwen qwen3 qwen3.5 moe abliterated uncensored sft dpo

Related

Total size
233 GB
Files
16
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-06-16 09:28

Files by quantization

Auxiliary files 16 files 233 GB
model-00003-of-00006.safetensors 45.3 GB 19311165 download
model-00005-of-00006.safetensors 45.1 GB d627f86d download
model-00001-of-00006.safetensors 45.1 GB 2e29d972 download
model-00004-of-00006.safetensors 43.6 GB 00b61f10 download
model-00002-of-00006.safetensors 43.6 GB 64ebd75d download
model-00006-of-00006.safetensors 5.51 GB 2309c03d download
model-mtp-official.safetensors 4.70 GB 15c8ecba download
tokenizer.json 19.1 MB 639e352c download
model.safetensors.index.json 3.88 MB cc8b932f download
OYM_banner.png 1.58 MB a714b89b download
README.md 9.57 KB 4565ae2e download
chat_template.jinja 7.57 KB a585dec8 download
config.json 3.35 KB 43e57ae6 download
.gitattributes 1.58 KB b6da95e2 download
tokenizer_config.json 1.20 KB b44c9480 download
generation_config.json 213 B 318011ae download

README current version from Hugging Face


license: mit
library_name: transformers
base_model: Qwen/Qwen3.5-122B-A10B
tags:

  • qwen
  • qwen3
  • qwen3.5
  • moe
  • abliterated
  • uncensored
  • sft
  • dpo
  • opus
  • qwopus
  • kimi
  • kimi-k2
  • distill
  • multimodal
  • vision
  • mtp
    pipeline_tag: image-text-to-text

OpenYourMind

Support & Community

☕ If these models are useful to you, consider supporting my work — it funds compute for more & larger abliterations.

Buy Me A Coffee

buymeacoffee.com/oym.kuato

💬 Discord: discord.gg/rhUZY5GEZr  ·  ₿ Bitcoin: bc1qsvfduzj9fjs9fugpc52yver3f2g8fp7xjxecdv


Qwopus3.5-122B-A10B-Kimi-K2.6-destill-healed-abliterated

Overview

Full BF16 weights of Qwopus3.5-122B-A10B-Kimi-K2.6-destill-healed-abliterated — the Kimi-K2.6-distilled, reasoning-DPO-healed evolution of OpenYourMind/Qwopus3.5-122B-A10B-abliterated-uncensored, itself an abliterated and supervised-finetuned variant of Qwen/Qwen3.5-122B-A10B (Mixture of Experts, ~10B active / 122B total). The model is uncensored, multimodal (image + text), and ships with the vision tower and MTP head intact so it is a drop-in replacement for the original base model at the architecture level.

The pipeline:

  1. Refusal Ablation — Residual-stream refusal directions (one per decoder layer, layers 19–45) were extracted via diff-in-means on a labeled prompt set and baked into the weights as a per-matrix delta — see the abliterix framework for the methodology.
  2. Healing — Stage A: Constrained-LoRA SFT on Opus reasoning data — Supervised finetuned on a curated set of Claude Opus reasoning traces (single-turn, ~8k rows). To keep the abliteration mathematically intact during training, a custom orthogonality projection is applied to every LoRA B-matrix on residual-write modules after each optimizer step (B := B − r·(rᵀB)), so the LoRA update is forbidden from re-introducing the refusal direction. LoRA rank 32, α 64, 54 protected modules across 27 decoder layers. Verified residual after training: max ‖rᵀB‖₂ = 8.5 × 10⁻¹⁰.
  3. Healing — Stage B: Unconstrained SFT on chosen completions — A second short SFT pass (LoRA r=16, α 32, no orthogonality constraint) on the chosen answers (including reasoning chains) from an internal preference dataset, to tighten on the deployment distribution and remove the last bits of drift introduced by Stage A.
  4. Kimi K2.6 Reasoning DPO — A targeted preference-optimization pass distilled from Kimi K2.6 to improve reasoning verbosity and eliminate degenerate looping. See the dedicated section below.
  5. Vision + MTP Restoration — The original Qwen3.5 vision tower (333 tensors, depth 27, hidden 1152) and MTP head (785 tensors, 1 hidden layer) were grafted back from the upstream Qwen/Qwen3.5-122B-A10B shards. Tensor names, shapes, and config.json schema (Qwen3_5MoeForConditionalGeneration, model_type: qwen3_5_moe) match the base model exactly — so this checkpoint loads anywhere the original loads.

Key Properties:

  • Uncensored across the standard refusal axes
  • Reasoning preserved and improved (Opus-style think-then-answer + Kimi K2.6 reasoning DPO)
  • Fewer looping / repetition failures on long conversations
  • Multimodal: vision (image / video) and MTP heads carried forward
  • Drop-in shape compatibility with Qwen/Qwen3.5-122B-A10B

Kimi K2.6 Reasoning DPO

On top of the base abliteration + Opus healing, this release adds a focused healing pass built from Kimi K2.6:

  • ~3,000 samples distilled from Kimi K2.6 were used for DPO (Direct Preference Optimization), alongside synthetic datasets also generated from Kimi K2.6.
  • Improved reasoning verbosity — the model now produces more complete, better-structured reasoning on the ~12% of requests where the previous release tended to under-explain or cut its chain-of-thought short.
  • Fixed looping / repetition — degenerate loops that appeared on 2–6% of long-tail conversations (long context, multi-turn) were largely eliminated.

The DPO pass targets the language model's reasoning behavior only; the abliteration, vision tower, and MTP head are unchanged by this step.

Evaluation

This model family outperforms the full-precision (BF16) Qwen/Qwen3.5-122B-A10B baseline across reasoning, coding, and tool-use benchmarks:

Benchmark Qwen3.5-122B-A10B (BF16, baseline) Qwopus3.5-122B-A10B
CTI 64.8 71.5
LiveCodeBench 78.9 79.9
BFCL 72.2 85.6

BFCL is the Berkeley Function-Calling Leaderboard (tool use); LiveCodeBench is contamination-controlled code generation.

The Qwopus figures above were measured on the NVFP4 build (4-bit weights); these full-precision BF16 weights match or exceed them. Even after 4-bit quantization the model stays ahead of the BF16 Qwen3.5-122B-A10B baseline.

Downloads / Other Formats

Format Repo Use it for
Full BF16 weights (this repo) Qwopus3.5-122B-A10B-Kimi-K2.6-destill-healed-abliterated Transformers / vLLM, fine-tuning, requantizing
NVFP4 (4-bit, ≈82 GB) Qwopus3.5-122B-A10B-Kimi-K2.6-destilled-abliterated-NVFP4 vLLM on a single ≥96 GB / Blackwell accelerator (vision + MTP included)
GGUF (Q4_K_M) …-Kimi-K2.6-destill-healed-abliterated-GGUF llama.cpp / LM Studio (text-only). MTP head included — see note below.
MLX 4-bit …-Kimi-K2.6-destill-healed-abliterated-MLX-4bit Apple Silicon / LM Studio (vision supported)

Files

File Description Size
model-0000{1..5}-of-00006.safetensors BF16 language + vision weights (48 decoder layers, MoE with 256 routed experts + shared expert per layer; Qwen3-VL vision tower folded into the shards) ~47–49 GB each
model-00006-of-00006.safetensors BF16 tail tensors ~5.9 GB
model-mtp-official.safetensors BF16 MTP head (785 tensors, 1 hidden layer) ~5.0 GB
model.safetensors.index.json Combined weight map —
config.json Multimodal config (Qwen3_5MoeForConditionalGeneration, model_type: qwen3_5_moe) —
tokenizer*, chat_template.jinja, generation_config.json Standard —

Total on disk: ~250 GB (233 GiB).

Usage

from transformers import AutoModelForImageTextToText, AutoProcessor

repo = "OpenYourMind/Qwopus3.5-122B-A10B-Kimi-K2.6-destill-healed-abliterated"
model = AutoModelForImageTextToText.from_pretrained(repo, dtype="bfloat16", device_map="auto")
processor = AutoProcessor.from_pretrained(repo)

messages = [{"role": "user", "content": [
    {"type": "image", "url": "path/to/image.jpg"},
    {"type": "text",  "text": "Describe this image in detail."},
]}]
inputs = processor.apply_chat_template(
    messages, add_generation_prompt=True, tokenize=True,
    return_tensors="pt", return_dict=True,
).to(model.device)
out = model.generate(**inputs, max_new_tokens=512)
print(processor.batch_decode(out, skip_special_tokens=True)[0])

Text-only inference works through the same class; if you don't need vision/MTP, you can also load just the language model with AutoModelForCausalLM.

Vision & MTP

Both the vision tower and the MTP (multi-token-prediction) head are included in these weights.

  • Vision works as expected (image / video → text).
  • MTP: the head is present and shape-compatible, but in our testing it produced no measurable speedup or quality gain on this checkpoint. It is shipped intact for completeness and forward-compatibility, but would need to be retrained to be useful — happy to do so if there is interest in the model.

Hardware

Full BF16 weights — fits comfortably on 2× H200 or 4× H100 (80 GB) with room for context. Single-node inference targets ≥ 130 GB total accelerator memory. For Apple Silicon, use the MLX 4-bit build linked above.

Notes

  • License: Other (inherits from the Qwen3.5 base license)
  • Base Model: Qwen/Qwen3.5-122B-A10B
  • Healing: Opus reasoning SFT + Kimi K2.6 reasoning DPO (≈3,000 distilled samples + synthetic data)
  • Modality: Text + Vision (image / video) + MTP
  • Architecture: Qwen3 MoE (~10B active / 122B total) + Qwen3-VL vision tower + MTP head

Thanks

  • Jackrong — for the idea of Qwopus merges (Opus distillations on Qwen models).
  • wangzhang — for the wonderful abliterix framework, which was customized to do this abliteration.

Disclaimer

Use is the responsibility of the user. Ensure your usage complies with applicable laws, platform rules, and deployment requirements.

README history 6 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-06-16Move Support & Community section to top (below banner); unify across models22711d79.6 KB
    Loading...
  2. 2026-06-08Add OYM banner to top of model card66ddf8a9.5 KB
    Loading...
  3. 2026-06-08Add highlighted Buy Me a Coffee support section86160bd9.4 KB
    Loading...
  4. 2026-05-24Add NVFP4 quant to Downloads + Evaluation table3253c5c9 KB
    Loading...
  5. 2026-05-21Change license to MITaf86dc18 KB
    Loading...
  6. 2026-05-20Add model card5acc7138 KB
    Loading...

Discussions 2 threads

  1. 2026-05-26Functional MTPopen3 💬#2
    Loading...
  2. 2026-05-20Changing the license from other to anything elseopen3 💬#1
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration