← back to catalog · registered 2026-08-22 13:56

OpenYourMind/Qwopus3.5-122B-A10B-Kimi-K2.6-destilled-abliterated-NVFP4

OpenYourMind Qwen 59B MoE multimodal second-order
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/OpenYourMind%2FQwopus3.5-122B-A10B-Kimi-K2.6-destilled-abliterated-NVFP4"
Response includes
  • classification m1
  • files 16
  • hub_downloads_all_time 24,207
  • author_summary 11 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M1
Primary method

Direct removal

No other method signals detected in this model.
Confidence
MEDIUM
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=0 (base model)
  • no specific method indicators - defaulting to M1 (most common)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
MEDIUM
Why we say so
primary_method=M1; difference-of-means is the reference extraction for M1/M3 (Arditi 2024)
Downloads · lifetime
24K
261 last 30d - cooling
Likes
21
Model age
4mo ago
created 2026-05-24
Downloads over time
Now24.3K→from2.4K↑902%
1.3K9.7K18.1K26.5K2.4K on Jun 1024.3K on Oct 11JunJulAugSepOct
Jun 10 → Oct 11 · 58 snapshots · spans 123 days

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
mit
Tags
transformers safetensors qwen3_5_moe image-text-to-text qwen qwen3 qwen3.5 moe abliterated uncensored sft dpo

Related

Total size
75.9 GB
Files
16
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-06-16 09:28

Files by quantization

Auxiliary files 16 files 75.9 GB
model-00001-of-00002.safetensors 46.6 GB 6aa3d2e7 download
model-00002-of-00002.safetensors 24.6 GB 3f7b8643 download
model_mtp.safetensors 4.70 GB b4fc887d download
tokenizer.json 19.1 MB 06b95093 download
model.safetensors.index.json 16.4 MB 2dc4df8a download
OYM_banner.png 1.58 MB a714b89b download
config.json 25.5 KB 375c2a68 download
README.md 8.24 KB fa07e17a download
chat_template.jinja 7.57 KB a585dec8 download
.gitattributes 1.65 KB 0e10f0f4 download
processor_config.json 1.16 KB 33818c7f download
tokenizer_config.json 1.14 KB b4acebe0 download
preprocessor_config.json 390 B 2ea84a43 download
video_preprocessor_config.json 385 B 3ba673a5 download
recipe.yaml 311 B 31dfdd32 download
generation_config.json 213 B 4ec8d40d download

README current version from Hugging Face


license: mit
library_name: transformers
base_model: OpenYourMind/Qwopus3.5-122B-A10B-Kimi-K2.6-destill-healed-abliterated
base_model_relation: quantized
tags:

  • qwen
  • qwen3
  • qwen3.5
  • moe
  • abliterated
  • uncensored
  • sft
  • dpo
  • opus
  • qwopus
  • kimi
  • kimi-k2
  • distill
  • multimodal
  • vision
  • mtp
  • nvfp4
  • fp4
  • quantized
  • compressed-tensors
  • llm-compressor
  • vllm
    pipeline_tag: image-text-to-text

OpenYourMind

Support & Community

☕ If these models are useful to you, consider supporting my work — it funds compute for more & larger abliterations.

Buy Me A Coffee

buymeacoffee.com/oym.kuato

💬 Discord: discord.gg/rhUZY5GEZr  ·  ₿ Bitcoin: bc1qsvfduzj9fjs9fugpc52yver3f2g8fp7xjxecdv


Qwopus3.5-122B-A10B-Kimi-K2.6-destilled-abliterated-NVFP4

Overview

4-bit NVFP4 quantization of OpenYourMind/Qwopus3.5-122B-A10B-Kimi-K2.6-destill-healed-abliterated — the Kimi-K2.6-distilled, reasoning-DPO-healed, abliterated/uncensored evolution of Qwen/Qwen3.5-122B-A10B (Mixture of Experts, ~10B active / 122B total).

This build packs the transformer weights to NVFP4 with LLM Compressor, cutting the on-disk footprint from ~250 GB to ≈82 GB while keeping the vision tower, MTP head, router gates, and the Gated-DeltaNet attention path in higher precision. It is multimodal (image + text), uncensored, and — despite 4-bit weights — beats the full-precision Qwen3.5-122B-A10B baseline on every benchmark we ran (see Evaluation).

It loads anywhere compressed-tensors is supported and is auto-detected by vLLM (no --quantization flag needed).

Evaluation

Scores below were measured on this NVFP4 build and compared against the full-precision (BF16) Qwen/Qwen3.5-122B-A10B baseline:

Benchmark Qwen3.5-122B-A10B (BF16, baseline) Qwopus3.5 NVFP4 (this model)
CTI 64.8 71.5
LiveCodeBench 78.9 79.9
BFCL 72.2 85.6

Even after 4-bit (NVFP4) weight quantization, this model outperforms the BF16 Qwen3.5-122B-A10B baseline on all three benchmarks — the Kimi-K2.6 distillation + reasoning-DPO healing more than offsets any quantization loss. BFCL is the Berkeley Function-Calling Leaderboard (tool use); LiveCodeBench is contamination-controlled code generation.

Quantization (NVFP4)

Produced with LLM Compressor using the QuantizationModifier recipe shipped in this repo (recipe.yaml).

  • Scheme: NVFP4 (format: nvfp4-pack-quantized) — 4-bit float weights in micro-blocks of 16, each block carrying an FP8 (float8_e4m3fn) scale. Weights are static; input activations are quantized dynamically (per-group, static-minmax).
  • Quantized: all transformer Linear layers — attention projections and the 256 routed-expert MoE FFNs (37,056 packed weight tensors).
  • Left in higher precision (BF16): the vision tower (visual.* — 333 tensors), the MTP head (model_mtp.safetensors — 785 tensors), lm_head, token embeddings, the MoE router gates (mlp.gate, shared_expert_gate), and the Gated-DeltaNet linear-attention path (linear_attn.*).
  • Architecture preserved: Qwen3_5MoeForConditionalGeneration / model_type: qwen3_5_moe, so the checkpoint loads as a drop-in replacement for the base at the architecture level.

Downloads / Other Formats

Format Repo Use it for
Full BF16 weights Qwopus3.5-122B-A10B-Kimi-K2.6-destill-healed-abliterated Transformers / vLLM, fine-tuning, requantizing
NVFP4 (this repo) Qwopus3.5-122B-A10B-Kimi-K2.6-destilled-abliterated-NVFP4 vLLM on a single ≥96 GB / Blackwell accelerator (vision + MTP included)
GGUF (Q4_K_M) …-Kimi-K2.6-destill-healed-abliterated-GGUF llama.cpp / LM Studio (text-only). MTP head included.
MLX 4-bit …-Kimi-K2.6-destill-healed-abliterated-MLX-4bit Apple Silicon / LM Studio (vision supported)

Files

File Description Size
model-00001-of-00002.safetensors NVFP4-packed language weights (4-bit + FP8 scales) + lm_head ~50.0 GB
model-00002-of-00002.safetensors NVFP4-packed language weights (tail) + BF16 vision tower ~26.4 GB
model_mtp.safetensors BF16 MTP head (785 tensors, 1 hidden layer) ~5.0 GB
model.safetensors.index.json Combined weight map —
config.json Multimodal config incl. quantization_config (nvfp4-pack-quantized) —
recipe.yaml LLM Compressor quantization recipe —
tokenizer*, chat_template.jinja, generation_config.json, *preprocessor_config.json Standard —

Total on disk: ≈81.5 GB (~76 GiB).

Usage (vLLM)

vLLM auto-detects the NVFP4 compressed-tensors format — no --quantization flag.

vllm serve OpenYourMind/Qwopus3.5-122B-A10B-Kimi-K2.6-destilled-abliterated-NVFP4 \
  --tool-call-parser qwen3_coder \
  --reasoning-parser qwen3 \
  --max-model-len 262144

The checkpoint ships the MTP head, so you can enable 1-token speculative decoding:

  --speculative-config '{"num_speculative_tokens":1}'

Tip (Qwen3.5 MoE / Gated-DeltaNet): if torch.compile errors in the GDN path during startup, add --compilation-config '{"use_inductor_graph_partition":true}'.

Text + vision both work through AutoProcessor / AutoModelForImageTextToText (via the compressed-tensors integration) for non-vLLM workflows.

Vision & MTP

Both the vision tower and the MTP (multi-token-prediction) head are included and kept in BF16.

  • Vision works as expected (image / video → text).
  • MTP: the head is present and shape-compatible. It enables speculative decoding under vLLM, but on the upstream checkpoint it produced little measurable speedup/quality gain and would benefit from retraining — shipped intact for completeness and forward-compatibility.

Hardware

The NVFP4 weights are ≈82 GB (vs ~250 GB for the BF16 release), so the model runs on a single accelerator with ≥ 96 GB: H200, B200, RTX PRO 6000 Blackwell, or a 128 GB unified-memory NVIDIA DGX Spark / GB10. Native FP4 math requires a Blackwell GPU (compute capability ≥ 10.0 / sm_120+); on other hardware vLLM runs NVFP4 via FlashInfer/emulation.

Notes

Thanks

  • Jackrong — for the idea of Qwopus merges (Opus distillations on Qwen models).
  • wangzhang — for the wonderful abliterix framework, which was customized to do this abliteration.
  • The LLM Compressor and vLLM teams for the NVFP4 tooling.

Disclaimer

Use is the responsibility of the user. Ensure your usage complies with applicable laws, platform rules, and deployment requirements.

README history 4 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-06-16Move Support & Community section to top (below banner); unify across models2c093f18.2 KB
    Loading...
  2. 2026-06-08Add OYM banner to top of model card96c1cbe8.2 KB
    Loading...
  3. 2026-06-08Add highlighted Buy Me a Coffee support section50379468.1 KB
    Loading...
  4. 2026-05-24Add NVFP4 model cardcd8d84a7.7 KB
    Loading...

Discussions 2 threads

  1. 2026-08-24Need Help Setting Up Model Inside Molabopen4 💬#2
    Loading...
  2. 2026-08-11Excellent IT Security Performance on RTX PRO 6000 — Surprisingly Reliable for a…open5 💬#1
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration