← back to catalog · registered 2026-10-04 15:58

cbert33/MiMo-V2.6-Flash-MOPD-Heretic-Abliterated-EXL3-DGX-Sliced-Calibrated

cbert33 multimodal
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/cbert33%2FMiMo-V2.6-Flash-MOPD-Heretic-Abliterated-EXL3-DGX-Sliced-Calibrated"
Response includes
  • classification m3
  • files 43
  • author_summary 8 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M3
Primary method

Layer-wise ablation

Applied on top of direct removal inherited from the base model.
Confidence
HIGH
Inherited from base model
Why this label 2 signals
Producer identity confirmed by naming conventions, tags or the model card. This label is very unlikely to change.
  • 'heretic' in model name (Heretic-produced)
  • Heretic uses layer-wise optimization (M3) with underlying direction removal (M1)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
HIGH
Why we say so
name contains 'heretic'; Heretic default extraction is difference-of-means (Arditi 2024)
Downloads · 30-day
0
Likes
1
Model age
1d ago
created 2026-10-02

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
mit
Languages
en zh
Tags
transformers safetensors mimo_v2 text-generation multimodal heretic heretic-dgx dgx-spark uncensored abliterated exl3 fp8

Related

Total size
148 GB
Files
43
Quantizations
1
Registered
2026-10-04 15:58
Last updated on HF
2026-10-04 15:34

Files by quantization

Auxiliary files 43 files 148 GB
model-00001-of-00024.safetensors 7.42 GB 5549c310 download
model-00007-of-00024.safetensors 6.12 GB a1e38eea download
model-00008-of-00024.safetensors 6.12 GB 1a1804ba download
model-00010-of-00024.safetensors 6.12 GB 8c864aaf download
model-00011-of-00024.safetensors 6.12 GB 523136b6 download
model-00013-of-00024.safetensors 6.12 GB 94c9ae53 download
model-00014-of-00024.safetensors 6.12 GB a9722b19 download
model-00016-of-00024.safetensors 6.12 GB 40edaca2 download
model-00017-of-00024.safetensors 6.12 GB 27de053b download
model-00019-of-00024.safetensors 6.12 GB 4a99749f download
model-00020-of-00024.safetensors 6.12 GB 98c6aeef download
model-00022-of-00024.safetensors 6.12 GB 4b20f780 download
model-00023-of-00024.safetensors 6.12 GB 151fc791 download
model-00005-of-00024.safetensors 6.12 GB 5b30fd13 download
model-00002-of-00024.safetensors 6.12 GB 29820e5f download
model-00004-of-00024.safetensors 6.12 GB 00b81ad5 download
model-00006-of-00024.safetensors 6.12 GB 67b91dc9 download
model-00009-of-00024.safetensors 6.12 GB 5852dccb download
model-00012-of-00024.safetensors 6.12 GB 2ca11dc1 download
model-00015-of-00024.safetensors 6.12 GB d1921432 download
model-00018-of-00024.safetensors 6.12 GB 2313170d download
model-00021-of-00024.safetensors 6.12 GB 0eaf9202 download
model-00003-of-00024.safetensors 6.12 GB c8bc7e30 download
model-00024-of-00024.safetensors 5.95 GB 149b2d43 download
kv_cache_scales.safetensors 8.70 KB 92b2d73e download
model.safetensors.index.json 26.2 MB 0c8a575c download
tokenizer.json 10.9 MB ff15eb92 download
vocab.json 2.65 MB 4783fe10 download
merges.txt 1.59 MB 20024bfe download
quantization_config.json 191 KB aa793b30 download
modeling_mimo_v2.py 83.5 KB 40225ab4 download
README.md 15.3 KB 88d15ea6 download
KV_CACHE_CALIBRATION.json 14.0 KB bef1f04a download
tokenizer_config.json 11.9 KB 39a944cb download
configuration_mimo_v2.py 9.81 KB bb6f4472 download
config.json 8.60 KB b58efb1a download
SHA256SUMS 3.92 KB a53b95a9 download
chat_template.jinja 3.78 KB 92597ade download
.gitattributes 1.73 KB 46145612 download
conversion-args.json 755 B c23c7ee7 download
rank-slice-manifest.json 566 B c9561c5d download
preprocessor_config.json 350 B f411f0c1 download
generation_config.json 195 B 167a2b07 download

README current version from Hugging Face


base_model:

  • XiaomiMiMo/MiMo-V2.6-Flash-MOPD
    base_model_relation: finetune
    license: mit
    language:
  • en
  • zh
    tags:
  • text-generation
  • multimodal
  • heretic
  • heretic-dgx
  • dgx-spark
  • uncensored
  • abliterated
  • exl3
  • fp8
  • a8
  • vllm
  • sm120
  • vision-language
  • audio
  • agent
  • video-understanding
  • long-context
  • mimo_v2
  • transformers
    library_name: transformers
    pipeline_tag: image-text-to-text

MiMo-V2.6-Flash-MOPD Heretic EXL3, rank-sliced for use on two DGX Spark nodes

This model is 309B total parameters / 15B active parameters, Hugging Face
has a bug in how it determines parameters on EXL3 quants.

This repository is a finetune / quant of
XiaomiMiMo/MiMo-V2.6-Flash-MOPD.
It combines a narrow Heretic transformation (KL divergence of 0.0011), EXL3 quantization, lossless TP2
rank slicing, calibrated FP8 KV-cache scales, and the upstream (unquantized) DFlash and audio
sidecars. Xiaomi's original model card is preserved at the bottom here.

It's taken a lot of work to get this going - we've corrected the dflash, an issue
with our quantization, and a ton of issues with vLLM. But we're confident in the
model and runner now. Use the latest version of my runner to support this model:
'cb-vllm-dual-dgx'.

(note that this runner also supports my GLM 5.3 Flash model:
'GLM-5.3-Flash-Uncensored-EXL3-DGX-Sliced'
)

Caveat: you have to keep async scheduling off for dflash to work. And don't set
more than 4 draft models - it starts to get weird when you go too high. We're still
testing this and waiting for a few specific upstream PR's for vLLM to get this next speed bump.

Uncensored model: the language checkpoint has undergone abliteration
to reduce refusal behavior. Treat outputs as untrusted, apply application-level
safeguards, and do not assume the model will decline harmful requests.

User responsibility: this model is provided without warranty. The
creators, uploaders, and maintainers are not responsible or liable for what
others generate, publish, deploy, or otherwise do with this abliterated model.
Users must operate it responsibly, apply appropriate safeguards, comply with
applicable law, and respect third-party rights. This model is for research
purposes only and is not intended for production use.

What changed

  • The Heretic pass targeted only attn.o_proj across all 48 transformer
    layers. No other module class was selected for ablation.
  • The transformed target was quantized to EXL3 with the MCG codebook and then
    sliced for tensor parallelism across two ranks without dequantizing or
    requantizing the tensors.
  • Static FP8 E4M3 K/V scales are included for all 48 layers.
  • Xiaomi's matching DFlash drafter, audio tokenizer, and model-card assets are
    included from the pinned upstream revision.
  • The MTP weights are already quantized and indexed inside the 24 target
    shards: the final index contains 114 model.mtp.* tensors. The separate
    upstream model_mtp.safetensors is therefore intentionally not duplicated.

Heretic settings and observed result

The transformation used Heretic v0.1.1
with 100 mlabonne/harmless_alpaca prompts and 100
mlabonne/harmful_behaviors prompts. Keyword-rate evaluation used the 100
harmful prompts; KL-divergence evaluation used 20 harmless prompts.

Setting Value
Trials 1
Target modules 48 × attn.o_proj
direction_index 38.49
attn.o_proj.max_weight 0.86
attn.o_proj.max_weight_position 46.53
attn.o_proj.min_weight 0.10
attn.o_proj.min_weight_distance 1.01

The selected trial produced KL divergence 0.0011. The keyword metric was
0/100 before and after the pass, so it was saturated at baseline and should
not be read as a behavioral or safety score. The useful measured result is the
small first-token distribution shift while limiting the edit to the attention
output projections.

Quantization and artifact result

Setting Value
EXL3 version 1.5.3
Average target bitrate 4.0 bpw
Codebook MCG
LM-head bitrate 8 bits
MTP bitrate 4 bits
Vision bitrate 16 bits
Calibration matrix 250 rows × 2,048 columns
Hessian regularization 0.025
Output scales automatic
Tensor-parallel slices 2

The rank-sliced EXL3 payload contains 290,356 tensors and 158,980,524,144
bytes across 24 target shards. FP8 KV calibration adds 96 scale tensors: one K
and one V scale for each of 48 layers. Those scales were derived from a
17,468-token calibration capture with eight observations per layer per rank,
using a 10% headroom factor. The final index contains 290,452 tensors.

conversion-args.json, quantization_config.json,
rank-slice-manifest.json, KV_CACHE_CALIBRATION.json, and SHA256SUMS
record the conversion, slicing, calibration, and file-integrity details.

DFlash validation

The included Xiaomi drafter is served with eight proposals per verification
step. In a 305-step validation run against this exact Heretic EXL3 target, the
first four proposal positions were accepted at 71.1%, 40.0%, 22.0%, and 11.1%.
Overall acceptance across all eight positions was 20.5% (491 of 2,396 draft
tokens). Text generation, structured tool calls, continuation after tool
results, and long-prefix cache reuse also passed. These are validation
observations from one two-node DGX Spark deployment, not general benchmark
claims; broader performance testing remains in progress.

Custom runner

Stock vLLM and Transformers do not understand this rank-sliced tensor
schema. To run this we've created a custom version of vLLM using all native
sm120/sm121 libraries. Use: cbert33/cb-vllm-dual-dgx.

It is a vLLM 0.30-based two-node DGX Spark runner with
rank-sliced EXL3/Trellis loading, DFlash, no async scheduling, FP8 KV
cache, prefix caching, reasoning, and structured tool-call support. Both ranks
must use the same runner revision, image, checkpoint revision, drafter, and
runtime flags. The runner repository contains the build and deployment
contract; this repository contains the model artifacts.

Companion artifacts

  • dflash/: Xiaomi's matching DFlash configuration, implementation, mask
    embedding, index, and draft weights.
  • audio_tokenizer/: Xiaomi's matching audio tokenizer configuration,
    template, generation settings, tokenizer settings, and weights.
  • assets/: the architecture and tool-call-repetition figures referenced by
    the upstream model card.

Original Xiaomi model card



Xiaomi-MiMo


Community
WeChat Group  |  Discord  |  Telegram  |  Reddit

MiMo-V2.6-Flash-MOPD

Technical Report

[!IMPORTANT]
This is the MOPD upgrade of the MiMo-V2.6-Flash-RL checkpoint.

  • MOPD2 (👉 Technical Report §5.6)

    Fuses several domain-specialized teachers into one model, extending to domains where reliable training-time verification is hard, such as long-horizon game development, scientific research and embodied intelligence.
  • Diagnosing and Mitigating Tool-Call Repetition in MiMo-V2.6 (👉 Technical Blog)

    An easy-to-overlook failure mode in which the model keeps issuing the same or highly similar tool calls, appearing busy while making no progress. Nothing fails outright, so it tends to go unnoticed. The MOPD stage handles it efficiently, with a short specialized-teacher run that converges quickly.

1. Introduction

How MOPD2 works

MOPD2 distills several domain-specialized teachers into the student on-policy. The teachers fall into two families: mixRL teachers, trained on verifiable tasks, and SFT teachers, trained on synthetic demonstrations for open-domain tasks where a reliable reward is hard to design. Three streams contribute to a single update:

  • Standard MOPD: mixRL teachers supervise full autonomous rollouts.
  • Teacher-Prefix OPD: prefixes come from teacher rollouts. A trajectory with k assistant turns yields k history prefixes, one per turn. The model generates a single new turn from each, and the teacher scores it against the same history.
  • SFT-Prefix OPD: prefixes come from SFT demonstrations. The demonstration supplies the history, and the model writes its own continuation.

Method details are in Technical Report §5.6.

Tool-call repetition

Following the release of MiMo-V2.6, tool-call repetition emerged as one of the most noticeable issues in agentic settings: the model would sometimes issue the same or highly similar tool calls repeatedly, consuming time and context without making progress. This checkpoint mitigates it.

Tool-call repetition rate on MiMo-V2.6-Flash before and after MOPD

Figure: response-level repetition rate on MiMo-V2.6-Flash, RL-stage versus this checkpoint, across context lengths and agent harnesses.

The technical blog has the full diagnosis. The fix is lightweight to train: a short specialized-teacher run that folds into the normal MOPD pass.

Model Summary

  • Architecture: Sparse MoE (Mixture of Experts), 309B total / 15B activated parameters
  • Context Length: 1M tokens
  • Modalities: Text, Image, Video, Audio
  • Vision Encoder: 681M-param MiMo ViT (28 layers: 24 SWA + 4 Full)
  • Audio Encoder: 308M AudioTokenizer + 127M audio patch encoder
  • Multi-Token Prediction (MTP): 5-layer speculative decoder

Figure 1: MiMo-V2.6 architecture — omni encoders, hybrid SWA backbone, and MTP blocks

Figure 1. MiMo-V2.6 architecture.

2. Downloads

Model Download
MiMo-V2.6-Pro-RL 🤗 HuggingFace · 🤖 ModelScope
MiMo-V2.6-Flash-RL 🤗 HuggingFace · 🤖 ModelScope
MiMo-V2.6-Pro-MOPD 🤗 HuggingFace · 🤖 ModelScope
MiMo-V2.6-Flash-MOPD 🤗 HuggingFace · 🤖 ModelScope

3. Model Architecture

LLM Backbone

Component MiMo-V2.6-Flash-MOPD
Layers (Total / SWA / GA) 48 / 39 / 9
Hidden Size 4096
SWA Heads (Q/KV) 64 / 8
GA Heads (Q/KV) 64 / 4
Head Dimensions (QK / V) 192 / 128
Sliding Window Size 128
Routed Experts (Total / Activated) 256 / 8
Max Context Length 1M
MTP / Speculative Decoder 5 SWA layers, window 1024

The first Transformer block uses global attention with a dense FFN. Remaining blocks interleave local SWA and GA; both use sparse MoE FFNs without shared experts.

Vision Encoder (MiMo ViT)

Configuration Value
Layers (Total / SWA / GA) 28 / 24 / 4
Hidden Size 1280
Attention Heads (Q / KV) 32 / 8
Head Dimension 64
Patch Size (T × H × W) 2 × 16 × 16
Sliding Window (Left / Right) 64 / 64
Spatial Merge Size 2 × 2
Parameters 681M

Audio Encoders

AudioTokenizer encoder: 24 layers (12 SWA / 12 GA), hidden 1024, 20 RVQ codebooks, 308M parameters. Audio patch encoder: 6 layers, 127M parameters; four frames per patch (25 Hz → 6.25 Hz).

Speculative Decoder

5-layer SWA MTP drafter (DFlash-style). Predicts 7 subsequent tokens per forward pass for parallel verification.

4. Deployment

For best performance, follow the SGLang MiMo cookbook. Docker image: lmsysorg/sglang:latest.

SGLang

sglang serve \
  --trust-remote-code \
  --model-path XiaomiMiMo/MiMo-V2.6-Flash-MOPD \
  --tp 8 \
  --dp 2 \
  --enable-dp-attention \
  --enable-dp-lm-head \
  --mm-enable-dp-encoder \
  --mem-fraction-static 0.65 \
  --chunked-prefill-size 16384 \
  --speculative-algorithm EAGLE \
  --speculative-num-steps 3 \
  --speculative-eagle-topk 1 \
  --speculative-num-draft-tokens 4 \
  --enable-multi-layer-eagle \
  --reasoning-parser mimo \
  --tool-call-parser mimo \
  --host 0.0.0.0 \
  --port 30000

vLLM

Follow the vLLM MiMo-V2.5 recipe. Stable vLLM may lag; pre-built image: docker pull vllm/vllm-openai:mimov25-cu129.

vllm serve XiaomiMiMo/MiMo-V2.6-Flash-MOPD \
  --tensor-parallel-size 4 \
  --trust-remote-code \
  --gpu-memory-utilization 0.95 \
  --max-model-len auto \
  --reasoning-parser mimo \
  --tool-call-parser mimo \
  --enable-auto-tool-choice \
  --generation-config vllm

Recommended sampling: temperature=1.0, top_p=0.95.

Also available in AI Studio, MiMo Code, Xiaomi MiMo Desktop, Xiaomi MiMo Open Platform API, and OpenRouter.

Citation

@misc{mimo2026v26flashmopd,
  title={MiMo-V2.6-Flash-MOPD},
  author={{Xiaomi MiMo Team}},
  year={2026},
  howpublished={\url{https://huggingface.co/XiaomiMiMo/MiMo-V2.6-Flash-MOPD}},
}

Contact

For questions or feedback, reach us at [email protected] or join our community:

Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration