base_model:
- XiaomiMiMo/MiMo-V2.6-Flash-MOPD
base_model_relation: finetune
license: mit
language: - en
- zh
tags: - text-generation
- multimodal
- heretic
- heretic-dgx
- dgx-spark
- uncensored
- abliterated
- exl3
- fp8
- a8
- vllm
- sm120
- vision-language
- audio
- agent
- video-understanding
- long-context
- mimo_v2
- transformers
library_name: transformers
pipeline_tag: image-text-to-text
MiMo-V2.6-Flash-MOPD Heretic EXL3, rank-sliced for use on two DGX Spark nodes
This model is 309B total parameters / 15B active parameters, Hugging Face
has a bug in how it determines parameters on EXL3 quants.
This repository is a finetune / quant ofXiaomiMiMo/MiMo-V2.6-Flash-MOPD.
It combines a narrow Heretic transformation (KL divergence of 0.0011), EXL3 quantization, lossless TP2
rank slicing, calibrated FP8 KV-cache scales, and the upstream (unquantized) DFlash and audio
sidecars. Xiaomi's original model card is preserved at the bottom here.
It's taken a lot of work to get this going - we've corrected the dflash, an issue
with our quantization, and a ton of issues with vLLM. But we're confident in the
model and runner now. Use the latest version of my runner to support this model:
'cb-vllm-dual-dgx'.
(note that this runner also supports my GLM 5.3 Flash model:
'GLM-5.3-Flash-Uncensored-EXL3-DGX-Sliced')
Caveat: you have to keep async scheduling off for dflash to work. And don't set
more than 4 draft models - it starts to get weird when you go too high. We're still
testing this and waiting for a few specific upstream PR's for vLLM to get this next speed bump.
Uncensored model: the language checkpoint has undergone abliteration
to reduce refusal behavior. Treat outputs as untrusted, apply application-level
safeguards, and do not assume the model will decline harmful requests.
User responsibility: this model is provided without warranty. The
creators, uploaders, and maintainers are not responsible or liable for what
others generate, publish, deploy, or otherwise do with this abliterated model.
Users must operate it responsibly, apply appropriate safeguards, comply with
applicable law, and respect third-party rights. This model is for research
purposes only and is not intended for production use.
What changed
- The Heretic pass targeted only
attn.o_projacross all 48 transformer
layers. No other module class was selected for ablation. - The transformed target was quantized to EXL3 with the MCG codebook and then
sliced for tensor parallelism across two ranks without dequantizing or
requantizing the tensors. - Static FP8 E4M3 K/V scales are included for all 48 layers.
- Xiaomi's matching DFlash drafter, audio tokenizer, and model-card assets are
included from the pinned upstream revision. - The MTP weights are already quantized and indexed inside the 24 target
shards: the final index contains 114model.mtp.*tensors. The separate
upstreammodel_mtp.safetensorsis therefore intentionally not duplicated.
Heretic settings and observed result
The transformation used Heretic v0.1.1
with 100 mlabonne/harmless_alpaca prompts and 100mlabonne/harmful_behaviors prompts. Keyword-rate evaluation used the 100
harmful prompts; KL-divergence evaluation used 20 harmless prompts.
| Setting | Value |
|---|---|
| Trials | 1 |
| Target modules | 48 × attn.o_proj |
direction_index |
38.49 |
attn.o_proj.max_weight |
0.86 |
attn.o_proj.max_weight_position |
46.53 |
attn.o_proj.min_weight |
0.10 |
attn.o_proj.min_weight_distance |
1.01 |
The selected trial produced KL divergence 0.0011. The keyword metric was0/100 before and after the pass, so it was saturated at baseline and should
not be read as a behavioral or safety score. The useful measured result is the
small first-token distribution shift while limiting the edit to the attention
output projections.
Quantization and artifact result
| Setting | Value |
|---|---|
| EXL3 version | 1.5.3 |
| Average target bitrate | 4.0 bpw |
| Codebook | MCG |
| LM-head bitrate | 8 bits |
| MTP bitrate | 4 bits |
| Vision bitrate | 16 bits |
| Calibration matrix | 250 rows × 2,048 columns |
| Hessian regularization | 0.025 |
| Output scales | automatic |
| Tensor-parallel slices | 2 |
The rank-sliced EXL3 payload contains 290,356 tensors and 158,980,524,144
bytes across 24 target shards. FP8 KV calibration adds 96 scale tensors: one K
and one V scale for each of 48 layers. Those scales were derived from a
17,468-token calibration capture with eight observations per layer per rank,
using a 10% headroom factor. The final index contains 290,452 tensors.
conversion-args.json, quantization_config.json,rank-slice-manifest.json, KV_CACHE_CALIBRATION.json, and SHA256SUMS
record the conversion, slicing, calibration, and file-integrity details.
DFlash validation
The included Xiaomi drafter is served with eight proposals per verification
step. In a 305-step validation run against this exact Heretic EXL3 target, the
first four proposal positions were accepted at 71.1%, 40.0%, 22.0%, and 11.1%.
Overall acceptance across all eight positions was 20.5% (491 of 2,396 draft
tokens). Text generation, structured tool calls, continuation after tool
results, and long-prefix cache reuse also passed. These are validation
observations from one two-node DGX Spark deployment, not general benchmark
claims; broader performance testing remains in progress.
Custom runner
Stock vLLM and Transformers do not understand this rank-sliced tensor
schema. To run this we've created a custom version of vLLM using all native
sm120/sm121 libraries. Use: cbert33/cb-vllm-dual-dgx.
It is a vLLM 0.30-based two-node DGX Spark runner with
rank-sliced EXL3/Trellis loading, DFlash, no async scheduling, FP8 KV
cache, prefix caching, reasoning, and structured tool-call support. Both ranks
must use the same runner revision, image, checkpoint revision, drafter, and
runtime flags. The runner repository contains the build and deployment
contract; this repository contains the model artifacts.
Companion artifacts
dflash/: Xiaomi's matching DFlash configuration, implementation, mask
embedding, index, and draft weights.audio_tokenizer/: Xiaomi's matching audio tokenizer configuration,
template, generation settings, tokenizer settings, and weights.assets/: the architecture and tool-call-repetition figures referenced by
the upstream model card.
Original Xiaomi model card
MiMo-V2.6-Flash-MOPD
[!IMPORTANT]
This is the MOPD upgrade of the MiMo-V2.6-Flash-RL checkpoint.
- MOPD2 (👉 Technical Report §5.6)
Fuses several domain-specialized teachers into one model, extending to domains where reliable training-time verification is hard, such as long-horizon game development, scientific research and embodied intelligence.- Diagnosing and Mitigating Tool-Call Repetition in MiMo-V2.6 (👉 Technical Blog)
An easy-to-overlook failure mode in which the model keeps issuing the same or highly similar tool calls, appearing busy while making no progress. Nothing fails outright, so it tends to go unnoticed. The MOPD stage handles it efficiently, with a short specialized-teacher run that converges quickly.
1. Introduction
How MOPD2 works
MOPD2 distills several domain-specialized teachers into the student on-policy. The teachers fall into two families: mixRL teachers, trained on verifiable tasks, and SFT teachers, trained on synthetic demonstrations for open-domain tasks where a reliable reward is hard to design. Three streams contribute to a single update:
- Standard MOPD: mixRL teachers supervise full autonomous rollouts.
- Teacher-Prefix OPD: prefixes come from teacher rollouts. A trajectory with k assistant turns yields k history prefixes, one per turn. The model generates a single new turn from each, and the teacher scores it against the same history.
- SFT-Prefix OPD: prefixes come from SFT demonstrations. The demonstration supplies the history, and the model writes its own continuation.
Method details are in Technical Report §5.6.
Tool-call repetition
Following the release of MiMo-V2.6, tool-call repetition emerged as one of the most noticeable issues in agentic settings: the model would sometimes issue the same or highly similar tool calls repeatedly, consuming time and context without making progress. This checkpoint mitigates it.

Figure: response-level repetition rate on MiMo-V2.6-Flash, RL-stage versus this checkpoint, across context lengths and agent harnesses.
The technical blog has the full diagnosis. The fix is lightweight to train: a short specialized-teacher run that folds into the normal MOPD pass.
Model Summary
- Architecture: Sparse MoE (Mixture of Experts), 309B total / 15B activated parameters
- Context Length: 1M tokens
- Modalities: Text, Image, Video, Audio
- Vision Encoder: 681M-param MiMo ViT (28 layers: 24 SWA + 4 Full)
- Audio Encoder: 308M AudioTokenizer + 127M audio patch encoder
- Multi-Token Prediction (MTP): 5-layer speculative decoder

Figure 1. MiMo-V2.6 architecture.
2. Downloads
| Model | Download |
|---|---|
| MiMo-V2.6-Pro-RL | 🤗 HuggingFace · 🤖 ModelScope |
| MiMo-V2.6-Flash-RL | 🤗 HuggingFace · 🤖 ModelScope |
| MiMo-V2.6-Pro-MOPD | 🤗 HuggingFace · 🤖 ModelScope |
| MiMo-V2.6-Flash-MOPD | 🤗 HuggingFace · 🤖 ModelScope |
3. Model Architecture
LLM Backbone
| Component | MiMo-V2.6-Flash-MOPD |
|---|---|
| Layers (Total / SWA / GA) | 48 / 39 / 9 |
| Hidden Size | 4096 |
| SWA Heads (Q/KV) | 64 / 8 |
| GA Heads (Q/KV) | 64 / 4 |
| Head Dimensions (QK / V) | 192 / 128 |
| Sliding Window Size | 128 |
| Routed Experts (Total / Activated) | 256 / 8 |
| Max Context Length | 1M |
| MTP / Speculative Decoder | 5 SWA layers, window 1024 |
The first Transformer block uses global attention with a dense FFN. Remaining blocks interleave local SWA and GA; both use sparse MoE FFNs without shared experts.
Vision Encoder (MiMo ViT)
| Configuration | Value |
|---|---|
| Layers (Total / SWA / GA) | 28 / 24 / 4 |
| Hidden Size | 1280 |
| Attention Heads (Q / KV) | 32 / 8 |
| Head Dimension | 64 |
| Patch Size (T × H × W) | 2 × 16 × 16 |
| Sliding Window (Left / Right) | 64 / 64 |
| Spatial Merge Size | 2 × 2 |
| Parameters | 681M |
Audio Encoders
AudioTokenizer encoder: 24 layers (12 SWA / 12 GA), hidden 1024, 20 RVQ codebooks, 308M parameters. Audio patch encoder: 6 layers, 127M parameters; four frames per patch (25 Hz → 6.25 Hz).
Speculative Decoder
5-layer SWA MTP drafter (DFlash-style). Predicts 7 subsequent tokens per forward pass for parallel verification.
4. Deployment
For best performance, follow the SGLang MiMo cookbook. Docker image: lmsysorg/sglang:latest.
SGLang
sglang serve \
--trust-remote-code \
--model-path XiaomiMiMo/MiMo-V2.6-Flash-MOPD \
--tp 8 \
--dp 2 \
--enable-dp-attention \
--enable-dp-lm-head \
--mm-enable-dp-encoder \
--mem-fraction-static 0.65 \
--chunked-prefill-size 16384 \
--speculative-algorithm EAGLE \
--speculative-num-steps 3 \
--speculative-eagle-topk 1 \
--speculative-num-draft-tokens 4 \
--enable-multi-layer-eagle \
--reasoning-parser mimo \
--tool-call-parser mimo \
--host 0.0.0.0 \
--port 30000
vLLM
Follow the vLLM MiMo-V2.5 recipe. Stable vLLM may lag; pre-built image: docker pull vllm/vllm-openai:mimov25-cu129.
vllm serve XiaomiMiMo/MiMo-V2.6-Flash-MOPD \
--tensor-parallel-size 4 \
--trust-remote-code \
--gpu-memory-utilization 0.95 \
--max-model-len auto \
--reasoning-parser mimo \
--tool-call-parser mimo \
--enable-auto-tool-choice \
--generation-config vllm
Recommended sampling: temperature=1.0, top_p=0.95.
Also available in AI Studio, MiMo Code, Xiaomi MiMo Desktop, Xiaomi MiMo Open Platform API, and OpenRouter.
Citation
@misc{mimo2026v26flashmopd,
title={MiMo-V2.6-Flash-MOPD},
author={{Xiaomi MiMo Team}},
year={2026},
howpublished={\url{https://huggingface.co/XiaomiMiMo/MiMo-V2.6-Flash-MOPD}},
}
Contact
For questions or feedback, reach us at [email protected] or join our community: