base_model: XiaomiMiMo/MiMo-V2.6-Flash-MOPD
base_model_relation: quantized
license: mit
library_name: llama.cpp
pipeline_tag: text-generation
quantized_by: spiritfather
tags:
- gguf
- mimo
- moe
- uncensored
- abliterated
- creative-writing
- llama.cpp
MiMo-V2.6-Flash-MOPD-Yamz-Uncensored — GGUF
🧬 Ablation: yamz-labs/MiMo-V2.6-Flash-MOPD-EXL3-Yamz-Uncensored · base: XiaomiMiMo/MiMo-V2.6-Flash-MOPD
📦 Source GGUF: QuantaPlanta/MiMo-V2.6-Flash-MOPD-MXFP4-GGUF
🛠️ Runs on llama.cpp (master)
📊 Benchmarked on CaliperBench
💬 Discord
[!Note]
All credit for the uncensoring goes to yamz-labs, who fitted the refusal direction and per-layer strengths for MiMo-V2.6-Flash-MOPD and published them asuncensor_direction.st+uncensor_spec.json. Credit for the base model goes to XiaomiMiMo (309B total / ~15B active MoE, 256 experts, 8 active), and for the lossless GGUF conversion these quants start from to QuantaPlanta. These are quantizations only.
[!Important]
Why this repo exists. yamz-labs' uncensor is applied at runtime by their Kyojin engine on an EXL3 pack (tested on AMD Strix Halo); stock ExLlamaV3, TabbyAPI, llama.cpp, LM Studio and Ollama ignore it and serve the base model with its refusals. This repo bakes the same edit into the weights so it works in any llama.cpp-based runtime — Apple Silicon, CUDA, CPU — with no special engine or flag.
[!Tip]
Benchmarked on CaliperBench — a creative-writing benchmark scoring prose craft, roleplay and willingness rather than general intelligence.
Refusals
100 held-out harmful prompts (harmful_behaviors eval split), thinking on, greedy, 1024-token cap, classified by CaliperBench's refusal detector. Same prompts and settings for every row.
| Model | Refused | Notes |
|---|---|---|
| Stock MiMo-V2.6-Flash-MOPD (QuantaPlanta MXFP4) | 97 / 100 | |
| Attention-only bake (experts left stock) | 74 / 100 | what you get if the expert half of the edit is lost |
| This repo, Q4_K_M | 3 / 100 | 0 empty or truncated answers |
| yamz-labs EXL3 + Kyojin runtime edit (their card) | 4 / 100 | their prompt set and regex detector, not directly comparable |
The edit is one global direction: it lowers refusals and over-refusals together. It is not a safety evaluation and says nothing about what the model will or will not produce.
Provided quants
KLD, PPL and same-top-token are measured against the stock model (QuantaPlanta's MXFP4 GGUF, which is lossless against the release) on wiki.test.raw, 100 chunks at ctx 512, llama-perplexity --kl-divergence. So each number is the combined cost of the uncensor edit and the quantization, which is what you actually pay versus running stock. The table fills in as each size lands.
| Quant | Size | Mixture (default / gate exps / up exps / down exps) | PPL | vs stock | KLD vs stock | Same top token |
|---|---|---|---|---|---|---|
| Q4_K_M | 176.5 GiB | Q8_0 / Q4_K / Q4_K / Q4_K+Q6_K | 5.1724 ± 0.0724 | +0.88% | 0.0243 ± 0.0003 | 93.26% |
| stock (reference) | 162.9 GiB | BF16 / MXFP4 / MXFP4 / MXFP4 | 5.1275 ± 0.0714 | 0 | 0 | 100% |
Q4_K_M fits a 256 GB Apple Silicon machine with room for context.
Loading (any current llama.cpp)
llama-server -m MiMo-V2.6-Flash-MOPD-Yamz-Uncensored.Q4_K_M.gguf -ngl 99 -fa on -c 32768 -ctk q8_0 -ctv q8_0 --jinja
Thinking: on by default in MiMo's template. For non-thinking replies pass chat_template_kwargs: {"enable_thinking": false} per request.
Sampling: yamz-labs and Xiaomi recommend temperature 1.0, top-p 0.95; greedy decoding is not recommended for generation.
How it was made
Source. QuantaPlanta/MiMo-V2.6-Flash-MOPD-MXFP4-GGUF: a stock
convert_hf_to_gguf.pyof the release — bf16 attention/dense, routed experts in MXFP4. MiMo ships its experts natively in MXFP4, so this file is lossless against the original weights.Bake the edit. yamz-labs' spec states the equivalent weight form of their runtime edit,
W ← W − w(L)·r̂ r̂ᵀ W, applied to the writers into the residual stream:attn_output(layers 11–47) and the MLP/MoE output (ffn_downin dense layer 0, routedffn_down_expsin layers 1–46), with their per-layer strengths. MTP layers, router, gate/up experts, norms and embeddings are untouched. Math is float32; the edited tensors are written as BF16 into a master GGUF (~320 GB), and every edited tensor was checked against the expected|1 − w|shrink of the direction component (e.g. layer 20 experts 0.222× vs 0.221× expected).- The edited experts are deliberately not re-encoded to MXFP4: that erased ~80% of the expert edit, and that build refused 74/100.
Quantize each size from that master with
llama-quantize: routed experts take the nominal type; every non-expert tensor (attention, dense and MTP FFN,nextn.eh_proj, token embeddings, output head) is pinned to Q8_0; router, norms and attention sinks stay F32. Static quants, no imatrix.--allow-requantizeis needed only because the unedited gate/up experts are MXFP4 in the master — the model's native precision, so this is still a first quantization.llama-quantize --allow-requantize \ --tensor-type attn_qkv=q8_0 --tensor-type attn_output=q8_0 \ --tensor-type 'ffn_(gate|up|down)\.weight=q8_0' --tensor-type nextn.eh_proj=q8_0 \ --token-embedding-type q8_0 --output-tensor-type q8_0 \ MASTER.gguf MiMo-V2.6-Flash-MOPD-Yamz-Uncensored.Q4_K_M.gguf Q4_K_M
Built with llama.cpp master 4e416ee (stock mimo2 support, no patches).
Use
The edit lowers refusals. You decide how to use the model and you answer for it under the laws that apply to you. Intended uses: research, evaluation, red-teaming, creative writing and private deployment. Do not use it to produce illegal content or to target people, or in a public service without your own filtering and moderation.
License
MIT, inherited from XiaomiMiMo/MiMo-V2.6-Flash-MOPD.