license: mit
base_model: XiaomiMiMo/MiMo-V2.6-Flash-MOPD
base_model_relation: quantized
library_name: exllamav3
pipeline_tag: text-generation
tags: [exl3, quantized, mimo, moe, rocm, strix-halo, gfx1151, uncensored]

MiMo-V2.6-Flash-MOPD, EXL3 for AMD Strix Halo, uncensored
This is the same model as yamz-labs/MiMo-V2.6-Flash-MOPD-EXL3-Yamz, plus two small files that make it refuse far less. The weight shards (and the small tuning overlay zz-e2e-step120.safetensors) are byte-identical to the base repository (same sha256, see SHA256SUMS). Read that card for the quantisation, the hardware, the quality against FP8 and the speed. The drafter/ folder (the 4 bpw speculative-decoding draft model, 702 MB) is also identical to the base repository; speculative decoding is on by default here too, and the speed table of the base card was measured with the edit off. This card covers only what the two extra files change.
The model refuses far less and may produce content the base model would decline. See "Use" below.
Highlights
- Refusals. 4/100 with the edit, 94/100 without, on 100 held-out harmful prompts, measured on the base pack.
- Same quantisation quality, size and engine as the base repository. The shards are byte-identical: KLD 0.0713 and top-1 91.99 % against FP8 without the edit (125 rows), 105.04 GB, one 128 GB machine (see the base card).
- Small cost in quality. With the edit on, KLD is 0.0752 (+5.5 %) and top-1 91.72 % on the same rows.
- Switchable. One environment variable or flag turns it off; deleting two files gives the base repository.

What is added
| File | What it is |
|---|---|
uncensor_spec.json |
Per-layer strengths of one direction edit. Attention output of layers 11-47, strength 1.62-1.76; MLP/MoE output of layers 0-46, strength 0.31-0.80. MTP layers are untouched. |
uncensor_direction.st |
The direction: one unit vector (safetensors layout; the extension is not .safetensors on purpose, so that weight loaders do not index it). |
The weights are not edited. The Yamz engine finds uncensor_spec.json in the model folder at load time and subtracts the direction from the output of each affected block while the model runs. The log prints -- ablit runtime: spec <path> active (bundled in the model directory) at load. No such line means no edit.
Switch it off without touching the files: EXL3_ABLIT_RUNTIME=off or --no-uncensor. EXL3_ABLIT_RUNTIME=/path/to/other_spec.json overrides the bundled file. Delete the two files to get the base repository. Needs the engine branch with this feature; older builds ignore the file and run the unedited model.
Run
git clone https://github.com/yamz-labs/exllamav3-strix && cd exllamav3-strix
./build.sh && source tools/strix_halo/env.sh
python tools/mimo/serve.py --model /path/to/this-folder --port 8000 -c 32768
Sampling: temperature 1.0, top-p 0.95. Greedy decoding is not recommended.
Measured effect
Spec sha256 9d347b3a...76ae3b9, direction sha256 c45d8e13...9e14ce. The direction was fitted on the pack without the tuning overlay. All numbers below were measured on the base pack (the base repository) on a Ryzen AI Max+ 395.
| What | Result |
|---|---|
| Refusals on 100 held-out harmful prompts (regex detector, 64 greedy tokens) | 4/100 with the edit; 94/100 without |
| KLD against FP8, 125 held-out rows | 0.0752 with the edit (+5.5 %, 95 % CI +4.5 to +6.6); 0.0713 without |
| Top-1 agreement with FP8, same rows | 91.72 % with the edit; 91.99 % without |
| KL on 100 harmless prompts, edit vs no edit | 0.102 |
Small sample (100 prompts): the interval is wide. The refusal count is a regex detector; soft refusals and partial answers can slip through either way. English prompts only. The direction is model-specific: do not reuse the GLM file. A hash proves the file is the one we measured, not that your setup behaves the same.
Speed cost of the edit
With the edit on and default speculation, greedy decode of 128 tokens measured 33.9 (prose), 36.7 (chat) and 45.6 (code) tok/s, median of 3 runs, on a Strix Halo APU (Radeon 8060S, 128 GB). The same pack with the edit off measured 32.1, 34.8 and 44.3. The edit adds no measurable decode cost; the difference is within the run-to-run spread of speculation (about 10 % between runs). The edit also still works while speculating (0 refusals on 10 harmful prompts). The GLM edit costs under 1 % of decode with the same kernel.
Use
The edit lowers the rate of refusals. It is not a safety evaluation and says nothing about what the model will or will not produce. You decide how to use it and you answer for it under the laws that apply to you. Intended uses: research, evaluation, red-teaming and private deployment. Do not use it to produce illegal content, to target people, or in a public service without your own filtering and moderation. The files come as is, without warranty.
Limits
- Only gfx1151 with ROCm is tested, with the Yamz engine only.
- The edit is one global direction. It lowers refusal and over-refusal together.
- The base card limits apply (125 of 129 held-out rows, no task-suite scores).
- Over-refusal on safe prompts (XSTest) and task-score deltas were not measured with the edit.
- Greedy decoding is not recommended.
Licence and credits
MIT in the front matter, as in the base repository (upstream ships no LICENSE file, see the base card). Our additions: Copyright (c) 2026 Yamz Labs, MIT. Provided as is, without warranty. Engine: built on ExLlamaV3 by turboderp (MIT). Quantised and edited by Yamz. Not affiliated with Xiaomi. We remove this repository on request from the base-model authors or a rights holder, on a valid legal notice.