← back to catalog · registered 2026-10-02 20:58

yamz-labs/MiMo-V2.6-Flash-MOPD-EXL3-Yamz-Uncensored

Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/yamz-labs%2FMiMo-V2.6-Flash-MOPD-EXL3-Yamz-Uncensored"
Response includes
  • classification m-uncensored
  • files 32
  • author_summary 2 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M-U
Primary method

Uncensored (method unknown)

No other method signals detected in this model.
Confidence
LOW
Why this label 3 signals
Weak or ambiguous signals. Best guess based on catalog patterns; treat as tentative and check the evidence below.
  • 'uncensored' in name/tags but no 'abliterated' marker
  • method not identifiable from author declaration alone
  • may be DPO fine-tune, prompt engineering, or unknown technique
Refusal direction extraction

No specific extraction method could be identified for this model. The producer either did not document it or used a proprietary pipeline.

What is a refusal direction? →
Downloads · 30-day
0
Likes
0
Model age
today
created 2026-10-02

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
mit
Tags
exllamav3 safetensors mimo_v2 exl3 quantized mimo moe rocm strix-halo gfx1151 uncensored text-generation

Related

Total size
97.8 GB
Files
32
Quantizations
1
Registered
2026-10-02 20:58
Last updated on HF
2026-10-02 21:45

Files by quantization

Auxiliary files 32 files 97.8 GB
model-00003-of-00014.safetensors 7.81 GB f5da0781 download
model-00001-of-00014.safetensors 7.56 GB 564b4127 download
model-00014-of-00014.safetensors 7.23 GB a94cc8a4 download
model-00004-of-00014.safetensors 7.00 GB f42c1571 download
model-00007-of-00014.safetensors 6.94 GB f6ae7617 download
model-00009-of-00014.safetensors 6.94 GB 37ff95d7 download
model-00011-of-00014.safetensors 6.94 GB c09d9e5f download
model-00006-of-00014.safetensors 6.94 GB 9085015c download
model-00008-of-00014.safetensors 6.94 GB cf389d0c download
model-00010-of-00014.safetensors 6.94 GB 2af5cc98 download
model-00012-of-00014.safetensors 6.94 GB 2fdbcb24 download
model-00002-of-00014.safetensors 6.75 GB 03f69d2e download
model-00013-of-00014.safetensors 6.69 GB 96b7d9e3 download
model-00005-of-00014.safetensors 6.19 GB 71b35285 download
zz-e2e-step120.safetensors 2.22 MB 809b6bbf download
model.safetensors.index.json 12.1 MB 0175c938 download
tokenizer.json 10.9 MB ff15eb92 download
vocab.json 2.65 MB 4783fe10 download
merges.txt 1.59 MB 20024bfe download
modeling_mimo_v2.py 83.5 KB 40225ab4 download
uncensor_direction.st 16.1 KB 62aed313 download
tokenizer_config.json 11.9 KB 39a944cb download
LICENSE-APACHE-2.0 11.1 KB d6456956 download
configuration_mimo_v2.py 9.81 KB bb6f4472 download
config.json 7.44 KB 662f8302 download
README.md 5.88 KB d3bab10c download
chat_template.jinja 3.78 KB 92597ade download
SHA256SUMS 3.15 KB 5880dbb0 download
uncensor_spec.json 3.02 KB e0b3773c download
.gitattributes 1.66 KB 160883bb download
LICENSE 1.29 KB 64262ce6 download
generation_config.json 195 B 167a2b07 download

README current version from Hugging Face


license: mit
base_model: XiaomiMiMo/MiMo-V2.6-Flash-MOPD
base_model_relation: quantized
library_name: exllamav3
pipeline_tag: text-generation
tags: [exl3, quantized, mimo, moe, rocm, strix-halo, gfx1151, uncensored]

Yamz Labs

MiMo-V2.6-Flash-MOPD, EXL3 for AMD Strix Halo, uncensored

This is the same model as yamz-labs/MiMo-V2.6-Flash-MOPD-EXL3-Yamz, plus two small files that make it refuse far less. The weight shards (and the small tuning overlay zz-e2e-step120.safetensors) are byte-identical to the base repository (same sha256, see SHA256SUMS). Read that card for the quantisation, the hardware, the quality against FP8 and the speed. The drafter/ folder (the 4 bpw speculative-decoding draft model, 702 MB) is also identical to the base repository; speculative decoding is on by default here too, and the speed table of the base card was measured with the edit off. This card covers only what the two extra files change.

The model refuses far less and may produce content the base model would decline. See "Use" below.

Highlights

  • Refusals. 4/100 with the edit, 94/100 without, on 100 held-out harmful prompts, measured on the base pack.
  • Same quantisation quality, size and engine as the base repository. The shards are byte-identical: KLD 0.0713 and top-1 91.99 % against FP8 without the edit (125 rows), 105.04 GB, one 128 GB machine (see the base card).
  • Small cost in quality. With the edit on, KLD is 0.0752 (+5.5 %) and top-1 91.72 % on the same rows.
  • Switchable. One environment variable or flag turns it off; deleting two files gives the base repository.

Refusals on 100 harmful prompts, with and without the edit

What is added

File What it is
uncensor_spec.json Per-layer strengths of one direction edit. Attention output of layers 11-47, strength 1.62-1.76; MLP/MoE output of layers 0-46, strength 0.31-0.80. MTP layers are untouched.
uncensor_direction.st The direction: one unit vector (safetensors layout; the extension is not .safetensors on purpose, so that weight loaders do not index it).

The weights are not edited. The Yamz engine finds uncensor_spec.json in the model folder at load time and subtracts the direction from the output of each affected block while the model runs. The log prints -- ablit runtime: spec <path> active (bundled in the model directory) at load. No such line means no edit.

Switch it off without touching the files: EXL3_ABLIT_RUNTIME=off or --no-uncensor. EXL3_ABLIT_RUNTIME=/path/to/other_spec.json overrides the bundled file. Delete the two files to get the base repository. Needs the engine branch with this feature; older builds ignore the file and run the unedited model.

Run

git clone https://github.com/yamz-labs/exllamav3-strix && cd exllamav3-strix
./build.sh && source tools/strix_halo/env.sh
python tools/mimo/serve.py --model /path/to/this-folder --port 8000 -c 32768

Sampling: temperature 1.0, top-p 0.95. Greedy decoding is not recommended.

Measured effect

Spec sha256 9d347b3a...76ae3b9, direction sha256 c45d8e13...9e14ce. The direction was fitted on the pack without the tuning overlay. All numbers below were measured on the base pack (the base repository) on a Ryzen AI Max+ 395.

What Result
Refusals on 100 held-out harmful prompts (regex detector, 64 greedy tokens) 4/100 with the edit; 94/100 without
KLD against FP8, 125 held-out rows 0.0752 with the edit (+5.5 %, 95 % CI +4.5 to +6.6); 0.0713 without
Top-1 agreement with FP8, same rows 91.72 % with the edit; 91.99 % without
KL on 100 harmless prompts, edit vs no edit 0.102

Small sample (100 prompts): the interval is wide. The refusal count is a regex detector; soft refusals and partial answers can slip through either way. English prompts only. The direction is model-specific: do not reuse the GLM file. A hash proves the file is the one we measured, not that your setup behaves the same.

Speed cost of the edit

With the edit on and default speculation, greedy decode of 128 tokens measured 33.9 (prose), 36.7 (chat) and 45.6 (code) tok/s, median of 3 runs, on a Strix Halo APU (Radeon 8060S, 128 GB). The same pack with the edit off measured 32.1, 34.8 and 44.3. The edit adds no measurable decode cost; the difference is within the run-to-run spread of speculation (about 10 % between runs). The edit also still works while speculating (0 refusals on 10 harmful prompts). The GLM edit costs under 1 % of decode with the same kernel.

Use

The edit lowers the rate of refusals. It is not a safety evaluation and says nothing about what the model will or will not produce. You decide how to use it and you answer for it under the laws that apply to you. Intended uses: research, evaluation, red-teaming and private deployment. Do not use it to produce illegal content, to target people, or in a public service without your own filtering and moderation. The files come as is, without warranty.

Limits

  • Only gfx1151 with ROCm is tested, with the Yamz engine only.
  • The edit is one global direction. It lowers refusal and over-refusal together.
  • The base card limits apply (125 of 129 held-out rows, no task-suite scores).
  • Over-refusal on safe prompts (XSTest) and task-score deltas were not measured with the edit.
  • Greedy decoding is not recommended.

Licence and credits

MIT in the front matter, as in the base repository (upstream ships no LICENSE file, see the base card). Our additions: Copyright (c) 2026 Yamz Labs, MIT. Provided as is, without warranty. Engine: built on ExLlamaV3 by turboderp (MIT). Quantised and edited by Yamz. Not affiliated with Xiaomi. We remove this repository on request from the base-model authors or a rights holder, on a valid legal notice.

Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration