license: mit
base_model: dealignai/MiMo-V2.6-Flash-RL-UNCENSORED
pipeline_tag: text-generation
library_name: gguf
quantized_by: cafonez
language:
- en
- zh
tags: - gguf
- mimo
- mimo2
- mimo-v2-6-flash
- uncensored
- mxfp4
- moe
- rocmfpx
- vulkan
- strix-halo
- gorgon-halo
- framework-desktop
MiMo-V2.6-Flash-RL UNCENSORED GGUF for the Framework Desktop (Gorgon Halo)
A 4.29 bpw GGUF of dealignai/MiMo-V2.6-Flash-RL-UNCENSORED, the refusal-removed edition of XiaomiMiMo/MiMo-V2.6-Flash-RL (309.8B total parameters, MoE with 256 experts, top-8). It is sized to run fully on the GPU of the 192 GB AMD Ryzen AI Max 400 ("Gorgon Halo") Framework Desktop, and was built and tested on that machine.
- Requires ROCmFPX
main(from PR #33 on). The routed experts are stored as one fusedffn_gate_up_expstensor per layer, which stock llama.cpp, LM Studio and Ollama do not load for MiMo yet. - Size: 154.6 GiB (166.0 GB) in 4 shards. Pass the first shard to
-m; the rest are found automatically. - Text only: the vision and audio encoders of the original are not included.
- Uncensored: refusals were removed by dealignai at the weight level (see their card). You are responsible for how you use it.
Files
| file | size |
|---|---|
MiMo-V2.6-Flash-RL-UNCENSORED-MXFP4-Q5K-00001-of-00004.gguf |
44.4 GB |
MiMo-V2.6-Flash-RL-UNCENSORED-MXFP4-Q5K-00002-of-00004.gguf |
44.2 GB |
MiMo-V2.6-Flash-RL-UNCENSORED-MXFP4-Q5K-00003-of-00004.gguf |
43.0 GB |
MiMo-V2.6-Flash-RL-UNCENSORED-MXFP4-Q5K-00004-of-00004.gguf |
34.4 GB |
Quantization recipe
| tensors | type | bpw | params | size |
|---|---|---|---|---|
| routed expert gate+up (fused, 47 MoE layers) | MXFP4 | 4.25 | 201.9B | 99.9 GiB |
| routed expert down (47 MoE layers) | MXFP4 | 4.25 | 100.9B | 49.9 GiB |
| attention (QKV, output), dense FFN (4 layers), MTP projections | Q5_K | 5.50 | 5.7B | 3.6 GiB |
| token embeddings | Q8_0 | 8.50 | 0.6B | 0.6 GiB |
| output head | Q6_K | 6.56 | 0.6B | 0.5 GiB |
| expert routers | F16 | 16 | 0.05B | 0.1 GiB |
| norms, biases | F32 | <0.01 GiB |
- The routed experts are the checkpoint's own MXFP4 weights, repacked into GGUF MXFP4 without requantization.
- Attention and the dense/MTP projections were quantized to Q5_K with an importance matrix computed on a mixed code/prose calibration set. This cut the bytes read per generated token and raised decode speed from about 20 to 23.8 tok/s.
- Each layer's gate and up experts are stored as one tensor (a byte copy), which lets the GPU run one expert matmul instead of two.
- The three MTP (next-token prediction) layers are included.
Quality
KL divergence of this file against the same model with attention at Q8_0 (experts identical in both): mean KLD 0.0185, same top-1 token 94.5% (llama-perplexity --kl-divergence, Vulkan). Fusing gate/up and storing the router in F16 left the KLD unchanged.
Speed on the Framework Desktop (Gorgon Halo)
AMD Ryzen AI Max+ PRO 495 / Radeon 8065S, 192 GB LPDDR5X, Ubuntu 26.04, kernel 7.0, Mesa 26.0.8 RADV, Vulkan backend, ROCmFPX main. llama-bench -fa 1 -b 2048 -ub 2048, 3 runs:
| test | t/s |
|---|---|
| prefill, 2048 tokens | 484.9 |
| decode | 23.6 |
| prefill, 2048 tokens at 32K context | 258.7 |
| decode at 32K context | 22.1 |
These are local measurements on one machine, not leaderboard results.
How to run
1. Let the GPU use the memory
The weights need about 155 GiB of GPU-visible memory plus the KV cache. On a 192 GB machine, raise the GTT limit with kernel parameters (this gives the GPU 176 GiB) and reboot:
ttm.pages_limit=46137344 ttm.page_pool_size=46137344
On Ubuntu, add them to GRUB_CMDLINE_LINUX_DEFAULT in /etc/default/grub, run sudo update-grub, and reboot. Check with:
cat /sys/class/drm/card*/device/mem_info_gtt_total # about 189000000000 bytes
Keep the BIOS GPU memory carve-out at its minimum: the GPU uses GTT, and memory carved out as VRAM is no longer available to it as GTT. Only run one model this size at a time; two will not fit in 192 GB.
2. Build ROCmFPX with Vulkan
sudo apt install build-essential cmake libvulkan-dev glslc mesa-vulkan-drivers
git clone https://github.com/ROCmFPX/ROCmFPX && cd ROCmFPX
cmake -B build -DGGML_VULKAN=ON -DCMAKE_BUILD_TYPE=Release
cmake --build build -j --target llama-server llama-cli
3. Download
hf download cafonez/MiMo-V2.6-Flash-RL-UNCENSORED-Gorgon-GGUF --local-dir MiMo-Gorgon-GGUF
4. Serve
./build/bin/llama-server \
-m MiMo-Gorgon-GGUF/MiMo-V2.6-Flash-RL-UNCENSORED-MXFP4-Q5K-00001-of-00004.gguf \
-dev Vulkan0 -ngl 999 -fa on \
-c 131072 -b 2048 -ub 2048 -np 1 \
--jinja --temp 0.7 --reasoning-budget 4096 \
--host 127.0.0.1 --port 8080
--jinjais required for the chat template (thinking and tool calls). Send"chat_template_kwargs": {"enable_thinking": false}in a request to turn thinking off.--temp 0.7: at the card default of 1.0 the model sometimes emitted hundreds of tool calls in one reply in agent use.--reasoning-budget 4096caps thinking per turn; without it, agent turns can spend 15K+ characters thinking.- For coding agents, also send
"parallel_tool_calls": falsewith each request so the model makes one tool call per reply. -c 131072fits comfortably with 176 GiB of GTT;-ub 2048gives the best prefill.
Then use any OpenAI-compatible client against http://127.0.0.1:8080/v1.
Credits
- Model: Xiaomi MiMo-V2.6-Flash-RL (MIT).
- Uncensored edition: dealignai (MIT).
- MiMo support in llama.cpp and its contributors; fused gate/up loading in ROCmFPX PR #33.