← back to catalog · registered 2026-10-02 12:58

cafonez/MiMo-V2.6-Flash-RL-UNCENSORED-Gorgon-GGUF

cafonez GGUF MoE second-order
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/cafonez%2FMiMo-V2.6-Flash-RL-UNCENSORED-Gorgon-GGUF"
Response includes
  • classification m-uncensored
  • files 6
  • author_summary 1 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M-U
Primary method

Uncensored (method unknown)

No other method signals detected in this model.
Confidence
LOW
Why this label 3 signals
Weak or ambiguous signals. Best guess based on catalog patterns; treat as tentative and check the evidence below.
  • 'uncensored' in name/tags but no 'abliterated' marker
  • method not identifiable from author declaration alone
  • may be DPO fine-tune, prompt engineering, or unknown technique
Refusal direction extraction

No specific extraction method could be identified for this model. The producer either did not document it or used a proprietary pipeline.

What is a refusal direction? →
Downloads · 30-day
0
Likes
0
Model age
today
created 2026-10-02

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
mit
Languages
en zh
Tags
gguf mimo mimo2 mimo-v2-6-flash uncensored mxfp4 moe rocmfpx vulkan strix-halo gorgon-halo framework-desktop
Total size
155 GB
Files
6
Quantizations
1
Registered
2026-10-02 12:58
Last updated on HF
2026-10-02 12:40

Files by quantization

Auxiliary files 6 files 155 GB
MiMo-V2.6-Flash-RL-UNCENSORED-MXFP4-Q5K-00001-of-00004.gguf 41.4 GB 1b67017a download
MiMo-V2.6-Flash-RL-UNCENSORED-MXFP4-Q5K-00002-of-00004.gguf 41.2 GB c7e9668c download
MiMo-V2.6-Flash-RL-UNCENSORED-MXFP4-Q5K-00003-of-00004.gguf 40.1 GB c2939340 download
MiMo-V2.6-Flash-RL-UNCENSORED-MXFP4-Q5K-00004-of-00004.gguf 32.0 GB f27c7a2b download
README.md 5.86 KB 5fb078c7 download
.gitattributes 1.86 KB 2e4c1679 download

README current version from Hugging Face


license: mit
base_model: dealignai/MiMo-V2.6-Flash-RL-UNCENSORED
pipeline_tag: text-generation
library_name: gguf
quantized_by: cafonez
language:

  • en
  • zh
    tags:
  • gguf
  • mimo
  • mimo2
  • mimo-v2-6-flash
  • uncensored
  • mxfp4
  • moe
  • rocmfpx
  • vulkan
  • strix-halo
  • gorgon-halo
  • framework-desktop

MiMo-V2.6-Flash-RL UNCENSORED GGUF for the Framework Desktop (Gorgon Halo)

A 4.29 bpw GGUF of dealignai/MiMo-V2.6-Flash-RL-UNCENSORED, the refusal-removed edition of XiaomiMiMo/MiMo-V2.6-Flash-RL (309.8B total parameters, MoE with 256 experts, top-8). It is sized to run fully on the GPU of the 192 GB AMD Ryzen AI Max 400 ("Gorgon Halo") Framework Desktop, and was built and tested on that machine.

  • Requires ROCmFPX main (from PR #33 on). The routed experts are stored as one fused ffn_gate_up_exps tensor per layer, which stock llama.cpp, LM Studio and Ollama do not load for MiMo yet.
  • Size: 154.6 GiB (166.0 GB) in 4 shards. Pass the first shard to -m; the rest are found automatically.
  • Text only: the vision and audio encoders of the original are not included.
  • Uncensored: refusals were removed by dealignai at the weight level (see their card). You are responsible for how you use it.

Files

file size
MiMo-V2.6-Flash-RL-UNCENSORED-MXFP4-Q5K-00001-of-00004.gguf 44.4 GB
MiMo-V2.6-Flash-RL-UNCENSORED-MXFP4-Q5K-00002-of-00004.gguf 44.2 GB
MiMo-V2.6-Flash-RL-UNCENSORED-MXFP4-Q5K-00003-of-00004.gguf 43.0 GB
MiMo-V2.6-Flash-RL-UNCENSORED-MXFP4-Q5K-00004-of-00004.gguf 34.4 GB

Quantization recipe

tensors type bpw params size
routed expert gate+up (fused, 47 MoE layers) MXFP4 4.25 201.9B 99.9 GiB
routed expert down (47 MoE layers) MXFP4 4.25 100.9B 49.9 GiB
attention (QKV, output), dense FFN (4 layers), MTP projections Q5_K 5.50 5.7B 3.6 GiB
token embeddings Q8_0 8.50 0.6B 0.6 GiB
output head Q6_K 6.56 0.6B 0.5 GiB
expert routers F16 16 0.05B 0.1 GiB
norms, biases F32 <0.01 GiB
  • The routed experts are the checkpoint's own MXFP4 weights, repacked into GGUF MXFP4 without requantization.
  • Attention and the dense/MTP projections were quantized to Q5_K with an importance matrix computed on a mixed code/prose calibration set. This cut the bytes read per generated token and raised decode speed from about 20 to 23.8 tok/s.
  • Each layer's gate and up experts are stored as one tensor (a byte copy), which lets the GPU run one expert matmul instead of two.
  • The three MTP (next-token prediction) layers are included.

Quality

KL divergence of this file against the same model with attention at Q8_0 (experts identical in both): mean KLD 0.0185, same top-1 token 94.5% (llama-perplexity --kl-divergence, Vulkan). Fusing gate/up and storing the router in F16 left the KLD unchanged.

Speed on the Framework Desktop (Gorgon Halo)

AMD Ryzen AI Max+ PRO 495 / Radeon 8065S, 192 GB LPDDR5X, Ubuntu 26.04, kernel 7.0, Mesa 26.0.8 RADV, Vulkan backend, ROCmFPX main. llama-bench -fa 1 -b 2048 -ub 2048, 3 runs:

test t/s
prefill, 2048 tokens 484.9
decode 23.6
prefill, 2048 tokens at 32K context 258.7
decode at 32K context 22.1

These are local measurements on one machine, not leaderboard results.

How to run

1. Let the GPU use the memory

The weights need about 155 GiB of GPU-visible memory plus the KV cache. On a 192 GB machine, raise the GTT limit with kernel parameters (this gives the GPU 176 GiB) and reboot:

ttm.pages_limit=46137344 ttm.page_pool_size=46137344

On Ubuntu, add them to GRUB_CMDLINE_LINUX_DEFAULT in /etc/default/grub, run sudo update-grub, and reboot. Check with:

cat /sys/class/drm/card*/device/mem_info_gtt_total   # about 189000000000 bytes

Keep the BIOS GPU memory carve-out at its minimum: the GPU uses GTT, and memory carved out as VRAM is no longer available to it as GTT. Only run one model this size at a time; two will not fit in 192 GB.

2. Build ROCmFPX with Vulkan

sudo apt install build-essential cmake libvulkan-dev glslc mesa-vulkan-drivers
git clone https://github.com/ROCmFPX/ROCmFPX && cd ROCmFPX
cmake -B build -DGGML_VULKAN=ON -DCMAKE_BUILD_TYPE=Release
cmake --build build -j --target llama-server llama-cli

3. Download

hf download cafonez/MiMo-V2.6-Flash-RL-UNCENSORED-Gorgon-GGUF --local-dir MiMo-Gorgon-GGUF

4. Serve

./build/bin/llama-server \
  -m MiMo-Gorgon-GGUF/MiMo-V2.6-Flash-RL-UNCENSORED-MXFP4-Q5K-00001-of-00004.gguf \
  -dev Vulkan0 -ngl 999 -fa on \
  -c 131072 -b 2048 -ub 2048 -np 1 \
  --jinja --temp 0.7 --reasoning-budget 4096 \
  --host 127.0.0.1 --port 8080
  • --jinja is required for the chat template (thinking and tool calls). Send "chat_template_kwargs": {"enable_thinking": false} in a request to turn thinking off.
  • --temp 0.7: at the card default of 1.0 the model sometimes emitted hundreds of tool calls in one reply in agent use.
  • --reasoning-budget 4096 caps thinking per turn; without it, agent turns can spend 15K+ characters thinking.
  • For coding agents, also send "parallel_tool_calls": false with each request so the model makes one tool call per reply.
  • -c 131072 fits comfortably with 176 GiB of GTT; -ub 2048 gives the best prefill.

Then use any OpenAI-compatible client against http://127.0.0.1:8080/v1.

Credits

Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration