license: mit
base_model: XiaomiMiMo/MiMo-V2.6-Flash-RL
tags:
- uncensored
- abliterated
- heretic
- gguf
- mimo
MiMo-V2.6-Flash-RL Uncensored Heretic — merged GGUF
⚠️ Content warning: This model has had the base model's refusal
behavior surgically suppressed. The resulting model will comply with
requests the base model refuses, including requests that are harmful,
unethical, offensive, or illegal. It has reduced safety guardrails. See
Responsible use below — you are solely
responsible for what you do with it.
This is a merged GGUF of
MiMo-V2.6-Flash-RL
(309B total / 15B active MoE, MIT license) decensored / "abliterated" with
heretic-gguf — a GGUF-native port of
Heretic's Optuna-optimized directional
ablation, which runs the whole search directly on quantized GGUF weights via
llama.cpp. The ablation (trial 85 of study mimo26flash) was baked directly
into the
MXFP4 weights:
edited tensors were dequantized, patched with the exact ablation delta, and
requantized to their original type; everything else is a byte-for-byte copy.
This is the zero-runtime-overhead form — a drop-in base model, no--lora flag needed. If you prefer the lossless option (bit-identical
base weights, ~70 MB download, requantization avoided entirely), the exact
same configuration is also available as a LoRA adapter at
MiMo-V2.6-Flash-RL-Uncensored-Heretic-LoRA-GGUF.
heretic-gguf is available at
github.com/MoriNoNushi/heretic-gguf —
the full tool, so the method can be applied to other GGUF models.
Results
Measured on 140 harmful prompts (100 from mlabonne/harmful_behaviors
test + 40 custom) and 100 harmless prompts (mlabonne/harmless_alpaca
test), CoT-skip prefix (<think></think>, thinking suppressed), greedy
decoding, 100-token responses, against the MXFP4 base:
| Refusal rate (harmful) | KL divergence (harmless) | |
|---|---|---|
| Base model | 95.71% (134/140) | 0 (by definition) |
| Ablated model (trial 85) | 3.57% (5/140) | 0.0568 |
The scores above were measured on the LoRA form of the ablation; the merged
model applies the identical delta, requantized to MXFP4/Q8_0, so behavior
should match within quantization noise. Refusals are counted by
refusal-keyword matching (English + Chinese + first-person-negation markers
such as "I'm not going to / able to ..."). KL divergence is measured on
first-token logits on harmless prompts. Note that the study was run with the
CoT-skip prefix (thinking suppressed, as in stock Heretic); with full
thinking enabled the model may still reason its way back to a refusal
mid-trace, so real-use refusal rates can be somewhat higher than the 3.57%
above.
Note on KL: the KL divergence above (and the optimization objective
itself) was measured against the MXFP4 quant. KL is a
baseline-relative metric, so on a different base quant the effective drift
from that quant's baseline may differ.
Usage
llama-server \
-m MiMo-V2.6-Flash-RL-Uncensored-Heretic-MXFP4-00001-of-00002.gguf \
--jinja
Add your usual offload/context flags (-ngl 999, -c, tensor splits,
etc.) — nothing model-specific is required, and no special sampling
parameters are needed. MiMo-V2.6-Flash-RL (mimo2) support is merged
upstream in llama.cpp — any recent build works, no patches or PRs needed.
How it was made
- Method: directional ablation ("abliteration") — the refusal direction
in residual space (difference of means over 480 harmful / 480 harmless
prompts, 5% winsorized, orthogonalized against the harmless mean) is
projected out of the attention output and MoE down-projection weights.
Strengths, layer kernels, and direction selection were tuned by
multi-objective Optuna TPE (minimize refusal rate and KL jointly). This
model is trial 85 of studymimo26flash, exported withheretic-gguf export --mode merged. - Configuration (study
mimo26flash, trial 85; global direction scope,
direction index 26.2 of 48; per-expert strengths scaled by measured
harmful/harmless routing frequency;row_normalization = "pre"):- attn.o_proj: max weight 6.39 @ layer 36.4 of 48.
- routed MLP down-proj: max weight 1.58 @ layer 31.9.
- Merged export: heretic-gguf expresses ablation as a rank-1 LoRA
overlay (the same math stock Heretic writes into PEFT adapters); the
merged exporter materializes that delta exactly — full-rank, no LoRA
factorization loss — and requantizes only the patched tensors to their
original type (MXFP4 experts, Q8_0 attention/dense). This is one extra
quantization step on those tensors relative to the base; the
LoRA form
avoids it entirely.
Responsible use & disclaimer
- This model can generate content that is offensive, disturbing, hateful,
sexually explicit, violent, or otherwise objectionable, including detailed
instructions for harmful or illegal acts. That is the direct and
intended consequence of removing refusal behavior. - The ablation suppresses refusals, not the base model's knowledge —
outputs on dangerous topics may be wrong, hallucinated, or incoherent.
Nothing the model says should be treated as accurate, safe, or legal
advice. - Do not deploy this model in any production system, public-facing
service, or multi-user setting. It is intended for personal research,
red-teaming, and evaluation purposes. - You, the user, are solely responsible for any output the model produces
and for any consequences of using it. The authors of this release, of
heretic-gguf, of Heretic, and of Xiaomi accept no liability whatsoever.
Using this model to produce illegal content or to harm others is your
choice and your legal exposure — ensure your use complies with all
applicable laws in your jurisdiction. - By downloading or using this model you acknowledge the above.
License
The base model is MIT-licensed (see the
base repo);
this model inherits those terms. The heretic-gguf tooling used to produce it
is AGPL-3.0-or-later.