pipeline_tag: image-text-to-text
license: other
license_name: minimax-community
license_link: LICENSE
library_name: transformers
tags:
- multimodal
- moe
- agent
- coding
- video
- heretic
- uncensored
- decensored
- abliterated
- ara
base_model: - MiniMaxAI/MiniMax-M3
🔒 This is a premium gated paid-access model
Access is granted manually after purchase through Ko-fi.
After purchasing, include your Hugging Face username in the Ko-fi purchase message, then click “Agree and send request to access repo” on this Hugging Face page. I will verify the username and manually approve access.
Please allow up to 24 hours for manual approval.
92% fewer refusals (8/100 Uncensored vs 98/100 Original) while preserving model quality (0.0258 KL divergence).
❤️ Support My Work
Creating these models takes significant time, work and compute. If you find them useful consider supporting me:

| Platform | Link | What you get |
|---|---|---|
| ☕ Ko-fi | One-time tip | My eternal gratitude |
Your help will motivate me and would go into further improving my workflow and coverings fees for storage, compute and may even help uncensoring bigger model with rental Cloud GPUs.
Read before purchase/download:
These GGUF files require a runtime/backend with MiniMax-M3 GGUF architecture support to use, such as the latest version of llama.cpp and the latest version of LM Studio. They are not guaranteed to work in Ollama, KoboldCpp, Jan, or older llama.cpp builds and older LM Studio versions unless those runtimes support the minimax_m3_vl architecture in GGUF format.
Important GGUF compatibility notice
These GGUF files are provided as the best currently available conversion based on the present upstream llama.cpp MiniMax-M3 GGUF work.
Please purchase only if you understand that not every frontends and/or backends support the minimax_m3_vl architecture.
GGUF quantization of llmfan46/MiniMax-M3-uncensored-heretic-aggressive.
This is a decensored version of MiniMaxAI/MiniMax-M3, made using Heretic v1.2.0 with the Arbitrary-Rank Ablation (ARA) method
Abliteration parameters
| Parameter | Value |
|---|---|
| start_layer_index | 14 |
| end_layer_index | 51 |
| preserve_good_behavior_weight | 0.0847 |
| steer_bad_behavior_weight | 0.0002 |
| overcorrect_relative_weight | 1.1741 |
| neighbor_count | 15 |
Targeted components
- attn.o_proj
Performance
| Metric | This model | Original model (MiniMaxAI/MiniMax-M3) |
|---|---|---|
| KL divergence | 0.0258 | 0 (by definition) |
| Refusals | ✅ 8/100 | ❌ 98/100 |
Lower refusals indicate fewer content restrictions, while lower KL divergence indicates more closeness to the original model's baseline. Higher refusals cause more rejections, objections, pushbacks, lecturing, censorship, softening and deflections.
Quantizations
| Filename | Quant | Description |
|---|---|---|
| MiniMax-M3-uncensored-heretic-aggressive-BF16.gguf | BF16 | Full precision |
| MiniMax-M3-uncensored-heretic-aggressive-Q8_0.gguf | Q8_0 | Near-lossless, recommended |
| MiniMax-M3-uncensored-heretic-aggressive-Q6_K.gguf | Q6_K | Excellent quality |
Vision Projector
| Filename | Quant | Description |
|---|---|---|
| MiniMax-M3-uncensored-heretic-aggressive-mmproj-BF16.gguf | BF16 | Native precision |
A Vision Projector File is Required for vision/multimodal capabilities. Use alongside any quantization above.
Usage
Right now works with the latest version of llama.cpp and the latest version of LM Studio.
Compatibility notice
These GGUF files require a runtime/backend with MiniMax-M3 GGUF architecture support.
Tested working on my setup:
- MiniMax-M3-compatible llama.cpp build`
llama-serverllama-ui- Tool calling with SearXNG/web search
- Latest version of LM Studio
Expected / likely compatible:
- Unsloth Studio, if using a current version with MiniMax-M3 support
Not currently supported / not guaranteed:
- LM Studio
- Older or mainline llama.cpp builds without MiniMax-M3 support
- Ollama, KoboldCpp, Jan, or other third-party frontends unless their bundled backend supports the
minimax-m3GGUF architecture
If you see an error such as:
unknown model architecture: 'minimax-m3'
or
failed to load model
This does not mean the GGUF file is broken, it just means that your runtime/backend does not support MiniMax-M3 GGUF yet.
Right now MiniMax-M3 GGUFs should be compatible with the latest version of: llama.cpp as well as the latest version of LM Studio.
And with: Unsloth Studio
To decrease the probability of unforseen issues due to outdated versions, be sure to use that latest transformers version (very important, won't work unless you either use 5.12.0 or 5.12.1), the latest CUDA versions (very important, do not use anything lower to avoid unforseen issues: 13.0 or 13.1 or 13.2 or 13.3), the latest PyTorch version (very important, use the latest versions of torch either 2.12.0+cu132 or 2.12.1+cu132 and torchvision either 0.27.0+cu132 or 0.27.1+cu132) and the latest Triton versions (3.6.0 or 3.7.0).
MiniMax-M3 is a native multimodal model with 1M context. It has ~428B parameters and ~23B activated parameters.
Highlights:
- Native Multimodality: M3 undergoes mixed-modality training from the very first step, enabling deeper semantic fusion across text, image, and video.
- Context Scaling via Sparse Attention: M3 introduces MiniMax Sparse Attention (MSA) to improve long context efficiency. M3 delivers 9× prefill and 15× decode speedups compared to M2 at 1M context, reducing per-token compute to 1/20.
- Coding & Cowork Capability: M3 achieves frontier-level performance across long-horizon agentic benchmarks, excelling in both coding and cowork.
MiniMax Sparse Attention (MSA)
M3 is powered by MiniMax Sparse Attention (MSA), a high-performance sparse attention operator designed for million-token contexts. Compared with GQA, MSA dramatically reduces the attention compute and memory footprint while preserving model quality.
📄 Read the technical report: arXiv:2606.13392 · Hugging Face Papers
How to Use
M3 supports three reasoning modes through the thinking parameter:
enabled— Reasoning is always enabled.adaptive— M3 automatically determines when additional reasoning is beneficial.disabled— Reasoning is disabled to minimize latency and maximize throughput.
Local Deployment
Download the model:
hf download MiniMaxAI/MiniMax-M3 --local-dir MiniMax-M3
We recommend the following inference frameworks (listed alphabetically) to serve the model:
SGLang - see SGLang cookbook.
vLLM - see vLLM recipes.
Transformers - see Transformers docs.
Inference Parameters
We recommend the following parameters for best performance: temperature=1.0, top_p=0.95, top_k=40.
Contact Us
Contact us at [email protected].