← back to catalog · registered 2026-08-22 13:56

mlx-community/MiniMax-M2.5-Uncensored-4bit

mlx-community Minimax 29B
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/mlx-community%2FMiniMax-M2.5-Uncensored-4bit"
Response includes
  • classification m-uncensored
  • files 37
  • hub_downloads_all_time 430
  • author_summary 207 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M-U
Primary method

Uncensored (method unknown)

No other method signals detected in this model.
Confidence
LOW
Why this label 3 signals
Weak or ambiguous signals. Best guess based on catalog patterns; treat as tentative and check the evidence below.
  • 'uncensored' in name/tags but no 'abliterated' marker
  • method not identifiable from author declaration alone
  • may be DPO fine-tune, prompt engineering, or unknown technique
Refusal direction extraction

No specific extraction method could be identified for this model. The producer either did not document it or used a proprietary pipeline.

What is a refusal direction? →
Downloads · lifetime
430
34 last 30d - cooling
Likes
1
Model age
7mo ago
created 2026-03-12
Downloads over time
Now449→from40↑1,023%
2017633349040 on Mar 11449 on Oct 11MarAprMayJunJulAugSepOct
Mar 11 → Oct 11 · 70 snapshots · spans 214 days

Metadata

Tags
safetensors minimax_m2 custom_code 4-bit region:us

Related

Total size
120 GB
Files
37
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-03-12 03:20

Files by quantization

Auxiliary files 37 files 120 GB
model-00001-of-00027.safetensors 4.82 GB 3f034a5b download
model-00009-of-00027.safetensors 4.50 GB 986a064f download
model-00015-of-00027.safetensors 4.50 GB e35ac7cb download
model-00012-of-00027.safetensors 4.50 GB c80104db download
model-00024-of-00027.safetensors 4.50 GB 092bb077 download
model-00018-of-00027.safetensors 4.50 GB 5df3d0ca download
model-00021-of-00027.safetensors 4.50 GB bb280348 download
model-00006-of-00027.safetensors 4.50 GB 39f51cf5 download
model-00003-of-00027.safetensors 4.50 GB 6e5f4c30 download
model-00023-of-00027.safetensors 4.48 GB 4884ec33 download
model-00011-of-00027.safetensors 4.48 GB 6f84b527 download
model-00017-of-00027.safetensors 4.48 GB 0cc30dc7 download
model-00020-of-00027.safetensors 4.48 GB 54c732a1 download
model-00022-of-00027.safetensors 4.48 GB 39ed83a4 download
model-00008-of-00027.safetensors 4.48 GB 61a575ae download
model-00019-of-00027.safetensors 4.48 GB 0920929c download
model-00016-of-00027.safetensors 4.48 GB a0b55d2c download
model-00005-of-00027.safetensors 4.48 GB cdbd0fbd download
model-00026-of-00027.safetensors 4.48 GB d7110e1b download
model-00010-of-00027.safetensors 4.48 GB 841f8866 download
model-00013-of-00027.safetensors 4.48 GB de44468d download
model-00025-of-00027.safetensors 4.48 GB d18d1765 download
model-00014-of-00027.safetensors 4.48 GB bdbe3699 download
model-00007-of-00027.safetensors 4.48 GB 39d47726 download
model-00004-of-00027.safetensors 4.48 GB 35ad6aca download
model-00002-of-00027.safetensors 4.48 GB 2c795076 download
model-00027-of-00027.safetensors 2.88 GB c52b567b download
tokenizer.json 14.8 MB 6b25b320 download
model.safetensors.index.json 167 KB c5d1f9ef download
modeling_minimax_m2.py 30.2 KB 8846d38a download
config.json 15.9 KB a103f651 download
configuration_minimax_m2.py 9.92 KB 7fcd9861 download
chat_template.jinja 6.37 KB 4623080a download
README.md 4.21 KB a7cdf094 download
.gitattributes 1.53 KB 52373fe2 download
tokenizer_config.json 534 B 64654235 download
generation_config.json 144 B 29d6840e download

README current version from Hugging Face

MiniMax-M2.5-Uncensored-4bit

A 4-bit quantized, uncensored variant of MiniMax-M2.5, built from vpyn/MiniMax-M2.5-CARVE-v1-BF16.

Optimized for the MLX ecosystem on Apple Silicon, this repo works especially well with mlx-openai-server, an OpenAI-compatible local inference server for MLX models with support for reasoning parsers, tool-call parsers, streaming, and multi-model serving.

Model description

This model is derived from the official MiniMax-M2.5 (MiniMaxAI/MiniMax-M2.5), a 229B-parameter frontier model strong in coding, tool use, search, and office tasks. The uncensored base MiniMax-M2.5-CARVE-v1-BF16 was created via CARVE-style uncensoring; this repository provides a 4-bit quantized version for lower memory use and faster inference while retaining the same architecture and chat format.

Intended use

This model is intended for users who want:

  • The capabilities of MiniMax-M2.5 in a smaller, faster 4-bit form.
  • An uncensored variant for research, creative, or uncensored-assistant use cases.

Disclaimer: This is an uncensored model. Outputs can be unfiltered. Use responsibly and in accordance with your local laws and policies.

Installation

The recommended way to run this model is with mlx-openai-server on Apple Silicon.

Install mlx-openai-server

python3 -m venv .venv
source .venv/bin/activate
pip install -U mlx-openai-server

To use the latest development version instead:

pip install -U git+https://github.com/cubist38/mlx-openai-server.git

Launch the model

mlx-openai-server launch --model-path MiniMax-M2.5-Uncensored-4bit --model-type lm --reasoning-parser minimax_m2 --tool-call-parser minimax_m2 --trust-remote-code

This exposes an OpenAI-compatible API at http://localhost:8000/v1, making the model easy to use from existing OpenAI SDKs, apps, and agent frameworks.

How to use

This is an MLX model for Apple Silicon. The recommended serving path is mlx-openai-server, and you can also run it directly with mlx_lm.

Python (mlx_lm)

from mlx_lm import load, generate

model_path = "MiniMax-M2.5-Uncensored-4bit"  # or your local path
model, tokenizer = load(
    model_path,
    tokenizer_config={"trust_remote_code": True},
)
response = generate(
    model,
    tokenizer,
    prompt="Hello, how are you?",
    max_tokens=256,
    temp=1.0,
    top_p=0.95,
    verbose=True,
)
print(response)

For chat, format messages with the model’s chat template (e.g. using the repo’s chat_template.jinja) before passing the resulting string as prompt.

Server (mlx-openai-server)

mlx-openai-server launch --model-path MiniMax-M2.5-Uncensored-4bit --model-type lm --reasoning-parser minimax_m2 --tool-call-parser minimax_m2 --trust-remote-code

mlx-openai-server is the best fit if you want OpenAI-compatible endpoints, streaming responses, structured outputs, reasoning/tool-call parsing, and easy integration with existing clients.

Acknowledgments

License

Follow the license terms of the original MiniMax-M2.5 model (Modified-MIT). See MiniMax-M2.5 and the MiniMax-M2.5 GitHub LICENSE for details.

README history 1 version

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-03-12Upload folder using huggingface_hub953415d4.2 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration