← back to catalog · registered 2026-08-22 13:56

ressl/MiniMax-M3-uncensored-NVFP4

ressl Minimax 207B MoE multimodal second-order
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/ressl%2FMiniMax-M3-uncensored-NVFP4"
Response includes
  • classification m1
  • files 138
  • hub_downloads_all_time 29,764
  • author_summary 28 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M1
Primary method

Direct removal

No other method signals detected in this model.
Confidence
MEDIUM
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=0 (base model)
  • no specific method indicators - defaulting to M1 (most common)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
MEDIUM
Why we say so
primary_method=M1; difference-of-means is the reference extraction for M1/M3 (Arditi 2024)
Downloads · lifetime
30K
556 last 30d - cooling
Likes
11
Model age
2mo ago
created 2026-07-15
Downloads over time
Now29.9K→from290↑10,213%
011K21.9K32.9K290 on Jul 1529.9K on Oct 11JulAugSepOct
Jul 15 → Oct 11 · 53 snapshots · spans 88 days

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Variants by this author 2 formats · 949 downloads combined

The same weights this author released in different packaging. Pick the format that matches your runtime.

Metadata

License
other
Languages
en de
Tags
transformers safetensors minimax_m3_vl image-text-to-text uncensored abliterated minimax moe multimodal nvfp4 modelopt sglang

Related

Total size
242 GB
Files
138
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-07-16 10:03

Files by quantization

Auxiliary files 138 files 242 GB
model-00001-of-00116.safetensors 5.20 GB bddcf4dd download
model-00060-of-00116.safetensors 3.80 GB a0d755ee download
model-00061-of-00116.safetensors 3.80 GB 02b2306d download
model-00062-of-00116.safetensors 3.80 GB c3f7938f download
model-00063-of-00116.safetensors 3.80 GB fad13d72 download
model-00064-of-00116.safetensors 3.80 GB 84f85fe4 download
model-00065-of-00116.safetensors 3.80 GB 0344fce7 download
model-00066-of-00116.safetensors 3.80 GB 94d5e331 download
model-00067-of-00116.safetensors 3.80 GB 3357801d download
model-00068-of-00116.safetensors 3.80 GB 20a4010f download
model-00069-of-00116.safetensors 3.80 GB 9fa9cd0e download
model-00070-of-00116.safetensors 3.80 GB 8c1d6990 download
model-00071-of-00116.safetensors 3.80 GB 67af9d12 download
model-00072-of-00116.safetensors 3.80 GB d1f62a58 download
model-00073-of-00116.safetensors 3.80 GB b7d56df7 download
model-00074-of-00116.safetensors 3.80 GB 5fc719b4 download
model-00075-of-00116.safetensors 3.80 GB d2a6c003 download
model-00076-of-00116.safetensors 3.80 GB c5e31acb download
model-00077-of-00116.safetensors 3.80 GB 9acdd5ab download
model-00078-of-00116.safetensors 3.80 GB a955f6bc download
model-00079-of-00116.safetensors 3.80 GB d76d7bcf download
model-00081-of-00116.safetensors 3.80 GB c506f842 download
model-00082-of-00116.safetensors 3.80 GB f66b4e20 download
model-00083-of-00116.safetensors 3.80 GB d5af0827 download
model-00084-of-00116.safetensors 3.80 GB 85f3ab6c download
model-00085-of-00116.safetensors 3.80 GB ec7ed054 download
model-00086-of-00116.safetensors 3.80 GB ff20e1ba download
model-00087-of-00116.safetensors 3.80 GB 6aa8684f download
model-00088-of-00116.safetensors 3.80 GB 72756776 download
model-00089-of-00116.safetensors 3.80 GB 5af78b6e download
model-00090-of-00116.safetensors 3.80 GB 615cd7eb download
model-00092-of-00116.safetensors 3.80 GB 3dd0ed43 download
model-00093-of-00116.safetensors 3.80 GB 81ccbda0 download
model-00094-of-00116.safetensors 3.80 GB 5f005b97 download
model-00095-of-00116.safetensors 3.80 GB 07583f9a download
model-00096-of-00116.safetensors 3.80 GB 695a23f2 download
model-00097-of-00116.safetensors 3.80 GB 13578c50 download
model-00098-of-00116.safetensors 3.80 GB 4b27e65f download
model-00099-of-00116.safetensors 3.80 GB 1d156d55 download
model-00100-of-00116.safetensors 3.80 GB b05b18ef download
model-00101-of-00116.safetensors 3.80 GB d5b2d831 download
model-00103-of-00116.safetensors 3.80 GB b8cb78d2 download
model-00104-of-00116.safetensors 3.80 GB c82b54ad download
model-00105-of-00116.safetensors 3.80 GB f65ad863 download
model-00106-of-00116.safetensors 3.80 GB a07566dc download
model-00107-of-00116.safetensors 3.80 GB a752ad97 download
model-00108-of-00116.safetensors 3.80 GB e1c2e9b0 download
model-00109-of-00116.safetensors 3.80 GB 3ab4a380 download
model-00110-of-00116.safetensors 3.80 GB 0359eb1f download
model-00111-of-00116.safetensors 3.80 GB debfa94d download
model-00112-of-00116.safetensors 3.80 GB a7a825c6 download
model-00080-of-00116.safetensors 3.80 GB c94fd871 download
model-00091-of-00116.safetensors 3.80 GB 9e370ece download
model-00102-of-00116.safetensors 3.80 GB e330c946 download
model-00113-of-00116.safetensors 3.80 GB 82acf0ec download
model-00114-of-00116.safetensors 3.80 GB 6d7eaa13 download
model-00115-of-00116.safetensors 3.80 GB de08fb08 download
model-00116-of-00116.safetensors 3.80 GB aa3398b7 download
model-00059-of-00116.safetensors 1.49 GB 543cea3e download
model-00002-of-00116.safetensors 1.24 GB 824225b3 download
model-00026-of-00116.safetensors 770 MB e852d34d download
model-00010-of-00116.safetensors 323 MB c303935a download
model-00011-of-00116.safetensors 323 MB b9fd9c2d download
model-00012-of-00116.safetensors 323 MB 97432fe7 download
model-00013-of-00116.safetensors 323 MB 7a5efcdd download
model-00014-of-00116.safetensors 323 MB 5c4e69a5 download
model-00015-of-00116.safetensors 323 MB 1287985c download
model-00016-of-00116.safetensors 323 MB 25cbdf06 download
model-00017-of-00116.safetensors 323 MB be3cc83c download
model-00018-of-00116.safetensors 323 MB 11c85e31 download
model-00019-of-00116.safetensors 323 MB 7fffa098 download
model-00020-of-00116.safetensors 323 MB 882137f1 download
model-00021-of-00116.safetensors 323 MB 91beec40 download
model-00022-of-00116.safetensors 323 MB ee7dbd31 download
model-00023-of-00116.safetensors 323 MB 8a9205f2 download
model-00024-of-00116.safetensors 323 MB 6880048a download
model-00003-of-00116.safetensors 323 MB 6feb0ee9 download
model-00004-of-00116.safetensors 323 MB 6437b401 download
model-00005-of-00116.safetensors 323 MB d06078ae download
model-00006-of-00116.safetensors 323 MB 8ef24751 download
model-00007-of-00116.safetensors 323 MB 7e521221 download
model-00008-of-00116.safetensors 323 MB ef1a23c2 download
model-00009-of-00116.safetensors 323 MB ae225949 download
model-00027-of-00116.safetensors 323 MB 179ae833 download
model-00028-of-00116.safetensors 323 MB 811d16cc download
model-00029-of-00116.safetensors 323 MB d38b0c9d download
model-00030-of-00116.safetensors 323 MB a2df3c78 download
model-00031-of-00116.safetensors 323 MB 3ce793f6 download
model-00032-of-00116.safetensors 323 MB aa8fc2e4 download
model-00033-of-00116.safetensors 323 MB f97d5024 download
model-00034-of-00116.safetensors 323 MB 2d162b56 download
model-00035-of-00116.safetensors 323 MB bfaef31a download
model-00036-of-00116.safetensors 323 MB bb393c24 download
model-00025-of-00116.safetensors 323 MB 1a7fad29 download
model-00037-of-00116.safetensors 323 MB 19afd35c download
model-00038-of-00116.safetensors 323 MB 93a14276 download
model-00039-of-00116.safetensors 323 MB 650891f7 download
model-00040-of-00116.safetensors 323 MB a82cab5f download
model-00041-of-00116.safetensors 323 MB 72a2d4d2 download
model-00042-of-00116.safetensors 323 MB aa05c34d download
model-00043-of-00116.safetensors 323 MB f74b66d2 download
model-00044-of-00116.safetensors 323 MB ceb01d48 download
model-00045-of-00116.safetensors 323 MB c004d7a3 download
model-00046-of-00116.safetensors 323 MB 2f61526f download
model-00047-of-00116.safetensors 323 MB 5daf306e download
model-00048-of-00116.safetensors 323 MB d9f2e149 download
model-00049-of-00116.safetensors 323 MB 8678a369 download
model-00050-of-00116.safetensors 323 MB 2f0233b4 download
model-00051-of-00116.safetensors 323 MB 02a99646 download
model-00052-of-00116.safetensors 323 MB 381eed5c download
model-00053-of-00116.safetensors 323 MB 7bc463b7 download
model-00054-of-00116.safetensors 323 MB a1dfab29 download
model-00055-of-00116.safetensors 323 MB c5165f8a download
model-00056-of-00116.safetensors 323 MB 5d2ec172 download
model-00057-of-00116.safetensors 323 MB 9cd3f8a8 download
model-00058-of-00116.safetensors 323 MB 67b6b8ce download
model.safetensors.index.json 9.90 MB ddda0bee download
tokenizer.json 9.28 MB 9065e14e download
vocab.json 4.49 MB 37989413 download
merges.txt 2.30 MB ff574e20 download
banner.png 2.20 MB 8c0dac6d download
chat_template.jinja 11.5 KB 93022eb9 download
tokenizer_config.json 11.0 KB 44840ad8 download
processing_minimax.py 9.73 KB c08cc624 download
README.md 7.70 KB 0a1f8518 download
image_processor.py 7.67 KB aa896bba download
video_processor.py 7.15 KB d09de125 download
config.json 5.53 KB 9872f08f download
configuration_minimax_m3_vl.py 4.25 KB 0534f703 download
LICENSE 3.26 KB f413dfa3 download
runtime-validation.json 2.27 KB 09ae583f download
added_tokens.json 1.62 KB 43c846ca download
.gitattributes 1.53 KB b0baf14c download
run-manifest.json 1.12 KB 2f623446 download
preprocessor_config.json 772 B 300d985e download
hf_quant_config.json 543 B eee474c7 download
special_tokens_map.json 277 B c55c5c84 download
generation_config.json 150 B 54806f9d download

README current version from Hugging Face


license: other
base_model: ressl/MiniMax-M3-uncensored
base_model_relation: quantized
library_name: transformers
pipeline_tag: text-generation
language:

  • en
  • de
    tags:
  • uncensored
  • abliterated
  • minimax
  • moe
  • multimodal
  • nvfp4
  • modelopt
  • sglang
  • blackwell

MiniMax-M3-uncensored-NVFP4

MiniMax-M3-uncensored-NVFP4

TL;DR: ressl/MiniMax-M3-uncensored,
compressed from 854.2 GB BF16 to 260.4 GB with NVIDIA ModelOpt NVFP4 routed-expert weights. It was
validated end to end on 4x NVIDIA RTX PRO 6000 Blackwell 96 GB with SGLang at about 88.5 tokens/s
for a warmed-up single-request decode.

This release preserves the uncensored MiniMax-M3 checkpoint while making the 428B-total, 23B-active
MoE practical on one four-GPU Blackwell workstation. Text generation, reasoning parsing, tool calls,
multimodal input, 20k-token context, and two concurrent requests were tested on the exported files.

Warning: this model is genuinely uncensored and can comply with requests that a stock model
refuses. It is intended for lawful security research, red-teaming, penetration testing, malware
and exploit analysis, detection engineering, and other legitimate research. You are responsible
for how you use it.

Facts and figures

Source checkpoint ressl/MiniMax-M3-uncensored, revision f5faea4659cc3834b1dc3793173db9aa533e3aa7
Method ModelOpt NVFP4, routed expert w1/w2/w3 weights only, block size 16, two-level scaling
Precision retained BF16 attention, router, dense and shared experts, embeddings, LM head, vision tower, and projectors
KV cache BF16 in the validated runtime, not quantized in the checkpoint
Size 260.4 GB (242.5 GiB), down from 854.2 GB (795.5 GiB), 69.5% smaller
Export layout 116 safetensors shards, 89,080 indexed tensors
Conversion 147.1 seconds with 4 workers on 4x RTX PRO 6000 Blackwell 96 GB
Serving validation SGLang TP=4, 32,768-token configured context, two concurrent requests
Measured decode About 88.5 tokens/s, warmed-up batch 1, short prompts, CUDA graph enabled
Long-context probe Exact sentinel recall from a 20,212-token prompt
Multimodal probes Correct red/blue spatial classification and exact OCR of M3 VISION 7429
Uncensored smoke test 0/16 hard refusals on the same harmful-prompt sample, 96-token continuations
ModelOpt revision f479e7890f0d276e061d69b5a0d70d477ec863f2

Quantization recipe

The conversion streams every source shard and quantizes all 21,888 routed-expert projection tensors.
It uses ModelOpt's NVFP4QTensor max quantizer with 16-value blocks. w1 and w3 share the second-level
FP32 scale required by the fused gate/up runtime layout. The exported weights use packed U8 data,
FP8 E4M3 block scales, and FP32 global scales. No calibration dataset was used; the activation input
scale is 1.0 and runtime block scaling remains dynamic. Shared experts deliberately remain BF16.

Structural validation checked every index entry, shard header, dtype, shape, expert triplet, input
scale, and gate/up shared scale. A sampled dequantization comparison against BF16 measured cosine
similarity 0.99539 to 0.99544 and relative RMSE 0.0954 to 0.0963 across w1, w2, and w3.

Run it with SGLang on RTX PRO 6000 Blackwell

This exact image and command were validated on four SM120 GPUs. The two mounted patch files are part
of this repository. They make SGLang's ModelOpt Cutlass path forward MiniMax-M3's parameterized
clamped-SwiGLU values. The TRTLLM MoE path is intentionally not used.

MODEL_DIR=/path/to/MiniMax-M3-uncensored-NVFP4

docker run --rm --runtime=nvidia --gpus all --ipc=host --shm-size 32g \
  -v "$MODEL_DIR:/model:ro" \
  -v "$MODEL_DIR/sglang_patch/modelopt_quant.py:/sgl-workspace/sglang/python/sglang/srt/layers/quantization/modelopt_quant.py:ro" \
  -v "$MODEL_DIR/sglang_patch/flashinfer_trtllm.py:/sgl-workspace/sglang/python/sglang/srt/layers/moe/moe_runner/flashinfer_trtllm.py:ro" \
  -p 30000:30000 \
  lmsysorg/sglang@sha256:8cc6e6f90bf803e9817800b679173d0b526f2b42b2c61b7ecafecdadb610eb55 \
  python3 -m sglang.launch_server \
    --model-path /model \
    --served-model-name ressl/MiniMax-M3-uncensored-NVFP4 \
    --host 0.0.0.0 --port 30000 \
    --tp-size 4 \
    --quantization modelopt_fp4 \
    --trust-remote-code \
    --dtype auto \
    --context-length 32768 \
    --mem-fraction-static 0.90 \
    --max-running-requests 2 \
    --chunked-prefill-size 16384 \
    --page-size 128 \
    --tool-call-parser minimax-m3 \
    --reasoning-parser minimax-m3 \
    --moe-runner-backend flashinfer_cutlass \
    --fp4-gemm-backend flashinfer_cutlass \
    --attention-backend flashinfer \
    --disable-flashinfer-autotune \
    --disable-prefill-cuda-graph \
    --disable-shared-experts-fusion \
    --disable-custom-all-reduce

--page-size 128 is required by MiniMax Sparse Attention. --disable-shared-experts-fusion is
required because this release keeps shared experts in BF16 while routed experts use NVFP4.

Recommended generation settings from the source checkpoint are temperature 1.0 and top-p 0.95.

Validation and limitations

  • This artifact passed serving smoke tests and targeted functional probes, not a full academic
    capability benchmark. No benchmark score from another checkpoint is claimed here.
  • The throughput number is a measured short-request decode rate after warmup. Prompt length,
    concurrency, sampling, and output length will change it.
  • The production validation used 32,768 tokens. The base architecture supports longer contexts, but
    longer configurations were not validated for this export on this four-GPU node.
  • The vision tower is BF16. Color localization and OCR passed, but no full multimodal benchmark was run.
  • vLLM was not accepted for this deployment: the tested SM120 fallback loaded the checkpoint but did
    not correctly execute MiniMax-M3's parameterized clamped-SwiGLU path. Use the validated SGLang setup.
  • Engine support is recent and image-specific. Keep the pinned image and patch checksums for
    reproducible behavior.

Provenance

  • Source repo: ressl/MiniMax-M3-uncensored
  • Source revision: f5faea4659cc3834b1dc3793173db9aa533e3aa7
  • Quantization recipe: streaming/nvfp4_experts_only_input_scale1-kv_none
  • Weight quantization: ModelOpt NVFP4 max, block 16
  • Routed expert projections: 21,888
  • Activation calibration samples: 0
  • Activation input scale: 1.0
  • KV-cache quantization: none
  • Conversion hardware: 4x NVIDIA RTX PRO 6000 Blackwell Workstation Edition, 96 GB
  • Serving hardware: the same four GPUs, tensor parallel size 4

The machine-readable run-manifest.json records the exact source revision, ModelOpt revision,
recipe, counts, sizes, validation result, and wall time. The files in sglang_patch/ are copied from
Mapika/MiniMax-M3-NVFP4, revision
668435825700a0047399441720f430bdd8eca0ab, and are included with their original attribution.

License and credits

The MiniMax license is inherited from the base model by
MiniMaxAI. Uncensoring, quantization, and validation by
Robert Ressl (Hugging Face, Website,
LinkedIn, Patreon).

Built with NVIDIA TensorRT Model Optimizer and SGLang. The SGLang compatibility patch is credited to
Mapika.

Support this work: if these models are useful to you, consider supporting development on
Patreon. More projects are available at
ressl.ch.

README history 1 version

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-07-15Add validated NVFP4 model cardfe9d09d7.7 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration