← back to catalog · registered 2026-08-22 13:56

random-robbie/MiniMax-M2.7-NVFP4-abliterated

random-robbie Minimax 112B MoE second-order
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/random-robbie%2FMiniMax-M2.7-NVFP4-abliterated"
Response includes
  • classification m1
  • files 27
  • hub_downloads_all_time 812
  • author_summary 1 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M1
Primary method

Direct removal

No other method signals detected in this model.
Confidence
MEDIUM
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=0 (base model)
  • no specific method indicators - defaulting to M1 (most common)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
MEDIUM
Why we say so
primary_method=M1; difference-of-means is the reference extraction for M1/M3 (Arditi 2024)
Downloads · lifetime
812
39 last 30d - cooling
Likes
0
Model age
3mo ago
created 2026-06-23
Downloads over time
Now823→from204↑303%
173410648885204 on Jun 24823 on Oct 11JunJulAugSepOct
Jun 24 → Oct 11 · 55 snapshots · spans 109 days

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
other
Languages
en zh
Tags
safetensors minimax_m2 minimax moe nvfp4 blackwell abliterated uncensored vllm modelopt text-generation conversational

Related

Total size
125 GB
Files
27
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-06-23 20:42

Files by quantization

Auxiliary files 27 files 125 GB
model-00001-of-00014.safetensors 10.3 GB 5ba5b843 download
model-00012-of-00014.safetensors 9.46 GB 3751048b download
model-00007-of-00014.safetensors 9.46 GB 06d55d33 download
model-00011-of-00014.safetensors 9.46 GB df7352b3 download
model-00005-of-00014.safetensors 9.46 GB 0700374c download
model-00006-of-00014.safetensors 9.46 GB ff314d5d download
model-00009-of-00014.safetensors 9.46 GB 58ad54d2 download
model-00010-of-00014.safetensors 9.46 GB cd695d83 download
model-00003-of-00014.safetensors 9.46 GB 8dc7ef3a download
model-00002-of-00014.safetensors 9.46 GB 77796b67 download
model-00008-of-00014.safetensors 9.40 GB aee6f116 download
model-00004-of-00014.safetensors 9.40 GB 214d57aa download
model-00013-of-00014.safetensors 9.40 GB f5a8ad55 download
model-00014-of-00014.safetensors 1.56 GB 9ed399b5 download
model.safetensors.index.json 18.6 MB 90e3353c download
tokenizer.json 14.8 MB 1497f696 download
merges.txt 2.30 MB a4449e69 download
modeling_minimax_m2.py 31.3 KB e7198d92 download
hf_quant_config.json 18.3 KB edd66284 download
tokenizer_config.json 10.9 KB d9ac4d99 download
configuration_minimax_m2.py 9.92 KB 7fcd9861 download
chat_template.jinja 6.37 KB a09ec0dd download
README.md 4.34 KB f9698b95 download
config.json 1.61 KB 47e1a7c5 download
.gitattributes 1.60 KB aa7aacd0 download
added_tokens.json 1.43 KB e5a619a7 download
generation_config.json 144 B 945acc24 download

README current version from Hugging Face


license: other
license_name: minimax-model-license
license_link: https://github.com/MiniMax-AI/MiniMax-M2.7/blob/main/LICENSE
language:

  • en
  • zh
    tags:
  • minimax
  • moe
  • nvfp4
  • blackwell
  • abliterated
  • uncensored
  • vllm
  • modelopt
    base_model:
  • llmfan46/MiniMax-M2.7-BF16-ultra-uncensored-heretic
    pipeline_tag: text-generation

MiniMax-M2.7 — Abliterated + NVFP4

MiniMax-M2.7 with refusal directions removed, requantized to NVIDIA NVFP4 for Blackwell GPUs.

What's changed

The source BF16 model has refusal behaviours removed at the weight level via abliteration. This NVFP4 version is a direct requantization of those weights — the architecture, tokenizer, and all other parameters are identical to the original MiniMax-M2.7.

Quantization details match nvidia's original NVFP4 format exactly:

  • MoE expert weights: NVFP4 (E2M1 FP4, group_size=16, FP8 per-block scales, FP32 per-tensor scale)
  • KV cache: FP8
  • Kept in BF16: embeddings (embed_tokens), attention projections (self_attn.*), MoE router (block_sparse_moe.gate), lm_head
  • input_scale scalars borrowed from nvidia/MiniMax-M2.7-NVFP4 calibration data

Model architecture

Parameter Value
Total parameters ~230B
Active parameters per token ~10B
Hidden size 3072
Layers 62
Experts (total / active) 256 / 8
Attention heads 48 (GQA, 8 KV heads)
Max context 204,800 tokens
Vocabulary 200,064 tokens

Hardware requirements

NVIDIA Blackwell only. NVFP4 (E2M1) is a Blackwell-native format and requires:

  • B200, GB200, or RTX 5090 (or later Blackwell GPUs)
  • Minimum 2× B200 (2× 180GB = 360GB HBM) recommended for tensor parallel serving
  • vLLM 0.22+ with modelopt quantization backend

This model will not run on Ampere or Hopper GPUs. If you need a version compatible with older hardware, use a GGUF or GPTQ quantization of the source BF16 model instead.

Usage

vLLM (recommended)

vllm serve random-robbie/MiniMax-M2.7-NVFP4-abliterated \
  --served-model-name minimax-m2.7 \
  --tensor-parallel-size 2 \
  --quantization modelopt \
  --kv-cache-dtype fp8 \
  --max-model-len 196608 \
  --gpu-memory-utilization 0.92 \
  --trust-remote-code

OpenAI-compatible API

from openai import OpenAI

client = OpenAI(base_url="http://localhost:8000/v1", api_key="none")

response = client.chat.completions.create(
    model="minimax-m2.7",
    messages=[{"role": "user", "content": "Hello!"}],
    max_tokens=512,
)
print(response.choices[0].message.content)

File structure

File Description
model-000XX-of-00014.safetensors 14 NVFP4 weight shards (~126GB total)
model.safetensors.index.json Tensor → shard index
hf_quant_config.json modelopt quantization metadata
config.json Model architecture config
modeling_minimax_m2.py Custom model code (trust_remote_code)
tokenizer.json / tokenizer_config.json Tokenizer

License

This model is derived from MiniMaxAI/MiniMax-M2.7 and is subject to the MiniMax Model License. Please review that license before use, particularly regarding commercial applications.

The abliteration technique (weight-level removal of refusal directions) was applied to the intermediate BF16 weights by llmfan46. This NVFP4 requantization was produced by random-robbie.

Credits

README history 2 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-06-23Fix license links to MiniMax-M2.7 repofa9379c4.3 KB
    Loading...
  2. 2026-06-23Add model card0ad06e04.3 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration