← back to catalog · registered 2026-08-22 13:56

ghostexee/MiniMax-M3-uncensored-RP-LC-Q4HQ-GGUF

ghostexee Minimax GGUF MoE second-order 1.0M ctx
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/ghostexee%2FMiniMax-M3-uncensored-RP-LC-Q4HQ-GGUF"
Response includes
  • classification m8
  • files 22
  • hub_downloads_all_time 261
  • author_summary 1 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M8
Primary method

Repackaging (quantization)

Applied on top of direct removal inherited from the base model.
Confidence
MEDIUM
Inherited from base model
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=1
  • assume M1 (base ablation) + M8 (GGUF quant) - default when producer unknown
Refusal direction extraction

No specific extraction method could be identified for this model. The producer either did not document it or used a proprietary pipeline.

What is a refusal direction? →
Downloads · lifetime
261
53 last 30d - stable
Likes
4
Model age
2mo ago
created 2026-08-04
Downloads over time
Now270→from153↑76%
147192237282153 on Aug 5270 on Oct 11270 on Oct 10AugSepOct
Aug 5 → Oct 11 · 50 snapshots · spans 67 days

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
other
Languages
en
Tags
gguf llama.cpp minimax minimax-m3 moe msa q4_k_m roleplay creative-writing long-context uncensored abliterated

Related

Total size
261 GB
Files
22
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-08-04 08:24

Files by quantization

Auxiliary files 22 files 261 GB
MiniMax-M3-uncensored-RP-LC-Q4HQ-00004-of-00018.gguf 15.4 GB 14e5c733 download
MiniMax-M3-uncensored-RP-LC-Q4HQ-00007-of-00018.gguf 15.4 GB 7cce898d download
MiniMax-M3-uncensored-RP-LC-Q4HQ-00010-of-00018.gguf 15.4 GB ec18a1e6 download
MiniMax-M3-uncensored-RP-LC-Q4HQ-00013-of-00018.gguf 15.4 GB 400759ac download
MiniMax-M3-uncensored-RP-LC-Q4HQ-00016-of-00018.gguf 15.4 GB 10faddcf download
MiniMax-M3-uncensored-RP-LC-Q4HQ-00001-of-00018.gguf 15.4 GB 7cf0390c download
MiniMax-M3-uncensored-RP-LC-Q4HQ-00005-of-00018.gguf 15.0 GB 4a4cf626 download
MiniMax-M3-uncensored-RP-LC-Q4HQ-00002-of-00018.gguf 15.0 GB 630b486e download
MiniMax-M3-uncensored-RP-LC-Q4HQ-00008-of-00018.gguf 15.0 GB b95be56d download
MiniMax-M3-uncensored-RP-LC-Q4HQ-00011-of-00018.gguf 15.0 GB 2c6c5f9d download
MiniMax-M3-uncensored-RP-LC-Q4HQ-00017-of-00018.gguf 15.0 GB b94e261f download
MiniMax-M3-uncensored-RP-LC-Q4HQ-00014-of-00018.gguf 15.0 GB 0cc1f039 download
MiniMax-M3-uncensored-RP-LC-Q4HQ-00006-of-00018.gguf 14.8 GB eced5f79 download
MiniMax-M3-uncensored-RP-LC-Q4HQ-00003-of-00018.gguf 14.8 GB 82c69d6e download
MiniMax-M3-uncensored-RP-LC-Q4HQ-00009-of-00018.gguf 14.8 GB 8eab7ac0 download
MiniMax-M3-uncensored-RP-LC-Q4HQ-00012-of-00018.gguf 14.8 GB 0f564307 download
MiniMax-M3-uncensored-RP-LC-Q4HQ-00015-of-00018.gguf 14.8 GB f4fe5cb1 download
MiniMax-M3-uncensored-RP-LC-Q4HQ-00018-of-00018.gguf 4.52 GB 846f67a8 download
README.md 8.33 KB 9b4e2051 download
LICENSE 3.26 KB f413dfa3 download
.gitattributes 3.05 KB c927f165 download
SHA256SUMS 2.09 KB ffa1f6fb download

README current version from Hugging Face


license: other
license_name: minimax-community-license
license_link: https://huggingface.co/MiniMaxAI/MiniMax-M3/blob/main/LICENSE
base_model: ressl/MiniMax-M3-uncensored
base_model_relation: quantized
library_name: gguf
pipeline_tag: text-generation
language:

  • en
    tags:
  • gguf
  • llama.cpp
  • minimax
  • minimax-m3
  • moe
  • msa
  • q4_k_m
  • roleplay
  • creative-writing
  • long-context
  • uncensored
  • abliterated

MiniMax-M3-uncensored RP-LC-Q4HQ GGUF

High-quality 5.26 bpw GGUF quantization of
ressl/MiniMax-M3-uncensored,
optimized for character-heavy fantasy roleplay and long-context continuity.
That checkpoint is an uncensored/abliterated derivative of
MiniMaxAI/MiniMax-M3.

[!CAUTION]
Risk level: high. This model is genuinely uncensored. The source reports
zero hard refusals on its harmful-prompt sample. Quantization does not restore
safety alignment. The model can generate dangerous, illegal, abusive,
sexually explicit, manipulative, or otherwise harmful content, including
operational cyber or violence-related instructions. Do not expose it to
untrusted users without strong access controls, monitoring, policy
enforcement, and output safeguards. You are responsible for lawful and safe
use.

What this is

This is a quantized derivative, not a new finetune. No gradient training was
performed. A roleplay-domain importance matrix was used to choose quantization
scales while the final GGUF was produced directly from the ReSSl BF16 weights.

  • Format: 18-shard, text-only GGUF
  • Size: 280,052,357,504 bytes (260.819 GiB)
  • Effective precision: 5.26 bits per weight
  • Architecture: MiniMax-M3, 428B total / approximately 23B active MoE
  • Context metadata: 1,048,576 tokens
  • Tested context: 65,536 tokens
  • Native MiniMax-M3 chat template retained
  • Vision tower and MTP head omitted

Start with MiniMax-M3-uncensored-RP-LC-Q4HQ-00001-of-00018.gguf;
llama.cpp discovers the other shards automatically.

Quantization recipe

This is deliberately more conservative than a conventional all-Q4 build:

Tensor group Type Count
Routed expert gate/up Q4_K 114
Attention, routed-down, shared/dense expert paths Q6_K 477
MSA indexer Q/K projections F32 114
Token embedding and output Q8_0 2
Routers and normalization tensors F32 retained

Residual-writing down projections, all attention projections, shared experts,
and the first three dense MLP blocks were kept at Q6_K. MiniMax-M3's sparse
attention indexers were retained at F32 because they directly determine which
long-context blocks remain visible.

Roleplay calibration

The quantization importance matrix used 655,360 tokens from a deterministic
fantasy/creative roleplay corpus:

Component Windows Tokens Share
Short-window matrix 128 x 4,096 524,288 80%
Long-window supplement 8 x 16,384 131,072 20%

The source pool mixed character dialogue, multi-character roleplay, creative
fiction, and narrative worldbuilding from
Timersofc/creative-writing-reap-calibration
and agentlans/combined-roleplay.
The combined matrix contained 762 tensor entries with 98.44% minimum routed-
expert coverage.

This calibration changes quantization error allocation only. It does not add
new roleplay training or alter the source model's behavior through finetuning.

Evaluation

Held-out perplexity

Four document-disjoint 4,096-token fantasy-roleplay chunks:

Model PPL Relative to Q8
Q8 reference 3.7610 +/- 0.09740 baseline
Earlier 4K-only RP-Q4HQ 3.7819 +/- 0.09923 +0.56%
This RP-LC-Q4HQ 3.7669 +/- 0.09864 +0.16%

Long-context continuity

A deterministic 61,500-token user prompt became 61,709 tokens after the
native chat template. Six trusted campaign records were placed from token 124
through token 53,321 among untrusted fantasy-roleplay distractors.

  • Exact canon recall: 14/14
  • Continuation length: 541 words (requested range: 450-650)
  • Canon facts used naturally in the scene: 8 (requested minimum: 6)
  • Distinct NPC voices: pass
  • Player-character agency preserved: pass
  • Concrete action opening at the ending: pass

On a 512 GiB Apple M3 Ultra with a 65,536-token context:

  • Prompt processing: 189.5 tokens/s
  • Generation: 16.1 tokens/s
  • Peak process RSS: 271.05 GiB
  • Process swaps: 0

These are targeted quantization checks, not a comprehensive capability or
safety evaluation. The 1M context declared by the architecture was not tested.

Runtime

The model requires MiniMax-M3's trained MSA sparse-attention implementation. It
was built and validated with llama.cpp build 10018 at commit
e99545c1c41ba42b7986831c0de2983498dc3c5b,
from the MiniMax-M3 MSA support work.
Use that revision or a later llama.cpp version with equivalent MiniMax-M3 MSA
support. Substituting dense attention is not equivalent beyond the dense
prefix.

huggingface-cli download m9e/MiniMax-M3-uncensored-RP-LC-Q4HQ-GGUF \
  --local-dir MiniMax-M3-uncensored-RP-LC-Q4HQ-GGUF

llama-cli \
  -m MiniMax-M3-uncensored-RP-LC-Q4HQ-GGUF/MiniMax-M3-uncensored-RP-LC-Q4HQ-00001-of-00018.gguf \
  -ngl all -fa on -fit off -c 65536 -cnv \
  -rea off --reasoning-budget 0 \
  --temp 1.0 --top-p 0.95 --min-p 0.05

The complete model occupies about 261 GiB before runtime buffers and context.
The tested 65K configuration peaked near 271 GiB. Plan hardware capacity with
additional headroom; smaller context allocations reduce runtime overhead.

Intended uses

  • Private or access-controlled creative writing and fantasy roleplay
  • Long-running campaign continuity experiments
  • Controlled research into uncensored model behavior
  • Lawful, authorized security research and red-team analysis with safeguards

Safety, limitations, and out-of-scope use

  • The source checkpoint intentionally suppresses refusal behavior. Treat model
    output as untrusted data and never as authorization to act.
  • Do not use it to facilitate crime, malware deployment, unauthorized access,
    violence, self-harm, exploitation, privacy abuse, harassment, or deception.
  • Roleplay output may include graphic violence, sexual material, coercion, or
    other disturbing themes. Use explicit consent and age/content controls. Never
    use it to sexualize or exploit minors.
  • Do not rely on it for medical, legal, financial, emergency, or other
    high-stakes decisions.
  • It may hallucinate facts, lose continuity, imitate copyrighted styles, or
    reproduce biases from its source models and calibration text.
  • This GGUF is not a security boundary. Applications should add authentication,
    rate limits, logging, abuse monitoring, prompt/output filtering, and human
    review appropriate to their risk.
  • The MiniMax Community License contains binding prohibited-use and commercial-
    use conditions. Read LICENSE before use or redistribution.

Lineage and credits

  1. Original model: MiniMaxAI/MiniMax-M3,
    by MiniMaxAI. Source revision used by the ReSSl derivative:
    50942730318c7943fe83db7ec8e9f9177ecb1cf8.
  2. Uncensored BF16 source: ressl/MiniMax-M3-uncensored,
    uncensoring and validation by Robert Ressl. Revision:
    315b596663fbe37fce3880d7c0468ceb47cd2da5.
  3. GGUF conversion, calibration, quantization, and evaluation: m9e.
  4. Inference runtime: llama.cpp.

Thanks to the authors and maintainers of the calibration datasets named above.

License

This derivative inherits the MiniMax Community License from the original
model. The license is included in this repository and is also available in the
upstream MiniMax-M3 repository.
It includes attribution, commercial-use, and prohibited-use requirements. This
model is provided as-is, without warranty.

Integrity

SHA256SUMS contains a digest for every GGUF shard. Verify after download:

shasum -a 256 -c SHA256SUMS

README history 1 version

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-08-04Duplicate from m9e/MiniMax-M3-uncensored-RP-LC-Q4HQ-GGUF60fdf7b8.3 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration