← back to catalog · registered 2026-08-22 13:56

lhca521/MiniMax-M2.7-abliterated-heretic-ara-AWQ

lhca521 Minimax 227B MoE second-order
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/lhca521%2FMiniMax-M2.7-abliterated-heretic-ara-AWQ"
Response includes
  • classification m3
  • files 39
  • hub_downloads_all_time 1,223
  • author_summary 23 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M3
Primary method

Layer-wise ablation

Applied on top of direct removal inherited from the base model.
Confidence
HIGH
Inherited from base model
Why this label 2 signals
Producer identity confirmed by naming conventions, tags or the model card. This label is very unlikely to change.
  • 'heretic' in model name (Heretic-produced)
  • Heretic uses layer-wise optimization (M3) with underlying direction removal (M1)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
HIGH
Why we say so
name contains 'heretic'; Heretic default extraction is difference-of-means (Arditi 2024)
Downloads · lifetime
1K
60 last 30d - cooling
Likes
1
Model age
5mo ago
created 2026-04-21
Downloads over time
Now1.2K→from22↑5,518%
04529051.4K22 on Apr 221.2K on Oct 11AprMayJunJulAugSepOct
Apr 22 → Oct 11 · 64 snapshots · spans 172 days

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
other
Tags
transformers safetensors minimax_m2 text-generation minimax moe mixture-of-experts abliterated heretic awq w4a16 quantized

Related

Total size
112 GB
Files
39
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-04-21 11:07

Files by quantization

Auxiliary files 39 files 112 GB
model-00015-of-00024.safetensors 4.66 GB 580c7422 download
model-00018-of-00024.safetensors 4.66 GB 9651f47a download
model-00007-of-00024.safetensors 4.66 GB 480a93f4 download
model-00021-of-00024.safetensors 4.66 GB b7597487 download
model-00010-of-00024.safetensors 4.66 GB ea4ec812 download
model-00013-of-00024.safetensors 4.66 GB 6a753628 download
model-00004-of-00024.safetensors 4.66 GB 484d0cc7 download
model-00012-of-00024.safetensors 4.66 GB 3dec9406 download
model-00023-of-00024.safetensors 4.66 GB 0f2317bf download
model-00009-of-00024.safetensors 4.66 GB d7aa22e5 download
model-00020-of-00024.safetensors 4.66 GB 621245de download
model-00006-of-00024.safetensors 4.66 GB e41d8bab download
model-00017-of-00024.safetensors 4.66 GB c6bb9445 download
model-00014-of-00024.safetensors 4.66 GB a079f0d9 download
model-00008-of-00024.safetensors 4.66 GB c9b9d3bd download
model-00011-of-00024.safetensors 4.66 GB 2db72f26 download
model-00016-of-00024.safetensors 4.66 GB d61433d0 download
model-00019-of-00024.safetensors 4.66 GB 1bad40b8 download
model-00022-of-00024.safetensors 4.66 GB e20d1860 download
model-00005-of-00024.safetensors 4.66 GB 8c7ce83d download
model-00003-of-00024.safetensors 4.66 GB 806478d6 download
model-00002-of-00024.safetensors 4.66 GB c445d98e download
model-00001-of-00024.safetensors 4.66 GB 42c4189a download
model-00024-of-00024.safetensors 4.49 GB 9f6895c7 download
tokenizer.json 14.8 MB fc7c5ebc download
model.safetensors.index.json 14.2 MB 91a0acae download
vocab.json 3.72 MB 3394578d download
merges.txt 2.30 MB a4449e69 download
modeling_minimax_m2.py 31.5 KB b8ba2586 download
README.md 13.7 KB be187822 download
tokenizer_config.json 11.0 KB 62024351 download
configuration_minimax_m2.py 9.96 KB fa618e64 download
chat_template.jinja 6.37 KB a09ec0dd download
config.json 5.39 KB e38c1e27 download
special_tokens_map.json 1.60 KB 86de0560 download
.gitattributes 1.60 KB aa7aacd0 download
added_tokens.json 1.43 KB e5a619a7 download
recipe.yaml 895 B e8dae0c5 download
generation_config.json 144 B f7ce0e8e download

README current version from Hugging Face


pipeline_tag: text-generation
license: other
license_name: other
license_link: https://github.com/MiniMax-AI/MiniMax-M2.7/blob/main/LICENSE
library_name: transformers
base_model: Youssofal/MiniMax-M2.7-abliterated-BF16
base_model_relation: quantized
tags:

  • minimax
  • minimax_m2
  • moe
  • mixture-of-experts
  • abliterated
  • heretic
  • awq
  • w4a16
  • quantized
  • compressed-tensors

MiniMax-M2.7-abliterated-heretic-ara-AWQ

AWQ W4A16 (group_size 128, symmetric) quantization of Youssofal/MiniMax-M2.7-abliterated-BF16
— a Heretic-ARA abliterated derivative of MiniMaxAI/MiniMax-M2.7.

⚠️ Decensored model. Safety guardrails have been deliberately removed.
Research and experimentation only. See full disclaimer below.

Quantization Details

Parameter Value
Method AWQ (Activation-aware Weight Quantization)
Scheme W4A16 (symmetric)
Weight Bits 4
Activation Bits 16
Group Size 128
Format compressed-tensors
Calibration Dataset HuggingFaceH4/ultrachat_200k
Calibration Samples 128
Max Sequence Length 512
Router Gates Unquantized (full precision)
LM Head Unquantized (full precision)
Experts Calibrated All 256 per layer (custom all-experts patch)
Compatible Inference Engine vLLM

Quantization Notes

  • Router gates kept full precision: MiniMax-M2's MoE uses sigmoid routing
    with an e_score_correction_bias. Quantizing the gate or this bias destroys
    routing decisions and produces multilingual gibberish. Both are explicitly
    excluded from quantization.
  • All 256 experts calibrated: With top-8 routing on a 256-expert model,
    naive AWQ leaves most experts with insufficient calibration data. This
    quantization uses a custom MoE forward patch during calibration that runs
    every expert on every batch (sparse for routed tokens, with a small dummy
    forward for unrouted experts to fire activation hooks).
  • QK-Norm absorbs Q/K smoothing: MiniMax-M2 has use_qk_norm=true. The
    RMSNorm weights of q_norm/k_norm absorb AWQ's per-channel scales applied
    to q_proj/k_proj outputs, preserving correctness.
  • GQA v→o smoothing skipped: With 8 KV heads and 48 query heads, the
    v_proj output dimensions don't match o_proj input dimensions for per-channel
    smoothing. llm-compressor correctly identifies this incompatibility and
    skips it.
  • MTP heads not preserved: The base checkpoint contains multi-token
    prediction (MTP) module weights that HF's MiniMaxM2ForCausalLM doesn't
    use. These are dropped from the quantized model.

Deployment

Recommended inference with vLLM:

vllm serve alonsoko/MiniMax-M2.7-abliterated-heretic-ara-AWQ \
    --trust-remote-code \
    --tensor-parallel-size 4 \
    --tool-call-parser minimax_m2 \
    --reasoning-parser minimax_m2 \
    --enable-auto-tool-choice

Recommended sampling: temperature=1.0, top_p=0.95, top_k=40 (per upstream MiniMax-M2 guidance).

MiniMax-M2 is an interleaved thinking model — when chaining assistant turns,
preserve <think>...</think> blocks from prior turns in the message history.

Hardware Requirements

Approximate VRAM for inference at this quantization (W4A16-G128):

  • Weights: ~115 GB
  • KV cache (per request, varies with context length): ~2-8 GB
  • Recommended: 2× 80GB GPUs (e.g., A100/H100) with --tensor-parallel-size 2,
    or 4× 48GB GPUs (e.g., L40S/A6000) with --tensor-parallel-size 4
  • Minimum: A single 141GB H200 should fit weights + modest context

This is a decensored version of MiniMaxAI/MiniMax-M2.7, made using Heretic v1.2.0+custom with the Arbitrary-Rank Ablation (ARA) method

⚠️ Disclaimer

This model is intended for research, experimentation, and testing purposes only.

  • This model may produce harmful, offensive, inappropriate, or otherwise objectionable content.
  • The abliteration process removes safety guardrails that were intentionally built into the original model.
  • Do not use this model in production systems, consumer-facing applications, or any context
    where harmful outputs could cause real-world harm.
  • The authors and contributors of this toolkit bear no responsibility for any misuse of this model
    or any harm caused by outputs generated by this model.
  • By using this model, you agree that you are solely responsible for ensuring its use complies
    with all applicable laws and ethical guidelines.

This model is shared purely for academic and technical exploration of model internals.

Abliteration parameters

Parameter Value
start_layer_index 30
end_layer_index 51
preserve_good_behavior_weight 0.4512
steer_bad_behavior_weight 0.0037
overcorrect_relative_weight 0.8804
neighbor_count 14

About the Base Model

Original model: MiniMaxAI/MiniMax-M2.7

MiniMax-M2.7 is a 230B-parameter sparse MoE (10B active) built for agentic
workflows, coding, and tool use. It uses interleaved thinking with <think>...</think>
blocks. See the base model card
for capabilities, benchmarks, and deployment details.

README history 1 version

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-04-21Duplicate from alonsoko/MiniMax-M2.7-abliterated-heretic-ara-AWQ70adc7f13.7 KB
    Loading...

Discussions 1 thread

  1. 2026-04-23Reportopen1 💬#1
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration