For privacy reasons a browser tells us at most
"≥ 8 GB RAM, 8 cores" - same reading whether you have 8 GB
or 128 GB. It has no idea how much RAM is free right now, which
apps are open, or whether you have a GPU.
The Abliteration app is integrated with your machine
It reads your exact RAM, GPU model and VRAM,
free memory right now, and picks the sharpest quant that
still fits. Every model page lights up precisely for your rig.
And you can chat with any model, right now
The app is a full local runtime - no API keys, no subscription,
everything runs on your machine. Click any model on this site and
start a conversation in seconds.
Streaming per-shard quantization script: for each Linear weight W, compute per-output-channel scale = |W|.amax(dim=1) / 448.0, then W_fp8 = (W / scale).to(fp8_e4m3fn). No calibration data required (FP8_DYNAMIC scheme).
License
Inherits the non-commercial MiniMax M-Series license from the base model.
README history
2 versions
The author's README evolved over time. Click a version to see its content at that point.
2026-04-17Add model card with proper lineage (Youssofal/BF16 → FP8)b4aff3a2 KB
Loading...
2026-04-17Add files using upload-large-folder tooleb7e2e6753 B
Loading...
Catalog is the map. Apps are the tools.
Run models on your own machine, not in the cloud.
Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.