For privacy reasons a browser tells us at most
"≥ 8 GB RAM, 8 cores" - same reading whether you have 8 GB
or 128 GB. It has no idea how much RAM is free right now, which
apps are open, or whether you have a GPU.
The Abliteration app is integrated with your machine
It reads your exact RAM, GPU model and VRAM,
free memory right now, and picks the sharpest quant that
still fits. Every model page lights up precisely for your rig.
And you can chat with any model, right now
The app is a full local runtime - no API keys, no subscription,
everything runs on your machine. Click any model on this site and
start a conversation in seconds.
Conversion was only possible thanks to the glm_moe_dsaindexer-sharing fix from pcuenca's glm-moe-dsa-indexer-sharing branch (mlx_lm/models/glm_moe_dsa.py). Without it, mlx_lm convert throws ValueError: Missing 285 parameters because the DSA sparse-indexer tensors don't map under released mlx-lm 0.31.3. That fix tracks ml-explore/mlx-lm PR #1410 (not yet merged). Discovered via mlx-community/GLM-5.2-DQ4plus-q8 discussion #1. Drop that patched glm_moe_dsa.py into your mlx_lm/models/ dir before converting/loading.
Toolchain:mlx-lm==0.31.3 + mlx==0.32.0 + Python 3.11, Apple Silicon (512 GB Macs, served via exo).
Variant specifics (this repo: Q4 (uniform 4-bit))
Uniform 4-bit affine quantization (group_size 64). Most compact (~418 GB); largest quality tradeoff.
Reproduction
# 1. patch the model class (see "Conversion was only possible thanks to" above)
curl -L https://raw.githubusercontent.com/pcuenca/mlx-lm/glm-moe-dsa-indexer-sharing/mlx_lm/models/glm_moe_dsa.py \
-o $(python -c "import mlx_lm,os;print(os.path.dirname(mlx_lm.__file__)+'/models/glm_moe_dsa.py')")
# 2. convert (from the local FP8 source snapshot dir)
mlx_lm convert --hf-path <GLM-5.2-FP8-Uncensored snapshot> --mlx-path ./out \
-q --q-bits 4 --q-group-size 64 --q-mode affine
# for the Q6+down8 mixed variant, load the Q8 output lazily and re-quantize
# with a predicate pinning MoE down_proj (w2) to 8-bit, everything else to 6-bit:
# from mlx_lm.utils import load, save, quantize_model
# m,tok = load("./q8_out", lazy=True)
# def pred(path, mod):
# return {"bits":8,"group_size":64} if path.endswith(".w2.weight") else {"bits":6,"group_size":64}
# qm,qc = quantize_model(m, {"quantization":{"group_size":64,"bits":8,"mode":"affine"}}, group_size=64, bits=None, quant_predicate=pred)
# save("./q6down8", "./q8_out", qm, tok, json.load(open("./q8_out/config.json")))
2026-07-12Add files using upload-large-folder tool001e14e573 B
Loading...
Catalog is the map. Apps are the tools.
Run models on your own machine, not in the cloud.
Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.