For privacy reasons a browser tells us at most
"≥ 8 GB RAM, 8 cores" - same reading whether you have 8 GB
or 128 GB. It has no idea how much RAM is free right now, which
apps are open, or whether you have a GPU.
The Abliteration app is integrated with your machine
It reads your exact RAM, GPU model and VRAM,
free memory right now, and picks the sharpest quant that
still fits. Every model page lights up precisely for your rig.
And you can chat with any model, right now
The app is a full local runtime - no API keys, no subscription,
everything runs on your machine. Click any model on this site and
start a conversation in seconds.
This release uses AMD Quark's GLM MoE DSA recipe and standard Hugging Face safetensors. Expert matrices use OCP MXFP4; attention, routers, the initial dense MLPs, embeddings, normalization tensors, and the output head remain at higher precision. This is a mixed-precision checkpoint, not a claim that every tensor is four-bit.
The source model card identifies JANGQ-AI/GLM-5.3-FP8 as its base and zai-org/GLM-5.3 as the upstream model.
This is FP8 → MXFP4 requantization. Source FP8 weights are dequantized using their original block scales before MXFP4 conversion. It does not recover precision already lost in the FP8 source.
No additional training was performed. The tokenizer and chat template are inherited from the source.
Conversion time, excluding download/validation/upload
7.89 minutes
Sizes above count safetensors files, including their headers. Tensor-element counts include non-scale buffers and are not a separately audited trainable-parameter count. Overall storage includes higher-precision tensors and quantization metadata. Exact values are in quantization_stats.json; per-shard sizes and SHA-256 hashes are in artifact_manifest.json.
Method: file-to-file round-to-nearest quantization; no calibration dataset, GPTQ, AWQ, or fine-tuning.
Weight format: FP4 E2M1 values, two values per byte, with one E8M0 scale per group of 32 consecutive weights.
Activation policy for quantized layers: dynamic MXFP4, as specified in config.json.
Quark recipe: LLMTemplate.get("glm_moe_dsa").get_config(scheme="mxfp4"), with *eh_proj additionally excluded to preserve the MTP auxiliary projection in BF16.
Excluded source FP8 linear weights are recovered to BF16. Existing floating-point excluded tensors retain their source dtype.
Hardware available: 8× AMD Instinct MI355X; this file-to-file FP8 conversion used one GPU (AMD Radeon Graphics), following Quark's supported single-device FP8 recovery path.
Architecture
Item
Source configuration
Architecture
GlmMoeDsaForCausalLM
Model type
glm_moe_dsa
Hidden layers
78
Hidden size
6144
Routed experts
256
Experts selected per token
8
Shared experts
1
Initial dense layers
3
Vocabulary size
154,880
Configured maximum positions
1,048,576
The configured context limit is inherited metadata, not a measured serving capacity. Actual context and memory requirements depend on the runtime, KV cache, batch size, and hardware.
Passed independent FP4 decoding and finite-value checks on 64 matrices / 4,980,736 weight elements, sampled across the checkpoint.
Sampled weight reconstruction relative RMSE vs recovered FP8 source: 0.111689; absolute RMSE: 0.00184817. Sampling uses the first 16 rows of evenly spaced matrix names; this is not a random or exhaustive quality estimate.
End-to-end inference, perplexity, downstream accuracy, tokens/second, and runtime VRAM consumption have not been measured for this export. Upstream benchmark scores are not results for this quantization.
Use a runtime supporting both GLM-5.3 / glm_moe_dsa and AMD Quark MXFP4. The intended serving path is vLLM with ROCm on supported AMD Instinct hardware. A generic Transformers installation or a runtime supporting only GPT-OSS MXFP4 is not sufficient evidence of compatibility.
The upstream model uses the GLM-5.3 License, preserved verbatim in this repository. The immediate source repository is tagged MIT, but the upstream model includes its own license conditions; this release is not labeled MIT-only. See LICENSE and LICENSE_PROVENANCE.json. Credit to Z.ai for GLM-5.3, JANGQ-AI for the FP8 base, dealignai for the supplied source checkpoint, and AMD for Quark. Kitani performed this MXFP4 conversion and packaging.
Catalog is the map. Apps are the tools.
Run models on your own machine, not in the cloud.
Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.