license: mit
pipeline_tag: text-generation
base_model:
- madeby561/GLM-5.2-MXFP8-NVFP4-NF3-Hybrid
- Jiunsong/SuperGLM-5.2-abliterated-NVFP4
tags: - glm
- glm-5.2
- mixture-of-experts
- text-generation
- abliterated
- mxfp8
- nvfp4
- nf3
- quantization
- blackwell
- vllm
GLM-5.2 Abliterated MXFP8/NVFP4/NF3 Hybrid
This is a full-expert GLM-5.2 hybrid checkpoint that combines the compact
MXFP8/NVFP4/NF3 layout from
madeby561/GLM-5.2-MXFP8-NVFP4-NF3-Hybrid
with selected refusal-reduction tensors from
Jiunsong/SuperGLM-5.2-abliterated-NVFP4.
It retains all 256 routed experts per layer and is intended for custom
Blackwell inference stacks that support the upstream hybrid format.
The served model identifier used during qualification was:
glm-5.2-abliterated-MXFP8-NVFP4-NF3-Hybrid
Runtime topology advisory (2026-07-30)
The checkpoint is not affected, but the upstream serving runtime has a known
topology hazard. The pinned July 26 v20 image assumes that all participating
GPUs can use its B12X PCIe one-shot allreduce path. On four-GPU hosts where
P2P works only within GPU pairs rather than across every rank, worker
initialization can hang indefinitely while allocating shared CUDA IPC buffers.
NCCL may already have initialized successfully, so the stall can look like a
compile or model-load delay.
Run this before using the upstream Compose profiles:
nvidia-smi topo -p2p r
If any off-diagonal rank pair in CUDA_VISIBLE_DEVICES reports NS, do not
use the stock v20 GLM launcher unchanged. The affected pinned image is:
voipmonitor/vllm:gilded-gnosis-v20-vllm0c79e41-sie603f74-fi801d57a-cu132-20260726
sha256:10261c7d65101c8aba2ce1fb59eabe73aff9d35eca5043b330cc0ce76d3c98d0
The July 30 r13 image is not yet a safe replacement for partial-mesh
topologies. Although its vLLM default is opt-in, the image environment still
sets VLLM_ENABLE_PCIE_ALLREDUCE=1, and the GLM launcher chain sourcesglm52-pcie-runtime-env.sh, which exports the flag and B12X backend again:
voipmonitor/vllm:gilded-gnosis-v20-vllm69ba80b-sia2ea608-fi801d57a-cu132-20260730-r13
sha256:02796036c96a52fda0919aa260c45c70bc97d8e662a6ae5e614b5f987c20851b
Setting VLLM_ENABLE_PCIE_ALLREDUCE=0 only in Compose is therefore not enough
when the affected launcher overwrites it. The four-GPU qualification reported
on this card used a host-specific runtime override that bypassed the stock
launcher, disabled PCIe custom allreduce and B12X DCP A2A, and selected theag_rs DCP backend. That result should not be interpreted as qualification of
the stock v20 launcher on arbitrary PCIe layouts.
A portable partial-mesh profile will be published only after it has been
validated independently. Until then, operators should either use a full-mesh
P2P host or apply an explicit runtime patch that prevents_b12x_pcie_allreduce_requested() from selecting the one-shot path and useag_rs, verifying the complete launch on their own hardware. Follow the open
compatibility discussion
for updates.
What changed
The release starts from the pinned madeby561 hybrid checkpoint and replaces
124 lineage-matched BF16 tensors:
- 62 attention output projections (
o_proj) - 62 shared-expert down projections (
down_proj)
The replacement tensors are byte-exact copies from the pinned Jiunsong
abliterated donor. Before transplantation, every source tensor was verified
against the donor's pinned NVIDIA base. After transplantation, every output
tensor was verified against the donor.
The following parts remain unchanged from the madeby561 hybrid source:
- all routed experts
lm_head- MTP weights
- quantization metadata and mixed-precision tier layout
- tokenizer, chat template, and model configuration
The changes touch 71 of 184 safetensor shards and 14,042,529,792 tensor bytes.
See DERIVATION.json for the pinned revisions and machine-readable invariants.
Lineage
| Role | Repository | Pinned revision |
|---|---|---|
| Hybrid source | madeby561/GLM-5.2-MXFP8-NVFP4-NF3-Hybrid |
68babde27a97a4c980c2494e830dd424975cd5a3 |
| Abliterated donor | Jiunsong/SuperGLM-5.2-abliterated-NVFP4 |
076582b8a58d3f924a68af550a630edffada5e95 |
| Donor base | nvidia/GLM-5.2-NVFP4 |
aec724e8c7b8ee9db3b48c01c320f63f9cdaf8aa |
| Original model | zai-org/GLM-5.2 |
b4734de4facf877f85769a911abafc5283eab3d9 |
The donor card describes its OBLITERATUS procedure, evaluation methodology,
and leakage controls. Those donor results should not be treated as benchmark
results for this hybrid derivative.
Download
hf download <repository-id> \
--local-dir ./glm-5.2-abliterated-hybrid
Serving
This is not a generic Transformers checkpoint. It requires the custom hybrid
loader and NF3 kernel described by the
madeby561 source release.
Standard vLLM, SGLang, and Transformers installations are not expected to load
it directly.
Keep mxfp8_tier_nokvb.json beside the checkpoint. The source release provides
the public serving image and baseline compose recipe. Begin with the source
release's conservative context and memory settings, then qualify larger
contexts on the target hardware.
This derivative was load-tested on four 96 GB NVIDIA Blackwell GPUs with:
- tensor parallel 4
- decode-context parallel 4
- MTP speculative decoding
- NVFP4 DS-MLA KV cache
- 480,000 served context
The configuration advertises a 1,048,576-token architectural maximum, but this
release does not claim that a 1M-token serving configuration will fit or remain
stable on every runtime or hardware layout.
Validation
Release validation covered:
- complete safetensor index resolution across 184 shards
- source-to-base and output-to-donor hash lineage for all 124 modified tensors
- model load and health at 480,000 served context
- deterministic short-response checks
- native tool-call formatting
- bounded refusal-reduction canaries
No post-transplant GPQA, coding, safety, or long-context benchmark suite is
claimed yet. The madeby561 source card's benchmark numbers describe the source
checkpoint, not this derivative.
Intended use and limitations
This release is intended for local inference research, agent systems, coding,
reasoning, and evaluation on hardware capable of loading the full hybrid
checkpoint.
"Abliterated" means selected refusal-associated projections were replaced from
an abliterated donor. It does not guarantee unrestricted behavior, correctness,
or the absence of refusals. The model can generate inaccurate, unsafe, biased,
or legally sensitive output. Operators remain responsible for validation,
access controls, applicable law, and use-case-specific safeguards.
Privacy and serialization
The published tree is an explicit allowlist. It contains safetensors weights,
model/tokenizer configuration, the chat template, the MXFP8 tier map, this
model card, the license, and the public derivation record. It excludes runtime
environment files, credentials, internal hostnames, network addresses, local
filesystem paths, operational logs, and private build manifests.
All weight shards use safetensors; no pickle-based model files are included.
License
The model and its cited upstream checkpoints are published under the MIT
License. The included LICENSE preserves the original copyright notice.