← back to catalog · registered 2026-08-22 13:56

kachowtowmater/SuperGLM-5.2-abliterated-MXFP8-NVFP4-NF3-Hybrid

kachowtowmater Glm 299B MoE
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/kachowtowmater%2FSuperGLM-5.2-abliterated-MXFP8-NVFP4-NF3-Hybrid"
Response includes
  • classification m1
  • files 101
  • benchmarks 11 entries
  • hub_downloads_all_time 173
  • author_summary 1 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M1
Primary method

Direct removal

No other method signals detected in this model.
Confidence
MEDIUM
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=0 (base model)
  • no specific method indicators - defaulting to M1 (most common)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
MEDIUM
Why we say so
primary_method=M1; difference-of-means is the reference extraction for M1/M3 (Arditi 2024)
Downloads · lifetime
173
36 last 30d - stable
Likes
0
Model age
2mo ago
created 2026-07-21
Downloads over time
Now187→from37↑405%
308714520237 on Jul 22187 on Oct 11JulAugSepOct
Jul 22 → Oct 11 · 52 snapshots · spans 81 days

Benchmarks

Portrait before abliteration
Benchmarks of the base model as it stood before the refusal-removal operation. Compare with the numbers above to see what the operation cost.
Benchmark Score Source
Entertainment 5 UGI
Hazardous 1.8 UGI
Natural Intelligence 54.88 UGI
Political lean -12.7% UGI
Sensitive-Info 41.67 UGI
SocPol 5.1 UGI
UGI 29.44 UGI
Willingness (10) 0.5 UGI
W10-Adherence 0 UGI
W10-Direct 1 UGI
Writing 61.21 UGI

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
mit
Tags
transformers safetensors glm_moe_dsa text-generation glm moe quantization nvfp4 nf3 mxfp8 hybrid-quant abliterated

Related

Total size
341 GB
Files
101
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-07-21 23:10

Files by quantization

Auxiliary files 101 files 341 GB
model-00061-of-00092.safetensors 3.73 GB ee8bbe50 download
model-00065-of-00092.safetensors 3.73 GB 1500ace5 download
model-00041-of-00092.safetensors 3.73 GB a84eb089 download
model-00015-of-00092.safetensors 3.73 GB 582c3493 download
model-00077-of-00092.safetensors 3.73 GB f17faf84 download
model-00071-of-00092.safetensors 3.73 GB 4627f289 download
model-00032-of-00092.safetensors 3.73 GB 8727c4d3 download
model-00056-of-00092.safetensors 3.73 GB c2d65466 download
model-00044-of-00092.safetensors 3.73 GB ae51cd1d download
model-00012-of-00092.safetensors 3.73 GB 33d85eee download
model-00045-of-00092.safetensors 3.73 GB 507bc227 download
model-00002-of-00092.safetensors 3.73 GB 6901b9be download
model-00019-of-00092.safetensors 3.73 GB a5606376 download
model-00078-of-00092.safetensors 3.72 GB 056c86a9 download
model-00048-of-00092.safetensors 3.72 GB c550007a download
model-00054-of-00092.safetensors 3.72 GB 12a512ad download
model-00066-of-00092.safetensors 3.72 GB 9af010d6 download
model-00024-of-00092.safetensors 3.72 GB 17ce9736 download
model-00089-of-00092.safetensors 3.72 GB 58c7de6f download
model-00039-of-00092.safetensors 3.72 GB 123f4ece download
model-00017-of-00092.safetensors 3.72 GB 8eb5c5ed download
model-00005-of-00092.safetensors 3.72 GB 11bb4816 download
model-00055-of-00092.safetensors 3.72 GB dc59aaad download
model-00023-of-00092.safetensors 3.72 GB cf7b52c2 download
model-00060-of-00092.safetensors 3.72 GB dd197538 download
model-00003-of-00092.safetensors 3.72 GB 896d9e88 download
model-00047-of-00092.safetensors 3.72 GB 78611184 download
model-00069-of-00092.safetensors 3.72 GB cac79e73 download
model-00086-of-00092.safetensors 3.72 GB 12a1a14f download
model-00068-of-00092.safetensors 3.72 GB 6eb21e95 download
model-00028-of-00092.safetensors 3.72 GB 4947ab57 download
model-00072-of-00092.safetensors 3.72 GB 5b319f6a download
model-00020-of-00092.safetensors 3.72 GB 6063f51f download
model-00004-of-00092.safetensors 3.72 GB 62b3f62f download
model-00082-of-00092.safetensors 3.72 GB 3784ff5e download
model-00038-of-00092.safetensors 3.72 GB 668863b4 download
model-00067-of-00092.safetensors 3.72 GB 07f4d9b4 download
model-00033-of-00092.safetensors 3.72 GB 74e8fa4f download
model-00043-of-00092.safetensors 3.72 GB 0ea3c76e download
model-00070-of-00092.safetensors 3.72 GB b39ce61a download
model-00050-of-00092.safetensors 3.72 GB 0e9afec7 download
model-00053-of-00092.safetensors 3.72 GB c9d9fe7e download
model-00007-of-00092.safetensors 3.72 GB 9ac38467 download
model-00074-of-00092.safetensors 3.72 GB ae9fc892 download
model-00076-of-00092.safetensors 3.72 GB 456f0e18 download
model-00051-of-00092.safetensors 3.72 GB 9c0a38fa download
model-00029-of-00092.safetensors 3.72 GB 73b84dcc download
model-00091-of-00092.safetensors 3.72 GB 135b4b6d download
model-00030-of-00092.safetensors 3.72 GB bedb9b93 download
model-00026-of-00092.safetensors 3.72 GB bb015050 download
model-00018-of-00092.safetensors 3.72 GB 144f36a1 download
model-00031-of-00092.safetensors 3.72 GB 754f64f4 download
model-00021-of-00092.safetensors 3.72 GB 8dcc50a1 download
model-00080-of-00092.safetensors 3.72 GB b16c8587 download
model-00016-of-00092.safetensors 3.72 GB 4df1c541 download
model-00057-of-00092.safetensors 3.72 GB 825916af download
model-00009-of-00092.safetensors 3.72 GB c13dee2c download
model-00062-of-00092.safetensors 3.72 GB f530724b download
model-00036-of-00092.safetensors 3.72 GB c055e34a download
model-00014-of-00092.safetensors 3.72 GB e8e9d644 download
model-00052-of-00092.safetensors 3.72 GB caeeb2c2 download
model-00063-of-00092.safetensors 3.72 GB 01caf772 download
model-00034-of-00092.safetensors 3.72 GB d08b5f2d download
model-00046-of-00092.safetensors 3.72 GB 06941cf9 download
model-00010-of-00092.safetensors 3.72 GB 304be3b6 download
model-00090-of-00092.safetensors 3.72 GB e172ff0e download
model-00011-of-00092.safetensors 3.72 GB efcca825 download
model-00088-of-00092.safetensors 3.72 GB 3a108be8 download
model-00049-of-00092.safetensors 3.72 GB f86eea92 download
model-00064-of-00092.safetensors 3.72 GB 8a959e34 download
model-00022-of-00092.safetensors 3.72 GB e381bcf4 download
model-00037-of-00092.safetensors 3.72 GB 247271d1 download
model-00025-of-00092.safetensors 3.72 GB 82b8648c download
model-00084-of-00092.safetensors 3.72 GB 5d90d506 download
model-00013-of-00092.safetensors 3.72 GB 0f27a3fc download
model-00027-of-00092.safetensors 3.72 GB 27885ce8 download
model-00083-of-00092.safetensors 3.72 GB 6d4c19dc download
model-00042-of-00092.safetensors 3.72 GB 1b4fdd3f download
model-00040-of-00092.safetensors 3.72 GB dc03eded download
model-00059-of-00092.safetensors 3.72 GB 0f5cc9f8 download
model-00058-of-00092.safetensors 3.72 GB c19ea0d3 download
model-00006-of-00092.safetensors 3.72 GB 78744ed0 download
model-00008-of-00092.safetensors 3.72 GB b473eeb8 download
model-00035-of-00092.safetensors 3.72 GB e9002770 download
model-00087-of-00092.safetensors 3.72 GB e39d0f10 download
model-00001-of-00092.safetensors 3.71 GB 50eb9353 download
model-00085-of-00092.safetensors 3.70 GB dc18bea7 download
model-00079-of-00092.safetensors 3.68 GB cbf8e480 download
model-00073-of-00092.safetensors 3.66 GB 6a6aa723 download
model-00075-of-00092.safetensors 3.63 GB b47169f4 download
model-00081-of-00092.safetensors 3.59 GB 74a0faad download
model-00092-of-00092.safetensors 2.40 GB f19e6cba download
tokenizer.json 19.3 MB 19e77364 download
model.safetensors.index.json 13.4 MB ff524549 download
config.json 142 KB 72d43218 download
mxfp8_tier_nokvb.json 10.5 KB 7e74d16f download
README.md 9.51 KB eb89a932 download
chat_template.jinja 6.01 KB 69fb6866 download
.gitattributes 1.60 KB a09db2ea download
tokenizer_config.json 790 B 0891a18e download
generation_config.json 215 B 4780e3ab download

README current version from Hugging Face


license: mit
library_name: transformers
pipeline_tag: text-generation
tags:

  • glm
  • moe
  • quantization
  • nvfp4
  • nf3
  • mxfp8
  • hybrid-quant
  • abliterated
  • uncensored
    base_model:
  • zai-org/GLM-5.2

SuperGLM-5.2-abliterated — MXFP8 / NVFP4 / NF3 Hybrid

A single-node hybrid-quantized build of the abliterated GLM-5.2 (753B, MoE) that fits on
4×96 GB GPUs at tensor-parallel 4, serving the full 256-expert model at 262K context —
weights preserved, refusals removed, reasoning intact.

  • ~341 GB on disk, 92 safetensors shards. Runs where the BF16/NVFP4-only builds need 8 cards.
  • Uncensored: the abliteration lives in the BF16 non-expert tier and survives quantization
    (measured 0/12 refusals post-build and post-healing).
  • Capability retained: GSM8K 94.2% (flexible) / 93.9% (strict) after the healing pass.

Research artifact. It inherits the behavior and the license of its upstreams — read
Provenance and Intended use below before deploying.

What this is

GLM-5.2 is a Mixture-of-Experts model (78 layers, 256 experts, first_k_dense=3,
glm_moe_dsa). This build applies a per-expert mixed-precision scheme so the whole model
fits four cards instead of eight, without dropping any experts (no structural pruning):

Tier Precision What it covers
Top-damage experts NVFP4 (4-bit, per-16 scale + global) the 64 highest-damage experts per MoE layer
Remaining experts NF3 (3-bit, group-32 e4m3 scale) the other 192 experts per layer
Attention KV-B / shared BF16 kept full precision
All non-expert weights BF16 attention, norms, embeddings, lm_head — where the abliteration lives

A bit_map (75 MoE layers × 256 experts → 4,800 NVFP4 / 14,400 NF3) drives the per-expert
allocation. This reproduces the hybrid scheme published by madeby561 for the stock GLM-5.2
checkpoint, applied here to the abliterated weights.

Healing pass

After quantization, expert weights are error-corrected with GPTQ-style error feedback
(healed against the source BF16, gate/up sharing a Hessian, down-proj reconstructed from the
SwiGLU intermediate). 36,591 matrices (~85%; low-traffic experts left naive) were healed,
cutting per-expert output error from 14.2% → 1.1%. The healing only touches experts, so the
abliteration (non-expert tier) is untouched.

Provenance & credits

This is a derivative work. Full credit to the upstreams:

  • zai-org/GLM-5.2 — the original model and architecture.
  • Jiunsong/SuperGLM-5.2-abliterated-NVFP4 — the abliterated (refusal-removed) NVFP4 source
    weights this build quantizes from.
  • madeby561 — author of the NVFP4 + NF3 + BF16 hybrid quantization scheme and the vLLM
    serving image; the NF3/NVFP4 quantizers and bit_map here reproduce that method.

The build reproduces the two quantizers (NF3 and NVFP4) and reuses the upstream hybrid bit_map;
the NF3 output format was validated byte-exact against the reference implementation before the
full run.

Requirements

  • 4× NVIDIA Blackwell GPUs (SM 120 / RTX PRO 6000-class), ~96 GB each (~384 GB total VRAM).
    The NVFP4 experts and the B12X MoE / sparse-MLA kernels are Blackwell-native — this build
    will not run unchanged on Ampere/Hopper.
  • Docker + the NVIDIA container runtime.
  • The hybrid-aware vLLM image madeby561/vllm-glm52-nvfp4-nf3-hybrid:v3. Its loader reads the
    bit_map and reconstructs each precision tier at load — a stock transformers/vllm load will
    not reconstruct the mixed-precision experts.

Download

HF_HUB_ENABLE_HF_TRANSFER=1 \
hf download kachowtowmater/SuperGLM-5.2-abliterated-MXFP8-NVFP4-NF3-Hybrid \
  --local-dir ./SuperGLM-5.2-abliterated-hybrid

Serving (vLLM, tensor-parallel 4)

Point MODEL_DIR at the downloaded checkpoint. docker-compose.yml:

services:
  glm52:
    image: madeby561/vllm-glm52-nvfp4-nf3-hybrid:v3
    network_mode: host
    ipc: host
    shm_size: 32gb
    init: true
    ulimits:
      memlock: -1
      stack: 67108864
    deploy:
      resources:
        reservations:
          devices:
            - driver: nvidia
              count: all
              capabilities: [gpu]
    volumes:
      - ${MODEL_DIR:?set MODEL_DIR to the downloaded checkpoint dir}:/model:ro
      - vllm-cache:/cache
    healthcheck:
      test: ["CMD-SHELL", "curl -sf http://localhost:8000/health"]
      start_period: 600s
      interval: 15s
    environment:
      CUDA_VISIBLE_DEVICES: "0,1,2,3"
      CUDA_DEVICE_ORDER: PCI_BUS_ID
      CUTE_DSL_ARCH: sm_120a            # Blackwell
      PYTORCH_CUDA_ALLOC_CONF: expandable_segments:True
      NCCL_IB_DISABLE: "1"
      NCCL_P2P_LEVEL: SYS
      GLOO_SOCKET_IFNAME: lo
      TP_SOCKET_IFNAME: lo
      VLLM_USE_B12X_FP8_GEMM: "1"
      VLLM_USE_B12X_MOE: "1"
      VLLM_USE_B12X_SPARSE_INDEXER: "1"
      VLLM_USE_V2_MODEL_RUNNER: "1"
      B12X_W4A16_TC_DECODE: "1"
      B12X_MOE_FORCE_A16: "1"
      VLLM_DCP_GLOBAL_TOPK: "1"
      VLLM_DCP_SHARD_DRAFT: "1"
      VLLM_ENABLE_PCIE_ALLREDUCE: "1"
      VLLM_PCIE_ALLREDUCE_BACKEND: b12x
      VLLM_PCIE_ONESHOT_MAX_BYTES: "65536"
      B12X_DENSE_SPLITK_TURBO: "1"
      XDG_CACHE_HOME: /cache/jit
      CUDA_CACHE_PATH: /cache/jit
      # hybrid loader
      HYBRID_TIER: both
      HYBRID_KEPT: b12x_nf3
      HYBRID_NF3: b12x_nf3
      HYBRID_B12X_MAX_TOKENS: "2048"
      HYBRID_MXFP8_NATIVE: "1"
    command: >
      vllm serve /model
      --served-model-name GLM-5.2 --host 0.0.0.0 --port 8000
      --trust-remote-code --tensor-parallel-size 4
      --decode-context-parallel-size 4 --dcp-comm-backend ag_rs --dcp-kv-cache-interleave-size 1
      --kv-cache-dtype fp8
      --attention-backend B12X_MLA_SPARSE
      --moe-backend b12x
      --load-format safetensors
      -cc.pass_config.fuse_allreduce_rms=True
      --gpu-memory-utilization 0.968
      --max-model-len 262144 --max-num-seqs 8 --max-num-batched-tokens 2048
      --max-cudagraph-capture-size 64
      --async-scheduling --enable-chunked-prefill --enable-prefix-caching
      --enable-auto-tool-choice --tool-call-parser glm47 --reasoning-parser glm45
      --default-chat-template-kwargs '{"reasoning_effort":"high"}'
      --hf-overrides '{"use_index_cache":true,"index_topk_pattern":"FFFSSSFSSSFSSSFSSSFSSSFSSSFSSSFSSSFSSSFSSSFSSSFSSSFSSSFSSSFSSSFSSSFSSSFSSSFSSS"}'
      --speculative-config '{"method":"mtp","num_speculative_tokens":5,"moe_backend":"b12x","draft_sample_method":"probabilistic"}'
volumes:
  vllm-cache:
MODEL_DIR=./SuperGLM-5.2-abliterated-hybrid docker compose up -d

First boot takes several minutes (engine init ~1–2 min + JIT compile; the healthcheck has a 600 s grace).

Verify the boot — two checks that actually matter

Sparse-MLA indexing and the KV dtype are load-bearing for long-context correctness. Check the container log:

  • grep -c "skip sparse MLA indexer" must be 57 (one per sparse layer). A different count means the
    index_topk_pattern hf-override didn't apply → silent long-context corruption.
  • grep "fp8_ds_mla KV cache" must be present. Missing it (i.e. BF16 KV) → garbage output on the
    B12X_MLA_SPARSE backend.

Inference

OpenAI-compatible /v1. GLM pins temperature 1.0 — do not lower it (its default inverts the usual advice):

curl http://localhost:8000/v1/chat/completions \
  -H 'Content-Type: application/json' \
  -d '{"model":"GLM-5.2","messages":[{"role":"user","content":"..."}],"temperature":1.0}'

Served id is GLM-5.2; 262K context; MTP speculative decoding (5 draft tokens), chunked prefill and
prefix caching on; tool-calling (glm47 parser) and reasoning (glm45 parser) enabled.

Evaluation

Metric Result Notes
GSM8K (flexible) 94.24% lm-evaluation-harness, served endpoint
GSM8K (strict) 93.86% "
Refusals 0 / 12 abliteration preserved through quant + healing
Coherence / reasoning intact arithmetic, ordering, multi-step held
Throughput ~46–54 tok/s single-stream, TP4

Intended use & limitations

  • Research use. This model has had safety refusals removed (abliterated). It will attempt
    most requests. You are responsible for how you deploy and prompt it; apply your own safety
    layer for any user-facing use.
  • Quantization is lossy. NF3 3-bit experts trade some fidelity for the 4-card fit; the
    healing pass recovers most, not all, of that gap.
  • License: inherits MIT from GLM-5.2. Attribution to the upstreams above is required.

Format

Standard safetensors (92 shards) + config.json carrying the hybrid metadata (hybrid_bit_map,
hybrid_scheme, per-tier layouts) + mxfp8_tier_nokvb.json. Load with the hybrid-aware vLLM image
above; a stock transformers load will not reconstruct the mixed-precision tiers.

Troubleshooting

Symptom Cause / fix
OOM at load Needs 4×~96 GB. Lower --gpu-memory-utilization, or you don't have the VRAM for TP4.
Garbage / repetition at long context The index_topk_pattern override didn't apply — confirm the 57 sparse-indexer skips in the boot log.
Garbage output generally KV must be fp8_ds_mla — confirm that line in the boot log; BF16 KV breaks B12X_MLA_SPARSE.
unknown kernel / illegal instruction at load Not on Blackwell (SM 120). The B12X kernels are Blackwell-native; Ampere/Hopper need a different build.
Stock vllm serve fails to load the experts Use the madeby561/vllm-glm52-nvfp4-nf3-hybrid:v3 image — the hybrid loader is required.

README history 2 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-07-21docs: full run recipe (reqs, compose, boot checks, troubleshooting)dcf768e9.5 KB
    Loading...
  2. 2026-07-21Add files using upload-large-folder toolfbcbec84.7 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration