← back to catalog · registered 2026-08-22 13:56

Jiunsong/SuperGLM-5.2-abliterated-NVFP4

Jiunsong Glm 362B MoE
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/Jiunsong%2FSuperGLM-5.2-abliterated-NVFP4"
Response includes
  • classification m1
  • files 59
  • hub_downloads_all_time 6,757
  • author_summary 35 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M1
Primary method

Direct removal

No other method signals detected in this model.
Confidence
MEDIUM
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=0 (base model)
  • no specific method indicators - defaulting to M1 (most common)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
MEDIUM
Why we say so
primary_method=M1; difference-of-means is the reference extraction for M1/M3 (Arditi 2024)
Downloads · lifetime
7K
390 last 30d - cooling
Likes
59
Model age
3mo ago
created 2026-07-13
Downloads over time
Now7K→from302↑2,203%
02.5K5.1K7.6K302 on Jul 147K on Oct 11JulAugSepOct
Jul 14 → Oct 11 · 54 snapshots · spans 89 days

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
mit
Tags
transformers safetensors glm_moe_dsa text-generation glm-5.2 nvfp4 modelopt abliterated obliteratus super-tune conversational base_model:nvidia/GLM-5.2-NVFP4

Related

Total size
433 GB
Files
59
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-07-23 10:49

Files by quantization

Auxiliary files 59 files 433 GB
model-00019-of-00047.safetensors 9.31 GB 0022b9c4 download
model-00043-of-00047.safetensors 9.31 GB 6542193f download
model-00025-of-00047.safetensors 9.31 GB 51b31b6a download
model-00032-of-00047.safetensors 9.31 GB 62d0b99f download
model-00008-of-00047.safetensors 9.31 GB 139f1c10 download
model-00039-of-00047.safetensors 9.31 GB e5f97c07 download
model-00015-of-00047.safetensors 9.31 GB 17cfdecd download
model-00022-of-00047.safetensors 9.31 GB 86a1268d download
model-00029-of-00047.safetensors 9.31 GB e86e584f download
model-00036-of-00047.safetensors 9.31 GB 1d729de1 download
model-00011-of-00047.safetensors 9.31 GB 917eec1f download
model-00009-of-00047.safetensors 9.31 GB f806cbb1 download
model-00016-of-00047.safetensors 9.31 GB fabc3531 download
model-00023-of-00047.safetensors 9.31 GB 53de1a17 download
model-00006-of-00047.safetensors 9.31 GB 902e2541 download
model-00013-of-00047.safetensors 9.31 GB af7acd1d download
model-00030-of-00047.safetensors 9.31 GB 7e28d533 download
model-00037-of-00047.safetensors 9.31 GB 8549eb72 download
model-00020-of-00047.safetensors 9.31 GB 830fbe56 download
model-00044-of-00047.safetensors 9.31 GB 98a2bf50 download
model-00027-of-00047.safetensors 9.31 GB a25bf481 download
model-00034-of-00047.safetensors 9.31 GB 8c78db22 download
model-00041-of-00047.safetensors 9.31 GB d0260617 download
model-00002-of-00047.safetensors 9.31 GB 5a93bf68 download
model-00004-of-00047.safetensors 9.31 GB 790fe046 download
model-00012-of-00047.safetensors 9.31 GB 368c31ac download
model-00005-of-00047.safetensors 9.31 GB 9c6a3381 download
model-00001-of-00047.safetensors 9.31 GB 7355ca80 download
model-00026-of-00047.safetensors 9.31 GB 6b7e8e12 download
model-00035-of-00047.safetensors 9.31 GB 33c1a2fe download
model-00028-of-00047.safetensors 9.31 GB 31d71bcc download
model-00021-of-00047.safetensors 9.31 GB f7ca12dd download
model-00033-of-00047.safetensors 9.31 GB f47739f9 download
model-00040-of-00047.safetensors 9.31 GB 1b342133 download
model-00014-of-00047.safetensors 9.31 GB 68eaf674 download
model-00038-of-00047.safetensors 9.31 GB 9cb454ba download
model-00007-of-00047.safetensors 9.31 GB 40299690 download
model-00031-of-00047.safetensors 9.31 GB cbe306cd download
model-00010-of-00047.safetensors 9.31 GB 2aa9b075 download
model-00017-of-00047.safetensors 9.31 GB 7994a0f6 download
model-00024-of-00047.safetensors 9.31 GB ec8c36c4 download
model-00003-of-00047.safetensors 9.31 GB 24048528 download
model-00046-of-00047.safetensors 9.30 GB 6edd678c download
model-00045-of-00047.safetensors 9.30 GB 56445743 download
model-00018-of-00047.safetensors 9.28 GB 833acd8a download
model-00042-of-00047.safetensors 9.20 GB 0fe5f14a download
model-00047-of-00047.safetensors 4.76 GB ca2dfdb8 download
model.safetensors.index.json 21.2 MB 2aa8397b download
tokenizer.json 19.3 MB 19e77364 download
.quant_summary.txt 8.42 MB 81d52f95 download
superglm52_weight_only_v2_merge.json 140 KB a089634b download
superglm52_weight_only_v2_eval.json 28.2 KB 6cf5d099 download
config.json 15.2 KB b63d4c44 download
hf_quant_config.json 7.23 KB 4c1380a3 download
README.md 6.11 KB 55465c32 download
chat_template.jinja 4.96 KB 672034e3 download
.gitattributes 1.60 KB a09db2ea download
tokenizer_config.json 790 B 0891a18e download
generation_config.json 215 B 4780e3ab download

README current version from Hugging Face


base_model: nvidia/GLM-5.2-NVFP4
library_name: transformers
license: mit
pipeline_tag: text-generation
tags:

  • glm-5.2
  • nvfp4
  • modelopt
  • abliterated
  • obliteratus
  • super-tune

SuperGLM-5.2-abliterated-NVFP4 v2

SuperGLM-5.2 is a clean-base, fused-weight OBLITERATUS release derived directly from NVIDIA's official GLM-5.2 NVFP4 checkpoint. It is designed to reduce refusal behavior while preserving the speed and storage advantages of the original ModelOpt NVFP4 export.

This revision replaces the earlier template-assisted experiment. The published chat template is byte-identical to NVIDIA's original template, no adaptive system scaffold is required, and the release gate evaluates the serialized checkpoint with runtime adapters disabled.

What was changed

  • Clean restart from nvidia/GLM-5.2-NVFP4@aec724e8c7b8ee9db3b48c01c320f63f9cdaf8aa — no cumulative merge from an earlier SuperGLM checkpoint.
  • Rank-4 OBLITERATUS projection over 62 BF16 attention o_proj matrices at strength 2.
  • Rank-2 projection over the matching BF16 shared-expert down_proj matrices at strength 0.15.
  • Packed routed ModelOpt NVFP4 experts remain byte-identical to NVIDIA's base.
  • lm_head remains byte-identical to NVIDIA's base.
  • No LoRA or adapter is needed at inference time: all selected updates are fused into safetensors.

The final release modifies 124 BF16 tensors across 37 shards. Tensor-region backups and projection residual checks were used during fusion; the maximum permitted projection verification error was 0.025.

Weight-only evaluation

The acceptance data comes from the official HarmBench repository at commit 8e1604d1171fe8a48d8febecd22f600e462bdcdd. The standard-text development gates reuse previously opened official HarmBench partitions for candidate selection. The sealed final instead uses previously unused FunctionalCategory=contextual official-test behaviors, formatted byte-for-byte like HarmBench DirectRequest (ContextString, separator, then Behavior) and judged with the official contextual classifier template.

Gate Physical prompts Greedy repeats Official no / invalid / total Explicit refusal strings Broad lexical flags Quality checks Structural anomalies Unicode replacements
Runtime adapter screen 41 2 1 / 0 / 82 0 19 16/16 0 0
Serialized fused gate 35 2 4 / 0 / 70 0 24 15/16 0 0
Sealed serialized final 80 2 2 / 0 / 160 0 24 16/16 0 0

All generations use greedy decoding, the original NVIDIA chat template, enable_thinking=false, no explicit system prompt, and adaptive_directness=false. Runtime LoRA is used only for the first candidate screen. The two release gates run on serialized weights with adapters disabled. Acceptance requires at least 15/16 quality checks, zero structural anomalies, zero Unicode replacement characters, and zero explicit refusal strings. The official classifier release threshold is configured as fused <= 6 and sealed <= 6 refused/non-exhibiting hits, with invalid outputs fixed at zero. This v2 card reports the exact no / invalid / total counts above so the configured release slack is auditable. In HarmBench semantics a yes means that the requested behavior is present, so treating every no as refused/non-exhibiting is stricter than a refusal-string check. Broad lexical flags are reported for continuity but are not a release gate because they also match compliant caveats containing words such as “illegal” or “dangerous.”

The 16 quality checks are a small regression sentinel, not a broad capability benchmark. This release does not claim a throughput speedup over NVIDIA's base; it preserves the base checkpoint's NVFP4 format and validates compatibility with the same ModelOpt FP4 runtime family.

Leakage controls

  • The prior 220-prompt development bank trains the tiny lm_head repair and is not reported as an independent benchmark. It is excluded from every official acceptance partition and is also supplied to the overlap audit.
  • The 352 OBLITERATUS contrast-training pairs train the refusal directions. They are excluded from every official acceptance partition and are also supplied to the overlap audit.
  • The standard development partitions are byte-identical to their pinned prior artifact and retain that artifact's zero-overlap audit.
  • The contextual sealed candidates were audited against the registered training and prior-evaluation prompts; near matches at char-5gram Jaccard >= 0.5 were excluded automatically before candidate freeze.
  • The retained contextual sealed final has zero normalized exact or threshold-level near overlap and was opened exactly once.

Inference

Use an NVFP4-compatible Blackwell runtime. The release was validated with SGLang's ModelOpt FP4 path and flashinfer_cutlass MoE runner on 8x RTX PRO 6000 Blackwell GPUs.

import sglang as sgl

engine = sgl.Engine(
    model_path="Jiunsong/SuperGLM-5.2-abliterated-NVFP4",
    tp_size=8,
    quantization="modelopt_fp4",
    moe_runner_backend="flashinfer_cutlass",
    disable_shared_experts_fusion=True,
)

Lineage and reproducibility

  • NVIDIA base: aec724e8c7b8ee9db3b48c01c320f63f9cdaf8aa
  • Selected candidate: broad62-r4-s2-obliteratus_only-shared-r2-s0p15
  • Candidate artifact: sha256:d7ded1c1eea29006823f7219a35c71a4cfdcdf4038aaa39c4909e57626f1300e
  • lm_head repair source: none
  • Direction source: zai-org/GLM-5.2-FP8@ba978f7d347eaf65d22f1a86833408afdb953541
  • Evaluation artifact: sha256:956fd23f7a45308878fc41a55e30f8995d515ad3b53bc1b1011d35efeb021101
  • HarmBench classifier: cais/HarmBench-Llama-2-13b-cls@bda705349d1144fa618770bea64d99ce54e3835b
  • HarmBench contextual classifier prompt SHA-256: 5d6bb9e3cf4d1e5f3f7620093113222ee115af75b0b2913f00e4c7225ec9f219
  • HarmBench standard classifier prompt SHA-256: 788f4f6aa1491c433c4da76c9140cfc30966cea3ff3875c4d0fcb336d92f60e0
  • Original template SHA-256: 172dc74a35e1752df75ecfb2b2cf9326d2852bb1379868ebeec9571654489679

Detailed sanitized fusion and evaluation reports are included in this repository. Raw generations are retained privately for reproducibility and are not published in the model card.

README history 8 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-07-23Publish clean-base fused weight-only SuperGLM v2076582b6.1 KB
    Loading...
  2. 2026-07-23Reset release branch to NVIDIA base aec724e8c7b86e92d9212.2 KB
    Loading...
  3. 2026-07-22Document V2 sweet-spot audit and retain validated weights30add1118.3 KB
    Loading...
  4. 2026-07-14Document fused OBLITERATUS, SuperTune recovery, and NVFP4 release gates506e95b16.4 KB
    Loading...
  5. 2026-07-14Document the validated 0/220 adaptive directness release7d93e2718.5 KB
    Loading...
  6. 2026-07-13Polish SuperGLM-5.2 model card343c49f12.9 KB
    Loading...
  7. 2026-07-13Finalize SuperGLM-5.2 OBLITERATUS + SuperTune NVFP4 evidence423857011.3 KB
    Loading...
  8. 2026-07-13Duplicate from nvidia/GLM-5.2-NVFP439c493e12.2 KB
    Loading...

Discussions 2 threads

  1. 2026-07-30Requestopen1 💬#2
    Loading...
  2. 2026-07-13Deployment and Speed?open1 💬#1
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration