← back to catalog · registered 2026-09-13 16:56

cbert33/GLM-5.3-Flash-Uncensored-EXL3-DGX-Sliced

cbert33 Glm MoE multimodal second-order
Abliteration classifier · v1.0.0
M1
Primary method

Direct removal

No other method signals detected in this model.
Confidence
MEDIUM
Why this label 3 signals
Method inferred from partial signals — repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=0 (base model)
  • no specific method indicators - defaulting to M1 (most common)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
MEDIUM
Why we say so
primary_method=M1; difference-of-means is the reference extraction for M1/M3 (Arditi 2024)
Downloads · 30-day
0
Likes
1
Model age
1d ago
created 2026-09-12

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
other
Languages
en
Tags
transformers safetensors glm5_next image-text-to-text glm exl3 tr3 rank-sliced tensor-parallel vllm dgx-spark sm120

Related

Total size
164 GB
Files
107
Quantizations
1
Registered
2026-09-13 16:56
Last updated on HF
2026-09-13 15:59

Files by quantization

Auxiliary files 107 files 164 GB
model-00083.safetensors 2.71 GB 02aac042 download
model-00089.safetensors 1.86 GB 87f24a59 download
model-00085.safetensors 1.82 GB 785fdf98 download
model-00090.safetensors 1.81 GB 5659a236 download
model-00086.safetensors 1.81 GB dc6e7c27 download
model-00087.safetensors 1.79 GB 8ecf0925 download
model-00088.safetensors 1.79 GB 99b6472a download
model-00091.safetensors 1.79 GB 0568d9ce download
model-00084.safetensors 1.79 GB e9832fed download
model-00028.safetensors 1.77 GB 95ac24bb download
model-00070.safetensors 1.77 GB 89d37efb download
model-00030.safetensors 1.77 GB c078afa9 download
model-00045.safetensors 1.77 GB 0370c429 download
model-00053.safetensors 1.77 GB d6b8ef43 download
model-00068.safetensors 1.77 GB ad2f32ec download
model-00049.safetensors 1.77 GB 05a3f690 download
model-00024.safetensors 1.77 GB 6c28f0be download
model-00026.safetensors 1.77 GB b7254516 download
model-00047.safetensors 1.77 GB 58cd57de download
model-00051.safetensors 1.77 GB 04c9d2ef download
model-00072.safetensors 1.77 GB 3dc32efa download
model-00074.safetensors 1.77 GB eddc8d3c download
model-00066.safetensors 1.77 GB 443f8f88 download
model-00032.safetensors 1.77 GB 4a7056cf download
model-00022.safetensors 1.77 GB f8ca2b7e download
model-00076.safetensors 1.77 GB ab8b5f00 download
model-00043.safetensors 1.77 GB 6e7498c2 download
model-00055.safetensors 1.77 GB ad0ffcfb download
model-00064.safetensors 1.77 GB 8fc7d738 download
model-00034.safetensors 1.77 GB 6d2448bc download
model-00020.safetensors 1.77 GB a226ab30 download
model-00078.safetensors 1.77 GB 6ad8fccd download
model-00041.safetensors 1.77 GB fdab403d download
model-00057.safetensors 1.77 GB 8e14a5a5 download
model-00062.safetensors 1.77 GB 829418cb download
model-00036.safetensors 1.77 GB 6254dc06 download
model-00015.safetensors 1.77 GB 640869d9 download
model-00018.safetensors 1.77 GB 3ff063f3 download
model-00080.safetensors 1.77 GB a21f89b7 download
model-00039.safetensors 1.77 GB 91315e6a download
model-00059.safetensors 1.77 GB 2744a230 download
model-00060.safetensors 1.77 GB 37c5f2ac download
model-00038.safetensors 1.77 GB a0413bd4 download
model-00081.safetensors 1.77 GB f6873c07 download
model-00016.safetensors 1.77 GB c2fa291b download
model-00017.safetensors 1.77 GB 4055610b download
model-00082.safetensors 1.77 GB 00bf5700 download
model-00037.safetensors 1.77 GB 5dc0453b download
model-00061.safetensors 1.77 GB e9d59622 download
model-00058.safetensors 1.77 GB e7556b4d download
model-00040.safetensors 1.77 GB 2643db45 download
model-00079.safetensors 1.77 GB 74df156e download
model-00019.safetensors 1.77 GB 582ff2ba download
model-00035.safetensors 1.77 GB ee937556 download
model-00063.safetensors 1.77 GB fbb7248a download
model-00056.safetensors 1.77 GB 4f28d80c download
model-00042.safetensors 1.77 GB 6aa446f8 download
model-00077.safetensors 1.77 GB b6389071 download
model-00021.safetensors 1.77 GB 7bad5157 download
model-00033.safetensors 1.77 GB 0e5422e7 download
model-00065.safetensors 1.77 GB f6082f2c download
model-00025.safetensors 1.77 GB 87804099 download
model-00073.safetensors 1.77 GB 01bee98b download
model-00048.safetensors 1.77 GB 748e5764 download
model-00050.safetensors 1.77 GB abf35638 download
model-00075.safetensors 1.77 GB 4bfb0061 download
model-00023.safetensors 1.77 GB 38cb3a69 download
model-00031.safetensors 1.77 GB 4fc93ae5 download
model-00046.safetensors 1.77 GB 73cb3371 download
model-00052.safetensors 1.77 GB cdbdfbbc download
model-00067.safetensors 1.77 GB 146bf418 download
model-00027.safetensors 1.77 GB 6612128b download
model-00029.safetensors 1.77 GB 14b81639 download
model-00044.safetensors 1.77 GB 17a29d91 download
model-00054.safetensors 1.77 GB faff5484 download
model-00069.safetensors 1.77 GB 6cd84503 download
model-00071.safetensors 1.77 GB 2d967194 download
model-00014.safetensors 1.77 GB 8cb51a91 download
model-00005.safetensors 1.77 GB 5bf7dee7 download
model-00007.safetensors 1.77 GB a8d45a63 download
model-00003.safetensors 1.77 GB 3796bdd3 download
model-00009.safetensors 1.77 GB e6ce6f1f download
model-00001.safetensors 1.77 GB a600a6d6 download
model-00011.safetensors 1.77 GB b14dcc73 download
model-00013.safetensors 1.77 GB 650a3be7 download
model-00012.safetensors 1.77 GB 4d4fd705 download
model-00010.safetensors 1.77 GB 4a7c46c8 download
model-00002.safetensors 1.77 GB e26bbb5e download
model-00004.safetensors 1.77 GB 2ff9dec4 download
model-00008.safetensors 1.77 GB 907a20f5 download
model-00006.safetensors 1.77 GB 52ec4fce download
model-00092.safetensors 1.18 GB 04d44a5b download
model.safetensors.index.json 28.7 MB 854966b2 download
tokenizer.json 19.3 MB 19e77364 download
ShapleyMcg-LICENSE 29.2 KB 03449050 download
SHA256SUMS 9.27 KB 1e1b65b9 download
chat_template.jinja 8.42 KB 15bf200e download
config.json 6.38 KB 9236288d download
README.md 5.42 KB 9ed613fa download
.gitattributes 1.72 KB 18a2c6ed download
RELEASE_MANIFEST.json 1.13 KB 865d8e14 download
exl3-mcg-storage-abi.json 943 B 40e1838a download
processor_config.json 909 B 3ec2a058 download
tokenizer_config.json 761 B e375fa0a download
quantization_config.json 544 B 443f2df6 download
rank-slice-manifest.json 502 B 037b9b70 download
generation_config.json 194 B 637ee6af download

README current version from Hugging Face


base_model: neko-legends/GLM-5.3-Flash-Uncensored-EXL3
base_model_relation: quantized
library_name: transformers
pipeline_tag: image-text-to-text
license: other
license_name: shapleymcg-license-1.0
license_link: https://github.com/brandonmmusic-max/shapleymcg/blob/main/LICENSE
tags:

  • glm
  • glm5_next
  • exl3
  • tr3
  • rank-sliced
  • tensor-parallel
  • vllm
  • dgx-spark
  • sm120
  • dflash2
  • multimodal
  • uncensored
  • abliterated
  • moe
  • safetensors
    quantized_by: neko-legends (Depths)
    language:
  • en

GLM-5.3 Flash Uncensored EXL3, rank-sliced for use on two DGX Spark nodes

What this is

This is a lossless TP2 storage transformation of
neko-legends/GLM-5.3-Flash-Uncensored-EXL3.
It stores each routed-expert EXL3 tensor as explicit rank0 and rank1 Trellis
tensors. The conversion does not dequantize, recalibrate, or requantize
Neko-Legends' version.

In plain terms, this uses the newer EXL3 standard of Rank-Slicing to split
across two nodes for dual DGX setups.

Uncensored model: the language checkpoint has undergone abliteration
to reduce refusal behavior. Treat outputs as untrusted, apply application-level
safeguards, and do not assume the model will decline harmful requests.

User responsibility: this model is provided without warranty. The
creators, uploaders, and maintainers are not responsible or liable for what
others generate, publish, deploy, or otherwise do with this abliterated model.
Users must operate it responsibly, apply appropriate safeguards, comply with
applicable law, and respect third-party rights. This model is for research
purposes only and is not intended for production use.

Why this exists

I tried a few different abliterated versions of GLM from different users here
and none of them worked well for me. Prone to looping, low tok/s, or a bit of
incoherence. I settled on Neko-Legends EXL3 version and really liked it, but the
only runner I could get it working with was the Entrpi vLLM one that uses
some older libraries and was pretty slow for me. (note that I think their builds
are awesome, I just wanted to take advantage of some newer updates.)

We created a new build of vLLM with native DGX libraries to run the Neko-Legends model
and then realized it wouldn't work unless we rank-sliced it. So we updated the model.

Our custom runner saw an average of 15-25% tok/s speed increase in agentic usage
over the previous EXL3 one we used.

Required runner

Stock vLLM 0.29 and Transformers do not understand this rank-sliced tensor
schema. To run this we've created a custom version of vLLM using all native
sm120/sm121 libraries. Use:

cbertucci33/vllm-v29-glm53flash-exl3-dgx

The live qualification image was built from runtime commit 974329d88. Later
commits are compatible only when they preserve the same exl3-trellis schema.

The tested runner profile uses:

vLLM: 0.29.0 base
tensor parallelism: 2
quantization: exl3
load format: instanttensor
attention backend: FLASHINFER_MLA_SPARSE_SM120
target KV cache: fp8_ds_mla
maximum model length: 800000
maximum sequences: 2

The DFlash2 draft checkpoint is separate from this repository. The qualified
production profile used:

local-inference-lab/GLM-5.3-Flash-DFlash2-MXFP8

with 7 speculative tokens. Note that this Dflash drafter was created to
accept 7 tokens. However we customized our runner to accept any number of draft
tokens between 1-7. We recommend using 5 as the best balance for performance.

Also note - in the repository for our runner is a python script to slice any
other EXL3 model for this type of multi-node inference if needed.

Runtime qualification

This exact checkpoint completed:

  • a full two-node, 164 GiB checkpoint load on two NVIDIA DGX Sparks;
  • Sparkinfer Trellis execution for the rank-sliced experts;
  • native SM120 FlashInfer attention;
  • warmup and CUDA graph capture;
  • a bounded API canary;
  • internal smoke test with 15 model turns and 17 tool calls, scored 8/8.

The accepted workload processed 335,046 prompt tokens and 18,672 completion
tokens in 521.94 seconds. Both ranks stayed at zero restarts with no OOM or
runtime error signature. The runner reported 1,152,606 logical KV-cache tokens
with 11,000,000,000 cache bytes per rank.

Licensing and attribution

This checkpoint remains a mixed-license artifact:

  • Z.ai's GLM-5.3 Flash base is MIT licensed.
  • The uncensored weight lineage is from orcarouter.
  • The EXL3 quantization was produced by neko-legends (Depths).
  • Per-expert suh, svh, and mcg scale tensors derive from the
    ShapleyMcg-calibrated checkpoint and remain under the ShapleyMcg License
    v1.0 included as ShapleyMcg-LICENSE.
  • This repository changes storage layout only. It does not claim authorship
    of the source model or quantization.

ShapleyMcg was created by Brandon M. Music. Its attribution-required license
grants no rights to the person known as 0xSero. Use of ShapleyMcg without the
required attribution is unlicensed.

@misc{music2026shapleymcg,
  author = {Music, Brandon M.},
  title  = {ShapleyMCG: An Auditable Calibration-to-Encoding Pipeline for
            Low-Bit Mixture-of-Experts Models},
  year   = {2026},
  url    = {https://github.com/brandonmmusic-max/shapleymcg},
  note   = {Licensed under the ShapleyMcg License v1.0}
}
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in app" button that hands off directly to a local runtime of your choice - Infrahuman, LM Studio, or Ollama. No API keys, no subscription, no prompt leakage.