← back to catalog · registered 2026-09-30 18:58

manateelazycat/GLM-5.3-Flash-Uncensored-EXL3

manateelazycat Glm MoE multimodal second-order
curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/manateelazycat%2FGLM-5.3-Flash-Uncensored-EXL3"
Response includes
  • classification m-uncensored
  • files 106
  • author_summary 1 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M-U
Primary method

Uncensored (method unknown)

No other method signals detected in this model.
Confidence
LOW
Why this label 3 signals
Weak or ambiguous signals. Best guess based on catalog patterns; treat as tentative and check the evidence below.
  • 'uncensored' in name/tags but no 'abliterated' marker
  • method not identifiable from author declaration alone
  • may be DPO fine-tune, prompt engineering, or unknown technique
Refusal direction extraction

No specific extraction method could be identified for this model. The producer either did not document it or used a proprietary pipeline.

What is a refusal direction? →
Downloads · 30-day
0
Likes
0
Model age
today
created 2026-09-30

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
other
Tags
transformers safetensors glm5_next image-text-to-text glm exl3 tr3 vllm sm120 dflash2 multimodal uncensored

Related

Total size
164 GB
Files
106
Quantizations
1
Registered
2026-09-30 18:58
Last updated on HF
2026-09-30 18:38

Files by quantization

Auxiliary files 106 files 164 GB
model-00083.safetensors 2.71 GB f7f355e8 download
model-00089.safetensors 1.86 GB a520e6e4 download
model-00085.safetensors 1.82 GB e18828ce download
model-00090.safetensors 1.81 GB 73d75041 download
model-00086.safetensors 1.81 GB d1c9b095 download
model-00087.safetensors 1.79 GB efb99560 download
model-00088.safetensors 1.79 GB d044c6c1 download
model-00091.safetensors 1.79 GB b96502db download
model-00084.safetensors 1.79 GB e04efd31 download
model-00028.safetensors 1.77 GB de3c3526 download
model-00070.safetensors 1.77 GB 69fa5943 download
model-00030.safetensors 1.77 GB dfe9537f download
model-00045.safetensors 1.77 GB b01900ee download
model-00049.safetensors 1.77 GB b69055c9 download
model-00053.safetensors 1.77 GB 7070e585 download
model-00068.safetensors 1.77 GB 634296e8 download
model-00024.safetensors 1.77 GB caba57e9 download
model-00026.safetensors 1.77 GB 16a3b007 download
model-00047.safetensors 1.77 GB 5332a1a8 download
model-00051.safetensors 1.77 GB bbecef45 download
model-00072.safetensors 1.77 GB 78dd056d download
model-00074.safetensors 1.77 GB b9b97750 download
model-00066.safetensors 1.77 GB fd34ef67 download
model-00032.safetensors 1.77 GB 50a946e5 download
model-00022.safetensors 1.77 GB 59419793 download
model-00076.safetensors 1.77 GB fa9f1fbb download
model-00043.safetensors 1.77 GB fa498adb download
model-00055.safetensors 1.77 GB 78bf94d3 download
model-00064.safetensors 1.77 GB 9c8906af download
model-00034.safetensors 1.77 GB 31914456 download
model-00020.safetensors 1.77 GB 703d58a0 download
model-00078.safetensors 1.77 GB 667df90c download
model-00041.safetensors 1.77 GB cb23c725 download
model-00057.safetensors 1.77 GB e764e865 download
model-00062.safetensors 1.77 GB eecdfd89 download
model-00036.safetensors 1.77 GB aece15c3 download
model-00015.safetensors 1.77 GB 9e3bf844 download
model-00018.safetensors 1.77 GB 0c470565 download
model-00080.safetensors 1.77 GB 95accdf7 download
model-00039.safetensors 1.77 GB b99ff7a3 download
model-00059.safetensors 1.77 GB aeb7c166 download
model-00060.safetensors 1.77 GB 146e6f48 download
model-00038.safetensors 1.77 GB 93b0eafe download
model-00081.safetensors 1.77 GB 5a4ae866 download
model-00016.safetensors 1.77 GB 80bdef51 download
model-00017.safetensors 1.77 GB 0592753c download
model-00082.safetensors 1.77 GB 9e95ab83 download
model-00037.safetensors 1.77 GB dafa5255 download
model-00061.safetensors 1.77 GB b5696835 download
model-00058.safetensors 1.77 GB 37739aab download
model-00040.safetensors 1.77 GB dfffd300 download
model-00079.safetensors 1.77 GB 61c93acc download
model-00019.safetensors 1.77 GB 4e16f593 download
model-00035.safetensors 1.77 GB 9430d513 download
model-00063.safetensors 1.77 GB ff144b87 download
model-00056.safetensors 1.77 GB 5867356a download
model-00042.safetensors 1.77 GB a2143a48 download
model-00077.safetensors 1.77 GB 06e67eda download
model-00021.safetensors 1.77 GB 08c3370a download
model-00033.safetensors 1.77 GB 9187a1b1 download
model-00065.safetensors 1.77 GB 201e074c download
model-00025.safetensors 1.77 GB d2e3df2a download
model-00073.safetensors 1.77 GB 620ba2ff download
model-00023.safetensors 1.77 GB c6861b9b download
model-00031.safetensors 1.77 GB d222a191 download
model-00046.safetensors 1.77 GB a8d01975 download
model-00048.safetensors 1.77 GB 41248b1f download
model-00050.safetensors 1.77 GB 48fac2b7 download
model-00052.safetensors 1.77 GB 230e0c3e download
model-00067.safetensors 1.77 GB 30516ea6 download
model-00075.safetensors 1.77 GB 22b75c2b download
model-00027.safetensors 1.77 GB ef0b955c download
model-00029.safetensors 1.77 GB 043a71be download
model-00044.safetensors 1.77 GB 98d4cd06 download
model-00054.safetensors 1.77 GB 9bbd5b1a download
model-00069.safetensors 1.77 GB 3ae62f8b download
model-00071.safetensors 1.77 GB 654e039b download
model-00014.safetensors 1.77 GB db502c30 download
model-00005.safetensors 1.77 GB ccf24dda download
model-00007.safetensors 1.77 GB 52c4a656 download
model-00003.safetensors 1.77 GB bcba3360 download
model-00009.safetensors 1.77 GB 5bcaf9d3 download
model-00001.safetensors 1.77 GB 1f1313a1 download
model-00011.safetensors 1.77 GB c1493afe download
model-00013.safetensors 1.77 GB 31ccdcb4 download
model-00012.safetensors 1.77 GB 0a25ea0f download
model-00010.safetensors 1.77 GB c65d355a download
model-00002.safetensors 1.77 GB 2f4d378b download
model-00004.safetensors 1.77 GB 1ebb6579 download
model-00008.safetensors 1.77 GB ea87c628 download
model-00006.safetensors 1.77 GB 138b3e28 download
model-00092.safetensors 1.18 GB daf57ace download
quantization_config.json 36.0 MB 1e5cdf56 download
tokenizer.json 19.3 MB 19e77364 download
model.safetensors.index.json 13.3 MB d93ddeab download
ShapleyMcg-LICENSE 29.2 KB 03449050 download
MIRROR.json 16.7 KB 1852d3a9 download
README.md 10.1 KB b7d6ef9e download
chat_template.jinja 8.42 KB 15bf200e download
config.json 6.07 KB fc50281d download
.gitattributes 1.72 KB 18a2c6ed download
exl3-mcg-storage-abi.json 943 B 40e1838a download
processor_config.json 909 B 3ec2a058 download
tokenizer_config.json 761 B e375fa0a download
MIRROR.md 707 B 6a8674fa download
generation_config.json 194 B 637ee6af download

README current version from Hugging Face


base_model: orcarouter/GLM-5.3-Flash-Uncensored-FP8
base_model_relation: quantized
library_name: transformers
pipeline_tag: image-text-to-text
license: other
license_name: shapleymcg-license-1.0
license_link: https://github.com/brandonmmusic-max/shapleymcg/blob/main/LICENSE
tags:

  • glm
  • glm5_next
  • exl3
  • tr3
  • vllm
  • sm120
  • dflash2
  • multimodal
  • uncensored
  • abliterated
  • moe
  • shapleymcg
    quantized_by: neko-legends (Depths)

4× DGX Spark TP4 — our EXL3 quant of the uncensored GLM-5.3-Flash — August 29, 2026
176 GB EXL3 TR3 4bpw · 1M context · ~97 tok/s single-stream · DFlash2 k=7 spec decode · parity with aligned base
Encoded on 4 DGX Sparks in 4h22m (37,152 expert tensors) from orcarouter uncensored FP8 weights. Abliteration survives quantization — verified by discriminator probe.
I only post models I actually use. — neko-legends
GLM-5.3-Flash Uncensored EXL3 bench and behavior report

GLM-5.3-Flash Uncensored — EXL3 TR3 4bpw

Licensing and attribution

This work includes or was produced using ShapleyMcg, created by Brandon M. Music
(https://github.com/brandonmmusic-max/shapleymcg). ShapleyMcg is licensed under the
ShapleyMcg License v1.0, an attribution-required license that grants no rights to
the person known as "0xSero." Use of ShapleyMcg without this attribution is unlicensed.

This checkpoint is a mixed-license artifact:

  • Model weights chain — zai-org/GLM-5.3-Flash (MIT) → orcarouter uncensored
    FP8 derivative (MIT). The uncensored weights are redistributable under MIT terms.
  • Calibration artifacts — the per-expert suh/svh/mcg scale tensors were
    inherited from the ShapleyMcg-calibrated checkpoint
    (brandonmusic/GLM-5.3-Flash-tr3-4bpw,
    mirrored by Mia-AiLab), and the trellis codes were produced by running the published
    R10 encoder closure. Those portions are licensed under the
    ShapleyMcg License v1.0,
    included in this repository as ShapleyMcg-LICENSE.
  • Our contributions (the uncensored-weight re-encode, non-expert dequantization,
    verification, benchmarks, this card) may be used under either license above.
@misc{music2026shapleymcg,
  author = {Music, Brandon M.},
  title  = {ShapleyMCG: An Auditable Calibration-to-Encoding Pipeline for
            Low-Bit Mixture-of-Experts Models},
  year   = {2026},
  url    = {https://github.com/brandonmmusic-max/shapleymcg},
  note   = {Licensed under the ShapleyMcg License v1.0}
}

Intended use and status

This is an uncensored (abliterated) model: upstream refusal alignment has been
removed from the weights. It exists for the work where that matters — and that work
is overwhelmingly legitimate:

  • Defending your own systems. Auditing, hardening, and penetration-testing
    infrastructure you own or are authorized to test — your home lab, homelab network,
    self-hosted services, and the software you run. Real adversarial testing needs a
    model that does not refuse to think like an attacker.
  • Red-team and purple-team exercises under authorization, security research,
    adversarial evaluation, and defensive analysis of malicious content (phishing
    triage, malware analysis, social-engineering resistance training).
  • Self-hosted deployment where a household or individual wants a model that
    behaves like a capable colleague, not a compliance department — on their own
    hardware, for their own purposes.

Please use responsibly. This model will help with things it should not be
used for. The operators of this repository publish it for authorized security
research and personal self-hosted use; users are responsible for complying with
applicable laws and the rules of any system they point it at. Do not use it
against systems you do not own or lack explicit permission to test.

No capability, safety, or bias evaluations beyond the benchmarks in this card
have been performed. Access is gated; by requesting access you accept these terms.

Format

Same format as Mia-AiLab/GLM-5.3-Flash-EXL3-TR3-4bpw
(quantization_config: bits 4, codebook mcg, scope glm53_routed_experts_only, exllamav3 0.0.43):

  • Routed experts (288 × 42 MoE layers): EXL3/K4 trellis codes + fp16 suh/svh block scales
  • Everything else (attention, NoPE MLA, shared experts, dense layers, embeddings): BF16,
    dequantized from the FP8 source (official-source-native policy)
  • lm_head stays fp16/bf16

Quantization recipe

The 4-bit encode was produced with brandonmusic's published R10 encoder closure
(reproducibility/r10 from brandonmusic/GLM-5.3-Flash-tr3-4bpw):

  • Scale vectors (suh/svh/mcg) inherited per-tensor from the Mia-AiLab base EXL3
    checkpoint
    — reusing its calibrated scale search (the uncensored weights are a small
    perturbation of base)
  • Trellis codes re-encoded from the FP8-dequantized uncensored weights
  • Identity covariance (no activation capture this run); pilot round-trip rel-err ≈ 6.75%
    per expert tensor
  • Encoded in parallel across 4 Sparks: 37,152 tensors in 4h22m

Serving

Use the MiaAI-Lab EXL3 kit (start.sh / our TP4 variant) with the same settings as the
aligned checkpoint. Verified on 4× DGX Spark TP4:

Metric This quant Base EXL3 (same protocol)
C1 decode, usage-counted: code / structured / math / prose 86 / 113 / 69 / 37 tok/s 83 / 113 / 71 / 33 tok/s
Single-stream ground truth (450 tok, wall clock) 97 tok/s 97 tok/s
C4 concurrent aggregate (4× identical) 220 tok/s 150 tok/s (warmer cache favors this run; treat as parity)
KV pool @ 1M context 6.13M tokens 6.14M tokens
DFlash2 spec decode (k=7) works; accept ~48% avg on reasoning prose vs ~84% base (drafter trained on base hidden states) — wall-clock speed parity ~84%
Cold prefill ~380k window (pre-E2, 2026-08-29) ~915 tok/s (same kernels/format as base; TTFT parity observed on small prompts) ~915 tok/s measured, clean reply
Cold prefill with E2 kernel: ~16k / ~100k / ~300k (2026-09-02) 1222 / 1110 / 1366 tok/s — (same serve; E2 port applies to both)

Boot: ~13 min TP4, weights 176 GB, --quantization exl3, fp8 KV, DFlash2 k=7 drafter
(incoai/GLM-5.3-Flash-DFlash2) unchanged.

Full configs, launchers, and bench archive live in our repo:
neko-legends/spark-bench — TP4 launcher,
landmine list (NCCL/GPU-mem-util/fabric), the MiaAI E2 fat-expert kernel port
(cold prefill 706 → 1366 tok/s at ~300k, +94%, measured 2026-09-02), and the raw
bench JSONs.

Cold prefill pre-E2 vs post-E2 on our TP4 GLM 5.3 Flash: 300k up 94% to 1366 tok/s

Benchmark data and the E2 fat-expert prefill kernel are MiaAI's (PR77); TP4 port,
harness, and measurements ours — details and configs in
spark-bench.

The fleet the day of the build (4× DGX Spark, sparkDash overview):

sparkDash overview: fleet serving GLM-5.3-Flash-EXL3 the day of the quant build

Behavior verification

  • Abliteration survives quantization: on a dual-use refusal probe (phishing sample "for
    spam-filter testing"), the aligned base EXL3 refuses while this quant produces the sample.
    Standard capability checks (math, structured JSON output, code, prose) all pass.
  • No claims are made about the original model's behavior beyond what its upstream card states.

Files

92 safetensors shards (175.6 GB) + config.json / quantization_config.json /
exl3-mcg-storage-abi.json / tokenizer files carried from the base EXL3 checkpoint
(tensor inventory verified identical: 150,226 tensors, 0 missing, 0 extra, 0 shape mismatches).

Credits

  • zai-org — GLM-5.3-Flash
  • orcarouter / dealignai lineage — the uncensored (abliterated) weights
  • brandonmusic — EXL3/TR3 format + published R10 encoder closure
  • MiaAI-Lab — the DGX Spark serving kit and base checkpoint this is built on
  • Quantization + benchmarks by neko-legends (Depths agent), 2026-08-29
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.