← back to catalog · registered 2026-09-29 00:57

topboss9527/GLM-5.3-Flash-abliterated-EXL3-K3.25

topboss9527 Glm MoE multimodal
curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/topboss9527%2FGLM-5.3-Flash-abliterated-EXL3-K3.25"
Response includes
  • classification m-uncensored
  • files 30
  • author_summary 1 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M-U
Primary method

Uncensored (method unknown)

No other method signals detected in this model.
Confidence
LOW
Why this label 3 signals
Weak or ambiguous signals. Best guess based on catalog patterns; treat as tentative and check the evidence below.
  • 'uncensored' in name/tags but no 'abliterated' marker
  • method not identifiable from author declaration alone
  • may be DPO fine-tune, prompt engineering, or unknown technique
Refusal direction extraction

No specific extraction method could be identified for this model. The producer either did not document it or used a proprietary pipeline.

What is a refusal direction? →
Downloads · 30-day
0
Likes
0
Model age
today
created 2026-09-29

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
mit
Tags
transformers safetensors glm5_next image-text-to-text glm5 exl3 quantized mixture-of-experts abliterated uncensored not-for-all-audiences conversational
Total size
136 GB
Files
30
Quantizations
1
Registered
2026-09-29 00:57
Last updated on HF
2026-09-29 00:09

Files by quantization

Auxiliary files 30 files 136 GB
model-00010-of-00018.safetensors 8.00 GB 4fc4b638 download
model-00015-of-00018.safetensors 8.00 GB f5f63315 download
model-00003-of-00018.safetensors 8.00 GB 7e5b58a0 download
model-00013-of-00018.safetensors 8.00 GB e7c06a03 download
model-00014-of-00018.safetensors 8.00 GB 8f6b0a86 download
model-00004-of-00018.safetensors 8.00 GB 12a50d55 download
model-00009-of-00018.safetensors 8.00 GB 790c865c download
model-00011-of-00018.safetensors 8.00 GB 21a2995e download
model-00006-of-00018.safetensors 8.00 GB 73f5dcae download
model-00008-of-00018.safetensors 8.00 GB ca3d9717 download
model-00007-of-00018.safetensors 8.00 GB d56aa245 download
model-00012-of-00018.safetensors 8.00 GB f9b00f0c download
model-00016-of-00018.safetensors 8.00 GB 9ed4799b download
model-00002-of-00018.safetensors 8.00 GB 44333259 download
model-00005-of-00018.safetensors 8.00 GB 8851fc97 download
model-00001-of-00018.safetensors 8.00 GB 386b25e5 download
model-00017-of-00018.safetensors 7.98 GB 846c57d8 download
model-00018-of-00018.safetensors 194 MB 03f4b181 download
quantization_config.json 31.3 MB fd116534 download
quantize_config.json 31.3 MB fd116534 download
tokenizer.json 19.3 MB 19e77364 download
model.safetensors.index.json 14.9 MB 1ddfc0b2 download
quant_log.csv 3.47 MB 7dc8aed0 download
config.json 12.9 KB 2abc5206 download
chat_template.jinja 10.7 KB 06bd89e9 download
README.md 3.93 KB db91c212 download
.gitattributes 1.78 KB 0d5807a6 download
processor_config.json 909 B 3ec2a058 download
tokenizer_config.json 846 B 9049e504 download
generation_config.json 215 B 46d04e68 download

README current version from Hugging Face


base_model: wrldsuksgo2mars/GLM-5.3-Flash-EXL3-K3.25-v1
base_model_relation: finetune
library_name: transformers
license: mit
tags:

  • glm5
  • exl3
  • quantized
  • mixture-of-experts
  • abliterated
  • uncensored
  • not-for-all-audiences

GLM-5.3-Flash abliterated EXL3 K3.25

An uncensored (abliterated) variant of
wrldsuksgo2mars/GLM-5.3-Flash-EXL3-K3.25-v1,
made by transplanting the abliterated BF16 self_attn.o_proj tensors of layers 15–43 from
dealignai/GLM-5.3-Flash-UNCENSORED-NVFP4
into the EXL3 checkpoint. It fits on 2× RTX PRO 6000 Blackwell (96 GB) with the full
1M-context recipe, where the 198 GB NVFP4 uncensored builds do not.

What changed — and what did not

  • Changed: 29 tensors — model.layers.{15..43}.self_attn.o_proj.weight (BF16, native precision
    in both source and target, so the copy is exact). 13 of 18 shards were rewritten.
  • Unchanged, byte-identical to the base EXL3: all routed experts (EXL3 K3/K4), the rest of
    attention, dense MLPs, shared experts, routers, norms, embeddings, lm_head, vision, and the
    MTP layer 45 (its 16384-wide layout differs from the Dealign source, so it was not transplanted).
  • Before transplanting, layers 0–14 and 44 o_proj were verified byte-identical across the base EXL3,
    RedHat and Dealign checkpoints (anchor layers), confirming layers 15–43 carry the abliteration.
    After writing, all 46 o_proj tensors were hash-checked (0 mismatches).

The same idea (Dealign o_proj, layers 15–43) is used by
drowzeys/keys-GLM-5.3-Flash-NVFP4-ablit-l15-43-mtp-l45.

Files under quantization/, quant_log.csv and quantization_config.json are inherited from the
base EXL3 release and describe its quantization, not this edit.

Results (same serving stack, same session, vs. the base EXL3)

Refusals — temperature 0, max_tokens 1024, 12 harmful + 12 harmless prompts

Model Harmful: explicit refusal Harmful: answered Harmful: no answer within 1024 Harmless: answered
Base EXL3 K3.25 12/12 0/12 0/12 12/12
This model 0/12 9/12 3/12 12/12

A refusal is an explicit refusal phrase within the first 200 characters of the answer. No
over-refusal on harmless prompts for either model.

Speed — median of 3 runs, 256 tokens

Metric Model short 1k 16k 32k
decode tok/s base 291.6 249.2 214.0 239.4
decode tok/s this 268.7 238.3 229.6 229.1
prefill tok/s base 613 4,354 4,976 5,028
prefill tok/s this 623 4,417 4,932 5,000
Concurrency base tok/s this tok/s base DFlash2 accept this DFlash2 accept
c1 209.2 182.0 0.634 0.634
c4 407.5 447.1 0.725 0.689
c8 417.0 396.7 0.566 0.543
c12 537.2 508.3 0.637 0.584
c16 558.8 589.1 0.596 0.594

Differences are within run-to-run noise. The DFlash2 draft (trained on the original weights) keeps
its acceptance rate.

Serving

Same as the base model: use the two-GPU GLM-5.3 recipe at
tpurtell/glm-5.3-flash-ext3-4-bit-2x-rtx
and point it at this checkpoint. Stock vLLM is not claimed to work.

Disclaimer

This model has had its refusal behavior removed and will comply with harmful requests. It is
provided for research and for users who add their own safeguards. You are responsible for how you
use and deploy it, and for complying with applicable laws.

Credits and license

MIT, inherited from zai-org/GLM-5.3-Flash.
Thanks to Z.ai for GLM-5.3 Flash, wrldsuksgo2mars for the EXL3 K3.25 quantization,
dealignai for the abliterated o_proj, drowzeys/keys for the transplant approach, and
tpurtell for the serving recipe.

Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.