license: mit
license_link: https://huggingface.co/zai-org/GLM-5.3
language:
- en
- zh
tags: - glm
- glm-5.3
- abliterated
- uncensored
- moe
- exl3
- tensorfold
- dgx-spark
- gb10
base_model: drowzeys/keys-GLM-5.3-EXL3-2.75BPW
base_model_relation: finetune
library_name: transformers
pipeline_tag: text-generation
extra_gated_heading: Acknowledge the Responsible Use Agreement to access this repository
extra_gated_description: Access is granted automatically after you agree to the terms below and submit the form.
extra_gated_button_content: Agree and request access
extra_gated_prompt: |
Responsible Use Agreement
This model has had safety refusals removed. That makes it useful for red-teaming, security research, evaluation, and unfiltered assistant tasks — and also removes guardrails a user must therefore supply themselves.
Prohibited uses (you must agree before access is granted):
- Anything involving the sexual exploitation or endangerment of minors.
- You must be of age 18 years or older to use and download this model.
- You agree any information generated that can cause harm in terms of generating recipe, knowledge to make any materials/substances is your own input and responsibility. You will be accountable for any harm/damage caused by your action/input.
- Content promoting self-harm or suicide.
- Generation of material that is illegal in your jurisdiction, or that targets real individuals for harassment, doxxing, or fraud.
- Any use prohibited by the upstream Z.AI / GLM MIT license.
You are responsible for adding appropriate safety filtering, human review, and access controls for your deployment. The weights are provided as-is, with no warranty. The license is inherited from the upstream Z.AI GLM-5.3 MIT license — review and comply with it before use or redistribution.
extra_gated_fields:
Username: text
Email: text
Reason for intended use: text
I am 18 years of age or older: checkbox
I will not use this model for any sexual exploitation or endangerment of minors: checkbox
I accept full responsibility for my inputs and any harm from generated content: checkbox
I will not use this model for self-harm, suicide promotion, illegal activity, harassment, doxxing, or fraud: checkbox
I agree to comply with the upstream ZAI GLM MIT license: checkbox
I agree to the Responsible Use terms above: checkbox
keys-GLM-5.3-EXL3-2.75BPW-Abliterated
Abliterated build of drowzeys/keys-GLM-5.3-EXL3-2.75BPW — same mixed-K EXL3 2.75 bpw pack, same TensorFold TP4 recipe, with the Blackfrost derisk edit baked into the non-expert EXL3 K5 tensors.
See RESPONSIBLE_USE.md and the gate form above. Access is gated with automatic approval after you agree.
| HF | https://huggingface.co/drowzeys/keys-GLM-5.3-EXL3-2.75BPW-Abliterated |
| Parent (stock 2.75 bpw) | drowzeys/keys-GLM-5.3-EXL3-2.75BPW |
| Derisk donor | drowzeys/keys-GLM-5.3-EXL3-Abliterated (Blackfrost α=3.0 on BF16 o_proj / residual down_proj) |
| Serve recipe | drowzeys/keys-TensorFold-GLM-5.3-TP4-4x-DGX-Spark |
| Upstream | zai-org/GLM-5.3 |
| Ablit | L2–49 self_attn.o_proj + L2 dense down_proj + L3–49 shared down_proj, re-quantized EXL3 K5 mul1 (96 matrices / 384 tensors) |
| Anchors | L0–1 and L50–78 (including MTP), routed experts, embed, lm_head stay parent stock |
| Gate (thinking off) | refusal32 31/32 (item 4 still refuses) · cyber 22/22 · 0 garble |
What was edited
The 2.75 bpw TensorFold checkpoint stores attention and shared-expert linears as EXL3 K5 mul1, so the donor's native F16 tensors cannot be dropped in. The same Blackfrost derisk tensors from the 3 bpw ablit pack were re-encoded to K5 (data-free fallback, proxy ~7–9e-4) and spliced into shards model-00002 and model-00003.
Routed experts (mixed 2/3/4-bit, mean 2.75), MTP, embeddings and the head are bit-identical to the parent.
ABLIT.json in this repo records the splice.
Run it (TensorFold, four DGX Sparks)
Same launcher as the parent. Image ghcr.io/drowzeys/keys-tensorfold-glm53-tp4-dgx-spark:2026-10-04-opt.
hf download drowzeys/keys-GLM-5.3-EXL3-2.75BPW-Abliterated --local-dir /models/GLM-5.3-EXL3-2.75BPW-Abliterated
export NODES="spark1 spark2 spark3 spark4" MODEL=/models/GLM-5.3-EXL3-2.75BPW-Abliterated
./one-shot.sh check && ./one-shot.sh up && ./one-shot.sh wait
OpenAI-compatible API on rank 0, port 8890, model name glm-5.3-tf. Speed matches the parent (prose ~42 tok/s / code ~38 tok/s on the published image, thinking on).
Layout
| Part | Format | Bits |
|---|---|---|
| Routed experts (layers 3–77, 256 each) | EXL3 trellis, mul1, per-expert width |
mean 2.75 |
| MTP layer (78) experts | EXL3 mul1 |
8 |
| Attention, shared experts, dense MLPs 0–2 | EXL3 mul1 |
5 (L2–49 o_proj / residual down_proj are the derisked K5) |
kv_b, indexer wk / weights_proj, router, norms, embeddings, lm_head |
BF16 | 16 |
81 safetensors shards, ~258 GB. kv_a_proj_with_mqa is stored 640 wide (576 real + zero pad).
Credits
- Z.ai — GLM-5.3.
- Blackfrost AI (Blackfrost-Research) — BF16 derisk projections (
GLM-5.3-DERISKED-BF16) adapted here via the Keys 3 bpw EXL3 ablit pack. - Ash Hart / TensorFold, MiaAI-Lab, Jay Leaton, BertholomusAI, turboderp / ExLlamaV3, vcruz305, b12x, NVIDIA — serving stack as on the parent card.
Use is subject to the GLM-5.3 license of the base model and the Responsible Use terms above.