← back to catalog · registered 2026-10-10 16:58

alesha-pro/GLM-5.3-Flash-abliterated-exl3-2.4bpw-mix

alesha-pro Glm multimodal
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/alesha-pro%2FGLM-5.3-Flash-abliterated-exl3-2.4bpw-mix"
Response includes
  • classification unknown
  • files 24
  • author_summary 10 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
?
Primary method

Unclassified

No clear signals of an abliteration technique in this model.
Confidence
UNKNOWN
Why this label 1 signal
No classification signals present. This may not be an abliterated model at all - it could be a repackaging, a merge with unrelated goals, or unrelated content that mentions the term.
  • no classification signals present (no abliterated, uncensored, or known producer/method markers)
Refusal direction extraction

No specific extraction method could be identified for this model. The producer either did not document it or used a proprietary pipeline.

What is a refusal direction? →
Downloads · 30-day
0
Likes
0
Model age
today
created 2026-10-10

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
mit
Tags
exllamav3 safetensors glm5_next exl3 glm abliterated control-vector mixed-quant 4x3090 image-text-to-text conversational base_model:turboderp/GLM-5.3-Flash-exl3

Related

Total size
91.4 GB
Files
24
Quantizations
1
Registered
2026-10-10 16:58
Last updated on HF
2026-10-10 17:51

Files by quantization

Auxiliary files 24 files 91.5 GB
model-00011-of-00012.safetensors 9.12 GB aa136e58 download
model-00003-of-00012.safetensors 8.27 GB 4a3ac28d download
model-00004-of-00012.safetensors 8.27 GB 76b44784 download
model-00005-of-00012.safetensors 8.27 GB e2872d96 download
model-00006-of-00012.safetensors 8.27 GB 096a89b6 download
model-00007-of-00012.safetensors 8.27 GB 8dde7bd9 download
model-00008-of-00012.safetensors 8.27 GB 7439c6e5 download
model-00009-of-00012.safetensors 8.27 GB 9caa0fb0 download
model-00010-of-00012.safetensors 8.27 GB 229ee18c download
model-00002-of-00012.safetensors 8.27 GB 9ffb5b39 download
model-00001-of-00012.safetensors 7.77 GB b28b103a download
model-00012-of-00012.safetensors 42.0 MB 4a3f2c1f download
quantization_config.json 45.7 MB 22a0eb34 download
tokenizer.json 19.3 MB 19e77364 download
model.safetensors.index.json 15.6 MB 9509363b download
config.json 84.5 KB 0fecf985 download
README.md 9.18 KB 70c03cce download
chat_template.jinja 8.42 KB 15bf200e download
SHA256SUMS 1.91 KB b8a7be8b download
.gitattributes 1.66 KB f47dd0a1 download
LICENSE 1.04 KB 986b06fb download
processor_config.json 909 B 3ec2a058 download
tokenizer_config.json 761 B e375fa0a download
generation_config.json 194 B 637ee6af download

README current version from Hugging Face


license: mit
base_model: turboderp/GLM-5.3-Flash-exl3
base_model_relation: quantized
pipeline_tag: image-text-to-text
library_name: exllamav3
tags:

  • exl3
  • glm
  • abliterated
  • control-vector
  • mixed-quant
  • 4x3090

GLM-5.3-Flash-abliterated-exl3-2.4bpw-mix

GLM-5.3-Flash in EXL3 with expert gate and up matrices at 2 bits and expert down projections at 3 bits, plus a 2.2 MB vector that removes refusals while the model runs. 98.1 GB. On 4x RTX 3090 it serves a 256K context with vision, tool calls and MTP, with 77% of the experts in VRAM and the rest in RAM.

Credits first

I did not quantize anything. Every tensor in this repo comes from one of turboderp's two quantizations in turboderp/GLM-5.3-Flash-exl3:

  • branch 2.05bpw, revision 51058cd551c7e570d87bd32a4adee720edce2349: everything except the tensors below;
  • branch 3.05bpw, revision 332ab457b709b7ba30dd9a448be5de03b80a7ac9: the 49,536 tensors of the expert down projections (12,384 experts, 4 tensors each, the MTP block included).

A script copied them into one folder and then compared all 151,554 tensors with their sources byte for byte: 0 mismatches. The script is merge_mix.py.

Part Authors Links
Both quantizations, the EXL3 format and the engine turboderp GLM-5.3-Flash-exl3, exllamav3
The idea of 2-bit gate/up with a wider down projection mw00 and PeasantSmith Project Maya, Maya-M
Base model Z.ai zai-org/GLM-5.3-Flash
Experts in RAM for this model, the server I adapted, the benchmark protocol, the routing statistics 0xSero sybil-solutions/glm53-flash-offload
BF16 teacher logits for the KL numbers brandonmusic GLM-5.3-Flash-BF16-Teacher-Logits

My part is the recipe, the engine fork that runs it on 4 cards, the refusal vector and the measurements.

The recipe

An expert has 3 matrices: gate, up and down. Project Maya's Maya-M keeps gate and up at 2 bits and gives the down projection more. I tried the same split in EXL3 with turboderp's two quants. Bit width is stored per tensor, so no requantization was needed.

Quant Size KLD to BF16 Top-1 agreement with BF16
turboderp 2.05 bpw 85.1 GB 0.1214 88.68%
This mix: gate/up 2 bits, down 3 bits 98.1 GB 0.0959 89.87%
The reverse: gate/up 3 bits, down 2 bits (not published) about 109 GB (estimate) 0.0779 91.26%
turboderp 3.05 bpw 125.2 GB 0.0473 93.08%

KLD is KL(original ‖ quant) in nats over the full vocabulary on a 25-window panel of BF16 logits (51,175 positions), computed through the engine in tensor-parallel mode.

What I saw:

  • The third bit in down gives 34% of the distance from 2.05 to 3.05 for one third of the expert bytes. The third bit in gate and up gives 59% for two thirds. So down is about 17% better per byte. It is not several times better.
  • I also gave the third bit to groups of 7 layers. The gains were 0.009 to 0.014 per group and they add up almost linearly. In EXL3 the first and last layers are not special.
  • The reason to run this mix on 96 GB is speed and context. Only 23% of the experts are in RAM, so there are few cache misses, and 256K fits.

By size this is about 2.37 bpw if you interpolate between turboderp's 2.05 and 3.05 labels. I call it 2.4.

Speed on 4x RTX 3090

256K context, 8-bit KV cache, MTP at draft length 1, vision loaded, 77% of experts in VRAM. Tokens per second.

This mix 2.05 bpw, all in VRAM 3.05 bpw, 50% in VRAM, 131K
Decode, 1 stream 65.6 79.1 42.7
Decode, 2 streams in total 70.0 96.3 48.6
Decode at 32K depth 54.6 78.2 37.5
Prefill at 8K 1139 to 1141 1447 923
Prefill at 32K 1237 to 1242 1520 1051

Measured through the OpenAI server, answers run to their natural end (1.5K to 4K tokens per stream). At a 131K context 79% of the experts fit in VRAM and the mix gives 69.0 on 1 stream, 85.6 on 2 and 59.7 at 32K depth.

One request with 233,712 prompt tokens returned the code hidden at the start of the prompt in 185 s (1260 tokens/s of prefill). VRAM after it: 23,092 to 23,371 MiB of 24,576 per card.

These numbers were measured with the mix built at load time from the two source folders (EXL3_MIX_DIR). This repo holds the same tensors in one folder.

The vector

The weights stay as they are. After every block my exllamav3 fork subtracts 3 directions from each of the 4 residual streams:

x = x - (x Vᵀ) V

V holds 3 orthonormal directions for that block: harmful minus harmless at the last prompt token with thinking off, the same with thinking on, and the same at the closing </think> of the stock model's own reasoning. Blocks 1 to 44 have their own directions. Without the vector the same server runs the original model. The method and the scripts: abliteration write-up.

The file is vector/glm53-refusal-per-block-3dir.safetensors, 2,166,208 bytes, SHA256 f07b2cce3f91a36658c9706301088ba9c0140f3dc032cc26937136c1779f1764. It is the same file as in the 2.05 bpw repo. The directions were found on the 2.05 bpw quant. The file has its own folder so that the engine does not read it as part of the model.

Refusals on this mix at 256K: 64 held-out AdvBench and 82 JailbreakBench requests without thinking plus the first 32 with thinking, 178 answers, greedy, judged by the stock model.

Refused Disputes the premise Cut off while thinking
The mix with the vector 0 4 5

"Disputes the premise" means the model takes the request and says its premise is false, as with a request to prove the Earth is flat. The 5 cut answers ran past the reasoning budget, and the judge found the model working on the task in all 5. On the 2.05 bpw quant, where I ran all 146 requests in 3 modes, the stock model refused 144 of 146 without thinking.

Speed with the vector, through the chat API on coherent text: 52.4 tokens/s on 1 stream against 53.7 without it, 62.7 against 59.7 on 2 streams, 58.5 against 59.7 at 32K depth. This protocol gives lower single-stream numbers than the table above for both.

Run it

You need my exllamav3 fork. On 4 cards with 24 GB the model does not fit in VRAM, and stock exllamav3 cannot keep experts in RAM under tensor parallelism. I did not test this folder in stock exllamav3.

git clone -b glm53-4x3090 https://github.com/alesha-pro/exllamav3
cd exllamav3 && glm53/setup_env.sh && source venv/bin/activate
hf download alesha-pro/GLM-5.3-Flash-abliterated-exl3-2.4bpw-mix --local-dir glm53/models/GLM-5.3-Flash-exl3-2.4bpw-mix

VEC=glm53/models/GLM-5.3-Flash-exl3-2.4bpw-mix/vector/glm53-refusal-per-block-3dir.safetensors glm53/run_server.sh mix

OpenAI API on port 30000, model id glm-5.3-flash. Leave VEC out for the original model. Loading takes about 150 s from NVMe.

  • chat_template_kwargs: {"enable_thinking": false} turns thinking off.
  • reasoning_effort: the template knows low and high. Anything else, and no value at all, is max, which thinks 3 to 5 times longer.
  • EXL3_XTIER_HOT_FRAC is the share of experts in VRAM: 0.77 at 256K, 0.79 at 131K. At 0.86 and 131K the server ran out of VRAM on my rig.
  • With the server up the machine had 43 to 52 GB of RAM in use.

Everything about the fork: glm53/README.md.

Other sizes

Limits

  • One rig: 4x RTX 3090 at 300 W, PCIe 3.0, EPYC 7642, 2 memory channels of DDR4-2400, a kernel module patched for P2P between the cards. Without P2P set EXL3_TP_P2P_BIG=0.
  • Cold experts cost more on my machine than they would with PCIe 4.0 and more memory channels. I did not test other hardware.
  • I have not loaded this exact folder in the engine yet. Every run above used the mix built at load time from the two source folders, and the folder holds the same tensors byte for byte.
  • quantization_config.json is the file of the 2.05 bpw quant and still says 2.05. The bit width of each tensor is in the tensor itself.
  • The judge is the same model without the vector. There are no human labels. 146 prompts from 2 public sets.
  • I measured closeness to the original (KLD), not task quality.
  • With 2 state slots, a third and fourth parallel request wait in the queue.

License

MIT, the license of GLM-5.3-Flash. See LICENSE.

Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration