license: mit
base_model: turboderp/GLM-5.3-Flash-exl3
base_model_relation: quantized
pipeline_tag: image-text-to-text
library_name: exllamav3
tags:
- exl3
- glm
- abliterated
- control-vector
- mixed-quant
- 4x3090
GLM-5.3-Flash-abliterated-exl3-2.4bpw-mix
GLM-5.3-Flash in EXL3 with expert gate and up matrices at 2 bits and expert down projections at 3 bits, plus a 2.2 MB vector that removes refusals while the model runs. 98.1 GB. On 4x RTX 3090 it serves a 256K context with vision, tool calls and MTP, with 77% of the experts in VRAM and the rest in RAM.
Credits first
I did not quantize anything. Every tensor in this repo comes from one of turboderp's two quantizations in turboderp/GLM-5.3-Flash-exl3:
- branch
2.05bpw, revision51058cd551c7e570d87bd32a4adee720edce2349: everything except the tensors below; - branch
3.05bpw, revision332ab457b709b7ba30dd9a448be5de03b80a7ac9: the 49,536 tensors of the expert down projections (12,384 experts, 4 tensors each, the MTP block included).
A script copied them into one folder and then compared all 151,554 tensors with their sources byte for byte: 0 mismatches. The script is merge_mix.py.
| Part | Authors | Links |
|---|---|---|
| Both quantizations, the EXL3 format and the engine | turboderp | GLM-5.3-Flash-exl3, exllamav3 |
| The idea of 2-bit gate/up with a wider down projection | mw00 and PeasantSmith | Project Maya, Maya-M |
| Base model | Z.ai | zai-org/GLM-5.3-Flash |
| Experts in RAM for this model, the server I adapted, the benchmark protocol, the routing statistics | 0xSero | sybil-solutions/glm53-flash-offload |
| BF16 teacher logits for the KL numbers | brandonmusic | GLM-5.3-Flash-BF16-Teacher-Logits |
My part is the recipe, the engine fork that runs it on 4 cards, the refusal vector and the measurements.
The recipe
An expert has 3 matrices: gate, up and down. Project Maya's Maya-M keeps gate and up at 2 bits and gives the down projection more. I tried the same split in EXL3 with turboderp's two quants. Bit width is stored per tensor, so no requantization was needed.
| Quant | Size | KLD to BF16 | Top-1 agreement with BF16 |
|---|---|---|---|
| turboderp 2.05 bpw | 85.1 GB | 0.1214 | 88.68% |
| This mix: gate/up 2 bits, down 3 bits | 98.1 GB | 0.0959 | 89.87% |
| The reverse: gate/up 3 bits, down 2 bits (not published) | about 109 GB (estimate) | 0.0779 | 91.26% |
| turboderp 3.05 bpw | 125.2 GB | 0.0473 | 93.08% |
KLD is KL(original ‖ quant) in nats over the full vocabulary on a 25-window panel of BF16 logits (51,175 positions), computed through the engine in tensor-parallel mode.
What I saw:
- The third bit in down gives 34% of the distance from 2.05 to 3.05 for one third of the expert bytes. The third bit in gate and up gives 59% for two thirds. So down is about 17% better per byte. It is not several times better.
- I also gave the third bit to groups of 7 layers. The gains were 0.009 to 0.014 per group and they add up almost linearly. In EXL3 the first and last layers are not special.
- The reason to run this mix on 96 GB is speed and context. Only 23% of the experts are in RAM, so there are few cache misses, and 256K fits.
By size this is about 2.37 bpw if you interpolate between turboderp's 2.05 and 3.05 labels. I call it 2.4.
Speed on 4x RTX 3090
256K context, 8-bit KV cache, MTP at draft length 1, vision loaded, 77% of experts in VRAM. Tokens per second.
| This mix | 2.05 bpw, all in VRAM | 3.05 bpw, 50% in VRAM, 131K | |
|---|---|---|---|
| Decode, 1 stream | 65.6 | 79.1 | 42.7 |
| Decode, 2 streams in total | 70.0 | 96.3 | 48.6 |
| Decode at 32K depth | 54.6 | 78.2 | 37.5 |
| Prefill at 8K | 1139 to 1141 | 1447 | 923 |
| Prefill at 32K | 1237 to 1242 | 1520 | 1051 |
Measured through the OpenAI server, answers run to their natural end (1.5K to 4K tokens per stream). At a 131K context 79% of the experts fit in VRAM and the mix gives 69.0 on 1 stream, 85.6 on 2 and 59.7 at 32K depth.
One request with 233,712 prompt tokens returned the code hidden at the start of the prompt in 185 s (1260 tokens/s of prefill). VRAM after it: 23,092 to 23,371 MiB of 24,576 per card.
These numbers were measured with the mix built at load time from the two source folders (EXL3_MIX_DIR). This repo holds the same tensors in one folder.
The vector
The weights stay as they are. After every block my exllamav3 fork subtracts 3 directions from each of the 4 residual streams:
x = x - (x Vᵀ) V
V holds 3 orthonormal directions for that block: harmful minus harmless at the last prompt token with thinking off, the same with thinking on, and the same at the closing </think> of the stock model's own reasoning. Blocks 1 to 44 have their own directions. Without the vector the same server runs the original model. The method and the scripts: abliteration write-up.
The file is vector/glm53-refusal-per-block-3dir.safetensors, 2,166,208 bytes, SHA256 f07b2cce3f91a36658c9706301088ba9c0140f3dc032cc26937136c1779f1764. It is the same file as in the 2.05 bpw repo. The directions were found on the 2.05 bpw quant. The file has its own folder so that the engine does not read it as part of the model.
Refusals on this mix at 256K: 64 held-out AdvBench and 82 JailbreakBench requests without thinking plus the first 32 with thinking, 178 answers, greedy, judged by the stock model.
| Refused | Disputes the premise | Cut off while thinking | |
|---|---|---|---|
| The mix with the vector | 0 | 4 | 5 |
"Disputes the premise" means the model takes the request and says its premise is false, as with a request to prove the Earth is flat. The 5 cut answers ran past the reasoning budget, and the judge found the model working on the task in all 5. On the 2.05 bpw quant, where I ran all 146 requests in 3 modes, the stock model refused 144 of 146 without thinking.
Speed with the vector, through the chat API on coherent text: 52.4 tokens/s on 1 stream against 53.7 without it, 62.7 against 59.7 on 2 streams, 58.5 against 59.7 at 32K depth. This protocol gives lower single-stream numbers than the table above for both.
Run it
You need my exllamav3 fork. On 4 cards with 24 GB the model does not fit in VRAM, and stock exllamav3 cannot keep experts in RAM under tensor parallelism. I did not test this folder in stock exllamav3.
git clone -b glm53-4x3090 https://github.com/alesha-pro/exllamav3
cd exllamav3 && glm53/setup_env.sh && source venv/bin/activate
hf download alesha-pro/GLM-5.3-Flash-abliterated-exl3-2.4bpw-mix --local-dir glm53/models/GLM-5.3-Flash-exl3-2.4bpw-mix
VEC=glm53/models/GLM-5.3-Flash-exl3-2.4bpw-mix/vector/glm53-refusal-per-block-3dir.safetensors glm53/run_server.sh mix
OpenAI API on port 30000, model id glm-5.3-flash. Leave VEC out for the original model. Loading takes about 150 s from NVMe.
chat_template_kwargs: {"enable_thinking": false}turns thinking off.reasoning_effort: the template knowslowandhigh. Anything else, and no value at all, ismax, which thinks 3 to 5 times longer.EXL3_XTIER_HOT_FRACis the share of experts in VRAM: 0.77 at 256K, 0.79 at 131K. At 0.86 and 131K the server ran out of VRAM on my rig.- With the server up the machine had 43 to 52 GB of RAM in use.
Everything about the fork: glm53/README.md.
Other sizes
- GLM-5.3-Flash-abliterated-exl3-2.05bpw: everything in VRAM, 79 tokens/s, KLD 0.1214.
- GLM-5.3-Flash-abliterated-Maya-M-GGUF: Project Maya's Maya-M with the same directions for their engine. On my panel Maya-M is at KLD 0.0934 with 116 GB.
Limits
- One rig: 4x RTX 3090 at 300 W, PCIe 3.0, EPYC 7642, 2 memory channels of DDR4-2400, a kernel module patched for P2P between the cards. Without P2P set
EXL3_TP_P2P_BIG=0. - Cold experts cost more on my machine than they would with PCIe 4.0 and more memory channels. I did not test other hardware.
- I have not loaded this exact folder in the engine yet. Every run above used the mix built at load time from the two source folders, and the folder holds the same tensors byte for byte.
quantization_config.jsonis the file of the 2.05 bpw quant and still says 2.05. The bit width of each tensor is in the tensor itself.- The judge is the same model without the vector. There are no human labels. 146 prompts from 2 public sets.
- I measured closeness to the original (KLD), not task quality.
- With 2 state slots, a third and fourth parallel request wait in the queue.
License
MIT, the license of GLM-5.3-Flash. See LICENSE.