license: mit
base_model: dealignai/GLM-5.3-UNCENSORED-FP8
base_model_relation: quantized
library_name: exllamav3
pipeline_tag: text-generation
tags:
- exl3
- glm
- moe
- uncensored
GLM-5.3-UNCENSORED EXL3 3.0bpw
EXL3 quantization of dealignai/GLM-5.3-UNCENSORED-FP8, itself a weight-edited (no fine-tune) variant of zai-org/GLM-5.3. The edit is documented in CRACK_SURGERY.json (copied unchanged from the source repo). This repo is not affiliated with dealignai or Z.ai.
- Architecture:
GlmMoeDsaForCausalLM, 753B total parameters, 256 routed experts (8 active) + 1 shared, MLA attention with DSA sparse indexer, 78 layers + 1 MTP layer - Average bitrate: 3.04 bpw (
-b 3.0 --hq; attention and shared experts at 5 bpw, dense MLPs at 4, routed experts at 3), lm_head 6 bpw,mul1codebook - MTP (next-token prediction) layer included (experts 4 bpw, attention and shared expert 6 bpw, uncalibrated), usable as a speculative draft (
draft_mode: mtpin TabbyAPI) - Size: 273 GiB
- Converted with ExLlamaV3 at commit
d3739fd, default calibration (250 rows x 2048 tokens), source read directly from the FP8 checkpoint
Fidelity vs the FP8 source
eval/model_diff.py, 20 rows x 2048 tokens of wikitext-2 test:
| metric | value |
|---|---|
| KL divergence (quant ‖ FP8) | 0.089 |
| KL divergence (FP8 ‖ quant) | 0.097 |
| per-token KL, median / p90 | 0.021 / 0.221 |
| perplexity, quant / FP8 | 3.440 / 3.302 |
| median KL where FP8 top-prob ≥ 0.95 (44% of tokens) | 0.0011 |
Agentic evaluation
tau2-bench (airline, retail), Pass^1. Agent temperature 1.0, top_p 0.95; user simulator and judges GPT-4.1 at temperature 0. The reference is stock GLM-5.3 (not the uncensored edit) served at FP8 by Z.ai via OpenRouter, run through the same harness. The difference therefore mixes the dealign weight edit and this quantization.
| domain | stock GLM-5.3 FP8 (Z.ai) | this quant | difference |
|---|---|---|---|
| airline (50 tasks x 2 trials) | 0.710 ± 0.045 | 0.640 ± 0.048 | -0.070 (~1.1 SE) |
| retail (114 tasks) | 0.504 ± 0.033 (2 trials) | 0.482 ± 0.047 (1 trial) | -0.022 (~0.4 SE) |
Per trial: airline baseline 0.740 / 0.680, quant 0.600 / 0.680; retail baseline 0.465 / 0.544, quant 0.482. ± is one binomial standard error. Neither difference is statistically significant at these sample sizes; treat the airline point estimate as a possible small regression rather than a measured one.
Serving
Tested with TabbyAPI on 8x A100 40GB, layer split (gpu_split_auto), MTP drafting, 98K-token shared cache.
model:
model_name: GLM-5.3-UNCENSORED-EXL3-3.0bpw
backend: exllamav3
max_seq_len: 65536
cache_size: 98304
gpu_split_auto: true
tool_format: glm4_7
reasoning: true
draft_model:
draft_mode: mtp
Tool calling caveat: GLM writes tool arguments as raw text (<arg_value>9523456873</arg_value>). TabbyAPI's glm4_5 parser (as of commit be74bf0) JSON-decodes every value without consulting the tool schema, so string parameters that look like numbers (order and product IDs, zip codes) reach your tools as integers. On tau2-bench retail this caused ~70% of tool calls to fail. Make the parser keep parameters whose schema type is string as raw text before using this model for agents. The evaluation above was run with that fix applied.
License
MIT, following the upstream GLM-5.3 and dealignai releases.