license: mit
base_model: dealignai/GLM-5.3-Flash-UNCENSORED-FP8
base_model_relation: quantized
library_name: mlx
pipeline_tag: image-text-to-text
language:
- en
- ko
- zh
tags: - mlx
- apple-silicon
- glm
- glm-5.3
- moe
- mixed-precision
- oq
- imatrix
- mtp
- uncensored
- abliterated
GLM-5.3-Flash-UNCENSORED-JK-oQ3.5e
JK is a custom mixed-precision MLX build of dealignai/GLM-5.3-Flash-UNCENSORED-FP8 (rev d5822b06, a weight-level uncensoring of zai-org/GLM-5.3-Flash, 320B total / about 18B active). It was quantized with oMLX oQ on an Apple M5 Ultra (256GB) to fit long-context work on that machine: 158 GiB on disk, about 15 GiB smaller than an oQ4e build of the same weights, at the same decode speed. Vision and the MTP draft head are kept.
Not affiliated with Z.AI or dealignai.
Layout
| Part | Bits (affine, group size 64) |
|---|---|
Routed experts gate_proj / up_proj, 14 most sensitive MoE layers (4, 5, 7, 8, 9, 11, 15, 16, 26, 27, 32, 35, 37, 39) |
4 |
Routed experts gate_proj / up_proj, the other 28 MoE layers |
3 |
Routed experts down_proj, all layers |
4 |
| MTP head experts | 4 |
Attention, shared experts, dense MLPs, embeddings, lm_head, DSA indexer |
8 |
About 4.22 bits per weight overall. Per-module bits are recorded in config.json (quantization).
How the 4-bit layers were chosen
On a uniform 4-bit copy of the base model, each MoE layer's expert gate_proj / up_proj was re-quantized to 3 bits with routing held fixed, and the relative squared error of that MoE block's output was measured (32 windows x 512 tokens). End-to-end KL against the reference was not used for ranking: it is dominated by routing changes, and even re-quantizing a layer at the same 4 bits moved it by 0.045 nats while NLL stayed within +-0.004.
Candidate sets were then compared by end-to-end NLL change (64 windows x 512 tokens, relative to the 4-bit reference):
4-bit gate_proj/up_proj layers |
NLL change |
|---|---|
| none (plain oQ3.5) | +0.041 |
| top 9 by measured error | +0.035 |
| first 7 + last 7 (position heuristic) | +0.033 |
| top 14 by measured error (this model) | +0.029 |
| top 19 by measured error | +0.026 |
Gains flatten after 14 layers, and 14 meets the size target.
Calibration (oQe imatrix)
512 windows x 512 tokens sampled from a 1.84M-token mix: the oMLX built-in oQe corpus (60%), private agent-session transcripts (20%), nightly session digests (10%) and technical notes (10%), with family-related content and secret-like lines removed. Only aggregate per-channel activation statistics enter the quantization; no text is stored in the weights. Every routed expert received tokens (5th percentile 2,179 tokens per expert). The MTP head has no imatrix entries and uses plain rounding.
Evaluation
Same runtime and settings for all rows. Percentages are relative to a uniform 4-bit MLX build of the base GLM-5.3-Flash.
| Decode NLL (1 token/step) | Decode NLL (3 tokens/step) | Korean held-out NLL | Tok/s greedy / sampled / 30K ctx | MTP acceptance | |
|---|---|---|---|---|---|
| Base GLM, uniform 4-bit | 0.99969 | 0.99341 | 2.1496 | 77.1 / 67.5 / 82.3 | 74.0% |
| Same uncensored weights, oQ4e (third-party build) | -0.7% | -0.8% | -1.5% | 69.2 / 63.5 / 75.2 | 73.9% |
| JK oQ3.5e | -0.4% | +1.0% | -0.7% | 69.6 / 64.2 / 75.7 | 72.9% |
- Functional checks passed: short summaries at low reasoning effort,
chat_template_kwargs, a sensitive request at low and high effort (answered), vision, and tool calling. - Six short reasoning problems at high effort: 6/6.
- No Chinese or Japanese characters leaked into Korean answers on ten everyday prompts (2,260 Hangul characters).
- Open question: unlike the other two builds, JK scores slightly worse with 3 tokens per step than with 1 token per step. This may be a numerical difference in the multi-token 3-bit expert kernels, which MTP verification uses.
Usage
GLM-5.3-Flash (glm5_next: hybrid linear and sparse MLA attention, hyper-connections, DSA indexer, MTP head) needs a runtime that supports this architecture. JK was built and tested with oMLX on Apple Silicon (GLM-5.3 branch, with MTP enabled and 3 draft tokens). Support in stock mlx-lm / mlx-vlm has not been verified.
The chat template is the one from the base GLM-5.3-Flash release. The dealignai template differs only in the clear_thinking default (true there, false here).
Responsible use
This model inherits dealignai's removal of safety refusals. It may produce content that the original model would refuse. You are responsible for how you use it and for complying with applicable laws.
License
MIT, as the upstream zai-org/GLM-5.3-Flash and dealignai/GLM-5.3-Flash-UNCENSORED-FP8. See LICENSE.