license: mit
license_link: LICENSE
base_model: orcarouter/GLM-5.3-Flash-Uncensored-FP8
base_model_relation: quantized
pipeline_tag: image-text-to-text
tags:
- gguf
- mixed-quant
- glm5-next
- moe
- dgx-spark
- ds4
- embedded-mtp
GLM-5.3-Flash Uncensored Mixed Quant GGUF
Mixed-precision GGUF derived from
orcarouter/GLM-5.3-Flash-Uncensored-FP8.
The recipe preserves higher precision in always-active paths and six boundary
routed layers, with IQ2_XXS as the minimum precision.
Status: build in progress. Model weights are not yet published. Sizes below
are conversion-plan estimates; runtime and quality results will be added after
the GGUF upload.
| Model property | Value |
|---|---|
| Architecture | 45 trunk layers; 34 KDA + 11 sparse-attention layers |
| Routed experts | 288 per routed layer, top-8 active |
| Native context | 1,048,576 tokens |
| MTP | One embedded block |
| Vision | Separate encoder, original BF16 weights |
| Planned main + MTP | 87.470 GiB |
| Planned total with vision | 88.520 GiB / 95.047 GB |
Quantization recipe
Layer indices are zero-based: trunk 0–44, MTP 45. Trunk layers 0–2 are dense.
| Tensor group | Type |
|---|---|
| Token embedding and output head | Q8_0 |
| Dense MLP layers 0–2 and shared experts | Q8_0 |
| KDA query/key projections | Q4_K |
| Remaining KDA and major sparse-attention/MLA projections | Q8_0 |
| Router weights/biases, norms and state controls | F32 |
| Indexer and mHC mixers; MTP hidden-state projection | BF16 |
| Routed gate/up, layers 3, 4, 5, 43, 44, 45 | IQ2_XS + imatrix |
| Routed gate/up, remaining routed layers | IQ2_XXS + imatrix |
| Routed down, layers 3, 4, 5, 43, 44, 45 | Q2_K (optional imatrix) |
| Routed down, remaining routed layers | IQ2_XS + imatrix |
| Vision encoder | BF16 |
Gate/up use the same format within each layer. Each tensor uses one block
format throughout; routed matrix widths remain divisible by 256. Calibration
sensitivity will be reported for this allocation.
Q2_K does not require an imatrix; optional importance weighting uses available
activation statistics without changing the format or file size.
Calibration uses GLM-native text inputs and relevant source references from
Inkling-Small Multimodal Calibration.
The imatrix is collected from text inputs. Original images are prepared for
separate image quality checks; image-conditioned activation calibration is not
claimed. MTP activation statistics are not collected by the current upstream
GLM graph; block 45 uses weight-based importance for IQ2_XS gate/up and standard
unweighted Q2_K down. Drafting behavior will be checked separately.
Runtime
This GGUF targets the native GLM format in
antirez/ds4. The mixed expert combinations need
the accompanying runtime extension; upstream llama.cpp compatibility is not
established. Exact runtime files and usage commands will accompany the weights.
The 1M context limit is preserved in metadata. Estimated KV/state/graph memory
at that limit is approximately 14.85 GiB, plus 1.15 GiB with MTP. Actual usable
context depends on available memory and runtime settings; DGX Spark maximum
context has not yet been measured. Image serving requires the BF16 encoder;
native video-input support is not established.
The source model's uncensored behavior and benchmark results are not new
measurements of this quantization. Quality results will identify the exact
tested build and settings.
License
MIT, inherited from the source checkpoint. See LICENSE.