license: mit
base_model: cantina-security/apex-flash-1-abliterated
base_model_relation: quantized
tags:
- glm5_next
- abliterated
- security-research
- gguf
- imatrix
pipeline_tag: image-text-to-text
library_name: gguf
apex-flash-1-abliterated-GGUF
GGUF quantization of cantina-security/apex-flash-1-abliterated (321B MoE, GLM-5.3-Flash architecture, llama.cpp arch glm5_next).
| Folder | Size | BPW | Notes |
|---|---|---|---|
UD-IQ3_XXS/ |
120.9 GB | 3.01 | Same per-tensor types as Unsloth's UD-IQ3_XXS of GLM-5.3-Flash |
How it was made
- Converted from the BF16 checkpoint with current llama.cpp
convert_hf_to_gguf.py. - Quantized with
llama-quantizeusing Unsloth's GLM-5.3-Flash importance matrix (imatrix_unsloth.gguf; apex is a fine-tune of the same base model) and a per-tensor type file reproducing unsloth/GLM-5.3-Flash-GGUFUD-IQ3_XXS:
routed expertsffn_gate/up_expsIQ2_S,ffn_down_expsIQ3_S (a few layers IQ4_XS / Q3_K / Q2_K), attention and shared experts Q6_K. - All 129 routed-expert tensors match Unsloth's types exactly. 332 small non-expert tensors (KDA
ssm_*,hc_*_fn, MLAattn_k_b/v_b/kv_a_mqa, indexer) stay BF16/F32 instead of Q8_0, because current llama.cpp does not quantize them. This is slightly larger and more precise. - The vision projector is not included; use
mmprojfrom the Unsloth repo (same vision tower).
Notes
- From the upstream card: this abliterated variant has not undergone a separate evaluation; it is intended for authorized security research. This quantization has not been benchmarked. Not affiliated with Cantina Security, Z.AI or Unsloth.
- Other formats: FP8, NVFP4.
License
MIT, inherited from the base model. Copyright (c) 2026 Z.AI Co., Ltd — see LICENSE.