← back to catalog · registered 2026-10-02 17:58

pqhaz/apex-flash-1-abliterated-FP8

pqhaz multimodal second-order
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/pqhaz%2Fapex-flash-1-abliterated-FP8"
Response includes
  • classification unknown
  • files 72
  • author_summary 1 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
?
Primary method

Unclassified

No clear signals of an abliteration technique in this model.
Confidence
UNKNOWN
Why this label 1 signal
No classification signals present. This may not be an abliterated model at all - it could be a repackaging, a merge with unrelated goals, or unrelated content that mentions the term.
  • no classification signals present (no abliterated, uncensored, or known producer/method markers)
Refusal direction extraction

No specific extraction method could be identified for this model. The producer either did not document it or used a proprietary pipeline.

What is a refusal direction? →
Downloads · 30-day
0
Likes
0
Model age
today
created 2026-10-02

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
mit
Tags
transformers safetensors glm5_next image-text-to-text abliterated security-research fp8 conversational base_model:cantina-security/apex-flash-1-abliterated base_model:quantized:cantina-security/apex-flash-1-abliterated license:mit endpoints_compatible

Related

Total size
306 GB
Files
72
Quantizations
1
Registered
2026-10-02 17:58
Last updated on HF
2026-10-02 17:42

Files by quantization

Auxiliary files 72 files 306 GB
model-00001-of-00062.safetensors 5.00 GB c5f64a01 download
model-00024-of-00062.safetensors 5.00 GB 3f67a1b1 download
model-00038-of-00062.safetensors 5.00 GB 6da87445 download
model-00045-of-00062.safetensors 5.00 GB 6d504eba download
model-00052-of-00062.safetensors 5.00 GB 6dbfcab9 download
model-00007-of-00062.safetensors 5.00 GB d496e7b0 download
model-00014-of-00062.safetensors 5.00 GB 2068aa87 download
model-00021-of-00062.safetensors 5.00 GB 0c1cbd99 download
model-00042-of-00062.safetensors 5.00 GB aaa41e79 download
model-00049-of-00062.safetensors 5.00 GB 5a77026a download
model-00008-of-00062.safetensors 5.00 GB 27a19e24 download
model-00015-of-00062.safetensors 5.00 GB 47c156f8 download
model-00018-of-00062.safetensors 5.00 GB 8e44ccac download
model-00025-of-00062.safetensors 5.00 GB 7bec56aa download
model-00029-of-00062.safetensors 5.00 GB f83f2384 download
model-00036-of-00062.safetensors 5.00 GB e9161449 download
model-00050-of-00062.safetensors 5.00 GB bb7bab05 download
model-00039-of-00062.safetensors 5.00 GB f08cb09c download
model-00043-of-00062.safetensors 5.00 GB d04c0679 download
model-00004-of-00062.safetensors 5.00 GB d1161cdc download
model-00012-of-00062.safetensors 5.00 GB 1e6a3e70 download
model-00019-of-00062.safetensors 5.00 GB 54a76a8b download
model-00026-of-00062.safetensors 5.00 GB 4500e438 download
model-00033-of-00062.safetensors 5.00 GB 2a4fb327 download
model-00054-of-00062.safetensors 5.00 GB 7c1ccc9a download
model-00031-of-00062.safetensors 5.00 GB 45cd3e60 download
model-00056-of-00062.safetensors 5.00 GB 26b1883d download
model-00047-of-00062.safetensors 5.00 GB 3e08d35e download
model-00060-of-00062.safetensors 5.00 GB 771ef372 download
model-00057-of-00062.safetensors 5.00 GB 062ac916 download
model-00009-of-00062.safetensors 5.00 GB e6bfb610 download
model-00017-of-00062.safetensors 5.00 GB ca84efb6 download
model-00028-of-00062.safetensors 5.00 GB b4a08be1 download
model-00035-of-00062.safetensors 5.00 GB 668eeeee download
model-00022-of-00062.safetensors 5.00 GB 3d73e7c5 download
model-00053-of-00062.safetensors 5.00 GB 176a0d9b download
model-00005-of-00062.safetensors 5.00 GB 29347de2 download
model-00011-of-00062.safetensors 5.00 GB f046f15b download
model-00040-of-00062.safetensors 5.00 GB 9934bf86 download
model-00046-of-00062.safetensors 5.00 GB f9a5432d download
model-00032-of-00062.safetensors 5.00 GB c1530ae3 download
model-00059-of-00062.safetensors 5.00 GB 10734f97 download
model-00003-of-00062.safetensors 5.00 GB fd0d7f82 download
model-00010-of-00062.safetensors 4.99 GB 0e0f73cc download
model-00048-of-00062.safetensors 4.99 GB 09c3ec0f download
model-00027-of-00062.safetensors 4.99 GB 078418ae download
model-00034-of-00062.safetensors 4.99 GB 3af0955f download
model-00041-of-00062.safetensors 4.99 GB 0202e2ac download
model-00020-of-00062.safetensors 4.99 GB 4ec28735 download
model-00013-of-00062.safetensors 4.99 GB d517b66b download
model-00006-of-00062.safetensors 4.99 GB 2d4abb55 download
model-00051-of-00062.safetensors 4.99 GB 22718be0 download
model-00044-of-00062.safetensors 4.99 GB 7104597a download
model-00030-of-00062.safetensors 4.99 GB 0d12849d download
model-00037-of-00062.safetensors 4.99 GB 2b0f0a80 download
model-00023-of-00062.safetensors 4.99 GB 8c1d8b71 download
model-00016-of-00062.safetensors 4.99 GB cd17bac2 download
model-00055-of-00062.safetensors 4.99 GB 09269681 download
model-00058-of-00062.safetensors 4.99 GB 86fbee8f download
model-00002-of-00062.safetensors 4.96 GB 40c5aed5 download
model-00061-of-00062.safetensors 4.94 GB fbc74677 download
model-00062-of-00062.safetensors 1.17 GB d2d39bb7 download
tokenizer.json 19.3 MB 19e77364 download
model.safetensors.index.json 8.02 MB 521dc5d1 download
config.json 67.8 KB 339bb6e6 download
chat_template.jinja 10.7 KB 06bd89e9 download
README.md 2.08 KB d3d242d9 download
.gitattributes 1.53 KB 52373fe2 download
LICENSE 1.04 KB 986b06fb download
processor_config.json 909 B 3ec2a058 download
tokenizer_config.json 761 B e375fa0a download
generation_config.json 194 B 637ee6af download

README current version from Hugging Face


license: mit
base_model: cantina-security/apex-flash-1-abliterated
base_model_relation: quantized
tags:

  • glm5_next
  • abliterated
  • security-research
  • fp8
    pipeline_tag: image-text-to-text
    library_name: transformers

apex-flash-1-abliterated-FP8

FP8 quantization of cantina-security/apex-flash-1-abliterated (321B MoE, GLM-5.3-Flash architecture), which is distributed in BF16 only (~643 GB).

This checkpoint is ~328 GB and fits on 4× 96 GB GPUs (e.g. RTX PRO 6000 Blackwell) with room left for KV cache.

Format

Identical layout to the official zai-org/GLM-5.3-Flash FP8 release:

  • quant_method: fp8, fmt: e4m3, activation_scheme: dynamic, weight_block_size: [128, 128]
  • Exactly the same set of tensors is quantized as in the official checkpoint (routed + shared experts, dense MLPs, MLA projections); everything in modules_to_not_convert (embeddings, lm_head, routers, norms, linear-attention/KDA, hyper-connections, vision tower) is left in BF16/F32 as in the source.
  • Each FP8 weight has a weight_scale_inv (F32, one scale per 128×128 block, scale = amax / 448). Round-to-nearest, no calibration data.
  • Tensor names and the quantization_config are copied from the official release, so any engine that serves zai-org/GLM-5.3-Flash should load this checkpoint the same way.

Serving

Use the same setup as for zai-org/GLM-5.3-Flash, e.g.

vllm serve pqhaz/apex-flash-1-abliterated-FP8 --tensor-parallel-size 4
python -m sglang.launch_server --model-path pqhaz/apex-flash-1-abliterated-FP8 --tp 4

Notes

  • From the upstream card: this abliterated variant has not undergone a separate evaluation. Results reported for apex-flash-1 apply to the standard checkpoint only. It is intended for authorized security research.
  • This quantization has not been separately benchmarked either.
  • Not affiliated with Cantina Security or Z.AI.

License

MIT, inherited from the base model. Copyright (c) 2026 Z.AI Co., Ltd — see LICENSE.

Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration