← back to catalog · registered 2026-10-02 22:58

pqhaz/apex-flash-1-abliterated-GGUF

pqhaz GGUF multimodal second-order
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/pqhaz%2Fapex-flash-1-abliterated-GGUF"
Response includes
  • classification unknown
  • files 3
  • author_summary 3 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
?
Primary method

Unclassified

No clear signals of an abliteration technique in this model.
Confidence
UNKNOWN
Why this label 1 signal
No classification signals present. This may not be an abliterated model at all - it could be a repackaging, a merge with unrelated goals, or unrelated content that mentions the term.
  • no classification signals present (no abliterated, uncensored, or known producer/method markers)
Refusal direction extraction

No specific extraction method could be identified for this model. The producer either did not document it or used a proprietary pipeline.

What is a refusal direction? →
Downloads · 30-day
0
Likes
0
Model age
today
created 2026-10-02

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Variants by this author 3 formats · 0 downloads combined

The same weights this author released in different packaging. Pick the format that matches your runtime.

Metadata

License
mit
Tags
gguf glm5_next abliterated security-research imatrix image-text-to-text base_model:cantina-security/apex-flash-1-abliterated base_model:quantized:cantina-security/apex-flash-1-abliterated license:mit endpoints_compatible region:us conversational

Related

Total size
0 B
Files
3
Quantizations
1
Registered
2026-10-02 22:58
Last updated on HF
2026-10-02 22:53

Files by quantization

Auxiliary files 3 files 4.80 KB
README.md 1.97 KB 70a7dcc7 download
.gitattributes 1.79 KB 20dd2c85 download
LICENSE 1.04 KB 986b06fb download

README current version from Hugging Face


license: mit
base_model: cantina-security/apex-flash-1-abliterated
base_model_relation: quantized
tags:

  • glm5_next
  • abliterated
  • security-research
  • gguf
  • imatrix
    pipeline_tag: image-text-to-text
    library_name: gguf

apex-flash-1-abliterated-GGUF

GGUF quantization of cantina-security/apex-flash-1-abliterated (321B MoE, GLM-5.3-Flash architecture, llama.cpp arch glm5_next).

Folder Size BPW Notes
UD-IQ3_XXS/ 120.9 GB 3.01 Same per-tensor types as Unsloth's UD-IQ3_XXS of GLM-5.3-Flash

How it was made

  • Converted from the BF16 checkpoint with current llama.cpp convert_hf_to_gguf.py.
  • Quantized with llama-quantize using Unsloth's GLM-5.3-Flash importance matrix (imatrix_unsloth.gguf; apex is a fine-tune of the same base model) and a per-tensor type file reproducing unsloth/GLM-5.3-Flash-GGUF UD-IQ3_XXS:
    routed experts ffn_gate/up_exps IQ2_S, ffn_down_exps IQ3_S (a few layers IQ4_XS / Q3_K / Q2_K), attention and shared experts Q6_K.
  • All 129 routed-expert tensors match Unsloth's types exactly. 332 small non-expert tensors (KDA ssm_*, hc_*_fn, MLA attn_k_b/v_b/kv_a_mqa, indexer) stay BF16/F32 instead of Q8_0, because current llama.cpp does not quantize them. This is slightly larger and more precise.
  • The vision projector is not included; use mmproj from the Unsloth repo (same vision tower).

Notes

  • From the upstream card: this abliterated variant has not undergone a separate evaluation; it is intended for authorized security research. This quantization has not been benchmarked. Not affiliated with Cantina Security, Z.AI or Unsloth.
  • Other formats: FP8, NVFP4.

License

MIT, inherited from the base model. Copyright (c) 2026 Z.AI Co., Ltd — see LICENSE.

Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration