license: mit
base_model: cantina-security/apex-flash-1-abliterated
base_model_relation: quantized
library_name: transformers
pipeline_tag: image-text-to-text
tags:
- glm5_next
- abliterated
- security-research
- nvfp4
- compressed-tensors
- llm-compressor
- quantized
apex-flash-1-abliterated-NVFP4
Preliminary model card. Quality evaluation is pending (see Evaluation).
This is an NVFP4 (W4A4) quantization of cantina-security/apex-flash-1-abliterated, in the compressed-tensors format. The checkpoint is about 205 GB, compared with about 643 GB for the BF16 source.
Base model
apex-flash-1-abliterated is an experimental derivative of apex-flash-1, Cantina Security's open-weights security model developed in partnership with Yeta (@yetalabs on X). apex-flash-1 is a reinforcement learning post-train of GLM-5.3-Flash (zai-org/GLM-5.3-Flash) for focused investigations: reading code, using tools, pursuing an exploit, and verifying its effect in a running target. The abliterated variant modifies refusal behavior broadly; the changes are not limited to security tasks.
Further reading from the base authors: Apex Flash release post, Explore Apex.
All credit for the model itself belongs to Cantina Security and Yeta. This repository only contains a quantized copy produced by a third party; it is not affiliated with or endorsed by them.
Intended use
As stated on the base card: authorized security research in environments the researcher owns or has permission to test.
Architecture
| Property | Value |
|---|---|
| Architecture | GLM-5.3-Flash (glm5_next, Glm5NextForConditionalGeneration) |
| Parameters | ~314B total |
| Layers | 45 decoder layers + 1 MTP layer (layer 45) |
| Experts | 288 routed (top-8) + 1 shared |
| Attention | Hybrid: 34 KDA linear-attention layers + 11 MLA/DSA sparse-attention layers (with lightning indexer) |
| Other | Manifold-constrained hyper-connections (mHC); 24-block ViT vision tower |
| Max positions | 1,048,576 |
| BF16 source size | ~643 GB |
Quantization
| Item | Value |
|---|---|
| Method | One-shot PTQ, QuantizationModifier(scheme="NVFP4") via llm-compressor oneshot |
| Format | nvfp4-pack-quantized (compressed-tensors) |
| Tooling | llm-compressor 0.14.0, compressed-tensors 0.19.0, transformers 5.17.0, torch 2.14.0 |
| Hardware | 8x NVIDIA A100-SXM4-80GB (Lambda), torchrun DDP |
| Calibration | moe_calibrate_all_experts=True (every expert sees every calibration token) |
| Checkpoint size | ~205 GB |
NVFP4 is W4A4. Weights are FP4 (E2M1) with group size 16, FP8 (E4M3) block scales and an FP32 per-tensor global scale. Activations use dynamic local FP8 block scales per 16 elements plus a static per-tensor input_global_scale calibrated from data.
What is quantized
Only the routed-expert gate/up/down projections in layers 3-44 are quantized (36,288 projections, the vast majority of parameters). Everything else stays BF16.
| Component | Precision |
|---|---|
| Routed experts, layers 3-44 (gate/up/down) | NVFP4 |
Vision tower, embeddings, lm_head |
BF16 |
| All attention (KDA, MLA/DSA, indexer) | BF16 |
MoE router mlp.gate and e_score_correction_bias |
BF16 |
| Shared experts; dense MLPs of layers 0-2 | BF16 |
| Hyper-connection parameters | BF16 |
| MTP layer 45 | BF16 (spliced from the source; llm-compressor 0.14.0 silently drops it for this architecture) |
Calibration data
512 samples, 501,470 tokens, every sample rendered through the model's own chat template.
| Domain | Source | Target share | Samples | Tokens |
|---|---|---|---|---|
| General chat | mlabonne/open-perfectblend | 30% | 154 | 98,634 |
| Cybersecurity instruction | Trendyol/Trendyol-Cybersecurity-Instruction-Tuning-Dataset | 20% | 102 | 69,059 |
| Vulnerable/fixed code review | CyberNative/Code_Vulnerability_Security_DPO | 15% | 77 | 27,616 |
Tool calling (re-rendered in GLM native <tool_call>/<arg_key> format) |
NousResearch/hermes-function-calling-v1 | 15% | 77 | 101,972 |
Long <think> code reasoning |
nvidia/OpenCodeReasoning | 20% | 102 | 204,189 |
| Total | 512 | 501,470 |
Context length used for calibration: at most 2048 tokens per sample (mean ~980). Calibration did not use long contexts. The impact is expected to be limited but is untested: weights are quantized independently of context length; activation block scales are computed dynamically at runtime; only the per-tensor activation global scale is static, and it was calibrated on sequences of 2048 tokens or fewer; attention and the KV path are BF16 and not quantized here. Long-context (>4K) quality has not been evaluated.
Fixes relative to the upstream llm-compressor GLM-5.3-Flash example
- MTP layer dropped: llm-compressor 0.14.0 only copies MTP tensors for configs with
num_mtp_layers/mtp_num_hidden_layersand anmtp.prefix, whereas this model usesnum_nextn_predict_layersandlayers.45.*. The layer is spliced back from the source (splice_mtp.py). - Target regex:
.*mlp\.experts\..*also matches layer 45, which would make loaders expect NVFP4 weights there. Targets are restricted to layers 3-44 with an explicit ignore for layer 45. - Missing dependencies in the upstream environment:
torchvisionandflash-linear-attention.
Structural verification
Result of verify_ckpt.py against the BF16 source: VERIFY OK. This confirms checkpoint structure only, not model quality.
| Check | Result |
|---|---|
| Tensors | 147,634 |
| Dtypes | bf16: 2,269; fp32: 72,789; uint8: 36,288; float8_e4m3fn: 36,288 |
| Quantized expert projections | 36,288 |
| Scales checked | 108,954 (all finite; global scales > 0) |
| MTP tensors | 889 |
| Size | 205.1 GB |
Evaluation
Pending. A BF16-vs-NVFP4 parity evaluation is in progress: a held-out set disjoint from the calibration data across the same five domains, plus 4096-token long samples and WikiText-2, measuring top-1 agreement, approximate KL divergence and perplexity. Results will be added to this card. No parity, accuracy-retention or benchmark claims are made at this time. The base model's reported results apply to the original apex-flash-1 BF16 checkpoint only.
Usage (untested on this checkpoint)
The compressed-tensors NVFP4 format is auto-detected by vLLM:
vllm serve Code4me2/apex-flash-1-abliterated-NVFP4 --tensor-parallel-size 2
MTP speculative decoding is available via:
--speculative-config '{"method":"mtp","num_speculative_tokens":1}'
Native NVFP4 kernels require Blackwell (SM100/SM12x). See the SM12x caveat below.
Known gaps and limitations
- Calibration length: limited to 2048 tokens per sample; long-context behavior is unevaluated.
- No multimodal calibration: no image or video data was used. The vision tower is BF16, but expert activations on image tokens were not calibrated. Image/video performance is unevaluated, as in the base.
- English-centric calibration: GLM is bilingual (zh/en); the calibration set is not.
- Plain round-to-nearest quantization with min-max observers: no GPTQ/AWQ-style error compensation and no MSE clipping search.
- All-expert calibration:
moe_calibrate_all_expertsexposes each expert to tokens it would not normally receive, which can make activation global scales conservative. - SM120/SM121 serving: stock vLLM 0.30 currently fails for all GLM-5.3-Flash checkpoints on SM120/SM121 (RTX PRO 6000, DGX Spark/GB10) with
pe_dim must be 64 for fp8_ds_mla(vllm-project/vllm#55773, #53963). Community patches exist (Libertai/vllm-sparse-mla-blackwell, tonyd2wild/DGX-Spark). - Inherited limitations: all limitations of the abliterated base apply. Refusal behavior is modified broadly, the variant was not separately evaluated, and image/video performance was not evaluated.
Reproducibility
The quantization recipe is in recipe.yaml. The logs quantize.log and verify.log and the calibration statistics calib_stats.json are included in this repository.