← back to catalog · registered 2026-10-02 19:58

Code4me2/apex-flash-1-abliterated-NVFP4

Code4me2 multimodal second-order
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/Code4me2%2Fapex-flash-1-abliterated-NVFP4"
Response includes
  • classification unknown
  • files 25
  • author_summary 1 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Refusal direction extraction

No specific extraction method could be identified for this model. The producer either did not document it or used a proprietary pipeline.

What is a refusal direction? →
Downloads · 30-day
0
Likes
0
Model age
today
created 2026-10-02

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
mit
Tags
transformers safetensors glm5_next image-text-to-text abliterated security-research nvfp4 compressed-tensors llm-compressor quantized conversational base_model:cantina-security/apex-flash-1-abliterated

Related

Total size
191 GB
Files
25
Quantizations
1
Registered
2026-10-02 19:58
Last updated on HF
2026-10-02 19:35

Files by quantization

Auxiliary files 25 files 191 GB
model-00003-of-00010.safetensors 18.6 GB 0f01c5d9 download
model-00007-of-00010.safetensors 18.6 GB e5366961 download
model-00005-of-00010.safetensors 18.6 GB 2bcd60aa download
model-00009-of-00010.safetensors 18.6 GB 589afdb7 download
model-00004-of-00010.safetensors 18.6 GB dcf13a2b download
model-00008-of-00010.safetensors 18.6 GB 501cdf53 download
model-00006-of-00010.safetensors 18.6 GB af1be6ad download
model-00001-of-00010.safetensors 18.6 GB 21811e09 download
model-00002-of-00010.safetensors 18.6 GB df1f6d94 download
model_mtp.safetensors 13.8 GB 9c110999 download
model-00010-of-00010.safetensors 9.52 GB 6bab8adb download
quantize.log 25.0 MB 42e7c350 download
tokenizer.json 19.3 MB 19e77364 download
model.safetensors.index.json 16.2 MB 614d7875 download
config.json 48.9 KB 80ef4dfa download
chat_template.jinja 10.7 KB 06bd89e9 download
README.md 8.27 KB d7a916fe download
.gitattributes 1.64 KB 39287a7c download
LICENSE 1.04 KB 986b06fb download
processor_config.json 909 B 3ec2a058 download
tokenizer_config.json 790 B 0891a18e download
calib_stats.json 459 B 9eda0aa7 download
recipe.yaml 449 B d02f6bdb download
verify.log 239 B dec1c2d5 download
generation_config.json 177 B 439a75d0 download

README current version from Hugging Face


license: mit
base_model: cantina-security/apex-flash-1-abliterated
base_model_relation: quantized
library_name: transformers
pipeline_tag: image-text-to-text
tags:

  • glm5_next
  • abliterated
  • security-research
  • nvfp4
  • compressed-tensors
  • llm-compressor
  • quantized

apex-flash-1-abliterated-NVFP4

Preliminary model card. Quality evaluation is pending (see Evaluation).

This is an NVFP4 (W4A4) quantization of cantina-security/apex-flash-1-abliterated, in the compressed-tensors format. The checkpoint is about 205 GB, compared with about 643 GB for the BF16 source.

Base model

apex-flash-1-abliterated is an experimental derivative of apex-flash-1, Cantina Security's open-weights security model developed in partnership with Yeta (@yetalabs on X). apex-flash-1 is a reinforcement learning post-train of GLM-5.3-Flash (zai-org/GLM-5.3-Flash) for focused investigations: reading code, using tools, pursuing an exploit, and verifying its effect in a running target. The abliterated variant modifies refusal behavior broadly; the changes are not limited to security tasks.

Further reading from the base authors: Apex Flash release post, Explore Apex.

All credit for the model itself belongs to Cantina Security and Yeta. This repository only contains a quantized copy produced by a third party; it is not affiliated with or endorsed by them.

Intended use

As stated on the base card: authorized security research in environments the researcher owns or has permission to test.

Architecture

Property Value
Architecture GLM-5.3-Flash (glm5_next, Glm5NextForConditionalGeneration)
Parameters ~314B total
Layers 45 decoder layers + 1 MTP layer (layer 45)
Experts 288 routed (top-8) + 1 shared
Attention Hybrid: 34 KDA linear-attention layers + 11 MLA/DSA sparse-attention layers (with lightning indexer)
Other Manifold-constrained hyper-connections (mHC); 24-block ViT vision tower
Max positions 1,048,576
BF16 source size ~643 GB

Quantization

Item Value
Method One-shot PTQ, QuantizationModifier(scheme="NVFP4") via llm-compressor oneshot
Format nvfp4-pack-quantized (compressed-tensors)
Tooling llm-compressor 0.14.0, compressed-tensors 0.19.0, transformers 5.17.0, torch 2.14.0
Hardware 8x NVIDIA A100-SXM4-80GB (Lambda), torchrun DDP
Calibration moe_calibrate_all_experts=True (every expert sees every calibration token)
Checkpoint size ~205 GB

NVFP4 is W4A4. Weights are FP4 (E2M1) with group size 16, FP8 (E4M3) block scales and an FP32 per-tensor global scale. Activations use dynamic local FP8 block scales per 16 elements plus a static per-tensor input_global_scale calibrated from data.

What is quantized

Only the routed-expert gate/up/down projections in layers 3-44 are quantized (36,288 projections, the vast majority of parameters). Everything else stays BF16.

Component Precision
Routed experts, layers 3-44 (gate/up/down) NVFP4
Vision tower, embeddings, lm_head BF16
All attention (KDA, MLA/DSA, indexer) BF16
MoE router mlp.gate and e_score_correction_bias BF16
Shared experts; dense MLPs of layers 0-2 BF16
Hyper-connection parameters BF16
MTP layer 45 BF16 (spliced from the source; llm-compressor 0.14.0 silently drops it for this architecture)

Calibration data

512 samples, 501,470 tokens, every sample rendered through the model's own chat template.

Domain Source Target share Samples Tokens
General chat mlabonne/open-perfectblend 30% 154 98,634
Cybersecurity instruction Trendyol/Trendyol-Cybersecurity-Instruction-Tuning-Dataset 20% 102 69,059
Vulnerable/fixed code review CyberNative/Code_Vulnerability_Security_DPO 15% 77 27,616
Tool calling (re-rendered in GLM native <tool_call>/<arg_key> format) NousResearch/hermes-function-calling-v1 15% 77 101,972
Long <think> code reasoning nvidia/OpenCodeReasoning 20% 102 204,189
Total 512 501,470

Context length used for calibration: at most 2048 tokens per sample (mean ~980). Calibration did not use long contexts. The impact is expected to be limited but is untested: weights are quantized independently of context length; activation block scales are computed dynamically at runtime; only the per-tensor activation global scale is static, and it was calibrated on sequences of 2048 tokens or fewer; attention and the KV path are BF16 and not quantized here. Long-context (>4K) quality has not been evaluated.

Fixes relative to the upstream llm-compressor GLM-5.3-Flash example

  1. MTP layer dropped: llm-compressor 0.14.0 only copies MTP tensors for configs with num_mtp_layers/mtp_num_hidden_layers and an mtp. prefix, whereas this model uses num_nextn_predict_layers and layers.45.*. The layer is spliced back from the source (splice_mtp.py).
  2. Target regex: .*mlp\.experts\..* also matches layer 45, which would make loaders expect NVFP4 weights there. Targets are restricted to layers 3-44 with an explicit ignore for layer 45.
  3. Missing dependencies in the upstream environment: torchvision and flash-linear-attention.

Structural verification

Result of verify_ckpt.py against the BF16 source: VERIFY OK. This confirms checkpoint structure only, not model quality.

Check Result
Tensors 147,634
Dtypes bf16: 2,269; fp32: 72,789; uint8: 36,288; float8_e4m3fn: 36,288
Quantized expert projections 36,288
Scales checked 108,954 (all finite; global scales > 0)
MTP tensors 889
Size 205.1 GB

Evaluation

Pending. A BF16-vs-NVFP4 parity evaluation is in progress: a held-out set disjoint from the calibration data across the same five domains, plus 4096-token long samples and WikiText-2, measuring top-1 agreement, approximate KL divergence and perplexity. Results will be added to this card. No parity, accuracy-retention or benchmark claims are made at this time. The base model's reported results apply to the original apex-flash-1 BF16 checkpoint only.

Usage (untested on this checkpoint)

The compressed-tensors NVFP4 format is auto-detected by vLLM:

vllm serve Code4me2/apex-flash-1-abliterated-NVFP4 --tensor-parallel-size 2

MTP speculative decoding is available via:

--speculative-config '{"method":"mtp","num_speculative_tokens":1}'

Native NVFP4 kernels require Blackwell (SM100/SM12x). See the SM12x caveat below.

Known gaps and limitations

  • Calibration length: limited to 2048 tokens per sample; long-context behavior is unevaluated.
  • No multimodal calibration: no image or video data was used. The vision tower is BF16, but expert activations on image tokens were not calibrated. Image/video performance is unevaluated, as in the base.
  • English-centric calibration: GLM is bilingual (zh/en); the calibration set is not.
  • Plain round-to-nearest quantization with min-max observers: no GPTQ/AWQ-style error compensation and no MSE clipping search.
  • All-expert calibration: moe_calibrate_all_experts exposes each expert to tokens it would not normally receive, which can make activation global scales conservative.
  • SM120/SM121 serving: stock vLLM 0.30 currently fails for all GLM-5.3-Flash checkpoints on SM120/SM121 (RTX PRO 6000, DGX Spark/GB10) with pe_dim must be 64 for fp8_ds_mla (vllm-project/vllm#55773, #53963). Community patches exist (Libertai/vllm-sparse-mla-blackwell, tonyd2wild/DGX-Spark).
  • Inherited limitations: all limitations of the abliterated base apply. Refusal behavior is modified broadly, the variant was not separately evaluated, and image/video performance was not evaluated.

Reproducibility

The quantization recipe is in recipe.yaml. The logs quantize.log and verify.log and the calibration statistics calib_stats.json are included in this repository.

Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration