license: mit
base_model: cantina-security/apex-flash-1-abliterated
base_model_relation: quantized
pipeline_tag: text-generation
tags:
- gguf
- ds4
- mixed-quantization
- abliterated
Apex Flash 1 abliterated — DS4 mixed Q2 GGUF
Text-only conversion of Cantina Security Apex Flash 1 abliterated,
revision cecb5eeb9c6b32404a0dd930df81de2c239bdd84.
Converted directly from BF16 using IQ2_XXS routed gate/up, Q2_K routed down and
higher-precision non-expert tensors, matching the stock counterpart's recipe.
Q2 names a mixed recipe, not a uniform tensor format. No activation-calibrated
importance matrix was used; IQ2 uses a weight-column-energy fallback. MTP tensors
are retained, but speculative decoding is not qualified. Vision weights are omitted.
File: apex-flash-1-abliterated-DS4-Q2.gguf, 96,505,818,944 bytes (89.88 GiB),
1,412 tensors.
Compatibility and verification
Requires a DS4 runtime supporting the GLM architecture and these mixed expert
types. GGUF alone does not establish compatibility with llama.cpp, Ollama, MLX,
LM Studio or their embedded engines. Memory needs depend on the exact runtime,
expert cache and context; file size alone is not a RAM requirement.
Full output-payload integrity verification passed. The conversion does not
inherit upstream benchmark scores. Abliteration targets refusal behavior;
it does not establish improved reasoning, reliability or capability.
Local qualification
Experimental, not a stable-agent release. Tested on 2026-10-04 using an
Apple M5 Max MacBook Pro (Mac17,6), 128 GiB RAM, macOS 27.0.1 and the pinned DS4
runtime below. Context was 8,192 tokens; expert-cache target 16 GiB, zero full
resident layers, SSD streaming, no MTP or draft model. The native allocation
plan was 27.26 GiB; this is not a measured total-RAM requirement.
| Measurement | Result |
|---|---|
| Decode, output tokens/s | Median 11.13; range 10.98–11.40, n=3 |
| Prefill latency | Median 8.222 s |
| First visible streamed fragment | Median 8.410 s |
| Bounded quality tests | 3/4 pass |
| Strict typed-tool tests | 34/36 pass |
Throughput used 535 input / 128 output tokens, thinking off, temperature 0,
top-p 1, top-k 0, seed 36 and zero cached prompt tokens. One tiny warm-up preceded
the samples; the OS file cache was not cleared. This is a sequential pilot,
not an ABBA comparison or a publisher benchmark.
Reasoning with thinking enabled, state replay and the parameterization fixture
passed. Quality grades check extracted answers/correctness, not JSON-only
compliance; state replay and parameterization appended prose despite JSON-only
requests. The remaining quality case produced invalid JSON (None rather thannull). Two array-delimiter tool failures occurred on Responses, streaming and
nonstreaming: delimiter text changed and a string was returned instead of the
required array. Chat Completions and Anthropic each passed 12/12.
These are exact-content/type tests, not security scores. Actual Prime/OpenCode
tests hit harness usage-accounting/alias/receipt issues; agent delivery and
executed coding are unqualified, not proven incompatible. The native server
shut down cleanly and artifact identities were rechecked. Requests used a
600-second socket limit; correctness grading matches the stock test.
Larger contexts, vision, speculation, refusal behavior and quality parity with
BF16 are untested. This small suite does not establish general coding ability
or superiority to the stock checkpoint.
Run on Apple Silicon
Use antirez/ds4 at4bd088c20da0b905771fb94142cf12d669103c20, plus the included runtime patch.
From this downloaded model folder, with Xcode Command Line Tools installed:
git clone https://github.com/antirez/ds4.git ds4-apex
cd ds4-apex
git checkout --detach 4bd088c20da0b905771fb94142cf12d669103c20
git apply --check ../runtime/schema-aware-glm-v2.patch
git apply ../runtime/schema-aware-glm-v2.patch
make -j4 ds4-server CC=clang
./ds4-server --metal -m ../apex-flash-1-abliterated-DS4-Q2.gguf \
--ctx 8192 --tokens 32768 --host 127.0.0.1 --port 18890 \
--ssd-streaming --ssd-streaming-cold \
--ssd-streaming-cache-experts 16GB --ssd-streaming-full-layers 0
Keep DS4_QUALIFICATION_RAW_CAPTURE undefined. Patch SHA256:2ba7e633e0be1595d82b0b8191cc73c66292fbcf046b1b663e5d18fb88bed992.
Expected patched ds4_server.c SHA256:bffea99f2bd5d122109cd64535a1b1781f8872dbbcd785420a1455d7cc6bff7f.
Binary hashes vary with SDK/compiler; reproducing source is not a new live pass.
The API is http://127.0.0.1:18890/v1. It advertises glm-5.3-flash, its
architecture alias, not the identity of the Apex fine-tune. Load only one
large model at a time, retain native memory guards, and allow long request
timeouts for reasoning. GUI integrations are not qualified by this upload.
provenance.json records pinned sources, recipe, verification scope, measured
results and the complete GGUF SHA256. Runtime code has its own license inruntime/LICENSE.
Credits and license
Cantina Security and Yeta for Apex; Z.AI for GLM-5.3-Flash.
Distributed under the upstream MIT license; see LICENSE.