← back to catalog · registered 2026-10-04 19:58

shamidk/apex-flash-1-abliterated-DS4-Q2-GGUF

shamidk GGUF second-order
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/shamidk%2Fapex-flash-1-abliterated-DS4-Q2-GGUF"
Response includes
  • classification unknown
  • files 5
  • author_summary 1 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
?
Primary method

Unclassified

No clear signals of an abliteration technique in this model.
Confidence
UNKNOWN
Why this label 1 signal
No classification signals present. This may not be an abliterated model at all - it could be a repackaging, a merge with unrelated goals, or unrelated content that mentions the term.
  • no classification signals present (no abliterated, uncensored, or known producer/method markers)
Refusal direction extraction

No specific extraction method could be identified for this model. The producer either did not document it or used a proprietary pipeline.

What is a refusal direction? →
Downloads · 30-day
0
Likes
1
Model age
today
created 2026-10-04

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
mit
Tags
gguf ds4 mixed-quantization abliterated text-generation base_model:cantina-security/apex-flash-1-abliterated base_model:quantized:cantina-security/apex-flash-1-abliterated license:mit endpoints_compatible region:us conversational

Related

Total size
89.9 GB
Files
5
Quantizations
1
Registered
2026-10-04 19:58
Last updated on HF
2026-10-04 19:35

Files by quantization

Auxiliary files 5 files 89.9 GB
apex-flash-1-abliterated-DS4-Q2.gguf 89.9 GB 2e5d4c7b download
provenance.json 17.9 KB 2a7a0b86 download
README.md 5.17 KB 12e23745 download
.gitattributes 1.55 KB 877f12ef download
LICENSE 1.04 KB 986b06fb download

README current version from Hugging Face


license: mit
base_model: cantina-security/apex-flash-1-abliterated
base_model_relation: quantized
pipeline_tag: text-generation
tags:

  • gguf
  • ds4
  • mixed-quantization
  • abliterated

Apex Flash 1 abliterated — DS4 mixed Q2 GGUF

Text-only conversion of Cantina Security Apex Flash 1 abliterated,
revision cecb5eeb9c6b32404a0dd930df81de2c239bdd84.

Converted directly from BF16 using IQ2_XXS routed gate/up, Q2_K routed down and
higher-precision non-expert tensors, matching the stock counterpart's recipe.
Q2 names a mixed recipe, not a uniform tensor format. No activation-calibrated
importance matrix was used; IQ2 uses a weight-column-energy fallback. MTP tensors
are retained, but speculative decoding is not qualified. Vision weights are omitted.

File: apex-flash-1-abliterated-DS4-Q2.gguf, 96,505,818,944 bytes (89.88 GiB),
1,412 tensors.

Compatibility and verification

Requires a DS4 runtime supporting the GLM architecture and these mixed expert
types. GGUF alone does not establish compatibility with llama.cpp, Ollama, MLX,
LM Studio or their embedded engines. Memory needs depend on the exact runtime,
expert cache and context; file size alone is not a RAM requirement.

Full output-payload integrity verification passed. The conversion does not
inherit upstream benchmark scores. Abliteration targets refusal behavior;
it does not establish improved reasoning, reliability or capability.

Local qualification

Experimental, not a stable-agent release. Tested on 2026-10-04 using an
Apple M5 Max MacBook Pro (Mac17,6), 128 GiB RAM, macOS 27.0.1 and the pinned DS4
runtime below. Context was 8,192 tokens; expert-cache target 16 GiB, zero full
resident layers, SSD streaming, no MTP or draft model. The native allocation
plan was 27.26 GiB; this is not a measured total-RAM requirement.

Measurement Result
Decode, output tokens/s Median 11.13; range 10.98–11.40, n=3
Prefill latency Median 8.222 s
First visible streamed fragment Median 8.410 s
Bounded quality tests 3/4 pass
Strict typed-tool tests 34/36 pass

Throughput used 535 input / 128 output tokens, thinking off, temperature 0,
top-p 1, top-k 0, seed 36 and zero cached prompt tokens. One tiny warm-up preceded
the samples; the OS file cache was not cleared. This is a sequential pilot,
not an ABBA comparison or a publisher benchmark.

Reasoning with thinking enabled, state replay and the parameterization fixture
passed. Quality grades check extracted answers/correctness, not JSON-only
compliance; state replay and parameterization appended prose despite JSON-only
requests. The remaining quality case produced invalid JSON (None rather than
null). Two array-delimiter tool failures occurred on Responses, streaming and
nonstreaming: delimiter text changed and a string was returned instead of the
required array. Chat Completions and Anthropic each passed 12/12.
These are exact-content/type tests, not security scores. Actual Prime/OpenCode
tests hit harness usage-accounting/alias/receipt issues; agent delivery and
executed coding are unqualified, not proven incompatible. The native server
shut down cleanly and artifact identities were rechecked. Requests used a
600-second socket limit; correctness grading matches the stock test.

Larger contexts, vision, speculation, refusal behavior and quality parity with
BF16 are untested. This small suite does not establish general coding ability
or superiority to the stock checkpoint.

Run on Apple Silicon

Use antirez/ds4 at
4bd088c20da0b905771fb94142cf12d669103c20, plus the included runtime patch.
From this downloaded model folder, with Xcode Command Line Tools installed:

git clone https://github.com/antirez/ds4.git ds4-apex
cd ds4-apex
git checkout --detach 4bd088c20da0b905771fb94142cf12d669103c20
git apply --check ../runtime/schema-aware-glm-v2.patch
git apply ../runtime/schema-aware-glm-v2.patch
make -j4 ds4-server CC=clang
./ds4-server --metal -m ../apex-flash-1-abliterated-DS4-Q2.gguf \
  --ctx 8192 --tokens 32768 --host 127.0.0.1 --port 18890 \
  --ssd-streaming --ssd-streaming-cold \
  --ssd-streaming-cache-experts 16GB --ssd-streaming-full-layers 0

Keep DS4_QUALIFICATION_RAW_CAPTURE undefined. Patch SHA256:
2ba7e633e0be1595d82b0b8191cc73c66292fbcf046b1b663e5d18fb88bed992.
Expected patched ds4_server.c SHA256:
bffea99f2bd5d122109cd64535a1b1781f8872dbbcd785420a1455d7cc6bff7f.
Binary hashes vary with SDK/compiler; reproducing source is not a new live pass.

The API is http://127.0.0.1:18890/v1. It advertises glm-5.3-flash, its
architecture alias, not the identity of the Apex fine-tune. Load only one
large model at a time, retain native memory guards, and allow long request
timeouts for reasoning. GUI integrations are not qualified by this upload.

provenance.json records pinned sources, recipe, verification scope, measured
results and the complete GGUF SHA256. Runtime code has its own license in
runtime/LICENSE.

Credits and license

Cantina Security and Yeta for Apex; Z.AI for GLM-5.3-Flash.
Distributed under the upstream MIT license; see LICENSE.

Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration