← back to catalog · registered 2026-08-22 13:56

apetersson/DeepSeek-V4-Flash-0731-Abliterated-DS4-Quality128

apetersson Deepseek GGUF MoE second-order 1.0M ctx
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/apetersson%2FDeepSeek-V4-Flash-0731-Abliterated-DS4-Quality128"
Response includes
  • classification m8
  • files 11
  • hub_downloads_all_time 36,700
  • author_summary 8 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M8
Primary method

Repackaging (quantization)

Applied on top of direct removal inherited from the base model.
Confidence
MEDIUM
Inherited from base model
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=1
  • assume M1 (base ablation) + M8 (GGUF quant) - default when producer unknown
Refusal direction extraction

No specific extraction method could be identified for this model. The producer either did not document it or used a proprietary pipeline.

What is a refusal direction? →
Downloads · lifetime
37K
1K last 30d - cooling
Likes
7
Model age
2mo ago
created 2026-08-01
Downloads over time
Now37K→from15.1K↑144%
14.1K22.4K30.8K39.2K15.1K on Aug 537K on Oct 11AugSepOct
Aug 5 → Oct 11 · 50 snapshots · spans 67 days

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
mit
Tags
gguf quantized deepseek deepseek-v4 deepseek-v4-flash moe mixture-of-experts abliterated mxfp4 iq2_xxs q2_k ds4

Related

Total size
109 GB
Files
11
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-08-04 19:03

Files by quantization

Auxiliary files 11 files 109 GB
DeepSeek-V4-Flash-0731-Abliterated-DS4-Quality128.gguf 95.8 GB 2cfc36b7 download
DeepSeek-V4-Flash-0731-Abliterated-DS4-Quality128-llamacpp-DSpark-support.gguf 6.80 GB 0582de4d download
DeepSeek-V4-Flash-0731-Abliterated-DS4-Quality128-DSpark-support.gguf 6.80 GB cd8593a2 download
BUILD_MANIFEST.json 35.0 KB b5fad153 download
repack_llamacpp_dspark.py 16.0 KB f12723a4 download
README.md 11.2 KB 05d1d7b9 download
ds4-upstream-issues.md 9.65 KB b42ab58c download
BUILD_PLAN.md 2.92 KB 00b7f5e4 download
.gitattributes 1.79 KB f3c1860b download
PROVENANCE.md 1.40 KB 9623eca9 download
SHA256SUMS 905 B f3c823ab download

README current version from Hugging Face


license: mit
library_name: gguf
pipeline_tag: text-generation
base_model: apetersson/DeepSeek-V4-Flash-0731-Abliterated-FP8
base_model_relation: quantized
quantized_by: apetersson
tags:

  • gguf
  • quantized
  • deepseek
  • deepseek-v4
  • deepseek-v4-flash
  • moe
  • mixture-of-experts
  • abliterated
  • mxfp4
  • iq2_xxs
  • q2_k
  • ds4
  • dspark
  • apple-silicon
  • metal

DeepSeek V4 Flash Abliterated — DS4 Quality128

Exact MXFP4 experts, maximum resident quality.

Runtime compatibility: the target GGUF runs with both DS4 and
llama.cpp on Metal. For DSpark speculative decoding, use the companion
whose filename identifies the runtime: DSpark-support for DS4 or
llamacpp-DSpark-support for a llama.cpp build with DeepSeek V4 DSpark
support.

This is a quality-first GGUF package built with the DS4 Quality128 quantization
policy. It is designed to keep DeepSeek V4 Flash resident on a 128 GB M1 Ultra
while preserving the most sensitive routed experts in their exact native
MXFP4 representation. DS4 in the package name identifies the quantization
profile. Use the same target GGUF with either runtime; only the optional DSpark
companions are runtime-specific.

The model is quantized directly from the abliterated FP8 checkpoint
apetersson/DeepSeek-V4-Flash-0731-Abliterated-FP8.
The abliteration affects only 36 attention wo_b tensors; routed-expert codes
and scales are unchanged from that checkpoint.

Artifacts

File Bytes GiB SHA-256
DeepSeek-V4-Flash-0731-Abliterated-DS4-Quality128.gguf 102,826,238,912 95.7644 2cfc36b761b59ea43531e7cdb02a690436a330e42ad57cb162726b385914df59
DeepSeek-V4-Flash-0731-Abliterated-DS4-Quality128-DSpark-support.gguf 7,297,737,120 6.7965 cd8593a232c9feebc4c91855d5ab486b17250fc8bc2f294bc80401f93b371566
DeepSeek-V4-Flash-0731-Abliterated-DS4-Quality128-llamacpp-DSpark-support.gguf 7,302,984,320 6.8014 0582de4d4f63c524651c61f06a5c9e3fc9deab697b44e5374cce409ed9181a92
Target + DS4 companion 110,123,976,032 102.5609 —
Target + llama.cpp companion 110,129,223,232 102.5658 —

The companions are runtime-specific alternatives. For DSpark, load only the
companion matching the runtime. Keeping all three GGUFs on disk uses
117,426,960,352 bytes (109.3624 GiB).

Quantization profile

  • Exact native MXFP4 gate/up/down routed experts on layers
    10, 14, 30, 34, 37, 38, 39, 40, 41, 42.
  • IQ2_XXS gate/up and Q2_K down routed experts on the other 33 MoE layers.
  • Q8 attention, shared-expert and output paths.
  • F16 protected indexer and auxiliary tensors.
  • The routed-expert imatrix applies only to genuinely requantized IQ2/Q2
    tensors; preserved MXFP4 needs no imatrix.
  • DSpark support uses IQ2_XXS gate/up and exact native MXFP4 down projections
    for target layers 40, 41 and 42.

Target GGUF type histogram (1,328 tensors): F32 492, F16 359, I32 3, Q8_0 345,
IQ2_XXS 66, Q2_K 33 and MXFP4 30. DSpark support histogram (81 tensors): F32
34, F16 7, Q8_0 31, IQ2_XXS 6 and MXFP4 3.

Runtime compatibility

The target GGUF is shared by both runtimes. The two companions use different
DSpark schemas and are not interchangeable:

Artifact DS4 stock llama.cpp b10210 llama.cpp fffbcbdb
Target GGUF Validated on Metal Validated on Metal Validated on Metal
DS4 companion Validated on Metal Not compatible Not compatible
llama.cpp companion Not compatible Not compatible Validated on Metal

Use the target GGUF directly for target-only inference in either runtime. For
DSpark speculative decoding, pair it with the companion listed for that
runtime. The DS4 companion requires
a recent ds4 version from the main branch,
with its DSpark generation support tracked in
antirez/ds4#642.

  • Quantizer SHA-256: f0a381f4ada808ea2afa740d964354fa327fc1235ba7cebf50874eb89fb97ac5
  • Runtime SHA-256: 2aaf20469b9918d6d6ab8787a02811c11228547cd979787879a03dba8a9e7824

The DS4 companion requires the recent ds4 main-branch version above. Stock
llama.cpp b10210 supports target-only inference but cannot use either
companion. llama.cpp DSpark requires the dflash companion and a build that
contains the DeepSeek V4 DSpark changes described below.

Validated llama.cpp configurations

Target-only with b10210

Target-only validation used a 128 GB M1 Ultra and stock Homebrew llama.cpp
build b10210 (000547513, 2026-07-31). The configuration used full Metal
offload, a 4,096-token context, 256-token batch and ubatch, Flash Attention, no
warmup and greedy decoding. A deterministic arithmetic probe returned the
correct answer, with preliminary measurements of 15.8 prompt tokens/s and 8.0
generation tokens/s. These figures are a single short smoke test, not a
sustained benchmark. This build supports target-only inference for this
package; neither companion is compatible with it.

DSpark with fffbcbdb

llama.cpp DSpark requires DeepSeek V4 MTP/DSpark, separate DSpark conversion,
sidecar discovery and Metal hyper-connection support. Validation used upstream
commit fffbcbdb9d5e56105a8842867a59bb9736520ca8 from 2026-08-02. Relevant
changes include
#25784,
#26458 and
#26459.

The llama.cpp sidecar contract uses architecture dflash, dflash.* metadata,
target tokenizer/model metadata and standardized tensor names such as blk.*,
markov_w1.weight and conf_proj.weight. The DS4 companion uses the DS4-native
deepseek4-dspark, dspark.* and mtp.* layout.

The llamacpp-DSpark-support.gguf companion is a container-only repack of the
DS4 companion. Its 81 tensor descriptors use llama.cpp names, and its header
contains the target tokenizer and standardized dflash metadata. It preserves
the complete 7,297,731,680-byte tensor-data region byte-for-byte without
requantization. The source and output payload SHA-256 is:

befbdb4a0f7e6b2626cfa9af0bc1560261a7a3235e0e6670b4f0c8df032a2e74

repack_llamacpp_dspark.py reproduces the conversion (SHA-256
9ffa5aedd0fa83b74846ff164aabd82097ee60e4610231493fe1ae5a702efb99)
without changing model weights or the DSpark quantization policy.

The validation configuration used a 128 GB M1 Ultra, full Metal offload, a
4,096-token context, 128-token batch and ubatch, Flash Attention, no warmup,
greedy decoding and a maximum DSpark block of five tokens. A short arithmetic
probe returned 42. A longer deterministic sequence produced the integers 1
through 40 correctly and reported:

Measurement Result
Generation throughput 27.0 tokens/s
Draft tokens generated 70
Draft tokens accepted 65
Draft acceptance 92.9%
Metal model allocation 98,057 MiB
Metal context allocation 111 MiB
Metal compute allocation 208 MiB
Free Metal working budget after allocation 3,467 MiB

These figures are a short correctness and compatibility probe, not a sustained
or cross-runtime benchmark. The tested build emits a nonfatal warning that the
Lightning Indexer for layer 2 is assigned to CPU and disabled; the warning does
not prevent DSpark from loading, drafting or accepting tokens.

Launch examples

llama.cpp b10210 target-only inference:

llama-cli \
  --model DeepSeek-V4-Flash-0731-Abliterated-DS4-Quality128.gguf \
  --ctx-size 4096 \
  --batch-size 256 --ubatch-size 256 \
  --gpu-layers 999 --flash-attn on \
  --reasoning off --reasoning-budget 0

DS4 target-only inference:

ds4 --metal \
  -m DeepSeek-V4-Flash-0731-Abliterated-DS4-Quality128.gguf

DS4 inference with its DSpark companion:

ds4 --metal \
  -m DeepSeek-V4-Flash-0731-Abliterated-DS4-Quality128.gguf \
  --mtp DeepSeek-V4-Flash-0731-Abliterated-DS4-Quality128-DSpark-support.gguf \
  --dspark --temp 0

llama.cpp DSpark inference with commit fffbcbdb:

llama-cli \
  --model DeepSeek-V4-Flash-0731-Abliterated-DS4-Quality128.gguf \
  --spec-type draft-dspark \
  --spec-draft-model \
    DeepSeek-V4-Flash-0731-Abliterated-DS4-Quality128-llamacpp-DSpark-support.gguf \
  --spec-draft-n-max 5 \
  --ctx-size 4096 \
  --batch-size 128 --ubatch-size 128 \
  --gpu-layers 999 --spec-draft-ngl 999 \
  --flash-attn on \
  --temp 0 --reasoning off --reasoning-budget 0

This command requires commit fffbcbdb or another build containing the
2026-08-02 DeepSeek V4 DSpark changes. Homebrew b10210 supports only the
target-only command above.

DS4 one-million-token context

The following residency estimates and --prefill-chunk recommendations apply
only to DS4; one-million-token context is unvalidated with llama.cpp.
--ctx 1048576 counts prompt and completion together. The safest
maximum-quality resident mode omits DSpark and uses a 2,048-token prefill
chunk:

ds4 --metal \
  -m DeepSeek-V4-Flash-0731-Abliterated-DS4-Quality128.gguf \
  --ctx 1048576 --prefill-chunk 2048

Estimated residency is at most 110.30 GiB, leaving at least 11.30 GiB
below Metal's approximately 121.60 GiB recommended working set. Omitting
DSpark does not reduce target-model quality; it only forgoes speculative decode.

Enable DSpark with the smaller chunk only after confirming peak memory on the
host:

ds4 --metal \
  -m DeepSeek-V4-Flash-0731-Abliterated-DS4-Quality128.gguf \
  --ctx 1048576 --prefill-chunk 1024 \
  --mtp DeepSeek-V4-Flash-0731-Abliterated-DS4-Quality128-DSpark-support.gguf \
  --dspark --temp 0

That mode is estimated at 114.02 GiB, about 7.58 GiB below the recommended
working set. DSpark with chunk 4096 is estimated at 123.24 GiB and is not a
reliably resident configuration.

Verification and provenance

The package passed strict target and DSpark planning, exact size/type/name-set
checks, strict imatrix coverage, source validation, and byte reproduction for
all 30 target plus three DSpark MXFP4 tensors. The llama.cpp companion passed its
81-tensor name contract, metadata contract, full payload size check,
byte-identical payload SHA-256 check and a real target-plus-draft inference
test. SHA256SUMS binds all three GGUFs, the repack script and the
documentation. See:

The sibling MLX package
DeepSeek-V4-Flash-0731-Abliterated-MLX-Quality128-suboptimal is available for
speed comparison; it is not the canonical quality artifact.

Treat the native MXFP4 Metal kernels and mixed DSpark path as experimental
relative to Q4_K. Benchmark correctness and throughput against DS4 v1 and the
MLX comparator before selecting an everyday launch configuration.

README history 6 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-08-04docs: point ds4 requirement to mainc0f4f5911.2 KB
    Loading...
  2. 2026-08-02Update README.md47c09c611.3 KB
    Loading...
  3. 2026-08-02Add files using upload-large-folder tool123954f11.3 KB
    Loading...
  4. 2026-08-01Update README.mdd9886665.5 KB
    Loading...
  5. 2026-08-01Add Hugging Face model card metadataf9051ba5.5 KB
    Loading...
  6. 2026-08-01Add files using upload-large-folder tool26e13235.2 KB
    Loading...

Discussions 1 thread

  1. 2026-08-02Vulkan mainstream llama.cpp ?open2 💬#1
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration