← back to catalog · registered 2026-08-22 13:56

neko-legends/DeepSeek-V4-Flash-0731-Abliterated-NVFP4

neko-legends Deepseek second-order
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/neko-legends%2FDeepSeek-V4-Flash-0731-Abliterated-NVFP4"
Response includes
  • classification m1
  • files 8
  • hub_downloads_all_time 1,928
  • author_summary 4 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M1
Primary method

Direct removal

No other method signals detected in this model.
Confidence
MEDIUM
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=0 (base model)
  • no specific method indicators - defaulting to M1 (most common)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
MEDIUM
Why we say so
primary_method=M1; difference-of-means is the reference extraction for M1/M3 (Arditi 2024)
Downloads · lifetime
2K
2 last 30d - cooling
Likes
5
Model age
2mo ago
created 2026-08-03
Downloads over time
Now1.9K→from1.8K↑4%
1.8K1.9K1.9K1.9K1.8K on Aug 51.9K on Oct 111.9K on Sep 23AugSepOct
Aug 5 → Oct 11 · 50 snapshots · spans 67 days

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
mit
Tags
deepseek-v4 nvfp4 fp8 modelopt quantized dgx-spark text-generation base_model:apetersson/DeepSeek-V4-Flash-0731-Abliterated-FP8 base_model:finetune:apetersson/DeepSeek-V4-Flash-0731-Abliterated-FP8 license:mit region:us

Related

Total size
0 B
Files
8
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-08-16 08:01

Files by quantization

Auxiliary files 8 files 342 KB
ledger-2026-08-16.png 105 KB ca20c95b download
c1-decode-journey-2026-08-16.png 100 KB 8814a3dd download
decode-at-depth-2026-08-15.png 69.1 KB 123bec50 download
c4-aggregate-2026-08-15.png 59.5 KB 6368c747 download
README.md 4.72 KB 841ef483 download
.gitattributes 1.67 KB 7488715a download
LICENSE 1.06 KB d62e3bef download
NOTICE 793 B 71acc765 download

README current version from Hugging Face


license: mit
base_model:

  • apetersson/DeepSeek-V4-Flash-0731-Abliterated-FP8
    pipeline_tag: text-generation
    tags:
  • deepseek-v4
  • nvfp4
  • fp8
  • modelopt
  • quantized
  • dgx-spark

DeepSeek V4 Flash 0731 Abliterated NVFP4 — archived conversion

This repo is now an informational page. The checkpoint weights that used
to live here have been removed (2026-08-16): the conversion is superseded
and there is nothing to download.

In plain English: this was a repackaging of DeepSeek V4 Flash (0731,
abliterated) with the routed-expert weights in NVIDIA NVFP4 format. Since
then, faster and better-supported paths won out — so instead of these
weights, use:


What four DGX Sparks do with the 32-32 checkpoint today

Single-stream decode (client wall, 2048-token completions, formal protocol):

C1 decode: TP2 baseline, broken boot, old record, and now

The full ledger — decode, prefill, and concurrency:

The ledger: TP2 vs broken vs record vs now

metric TP2 baseline (07-31) TP4 broken, no-spec (08-15) TP4 record (08-14) TP4 now (08-16)
C1 decode (tok/s) 67.7 33.5 103.4 136.25 median · 145.5 peak
Prefill cold (tok/s) 1576 ~950 ~940 2102 @32k · 2156 @8k
C4 aggregate (tok/s) 93.15 92.43 — 182.2

Decode at prompt depth — speculative-decoding acceptance, not depth, is the
variable:

decode at depth

Concurrency-4 aggregate:

c4 aggregate

Everything needed to reproduce — image, flags, fabric wiring, boot gates,
bench scripts, dated results: github.com/neko-legends/spark-bench.


About the removed conversion (reference)

  • Source:
    apetersson/DeepSeek-V4-Flash-0731-Abliterated-FP8 — routed experts were packed MXFP4 E2M1 with per-32 UE8M0 scales; we used NVIDIA ModelOpt's lossless MXFP4→NVFP4 cast (no dequantize/requantize) with activation scales calibrated on 2× GB10 DGX Spark (128 public prompts, batch 4, seq 512).
  • Mixed precision by design: routed experts NVFP4 E2M1; attention/dense FP8 E4M3; shared-expert and MTP tensors retained.
  • Export audit at release: 8,657,043,456 / 8,657,043,456 blocks cast, 33,024 tensors across 43 layers, 48/48 shards independently audited, zero structural errors.
  • Abliteration retention was verified at release against the upstream ablation manifest (all 36 named tensors bit-identical) plus a small behavioral suite (12/12 substantive, zero refusals).
  • The runtime bridge for the SM121 B12X path remains in
    runtime/, the reference inference code in
    inference/, and the DSV4 tokenizer encoding in
    encoding/.

Attribution

Review the upstream model card and license before deploying any descendant
checkpoint.

README history 14 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-08-16Archive: remove weights + old benchmarks; repo is now an informational page (...a74c3c84.7 KB
    Loading...
  2. 2026-08-16Rewrite model card: plain-English informational style, spark-bench records + ...8fc38c512.6 KB
    Loading...
  3. 2026-08-06Add responsible use gated access form734ac3a17.5 KB
    Loading...
  4. 2026-08-06Add DGX Spark recommendation notef82679915.8 KB
    Loading...
  5. 2026-08-03Document Forge and Anvil deployment validation7ad305515.2 KB
    Loading...
  6. 2026-08-03Publish SM121 B12X two-Spark benchmarks90d5d3214.6 KB
    Loading...
  7. 2026-08-03Publish SM121 B12X two-Spark benchmarksd4baa3114.5 KB
    Loading...
  8. 2026-08-03Publish SM121 B12X two-Spark benchmarksecfd83713.8 KB
    Loading...
  9. 2026-08-03Publish tuned two-Spark NVFP4 benchmarks189c6bb12.3 KB
    Loading...
  10. 2026-08-03Publish tuned two-Spark NVFP4 benchmarksf24670f11.8 KB
    Loading...
  11. 2026-08-03Add Neko Legends benchmark chart and model-card brandingb06526f7.3 KB
    Loading...
  12. 2026-08-03Add Neko Legends benchmark chart and model-card branding304615b6.9 KB
    Loading...
  13. 2026-08-03Document calibrated two-Spark NVFP4 conversion9e8b9665.4 KB
    Loading...
  14. 2026-08-03initial commit0c7dbff21 B
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration