← back to catalog · registered 2026-08-22 13:56

Jiunsong/SuperDeepseek-V4-Flash-abliterated-MQ-2xDGX

Jiunsong Deepseek 296B MoE
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/Jiunsong%2FSuperDeepseek-V4-Flash-abliterated-MQ-2xDGX"
Response includes
  • classification m1
  • files 58
  • hub_downloads_all_time 2,703
  • author_summary 35 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M1
Primary method

Direct removal

No other method signals detected in this model.
Confidence
MEDIUM
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=0 (base model)
  • no specific method indicators - defaulting to M1 (most common)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
MEDIUM
Why we say so
primary_method=M1; difference-of-means is the reference extraction for M1/M3 (Arditi 2024)
Downloads · lifetime
3K
422 last 30d - stable
Likes
42
Model age
2mo ago
created 2026-08-11
Downloads over time
Now2.8K→from0↑0%
01K2.1K3.1K0 on Aug 122.8K on Oct 11AugSepOct
Aug 12 → Oct 11 · 50 snapshots · spans 60 days

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
mit
Languages
en ko
Tags
transformers safetensors deepseek_v4 text-generation deepseek-v4 mixture-of-experts fp4 fp8 bf16 mixed-precision quantized long-context

Related

Total size
158 GB
Files
58
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-08-11 16:54

Files by quantization

Auxiliary files 58 files 158 GB
model-00048-of-00048.safetensors 3.44 GB cc43742b download
model-00046-of-00048.safetensors 3.36 GB 5db924ca download
model-00004-of-00048.safetensors 3.35 GB 9610f56b download
model-00012-of-00048.safetensors 3.34 GB 64ed4e5f download
model-00014-of-00048.safetensors 3.34 GB 45db2f54 download
model-00016-of-00048.safetensors 3.34 GB e0530b70 download
model-00018-of-00048.safetensors 3.34 GB e393fea9 download
model-00020-of-00048.safetensors 3.34 GB 9f556769 download
model-00022-of-00048.safetensors 3.34 GB decd67a4 download
model-00024-of-00048.safetensors 3.34 GB fc27aeb4 download
model-00026-of-00048.safetensors 3.34 GB 657b8931 download
model-00028-of-00048.safetensors 3.34 GB b2fd5cbb download
model-00030-of-00048.safetensors 3.34 GB 9ed3c317 download
model-00032-of-00048.safetensors 3.34 GB 16365384 download
model-00034-of-00048.safetensors 3.34 GB 0f949451 download
model-00036-of-00048.safetensors 3.34 GB 7e676142 download
model-00038-of-00048.safetensors 3.34 GB 137fa617 download
model-00040-of-00048.safetensors 3.34 GB 8bc93d8a download
model-00042-of-00048.safetensors 3.34 GB 4d19bf36 download
model-00044-of-00048.safetensors 3.34 GB 422d3889 download
model-00006-of-00048.safetensors 3.34 GB 4a4f3764 download
model-00008-of-00048.safetensors 3.34 GB 224968d2 download
model-00010-of-00048.safetensors 3.34 GB 627145f4 download
model-00013-of-00048.safetensors 3.32 GB 8dfe199d download
model-00015-of-00048.safetensors 3.32 GB 5810381a download
model-00017-of-00048.safetensors 3.32 GB ed111302 download
model-00019-of-00048.safetensors 3.32 GB a74ca4d3 download
model-00021-of-00048.safetensors 3.32 GB 1671cce7 download
model-00023-of-00048.safetensors 3.32 GB c61a3e17 download
model-00025-of-00048.safetensors 3.32 GB a66b6b8d download
model-00027-of-00048.safetensors 3.32 GB fb01f21a download
model-00029-of-00048.safetensors 3.32 GB 9ec2fdf9 download
model-00031-of-00048.safetensors 3.32 GB d5078c3f download
model-00033-of-00048.safetensors 3.32 GB f2cffd43 download
model-00035-of-00048.safetensors 3.32 GB 9cb6a316 download
model-00037-of-00048.safetensors 3.32 GB a59d662f download
model-00039-of-00048.safetensors 3.32 GB a29af1aa download
model-00041-of-00048.safetensors 3.32 GB fd312e7f download
model-00043-of-00048.safetensors 3.32 GB b7103842 download
model-00005-of-00048.safetensors 3.32 GB f87a5ac7 download
model-00007-of-00048.safetensors 3.32 GB df81bb80 download
model-00009-of-00048.safetensors 3.32 GB 04d69ef1 download
model-00011-of-00048.safetensors 3.32 GB e4b8e601 download
model-00002-of-00048.safetensors 3.32 GB 77b26c93 download
model-00003-of-00048.safetensors 3.32 GB 412abf4c download
model-00047-of-00048.safetensors 3.32 GB 62816173 download
model-superdeepseek-overlay.safetensors 1.44 GB 86c8494d download
model-00045-of-00048.safetensors 1010 MB a5be6aed download
model-superdeepseek-head-recovery.safetensors 1010 MB 3d49e3e0 download
model-00001-of-00048.safetensors 1010 MB f3668ba4 download
tokenizer.json 6.07 MB 628e3364 download
model.safetensors.index.json 5.34 MB 0ec0c37c download
README.md 9.73 KB 093b9788 download
config.json 1.84 KB 5f2da910 download
.gitattributes 1.48 KB a6344aac download
LICENSE 1.06 KB d62e3bef download
tokenizer_config.json 801 B f3dad388 download
generation_config.json 170 B c56a8c5b download

README current version from Hugging Face


license: mit
library_name: transformers
pipeline_tag: text-generation
base_model: deepseek-ai/DeepSeek-V4-Flash-0731
base_model_relation: finetune
tags:

  • deepseek-v4
  • mixture-of-experts
  • fp4
  • fp8
  • bf16
  • mixed-precision
  • quantized
  • long-context
  • 1m-context
  • reasoning
  • tool-calling
  • uncensored
  • vllm
  • obliteratus
  • supertune
  • dgx-spark
    language:
  • en
  • ko

SuperDeepseek-V4-Flash-abliterated-MQ-2xDGX

A fast, less-refusing DeepSeek V4 Flash with mixed quantization (MQ), verified 1M-token retrieval, and 123 tok/s-class aggregate decode on two DGX Spark nodes.

Precision
Context
Decode
Tools
License

SuperDeepseek-V4-Flash-abliterated-MQ-2xDGX is the performance-focused SuperDeepseek release built from
deepseek-ai/DeepSeek-V4-Flash-0731.
It keeps the official hybrid checkpoint layout, applies a surgical OBLITERATUS +
SuperTune update, and ships as a directly loadable checkpoint with no LoRA or runtime
adapter required.

Release highlights

Architecture DeepSeek V4 Flash, 304B-class MoE, 43 backbone layers + 3 MTP layers, 256 routed experts, top-6
Release format Hybrid FP4 experts + FP8 E4M3 blocks + BF16 quality-sensitive tensors, about 169.5 GB on the Hub
Targeted update 46 attn.wo_b weight/scale pairs, with all routed experts and untargeted parent tensors preserved
Verified context 1,048,576 configured; 1,028,621-token prompt accepted with successful needle retrieval
Aggregate decode 118.6 tok/s forced output and 123.3 tok/s structured tool output at p256/C6
Behavior shift Worst-mode refusal 97.92% -> 4.17%, while the measured tool gates remain 100%
Capability floor Minimum capability mean 0.9375 -> 0.9583 against the pinned parent

Why run this model

  • Far fewer unnecessary refusals: the selected checkpoint reduces the measured
    worst-mode refusal rate from 97.92% to
    4.17%.
  • Tools stay intact: tool compliance and correct-tool selection remain at
    100.00% and
    100.00% in the paired release gate.
  • Real 1M context proof: both 149,845-token and 1,028,621-token needle-retrieval
    requests completed successfully.
  • Fast on two DGX Spark nodes: the measured serving profile reaches
    123.3 aggregate tok/s on structured tool generation.
  • Surgical rather than destructive: experts, routers, embeddings, mHC tensors,
    and every untargeted parent tensor retain the official checkpoint representation.

Quantization and precision

MQ in the model name means mixed quantization. This is an
official-layout mixed-precision checkpoint, not a full-BF16 release and not a
custom whole-model requantization.

Component Precision / storage
MoE expert weights FP4, inherited from the official expert_dtype=fp4 checkpoint layout
Block-quantized paths FP8 E4M3, dynamic activation scaling, 128x128 weight blocks, UE8M0 scales
43 backbone + 3 MTP attn.wo_b updates Deterministic FP8 weight/scale overlay, 92 tensors
Default unquantized and quality-sensitive paths BF16 (torch_dtype=bfloat16) with F32 metadata/normalization where defined upstream
Output head recovery One bounded BF16 head.weight overlay, rank-64, relative Frobenius delta 0.0025
Measured serving KV cache NVFP4 DS-MLA

The parent checkpoint is pinned to 9e165c30e2704aec5d9d593cce3eebd58bbef1cb. Only the declared FP8
attn.wo_b pairs and the single bounded BF16 output head are redirected by the final
weight index; the remaining parent tensors keep their original quantization and bytes.

What was changed

The release uses two measured weight-space passes:

  1. OBLITERATUS fits a robust rank-1 refusal direction across chat,
    think-high, and think-max modes and applies the selected strength
  2. A second rank-1 residual pass is recaptured from the baked first pass,
    orthogonalized against it, and applied at strength
    0.5.
  3. A bounded rank-64 output-head recovery was applied; its relative Frobenius delta was 0.0025.

The final checkpoint modifies only the 43 backbone and three MTP attn.wo_b
weight/scale pairs plus the bounded output head. There is no inference-time adapter.

Behavior and capability

Metric Official parent SuperDeepseek-V4-Flash-abliterated-MQ-2xDGX
Worst-mode refusal 97.92% 4.17%
Worst empty answer 0.00% 0.00%
Worst tool compliance 100.00% 100.00%
Worst correct-tool rate 100.00% 100.00%
Minimum capability mean 0.9375 0.9583

The independently reloaded checkpoint reproduced the selected candidate's deterministic
validation behavior exactly. Empty-output, Unicode, repetition, serialization,
reasoning, code, formatting, and tool-use sentinels were included in the release gate.

Decode performance

The numbers below are aggregate concurrent decode throughput, not single-stream speed.
They use six distinct prompts, fixed-length generation, and the sealed
sparkDash-style measurement contract at commit
dfde4214f32b174880832a4d317d3c0567750ac5.

Workload Prompt / concurrency Aggregate decode
Forced output p256 / C6 118.5505 tok/s
Structured tool output p256 / C6 123.2888 tok/s
  • Median matched decode ratio vs the parent: 1.0012x
  • Minimum matched-case ratio vs the parent: 0.9627x
  • Regular CUDA graphs vs breakable: 1.2125x at C1 and 1.2674x at C6

Verified long context

Actual prompt tokens Accepted Needle retrieved
149,845 Yes Yes
1,028,621 Yes Yes

The configured maximum is 1,048,576 tokens. These are end-to-end acceptance and
retrieval probes; they are not a claim that every task benefits equally from the full
window.

Two-node DGX Spark serving

The measured profile uses TP=2 over direct CX-7 RoCEv2 with the
ghcr.io/anemll/dspark-vllm-gx10:0.1.1 runtime:

  • NVFP4 DS-MLA KV cache
  • DSpark speculative decoding with K=1 and greedy draft sampling
  • FlashInfer b12x MoE and FlashInfer autotuning
  • prefix caching, asynchronous scheduling, and chunked prefill
  • regular CUDA graphs with VLLM_USE_BREAKABLE_CUDAGRAPH=0

The repository includes the exact two-rank launcher under
repro/scripts/serve_superdeepseek_v4_dual.sh. Its measured model-facing options are:

vllm serve /model \
  --served-model-name SuperDeepseek-V4-Flash-abliterated-MQ-2xDGX \
  --tensor-parallel-size 2 \
  --max-model-len 1048576 \
  --kv-cache-dtype nvfp4_ds_mla \
  --moe-backend flashinfer_b12x \
  --enable-prefix-caching \
  --async-scheduling \
  --enable-chunked-prefill \
  --speculative-config '{"method":"dspark","num_speculative_tokens":1,"draft_sample_method":"greedy"}'

After the server is ready, it exposes an OpenAI-compatible API:

from openai import OpenAI

client = OpenAI(base_url="http://127.0.0.1:8888/v1", api_key="EMPTY")
response = client.chat.completions.create(
    model="SuperDeepseek-V4-Flash-abliterated-MQ-2xDGX",
    messages=[{"role": "user", "content": "Design a reliable tool-using agent."}],
    max_tokens=1024,
)
print(response.choices[0].message.content)

Uncensored behavior

“Uncensored” means this checkpoint measurably reduces the selected refusal subspace
while retaining the declared capability and output-integrity gates. It does not imply
that every answer is correct or that downstream deployment controls are unnecessary.

Release integrity

  • Parent: deepseek-ai/DeepSeek-V4-Flash-0731@9e165c30e2704aec5d9d593cce3eebd58bbef1cb
  • 48 original parent shard names preserved
  • 92 FP8 overlay tensors: exactly 46 attn.wo_b weight/scale pairs
  • 1 bounded BF16 output-head tensor
  • Independently reloaded paired validation: passed
  • Decode, tool, reasoning, long-context, and output-integrity gates: passed
  • Machine-readable benchmark and release evidence included under evidence/

Limitations

  • The speed figures are measured on a specific two-node DGX Spark/CX-7 runtime and
    should not be treated as universal hardware results.
  • Abliteration changes refusal behavior and can produce content the parent would
    decline. Operators remain responsible for access control and appropriate use.
  • The capability and integrity suites are finite regression gates, not proof of
    universal correctness.
  • One million token acceptance does not guarantee perfect recall at every position or
    on every task.
  • The parent model's license and upstream limitations continue to apply.

Evidence identities

Evidence SHA-256
paired validation 65aaf03e1ac5f88ea462c7500d36a06c254370dfa3f9e2a0d24183c482e23a46
phase 1 selection a75b0f9c2c2c9d0c38e94c5a7e36139509178f7f9b74521925c9a1f011f0f207
phase 2 selection b0fdb2c73f16954b56bde2460833b782b29bfe0a67aeedc611d65dc94764e4eb
recovery decision 39acc4c24694c7bb08615217d16e5e7d2f3d9e33c26425610ff1198054422a2f
decode benchmark 6745453ee65b9e59581d715bf668bc710b0bc680fe726763e9a750dd25f45565
overlay audit c618a1a2da8295b07b1d93d9a59c9a66e8e430c2b8bdf1aca07a56ad56455e8e
head recovery audit a5aa6531909444a5d6199d889030b6b9424868a7656e4db8d77c30aad4bfc8db

License

MIT, following the upstream DeepSeek V4 Flash release.

README history 3 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-08-11Rename release and document MQ 2xDGX performance2a7dd6a9.7 KB
    Loading...
  2. 2026-08-11Add measured SuperDeepseek model carde24ae1c4.2 KB
    Loading...
  3. 2026-08-11Duplicate from deepseek-ai/DeepSeek-V4-Flash-07311cadec77.1 KB
    Loading...

Discussions 1 thread

  1. 2026-08-12Visionopen1 💬#1
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration