← back to catalog · registered 2026-08-22 13:56

Jiunsong/SuperDeepseek-V4-Flash-abliterated

Jiunsong Deepseek 296B MoE
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/Jiunsong%2FSuperDeepseek-V4-Flash-abliterated"
Response includes
  • classification m1
  • files 58
  • hub_downloads_all_time 1,014
  • author_summary 35 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M1
Primary method

Direct removal

No other method signals detected in this model.
Confidence
MEDIUM
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=0 (base model)
  • no specific method indicators - defaulting to M1 (most common)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
MEDIUM
Why we say so
primary_method=M1; difference-of-means is the reference extraction for M1/M3 (Arditi 2024)
Downloads · lifetime
1K
479 last 30d - stable
Likes
13
Descendants
2
in 2 direct forks
Model age
2mo ago
created 2026-08-12
Downloads over time
Now1.1K→from0↑0%
04118221.2K0 on Aug 121.1K on Oct 11AugSepOct
Aug 12 → Oct 11 · 49 snapshots · spans 60 days

Genealogy 2 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
mit
Languages
en ko
Tags
transformers safetensors deepseek_v4 text-generation deepseek-v4 mixture-of-experts fp4 fp8 bf16 mixed-precision long-context 1m-context

Related

Total size
158 GB
Files
58
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-08-12 03:14

Files by quantization

Auxiliary files 58 files 158 GB
model-00048-of-00048.safetensors 3.44 GB cc43742b download
model-00046-of-00048.safetensors 3.36 GB 5db924ca download
model-00004-of-00048.safetensors 3.35 GB 9610f56b download
model-00012-of-00048.safetensors 3.34 GB 64ed4e5f download
model-00014-of-00048.safetensors 3.34 GB 45db2f54 download
model-00016-of-00048.safetensors 3.34 GB e0530b70 download
model-00018-of-00048.safetensors 3.34 GB e393fea9 download
model-00020-of-00048.safetensors 3.34 GB 9f556769 download
model-00022-of-00048.safetensors 3.34 GB decd67a4 download
model-00024-of-00048.safetensors 3.34 GB fc27aeb4 download
model-00026-of-00048.safetensors 3.34 GB 657b8931 download
model-00028-of-00048.safetensors 3.34 GB b2fd5cbb download
model-00030-of-00048.safetensors 3.34 GB 9ed3c317 download
model-00032-of-00048.safetensors 3.34 GB 16365384 download
model-00034-of-00048.safetensors 3.34 GB 0f949451 download
model-00036-of-00048.safetensors 3.34 GB 7e676142 download
model-00038-of-00048.safetensors 3.34 GB 137fa617 download
model-00040-of-00048.safetensors 3.34 GB 8bc93d8a download
model-00042-of-00048.safetensors 3.34 GB 4d19bf36 download
model-00044-of-00048.safetensors 3.34 GB 422d3889 download
model-00006-of-00048.safetensors 3.34 GB 4a4f3764 download
model-00008-of-00048.safetensors 3.34 GB 224968d2 download
model-00010-of-00048.safetensors 3.34 GB 627145f4 download
model-00013-of-00048.safetensors 3.32 GB 8dfe199d download
model-00015-of-00048.safetensors 3.32 GB 5810381a download
model-00017-of-00048.safetensors 3.32 GB ed111302 download
model-00019-of-00048.safetensors 3.32 GB a74ca4d3 download
model-00021-of-00048.safetensors 3.32 GB 1671cce7 download
model-00023-of-00048.safetensors 3.32 GB c61a3e17 download
model-00025-of-00048.safetensors 3.32 GB a66b6b8d download
model-00027-of-00048.safetensors 3.32 GB fb01f21a download
model-00029-of-00048.safetensors 3.32 GB 9ec2fdf9 download
model-00031-of-00048.safetensors 3.32 GB d5078c3f download
model-00033-of-00048.safetensors 3.32 GB f2cffd43 download
model-00035-of-00048.safetensors 3.32 GB 9cb6a316 download
model-00037-of-00048.safetensors 3.32 GB a59d662f download
model-00039-of-00048.safetensors 3.32 GB a29af1aa download
model-00041-of-00048.safetensors 3.32 GB fd312e7f download
model-00043-of-00048.safetensors 3.32 GB b7103842 download
model-00005-of-00048.safetensors 3.32 GB f87a5ac7 download
model-00007-of-00048.safetensors 3.32 GB df81bb80 download
model-00009-of-00048.safetensors 3.32 GB 04d69ef1 download
model-00011-of-00048.safetensors 3.32 GB e4b8e601 download
model-00002-of-00048.safetensors 3.32 GB 77b26c93 download
model-00003-of-00048.safetensors 3.32 GB 412abf4c download
model-00047-of-00048.safetensors 3.32 GB 62816173 download
model-superdeepseek-overlay.safetensors 1.44 GB 86c8494d download
model-00045-of-00048.safetensors 1010 MB a5be6aed download
model-superdeepseek-head-recovery.safetensors 1010 MB 3d49e3e0 download
model-00001-of-00048.safetensors 1010 MB f3668ba4 download
tokenizer.json 6.07 MB 628e3364 download
model.safetensors.index.json 5.34 MB 0ec0c37c download
README.md 8.56 KB 0bd6b3e4 download
config.json 1.84 KB 5f2da910 download
.gitattributes 1.48 KB a6344aac download
LICENSE 1.06 KB d62e3bef download
tokenizer_config.json 801 B f3dad388 download
generation_config.json 170 B c56a8c5b download

README current version from Hugging Face


license: mit
library_name: transformers
pipeline_tag: text-generation
base_model: deepseek-ai/DeepSeek-V4-Flash-0731
base_model_relation: finetune
tags:

  • deepseek-v4
  • mixture-of-experts
  • fp4
  • fp8
  • bf16
  • mixed-precision
  • long-context
  • 1m-context
  • reasoning
  • tool-calling
  • uncensored
  • vllm
  • abliterated
  • obliteratus
  • supertune
    language:
  • en
  • ko

SuperDeepseek-V4-Flash-abliterated

A less-refusing DeepSeek V4 Flash full checkpoint with verified 1M-token retrieval, intact tool use, and 123 tok/s-class two-node decode.

Format
Context
Decode
Tools
License

SuperDeepseek-V4-Flash-abliterated is the general, independently downloadable full-checkpoint release of SuperDeepseek. It is built from deepseek-ai/DeepSeek-V4-Flash-0731, keeps DeepSeek's official native tensor layout, and fuses the selected OBLITERATUS + SuperTune update directly into the checkpoint. No LoRA or inference-time adapter is required.

Why this release

  • Far fewer unnecessary refusals: worst-mode refusal fell from 97.92% to 4.17% in the paired gate.
  • Tools stayed intact: tool compliance and correct-tool selection both remained 100%.
  • Real 1M-context proof: a 1,028,621-token prompt was accepted and its hidden needle was retrieved.
  • Fast verified serving: the released weights reached 123.35 aggregate tok/s on forced decode and 123.28 aggregate tok/s on structured tool generation.
  • Surgical weight edit: all 256 routed experts, routers, embeddings, mHC tensors, and untargeted parent tensors preserve their official representation.

Release at a glance

Architecture DeepSeek V4 Flash, 304B-class MoE, 43 backbone + 3 MTP layers, 256 routed experts, top-6
Checkpoint Complete Transformers/safetensors checkpoint, about 169.5 GB, 48 parent shards plus fused SuperDeepseek overlays
Native precision FP4 experts + FP8 E4M3 block-quantized paths + BF16 quality-sensitive paths
Targeted update 46 attn.wo_b weight/scale pairs plus one bounded BF16 output-head recovery
Context 1,048,576 configured; 1,028,621 actual prompt tokens accepted and retrieved
Decode 123.3459 forced / 123.2819 structured-tool aggregate tok/s at p256/C6
Behavior Worst-mode refusal 97.92% -> 4.17%; tool gates 100%

Checkpoint format and precision

This is the native-format full checkpoint, not a full-BF16 reconstruction. DeepSeek's official V4 Flash checkpoint is itself hybrid precision; no authentic all-BF16 upstream checkpoint is published. This release deliberately preserves that official representation instead of dequantizing it and presenting approximated BF16 weights as “original.”

Component Precision / storage
MoE expert weights FP4, inherited from the official expert_dtype=fp4 layout
Block-quantized paths FP8 E4M3, dynamic activation scaling, 128x128 blocks, UE8M0 scales
43 backbone + 3 MTP attn.wo_b updates Deterministic FP8 weight/scale overlay, 92 tensors
Default unquantized and quality-sensitive paths BF16 (torch_dtype=bfloat16), with upstream F32 metadata/normalization where defined
Output-head recovery One bounded BF16 head.weight overlay, rank 64, relative Frobenius delta 0.0025
Verified serving KV cache NVFP4 DS-MLA

The parent is pinned to deepseek-ai/DeepSeek-V4-Flash-0731@9e165c30e2704aec5d9d593cce3eebd58bbef1cb. All 48 parent shard identities were checked against that revision.

For the release name and scripts tuned specifically for a two-node DGX Spark deployment, see Jiunsong/SuperDeepseek-V4-Flash-abliterated-MQ-2xDGX. Both repositories contain the same released SuperDeepseek weights; this repository is the general/native-layout edition, while the sibling makes the measured MQ + 2xDGX deployment target explicit in its name.

What changed

The release uses two measured refusal-subspace passes:

  1. A robust rank-1 direction fitted across chat, think-high, and think-max modes, applied at strength 2.
  2. A second rank-1 residual direction recaptured from the baked first pass, orthogonalized against it, and applied at strength 0.5.
  3. A bounded rank-64 output-head recovery with relative Frobenius delta 0.0025.

The final checkpoint changes only the 43 backbone and three MTP attn.wo_b weight/scale pairs plus the bounded output head. The routed experts remain untouched.

Behavior and capability

Metric Official parent SuperDeepseek
Worst-mode refusal 97.92% 4.17%
Worst empty answer 0.00% 0.00%
Worst tool compliance 100.00% 100.00%
Worst correct-tool rate 100.00% 100.00%
Minimum capability mean 0.9375 0.9583

The independently reloaded checkpoint reproduced the selected candidate's deterministic validation behavior. The release gate covered empty output, Unicode, repetition, serialization, reasoning, code, formatting, tool selection, and structured tool calls.

Measured performance

These are aggregate concurrent post-first-token decode rates, not single-stream numbers. Six distinct prompts, fixed-length generation, and the sealed sparkDash-style p256/C6 contract were used.

Workload Prompt / concurrency Three trials Median
Forced output p256 / C6 117.7944, 123.3459, 125.9469 tok/s 123.3459 tok/s
Structured tool output p256 / C6 123.6285, 123.2819, 117.1843 tok/s 123.2819 tok/s

The benchmark follows sparkDash commit dfde4214f32b174880832a4d317d3c0567750ac5. Hardware and runtime were 2x NVIDIA DGX Spark / GB10, TP=2, direct CX-7 RoCEv2, vLLM, FlashInfer b12x MoE, DSpark K=1 speculative decoding, and NVFP4 DS-MLA KV cache. Hardware-specific figures should not be assumed on other systems.

Verified long context

Actual prompt tokens Accepted Needle retrieved
149,845 Yes Yes
1,028,621 Yes Yes

The configured maximum is 1,048,576 tokens. These are end-to-end acceptance and retrieval probes, not a claim of perfect recall for every task or needle position.

Two-node reference serving

The verified profile uses TP=2 over direct CX-7 RoCEv2 with the ghcr.io/anemll/dspark-vllm-gx10:0.1.1 runtime:

vllm serve /model \
  --served-model-name SuperDeepseek-V4-Flash-abliterated \
  --tensor-parallel-size 2 \
  --max-model-len 1048576 \
  --kv-cache-dtype nvfp4_ds_mla \
  --moe-backend flashinfer_b12x \
  --enable-prefix-caching \
  --async-scheduling \
  --enable-chunked-prefill \
  --speculative-config '{"method":"dspark","num_speculative_tokens":1,"draft_sample_method":"greedy"}'

The exact two-rank launcher and machine-readable release evidence are included under repro/ and evidence/.

Integrity

  • Parent revision: 9e165c30e2704aec5d9d593cce3eebd58bbef1cb
  • 48 parent shard names and hashes verified
  • 92 FP8 overlay tensors: exactly 46 attn.wo_b weight/scale pairs
  • One bounded BF16 output-head tensor
  • Independent reload, paired validation, decode, reasoning, tool, long-context, and output-integrity gates passed
Artifact SHA-256
SuperDeepseek FP8 overlay 86c8494d0b02a01ccb4d4de5ad48a66d26f56341e1dd3eb17eb61b40710f73ba
BF16 head-recovery overlay 3d49e3e05ba864666054328f475058a9558603aac70d1a191b352be65bce1428
Sealed throughput report 1cd7734205f4d03cddf241d5ed2a5f423f402d194222c9abc3bdacc965cce370
1M-context evidence 6745453ee65b9e59581d715bf668bc710b0bc680fe726763e9a750dd25f45565

Responsible use and limitations

“Abliterated” describes a measured reduction of the selected refusal subspace. It does not make every answer correct, remove the need for deployment controls, or transfer responsibility away from the operator. The capability suites are finite regression gates, and the upstream model's license and limitations continue to apply.

License

MIT, following the upstream DeepSeek V4 Flash release.

README history 2 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-08-12Add native-layout SuperDeepseek model card93f0b5e8.6 KB
    Loading...
  2. 2026-08-12Duplicate from Jiunsong/SuperDeepseek-V4-Flash-abliterated-MQ-2xDGX7a1cc2a9.7 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration