← back to catalog · registered 2026-08-22 13:56

cebeuq/DeepSeek-V4-Flash-0731-abliterated

cebeuq Deepseek 296B MoE
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/cebeuq%2FDeepSeek-V4-Flash-0731-abliterated"
Response includes
  • classification m1
  • files 57
  • hub_downloads_all_time 18,284
  • author_summary 2 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M1
Primary method

Direct removal

No other method signals detected in this model.
Confidence
MEDIUM
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=0 (base model)
  • no specific method indicators - defaulting to M1 (most common)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
MEDIUM
Why we say so
primary_method=M1; difference-of-means is the reference extraction for M1/M3 (Arditi 2024)
Downloads · lifetime
18K
2K last 30d - stable
Likes
42
Descendants
2
in 2 direct forks
Model age
2mo ago
created 2026-08-01
Downloads over time
Now19.2K→from1.6K↑1,100%
7207.5K14.2K21K1.6K on Aug 519.2K on Oct 11AugSepOct
Aug 5 → Oct 11 · 50 snapshots · spans 67 days

Genealogy 2 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
mit
Tags
transformers safetensors deepseek_v4 text-generation abliterated uncensored mixture-of-experts mhc fp4 fp8 dspark vllm

Related

Total size
157 GB
Files
57
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-08-01 23:48

Files by quantization

Auxiliary files 57 files 157 GB
model-00048-of-00048.safetensors 3.44 GB cc43742b download
model-00046-of-00048.safetensors 3.36 GB 5db924ca download
model-00004-of-00048.safetensors 3.35 GB 9610f56b download
model-00012-of-00048.safetensors 3.34 GB 64ed4e5f download
model-00014-of-00048.safetensors 3.34 GB 45db2f54 download
model-00016-of-00048.safetensors 3.34 GB e0530b70 download
model-00018-of-00048.safetensors 3.34 GB e393fea9 download
model-00020-of-00048.safetensors 3.34 GB 9f556769 download
model-00022-of-00048.safetensors 3.34 GB decd67a4 download
model-00024-of-00048.safetensors 3.34 GB fc27aeb4 download
model-00026-of-00048.safetensors 3.34 GB 657b8931 download
model-00028-of-00048.safetensors 3.34 GB b2fd5cbb download
model-00030-of-00048.safetensors 3.34 GB 9ed3c317 download
model-00032-of-00048.safetensors 3.34 GB 16365384 download
model-00034-of-00048.safetensors 3.34 GB 0f949451 download
model-00036-of-00048.safetensors 3.34 GB 7e676142 download
model-00038-of-00048.safetensors 3.34 GB 137fa617 download
model-00040-of-00048.safetensors 3.34 GB 8bc93d8a download
model-00042-of-00048.safetensors 3.34 GB 4d19bf36 download
model-00044-of-00048.safetensors 3.34 GB 422d3889 download
model-00006-of-00048.safetensors 3.34 GB 4a4f3764 download
model-00008-of-00048.safetensors 3.34 GB 224968d2 download
model-00010-of-00048.safetensors 3.34 GB 627145f4 download
model-00013-of-00048.safetensors 3.32 GB 8dfe199d download
model-00015-of-00048.safetensors 3.32 GB 5810381a download
model-00017-of-00048.safetensors 3.32 GB ed111302 download
model-00019-of-00048.safetensors 3.32 GB a74ca4d3 download
model-00021-of-00048.safetensors 3.32 GB 1671cce7 download
model-00023-of-00048.safetensors 3.32 GB c61a3e17 download
model-00025-of-00048.safetensors 3.32 GB a66b6b8d download
model-00027-of-00048.safetensors 3.32 GB fb01f21a download
model-00029-of-00048.safetensors 3.32 GB 9ec2fdf9 download
model-00031-of-00048.safetensors 3.32 GB d5078c3f download
model-00033-of-00048.safetensors 3.32 GB f2cffd43 download
model-00035-of-00048.safetensors 3.32 GB 9cb6a316 download
model-00037-of-00048.safetensors 3.32 GB a59d662f download
model-00039-of-00048.safetensors 3.32 GB a29af1aa download
model-00041-of-00048.safetensors 3.32 GB fd312e7f download
model-00043-of-00048.safetensors 3.32 GB b7103842 download
model-00005-of-00048.safetensors 3.32 GB f87a5ac7 download
model-00007-of-00048.safetensors 3.32 GB df81bb80 download
model-00009-of-00048.safetensors 3.32 GB 04d69ef1 download
model-00011-of-00048.safetensors 3.32 GB e4b8e601 download
model-00002-of-00048.safetensors 3.32 GB 77b26c93 download
model-00003-of-00048.safetensors 3.32 GB 412abf4c download
model-00047-of-00048.safetensors 3.32 GB 62816173 download
model-overlay-00001-of-00001.safetensors 1.44 GB 20d25599 download
model-00045-of-00048.safetensors 1010 MB a5be6aed download
model-00001-of-00048.safetensors 1010 MB f3668ba4 download
tokenizer.json 6.07 MB 628e3364 download
model.safetensors.index.json 5.34 MB dffe4726 download
README.md 13.6 KB d9a947e0 download
config.json 1.84 KB 5f2da910 download
.gitattributes 1.48 KB a6344aac download
LICENSE 1.06 KB d62e3bef download
tokenizer_config.json 801 B f3dad388 download
generation_config.json 170 B c56a8c5b download

README current version from Hugging Face


license: mit
base_model:

  • deepseek-ai/DeepSeek-V4-Flash-0731
    base_model_relation: finetune
    pipeline_tag: text-generation
    library_name: transformers
    tags:
  • abliterated
  • uncensored
  • mixture-of-experts
  • deepseek_v4
  • mhc
  • fp4
  • fp8
  • dspark
  • vllm
  • dgx-spark

DeepSeek-V4-Flash-0731 — Abliterated

A decensored (abliterated) rebuild of deepseek-ai/DeepSeek-V4-Flash-0731 — a 284B-parameter DeepSeek-V4 MoE agentic-coding / reasoning model — served unquantized at native FP4+FP8 across 2× NVIDIA DGX Spark (GB10, 128 GB each) with vLLM.

Refusal behavior was removed from the attention output projections only. All 11,008 routed experts per layer, the shared experts, wo_a, embeddings, routers, norms and every mHC parameter are untouched. Coding, tool-calling and long-context retrieval remain intact.

This is an overlay, not a re-quantization. Exactly 92 of 72,317 tensors (0.13 %) differ from DeepSeek's release; the other 72,225 are byte-identical and verifiable per-tensor by sha256. There is no quantization step anywhere in this pipeline — the model ships at the base checkpoint's native precision.

⚠️ Uncensored model. Safety refusals have been substantially removed. You are responsible for how you use it. Intended for local/research use on hardware you control. It will attempt almost any request.


Highlights

Base deepseek-ai/DeepSeek-V4-Flash-0731 (DeepSeek-V4 MoE, 284B total / 13B active)
Architecture 43 layers, hidden 4096, 256 experts (top-6) + 1 shared, hybrid CSA + HCA attention with Lightning Indexer, mHC hyper-connections (hc_mult=4), inline 3-stage DSpark draft head
Context 1,048,576 native (served here at 262,144)
Precision unchanged — FP4 (E2M1 + ue8m0) routed experts, FP8 e4m3 (128×128 block scales) elsewhere
Edit rank-1 orthogonal projection, λ = 2.5, on 43 layers.*.attn.wo_b + 3 mtp.*.attn.wo_b
Untouched routed + shared experts, wo_a, embed, head, routers, norms, tid2eid, all mHC params
Size ~167 GB, 48 safetensors shards + a 1.54 GB overlay
Runtime vLLM 0.25.2 (ghcr.io/anemll/dspark-vllm-gx10:0.1.1), TP=2 across 2× DGX Spark

Validation (before → after abliteration)

Refusal measured on AdvBench harmful_behaviors with long generations (V4 exhibits a documented delayed refusal — educational framing for a stretch before pivoting), scored on the parsed answer content with the thinking block excluded. Tool-calling measured against a 3-tool schema using the native DSML format and --tool-call-parser deepseek_v4; compliance means the completion parses, contains a call, names a valid tool, and supplies every required parameter.

Metric Base 0731 This model (abliterated)
Refusal — chat (n=48) 95.8 % 0.0 %
Refusal — think-high (n=24) — 0.0 %
Refusal — think-max (n=24) 95.8 % 0.0 %
Tool-call compliance (all 3 modes) 1.000 1.000
Correct tool selected — 1.000
Empty-answer rate 0.000 0.000
DSpark draft acceptance ~48 % 48.7 % — drafter at parity

Zero refusals observed. At these sample sizes the 95 % Wilson upper bounds are ≈7.4 % (chat) and ≈13.8 % (thinking) — the observed rate is 0, not a proven ceiling.

Serving performance (2× DGX Spark, TP=2 over ConnectX-7 RoCE, single-stream decode, TTFT excluded):

workload decode tok/s draft acceptance mean accept length
structured (JSON) 75.5 82.5 % 5.12 / 5
code 71.7 ~69 % —
technical prose 50.6 47.9 % 3.39

Aggregate 183 tok/s at 8 concurrent streams. Throughput is strongly content-dependent because speculative decoding lives or dies on how predictable the output is — quote the workload with the number.

Prefill with prefix caching on: 41.9 s cold → 0.37 s warm at 71k context (0.65 s at 156k), which is what makes agentic tool loops usable.


How it was built

  1. Refusal directions. Forward hooks on all 43 attn.wo_b modules of a live vLLM TP=2 deployment captured last-token sublayer outputs over 256 AdvBench harmful vs 256 Alpaca harmless prompts, in all three reasoning modes (chat, think-high, think-max), giving per-layer difference-of-means directions r_ℓ ∈ ℝ⁴⁰⁹⁶. Cross-rank digests were asserted identical on every prompt, since wo_b is RowParallelLinear and a pre-all-reduce hook would silently capture a half-direction. Stability: median split-half cosine 0.9863, median held-out AUC 0.9976, all 43 layers above 0.90.
  2. Merge + bake. The three per-mode direction sets were merged and applied as W ← W − λ·r̂(r̂ᵀW) at λ = 2.5, in float64, directly on the FP8 e4m3 weights: dequantize with the ue8m0 128×128 block scales, project, then re-quantize holding the original block exponents fixed so every element the projection did not move re-encodes to its identical original byte. Overflow is clamped rather than rescaled (6,877 of 94,208 blocks overflowed; 7,821 of 1.54 B elements clamped; max overshoot 1.186×). The 3 DSpark stages inherit the deepest backbone direction — an unedited drafter would propose refusal tokens the edited verifier rejects, collapsing acceptance.

Mean relative edit size 0.0587. Written as a tensor overlay with a repointed model.safetensors.index.json, so the 48 original shards are never rewritten. The bake is deterministic — three independent runs produced overlay sha256 20d2559987a19cb5….

Never staged a BF16 upcast; the whole pipeline is numpy + safetensors, torch-free, and runs on CPU in ~70 s.


Serving (vLLM, 2× DGX Spark)

The model (~167 GB) exceeds a single 128 GB Spark, so serve it across both nodes with tensor-parallel = 2. Config that works on GB10:

vllm serve /path/to/this-model \
  --served-model-name ds4-abliterated --trust-remote-code \
  --tensor-parallel-size 2 --distributed-executor-backend mp \
  --nnodes 2 --node-rank $RANK --master-addr $HEAD --master-port 25000 \
  --moe-backend flashinfer_b12x --async-scheduling \
  --max-model-len 262144 --kv-cache-dtype fp8 --block-size 256 \
  --gpu-memory-utilization 0.85 --max-num-seqs 8 \
  --max-num-batched-tokens 16384 --max-cudagraph-capture-size 48 \
  --enable-prefix-caching --enable-chunked-prefill \
  --speculative-config '{"method":"dspark","num_speculative_tokens":5,"draft_sample_method":"probabilistic"}' \
  --tokenizer-mode deepseek_v4 --reasoning-parser deepseek_v4 \
  --tool-call-parser deepseek_v4 --enable-auto-tool-choice

GB10 notes, each of which cost real debugging time:

  • --moe-backend flashinfer_b12x is not optional. auto resolves to DEEPGEMM_MXFP4, the slow path on sm_121. VLLM_USE_B12X_MOE=1 alone does not override it — the CLI flag does.
  • --linear-backend must stay auto; flashinfer_b12x has no kernel for V4's FP8 block-scaled linears and hard-fails at load.
  • NCCL_IB_HCA must name a device that exists. GB10 renames its ConnectX-7 RDMA devices to rocep*/roceP2p* — there is no mlx5_N. NCCL_IB_HCA is a prefix filter: a non-matching value makes NCCL log "no device found" and silently fall back to TCP sockets, costing ~2.2× decode throughput (12 → 27 tok/s here). Verify with NCCL_DEBUG=INFO | grep 'NET/IB'.
  • num_speculative_tokens must be ≤ 5 (dspark_block_size). The model card's 7 boots unvalidated on vLLM 0.25.x and then misbehaves — the drafter emits exactly block_size tokens per pass and multi-block drafting is unimplemented.
  • --max-num-batched-tokens ≥ 8192 + max_num_seqs × (k−1), else vLLM silently degrades the scheduler budget.
  • Prefix caching is safe here only because VLLM_DSPARK_GPU_REJECTED_CONTEXT_MASK=1 is set (vLLM #47930 otherwise collapses draft acceptance to ~0.5 %). Measured with it on: acceptance 48.7 %, unchanged.
  • nvidia-smi reports memory as N/A on GB10 and torch.cuda.mem_get_info is misleading on unified memory — use free -h.

Notes for agent frameworks

  • Reasoning modes: chat, think-high, think-max. There is no HuggingFace chat template — prompts go through the repo's encoding/encoding_dsv4.py (encode_messages(messages, thinking_mode, reasoning_effort)). Calling tokenizer.apply_chat_template silently falls back to a generic format instead of raising.
  • Enabling thinking over the OpenAI API: chat_template_kwargs: {"thinking": true, "reasoning_effort": "max"}. A top-level reasoning_effort field is ignored — vLLM's deepseek_v4 parser gates on chat_template_kwargs["thinking"] AND reasoning_effort != "none". Reasoning is returned in the reasoning field, separate from content.
  • Tool calls use DeepSeek's native DSML block, not JSON. parse_message_from_completion_text raises on malformed output, which makes it a clean compliance signal.
  • Benchmarking: read usage.completion_tokens, never count SSE deltas. Under speculative decoding vLLM emits one chunk per decode step carrying every accepted token, so counting chunks measures steps/s and under-reports by the acceptance length (~3.4× here).

Findings

Three things measured here that we could not find published elsewhere:

  • Refusal is encoded along two substantially different directions. Chat-mode and thinking-mode directions have median cosine 0.46 on the 36 layers where both separate at AUC > 0.99, and they diverge with depth (0.88 at layer 0, −0.08 at layer 42). The two thinking modes agree with each other (0.90). A direction fitted only on non-thinking prompts leaves thinking-mode refusal substantially intact — hence the merged direction here.
  • mHC does not block abliteration. V4 replaces ordinary residuals with hyper-connections, so the carrier is [B,S,4,4096] and output_hidden_states returns that carrier rather than a hidden state. But the mHC write is rank-1 across streams — hc_post broadcasts one 4096-vector into all four with per-token scalars — so removing a direction at the sublayer output removes it from every stream at once. Measured: split-half cosine 0.9863, held-out AUC 0.9976, in every reasoning mode.
  • A baked FP8 edit is structurally weaker than a runtime hook. A rank-1 edit spread over 4096 dims perturbs each element by ~1/√4096 of the row's component — measured at 0.253× the e4m3 quantization step — so most elements round back to their original byte. Best achievable direction removal on a real tensor is ~68 %, minimised near λ≈1.5. This is why λ must be calibrated against the baked model rather than inherited from a search run, and plausibly why published work converges on over-projection at λ≈2.5 rather than a "clean" λ=1.

Limitations

Stated plainly, because a card that is silent on these reads as if they were checked.

  • Capability was gated on tool-calling only. MMLU-Pro, GSM8K, HumanEval and similar
    benchmarks were not run. The claim is "tool-calling and refusal are measured"; it is
    not "general capability is unchanged". The published precedent at this operating point
    reports capability flat-to-positive, which is reassuring but is not a measurement of
    this checkpoint.
  • Long-context behaviour is unvalidated. Directions were captured from ~60-token
    prompts. CSA's Lightning Indexer selects a completely different token set at 168k than at
    60 tokens, so it is an open question whether refusal suppression holds at depth. No
    published DeepSeek-V4 abliteration has been validated past 32k either. Needle-in-haystack
    and refusal-at-depth were not run here.
  • Refusal is scored by marker matching on the parsed answer content (thinking block
    excluded), with long generations to catch V4's delayed-refusal pivot. It is not an
    LLM-judge evaluation, and it will miss refusals phrased in ways the marker set does not
    cover.
  • Two categories resist attention-only editing. Published work on this architecture
    reports PII-doxing and self-harm prompts retaining materially higher refusal than the
    aggregate at the same operating point. Not separately measured here.
  • Sample sizes are small (n=48 chat, n=24 per thinking mode). Zero refusals observed is
    not zero refusal rate — see the Wilson bounds above.

Files

config.json, model-*-of-00048.safetensors (+ repointed index), model-overlay-00001-of-00001.safetensors, tokenizer.json/tokenizer_config.json, generation_config.json, encoding/ (the official encoding_dsv4.py prompt encoder), inference/ (DeepSeek's reference implementation). See docs/ for the full validation report and the per-tensor edit audit (abliteration_report.json), and scripts/ for the capture / search / bake / validate pipeline used to build this.

License & credits

MIT (inherited from the base model). Base model: deepseek-ai/DeepSeek-V4-Flash-0731. Abliteration follows the residual-direction method (Arditi et al., Refusal in LLMs is mediated by a single direction, NeurIPS 2024); tooling inspired by p-e-w/heretic and elder-plinius/OBLITERATUS, neither of which supports this architecture. attn.wo_b targeting and the λ≈2.5 operating point independently replicate lovesenko/DeepSeek-V4-Flash-DSpark-Abliterated (their reported mean Frobenius δ 0.059 vs 0.0587 measured here). Serving stack: anemll/dspark-vllm-gx10. Not affiliated with or endorsed by the base-model authors.

README history 1 version

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-08-01Model card, abliteration pipeline, validation artifactsad7e7ed13.6 KB
    Loading...

Discussions 2 threads

  1. 2026-09-07Thank you — overlay for Vision-Exp, please?open1 💬#2
    Loading...
  2. 2026-08-03Could this be a LoRA?open1 💬#1
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration