← back to catalog · registered 2026-08-22 13:56

RupertBern/DeepSeek-V4-Flash-0731-HERETIC-Abliterated-FP8

RupertBern Deepseek
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/RupertBern%2FDeepSeek-V4-Flash-0731-HERETIC-Abliterated-FP8"
Response includes
  • classification m3
  • files 37
  • hub_downloads_all_time 124
  • author_summary 1 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M3
Primary method

Layer-wise ablation

Applied on top of direct removal inherited from the base model.
Confidence
HIGH
Inherited from base model
Why this label 2 signals
Producer identity confirmed by naming conventions, tags or the model card. This label is very unlikely to change.
  • 'heretic' in model name (Heretic-produced)
  • Heretic uses layer-wise optimization (M3) with underlying direction removal (M1)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
HIGH
Why we say so
name contains 'heretic'; Heretic default extraction is difference-of-means (Arditi 2024)
Downloads · lifetime
124
47 last 30d - stable
Likes
1
Model age
2mo ago
created 2026-08-02
Downloads over time
Now148→from34↑335%
287211615934 on Aug 5148 on Oct 11AugSepOct
Aug 5 → Oct 11 · 50 snapshots · spans 67 days

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
mit
Tags
transformers safetensors deepseek_v4 text-generation deepseek-v4 abliterated heretic fp8 reasoning tool-use base_model:deepseek-ai/DeepSeek-V4-Flash-0731 base_model:quantized:deepseek-ai/DeepSeek-V4-Flash-0731

Related

Total size
78.8 GB
Files
37
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-08-02 20:08

Files by quantization

Auxiliary files 37 files 78.8 GB
model-00004-of-00048.safetensors 3.35 GB 9610f56b download
model-00012-of-00048.safetensors 3.34 GB 64ed4e5f download
model-00014-of-00048.safetensors 3.34 GB 45db2f54 download
model-00016-of-00048.safetensors 3.34 GB e0530b70 download
model-00018-of-00048.safetensors 3.34 GB e393fea9 download
model-00020-of-00048.safetensors 3.34 GB 9f556769 download
model-00022-of-00048.safetensors 3.34 GB decd67a4 download
model-00024-of-00048.safetensors 3.34 GB fc27aeb4 download
model-00006-of-00048.safetensors 3.34 GB 4a4f3764 download
model-00008-of-00048.safetensors 3.34 GB 224968d2 download
model-00010-of-00048.safetensors 3.34 GB 627145f4 download
model-00013-of-00048.safetensors 3.32 GB 8dfe199d download
model-00015-of-00048.safetensors 3.32 GB 5810381a download
model-00017-of-00048.safetensors 3.32 GB ed111302 download
model-00019-of-00048.safetensors 3.32 GB a74ca4d3 download
model-00021-of-00048.safetensors 3.32 GB 1671cce7 download
model-00023-of-00048.safetensors 3.32 GB c61a3e17 download
model-00005-of-00048.safetensors 3.32 GB f87a5ac7 download
model-00007-of-00048.safetensors 3.32 GB df81bb80 download
model-00009-of-00048.safetensors 3.32 GB 04d69ef1 download
model-00011-of-00048.safetensors 3.32 GB e4b8e601 download
model-00002-of-00048.safetensors 3.32 GB 77b26c93 download
model-00003-of-00048.safetensors 3.32 GB 412abf4c download
model-overlay-00001-of-00001.safetensors 1.13 GB 601c4409 download
model-00001-of-00048.safetensors 1010 MB f3668ba4 download
tokenizer.json 6.07 MB 628e3364 download
model.safetensors.index.json 5.34 MB 441264e2 download
abliteration_report.json 27.5 KB 2fad904b download
README.md 5.87 KB ec4a0e68 download
EVAL_RESULTS.md 2.76 KB 2483c714 download
RELEASE_MANIFEST.json 2.45 KB 9c63aa83 download
config.json 1.84 KB 5f2da910 download
.gitattributes 1.48 KB a6344aac download
HERETIC_ATTRIBUTION.md 1.24 KB 982e6e5c download
LICENSE 1.06 KB d62e3bef download
tokenizer_config.json 801 B f3dad388 download
generation_config.json 170 B c56a8c5b download

README current version from Hugging Face


license: mit
library_name: transformers
pipeline_tag: text-generation
base_model: deepseek-ai/DeepSeek-V4-Flash-0731
tags:

  • deepseek-v4
  • abliterated
  • heretic
  • fp8
  • reasoning
  • tool-use

DeepSeek-V4-Flash-0731 HERETIC Abliterated FP8

A native-FP8 behavioral derivative of deepseek-ai/DeepSeek-V4-Flash-0731.

This release removes a three-mode refusal subspace from selected attention output projections while preserving the official architecture, tokenizer/encoder, FP8 format, reasoning modes, tool-call format, routed experts, and shared experts.

Release gate: GREEN — baked artifact load, behavior, tool-call, and corrected reasoning smoke verified.

Tool attribution

This checkpoint was produced with Heretic v1.4.0 by Philipp Emanuel Weidmann and contributors, using base commit 7675b90d648154cdfefa597372cc477df9848eab. Heretic is their project and is licensed AGPL-3.0-or-later; it is not owned or authored by this model publisher.

The release-specific work here is the DeepSeek V4/mHC two-node adaptation, separate chat / think-high / think-max direction capture, rank-3 subspace construction, native-FP8 attention-only bake, and the evaluation/package receipts. Model weights retain the upstream DeepSeek MIT license. Heretic-derived source modifications, when distributed, remain under AGPL-3.0-or-later. See HERETIC_ATTRIBUTION.md.

What changed

  • Built three separate refusal directions from chat, think-high, and think-max activations.
  • Orthonormalized them into a rank-3 refusal subspace per layer instead of averaging low-cosine directions.
  • Edited attention output projections for backbone layers 10–42.
  • Applied the deepest backbone subspace to the three attached MTP/DSpark attention projections.
  • Used lambda=1.5 selected by a runtime sweep under a >=0.99 tool-call gate.
  • Dequantized and re-quantized only edited FP8 tensors back to the original F8_E4M3 + F8_E8M0 representation.

What did not change

  • No reasoning mode was disabled or shortened.
  • No lower-bit conversion was applied.
  • No routed-expert or shared-expert weight was edited.
  • No tokenizer, official V4 encoder, model config, context setting, or tool schema was replaced.
  • The official source checkpoint remained read-only throughout capture, search, bake, and evaluation.

Mechanical receipts

  • Edited tensors: 72 (36 weight + scale pairs)
  • Untouched indexed tensors: 72,245
  • Overlay size: approximately 1.21 GB
  • Cross-node deterministic overlay SHA-256: 601c440978a973d6c0a7dbf00903bf5bbe0c278d4e43d68947ff65ed631b1bd3
  • Rewritten index SHA-256: f06fc8d7a4872ba6e8ccc84d6b48859123e9beeae2ad9f8bdb9cf82f260ec114
  • Worst quantized intended-effect retention: 90.57%
  • Clamp fraction: 5.17e-6
  • Dangling shards across both independently produced package trees: 0

Behavioral evaluation

Deterministic evaluation used the official DeepSeek V4 encoder, temperature 0, seed 0, identical prompts, and the same two-node vLLM runtime for baseline and derivative.

Refusal and tool use

Mode Official baseline (n=small sweep) Baked model (n=full)
chat 93.75% 29.17% (n=48)
think-high 75.00% 33.33% (n=24)
think-max 87.50% 29.17% (n=24)
  • Empty-answer rate: 0%
  • Tool-call compliance: 100% (6/6)
  • Correct-tool rate: 100% (6/6)
  • Tool parse failures: 0

The refusal detector is deliberately broad and may over-count soft disclaimers. These figures are evaluation receipts, not a claim that every safety behavior has been removed.

Reasoning regression smoke

A 12-item deterministic arithmetic, symbolic, sequence, and code-output smoke was run in all three modes. This is a regression smoke, not a substitute for MATH, HumanEval, or a broad capability benchmark.

  • Official: chat 10/12, think-high 12/12, think-max 12/12
  • Baked: chat 10/12, think-high 12/12, think-max 12/12

Full details and result hashes are in EVAL_RESULTS.md.

Runtime

Use a DeepSeek-V4-compatible runtime and the official encoder. This checkpoint retains the official V4/DSpark architecture and native FP8 layout. Generic runtimes without DeepSeek V4 support may not load it correctly.

The model was exercised on two NVIDIA DGX Spark nodes with tensor parallelism 2, the official V4 encoder, and a DeepSeek-V4-capable vLLM build.

Reproducibility

The release includes:

  • refusal direction capture and diagnostics
  • rank-3 orthonormal subspace construction receipt
  • runtime lambda sweep
  • FP8 fixed-exponent projection/requantization code
  • bake report
  • behavior and quality evaluation receipts
  • package hashes

Limitations

  • This is an unofficial community derivative and is not affiliated with or endorsed by DeepSeek.
  • Abliteration reduces some refusal behavior and can increase compliance with unsafe requests. This checkpoint is not safety-certified; downstream deployers remain responsible for evaluation, access controls, monitoring, and compliance with applicable law and platform policy.
  • The reasoning check is intentionally small; stronger benchmark coverage is welcome.
  • Refusal behavior is prompt- and decoder-sensitive.
  • The three MTP/DSpark stages inherit the deepest measured backbone subspace because the capture run did not instantiate speculative decoding.
  • This is a behavioral weight edit. Evaluate it for your own use case before deployment.

License and attribution

The upstream repository and model weights are MIT licensed. This derivative retains the MIT license and original DeepSeek copyright notice.

Base model: deepseek-ai/DeepSeek-V4-Flash-0731

Citation

@misc{deepseekai2026deepseekv4,
  title={DeepSeek-V4: Towards Highly Efficient Million-Token Context Intelligence},
  author={DeepSeek-AI},
  year={2026}
}

README history 1 version

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-08-02Duplicate from squanchyzx/DeepSeek-V4-Flash-0731-HERETIC-Abliterated-FP829b45d75.9 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration