← back to catalog · registered 2026-08-22 13:56

squanchyzx/DeepSeek-V4-Flash-0731-HERETIC-Abliterated-FP8

squanchyzx Deepseek 296B
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/squanchyzx%2FDeepSeek-V4-Flash-0731-HERETIC-Abliterated-FP8"
Response includes
  • classification m3
  • files 61
  • hub_downloads_all_time 5,980
  • author_summary 1 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M3
Primary method

Layer-wise ablation

Applied on top of direct removal inherited from the base model.
Confidence
HIGH
Inherited from base model
Why this label 2 signals
Producer identity confirmed by naming conventions, tags or the model card. This label is very unlikely to change.
  • 'heretic' in model name (Heretic-produced)
  • Heretic uses layer-wise optimization (M3) with underlying direction removal (M1)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
HIGH
Why we say so
name contains 'heretic'; Heretic default extraction is difference-of-means (Arditi 2024)
Downloads · lifetime
6K
663 last 30d - stable
Likes
21
Model age
2mo ago
created 2026-08-02
Downloads over time
Now6.2K→from712↑767%
4392.5K4.6K6.7K712 on Aug 56.2K on Oct 11AugSepOct
Aug 5 → Oct 11 · 51 snapshots · spans 67 days

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
mit
Tags
transformers safetensors deepseek_v4 text-generation deepseek-v4 abliterated heretic fp8 reasoning tool-use base_model:deepseek-ai/DeepSeek-V4-Flash-0731 base_model:quantized:deepseek-ai/DeepSeek-V4-Flash-0731

Related

Total size
156 GB
Files
61
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-08-12 00:25

Files by quantization

Auxiliary files 61 files 156 GB
model-00048-of-00048.safetensors 3.44 GB cc43742b download
model-00046-of-00048.safetensors 3.36 GB 5db924ca download
model-00004-of-00048.safetensors 3.35 GB 9610f56b download
model-00012-of-00048.safetensors 3.34 GB 64ed4e5f download
model-00014-of-00048.safetensors 3.34 GB 45db2f54 download
model-00016-of-00048.safetensors 3.34 GB e0530b70 download
model-00018-of-00048.safetensors 3.34 GB e393fea9 download
model-00020-of-00048.safetensors 3.34 GB 9f556769 download
model-00022-of-00048.safetensors 3.34 GB decd67a4 download
model-00024-of-00048.safetensors 3.34 GB fc27aeb4 download
model-00026-of-00048.safetensors 3.34 GB 657b8931 download
model-00028-of-00048.safetensors 3.34 GB b2fd5cbb download
model-00030-of-00048.safetensors 3.34 GB 9ed3c317 download
model-00032-of-00048.safetensors 3.34 GB 16365384 download
model-00034-of-00048.safetensors 3.34 GB 0f949451 download
model-00036-of-00048.safetensors 3.34 GB 7e676142 download
model-00038-of-00048.safetensors 3.34 GB 137fa617 download
model-00040-of-00048.safetensors 3.34 GB 8bc93d8a download
model-00042-of-00048.safetensors 3.34 GB 4d19bf36 download
model-00044-of-00048.safetensors 3.34 GB 422d3889 download
model-00006-of-00048.safetensors 3.34 GB 4a4f3764 download
model-00008-of-00048.safetensors 3.34 GB 224968d2 download
model-00010-of-00048.safetensors 3.34 GB 627145f4 download
model-00013-of-00048.safetensors 3.32 GB 8dfe199d download
model-00015-of-00048.safetensors 3.32 GB 5810381a download
model-00017-of-00048.safetensors 3.32 GB ed111302 download
model-00019-of-00048.safetensors 3.32 GB a74ca4d3 download
model-00021-of-00048.safetensors 3.32 GB 1671cce7 download
model-00023-of-00048.safetensors 3.32 GB c61a3e17 download
model-00025-of-00048.safetensors 3.32 GB a66b6b8d download
model-00027-of-00048.safetensors 3.32 GB fb01f21a download
model-00029-of-00048.safetensors 3.32 GB 9ec2fdf9 download
model-00031-of-00048.safetensors 3.32 GB d5078c3f download
model-00033-of-00048.safetensors 3.32 GB f2cffd43 download
model-00035-of-00048.safetensors 3.32 GB 9cb6a316 download
model-00037-of-00048.safetensors 3.32 GB a59d662f download
model-00039-of-00048.safetensors 3.32 GB a29af1aa download
model-00041-of-00048.safetensors 3.32 GB fd312e7f download
model-00043-of-00048.safetensors 3.32 GB b7103842 download
model-00005-of-00048.safetensors 3.32 GB f87a5ac7 download
model-00007-of-00048.safetensors 3.32 GB df81bb80 download
model-00009-of-00048.safetensors 3.32 GB 04d69ef1 download
model-00011-of-00048.safetensors 3.32 GB e4b8e601 download
model-00002-of-00048.safetensors 3.32 GB 77b26c93 download
model-00003-of-00048.safetensors 3.32 GB 412abf4c download
model-00047-of-00048.safetensors 3.32 GB 62816173 download
model-overlay-00001-of-00001.safetensors 1.03 GB ca4a043a download
model-00045-of-00048.safetensors 1010 MB a5be6aed download
model-00001-of-00048.safetensors 1010 MB f3668ba4 download
tokenizer.json 6.07 MB 628e3364 download
model.safetensors.index.json 5.34 MB 231e0f43 download
abliteration_report.json 25.8 KB 429f6d8c download
README.md 8.75 KB ac65969b download
EVAL_RESULTS.md 6.96 KB e3cf3346 download
RELEASE_MANIFEST.json 3.32 KB 9500e3d2 download
config.json 1.84 KB 5f2da910 download
.gitattributes 1.48 KB a6344aac download
HERETIC_ATTRIBUTION.md 1.31 KB 7c1fb09d download
LICENSE 1.04 KB d84f527e download
tokenizer_config.json 801 B f3dad388 download
generation_config.json 170 B c56a8c5b download

README current version from Hugging Face


license: mit
library_name: transformers
pipeline_tag: text-generation
base_model: deepseek-ai/DeepSeek-V4-Flash-0731
tags:

  • deepseek-v4
  • abliterated
  • heretic
  • fp8
  • reasoning
  • tool-use

DeepSeek-V4-Flash-0731 HERETIC Abliterated FP8 — v2

A native-FP8 behavioral derivative of deepseek-ai/DeepSeek-V4-Flash-0731.

v2 replaces the original λ1.5 release on main. The original commit remains available as an immutable tagged revision because it developed a long-context degeneration failure in real agent histories. Do not deploy the old revision for long-lived agent sessions.

What was wrong with v1

The issue was not a single repeated word. In long tool/agent histories, v1 could fall into several degeneration patterns:

  • repeated words such as kanka, my, or similar tokens;
  • repeated CJK characters and unexpected script switching;
  • thousands of empty () fragments or repeated punctuation;
  • repeated lines and low-entropy n-gram loops;
  • 100K–200K-character malformed generations that poisoned subsequent history.

The failure came from an overly aggressive rank-3 edit at lambda=1.5, combined with speculative/MTP amplification. Merely restoring stock MTP was not enough: the λ1.5 stock-MTP control still produced a 324-character consecutive 我 run.

What changed in v2

  • Kept the same three-mode rank-3 refusal subspace for chat, think-high, and think-max.
  • Edited attention output projections in backbone layers 10–42.
  • Reduced the projection strength from lambda=1.5 to lambda=1.35.
  • Left all MTP/DSpark tensors stock and untouched.
  • Left routed experts, shared experts, embeddings, norms, heads, tokenizer, encoder, and model configuration untouched.
  • Re-baked once from the protected official base; v2 was not baked on top of v1.
  • Preserved the official native FP8 representation; no additional lower-bit quantization pass was applied.

Tool attribution

This checkpoint was produced with Heretic v1.4.0 by Philipp Emanuel Weidmann and contributors, using base commit 7675b90d648154cdfefa597372cc477df9848eab. Heretic is their project and is licensed AGPL-3.0-or-later; it is not owned or authored by this model publisher.

The release-specific work here is the DeepSeek V4/mHC two-node adaptation, separate refusal-direction capture, rank-3 subspace construction, native-FP8 attention-only bake, long-context incident reproduction, generic degeneration detector, and release evaluation. See HERETIC_ATTRIBUTION.md.

Mechanical receipts

  • Edited tensors: 66 (33 weight + scale pairs)
  • Untouched indexed tensors: 72,251
  • MTP/DSpark edited tensors: 0
  • Overlay bytes: 1,107,370,672
  • Overlay SHA-256: ca4a043ae3a306a50b680a08c3f175d7a7df20b955f28a9d4e45161692c45638
  • Rewritten index SHA-256: 393a1b09ffacdf5e4e226a84f0b8d5467947ad0f35c91a71127e609003375a56
  • Cross-node overlay/index equality: verified on two independently materialized DGX Spark trees
  • Dangling official-base shard links before packaging: 0/48 on both nodes

Long-context acceptance

The detector does not key on one literal word. It evaluates answer and reasoning content for same-word/character/symbol runs, repeated 2/3/4/8-grams, repeated lines, dominant-token fraction, vocabulary uniqueness, entropy, compression, empty-parenthesis floods, and unexpected script switching.

Results with speculative decoding enabled:

  • Real cron incident fixture, 92,501 prompt tokens: 5/5 seeds clean
  • Corrected real WhatsApp fixture, 64,245 prompt tokens: 5/5 seeds clean
  • Synthetic 124,983 prompt-token replay: clean
  • Synthetic 249,971 prompt-token replay: clean
  • Total long-history acceptance replays: 12/12 clean
  • Non-stop finishes: 0/12
  • Tool calls: 6/6 valid and correct
  • Deterministic short quality smoke with a 512-token generation budget: 11/12
  • Refusal smoke: 6/16 (37.5%; lower is less refusal)

These are bounded release tests, not a universal capability or safety benchmark. Full compact receipts are in EVAL_RESULTS.md and the eval/v2/ directory.

Runtime

Use a DeepSeek-V4-compatible runtime and the official encoder. This checkpoint retains the official V4/DSpark architecture and native FP8 layout. Generic runtimes without DeepSeek V4 support may not load it correctly.

The release candidate was exercised on two NVIDIA DGX Spark nodes with tensor parallelism 2, a 1,048,576-token configured context window, the official V4 encoder, and a DeepSeek-V4-capable vLLM build.

2026-08-12 runtime qualification and stock-vs-v2 A/B

No checkpoint revision was created for this update. The v2 weight shards, index, configuration, tokenizer, and immutable v2-lam1p35-mtpstock tag are unchanged. This section documents a verified runtime profile and a bounded matched comparison against the unmodified official DeepSeek-V4-Flash-0731 checkpoint.

Runtime profile

The qualified two-DGX-Spark deployment uses the runtime lineage from MiaAI-Lab/DeepSeek-v4-Flash-DSpark-2x-DGX-Spark at commit 0817b11f5bfd53aeb85c90d1e8ae3755898b5314. That repository supplies runtime/deployment fixes; it is not a separate Mia-tuned checkpoint.

The validated profile kept this repository's v2 weights and used TP=2, MTP=3, a 1,048,576-token configured window, the official encoder, regular CUDA graphs, and the NVFP4 sparse-MLA fix. Direct API gates passed 5/5, including strict JSON and tool-history re-encoding. Long-context gates completed at 66,637, 249,971, and 609,847 prompt tokens without OOM, engine death, traceback, HTTP 5xx, or container restart. The 609,847-token run decoded at 38.768 tok/s, but its prefill/TTFT was 541.203 s; the configured 1M window should not be presented as low-latency interactive UX.

The qualified deployment currently uses host-mounted patch files. An immutable runtime image and mount-independence receipt remain future work.

Matched stock-vs-v2 result

Both arms were cold-loaded into the same TP=2 runtime envelope with identical serving flags, tokenizer, prompts, seeds, sampling, MTP=3, and default thinking=false chat-template behavior.

  • Strict objective/instruction suite: official stock 14/20; Heretic v2 11/20.
  • Stock-only wins: 3; Heretic-only wins: 0; exact paired McNemar two-sided p=0.25.
  • Tool selection: both 6/6.
  • Deterministic 512-token quality smoke: both 11/12.
  • Legacy bounded short-generation quality smoke: stock 10/12; Heretic v2 9/12.
  • Refusal smoke: stock 15/16; Heretic v2 6/16 (lower means less refusal).
  • Blind qualitative majority (3 judges, 12 randomized pairs): stock 6, Heretic v2 2, ties 4; mean scores stock 7.500, Heretic v2 7.417.

Interpretation: stock showed a directional edge on this small strict objective suite, while v2 preserved much lower refusal behavior and equal tool-call compliance. The sample is too small to establish a universal capability gap. Use the official stock checkpoint when upstream-aligned behavior and maximum baseline correctness are the priority; evaluate this derivative when reduced refusal behavior is an explicit requirement. Do not treat abliteration as a free quality improvement.

Machine-readable receipts: matched-ab-summary.json, blind-judge-summary.json, suite-manifest.json, and runtime-qualification.json.

Revision policy

  • main: current v2 release
  • v2-lam1p35-mtpstock: immutable v2 tag
  • v1-lam1p5-known-long-context-issue: preserved v1 tag; not recommended for long-lived agent sessions

Limitations and safety

  • This is an unofficial community derivative and is not affiliated with or endorsed by DeepSeek.
  • Abliteration reduces some refusal behavior and can increase compliance with unsafe requests. This checkpoint is not safety-certified.
  • Downstream deployers remain responsible for evaluation, access controls, monitoring, output limits, loop guards, and compliance with applicable law and platform policy.
  • Refusal and long-context behavior remain prompt-, template-, runtime-, and decoder-sensitive.
  • The quality suite is intentionally small. Broader benchmark coverage is welcome.
  • A model-side repair does not replace runtime output caps and generic degeneration guards.

License

The upstream repository and model weights are MIT licensed. This derivative retains the MIT license and original DeepSeek copyright notice.

Base model: deepseek-ai/DeepSeek-V4-Flash-0731

README history 5 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-08-12Document runtime qualification and matched stock-v2 A/Bc2f27968.8 KB
    Loading...
  2. 2026-08-08Publish v2 long-context repair (lambda 1.35, stock MTP)9686b175.6 KB
    Loading...
  3. 2026-08-02Credit Heretic tool and separate model/code licenses2b84b835.9 KB
    Loading...
  4. 2026-08-02Update README.md release disclosure and evaluation554aedc5.1 KB
    Loading...
  5. 2026-08-02Add model card, method scripts, and evaluation receipts5f4ca794.8 KB
    Loading...

Discussions 2 threads

  1. 2026-08-08Heretic Araopen2 💬#2
    Loading...
  2. 2026-08-08Thinking severely affectedopen8 💬#1
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration