← back to catalog · registered 2026-08-22 13:56

gorbatjovy/DeepSeek-V4-Flash-0731-BahamutRU-t265-abliterated

gorbatjovy Deepseek 296B
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/gorbatjovy%2FDeepSeek-V4-Flash-0731-BahamutRU-t265-abliterated"
Response includes
  • classification m1
  • files 57
  • hub_downloads_all_time 139
  • author_summary 5 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M1
Primary method

Direct removal

No other method signals detected in this model.
Confidence
MEDIUM
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=0 (base model)
  • no specific method indicators - defaulting to M1 (most common)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
MEDIUM
Why we say so
primary_method=M1; difference-of-means is the reference extraction for M1/M3 (Arditi 2024)
Downloads · lifetime
139
49 last 30d - stable
Likes
0
Model age
8w ago
created 2026-08-16
Downloads over time
Now158→from41↑285%
358012517041 on Aug 19158 on Oct 11AugSepOct
Aug 19 → Oct 11 · 48 snapshots · spans 53 days

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
mit
Tags
safetensors deepseek_v4 deepseek deepseek-v4 abliterated uncensored heretic fp8 base_model:deepseek-ai/DeepSeek-V4-Flash-0731 base_model:quantized:deepseek-ai/DeepSeek-V4-Flash-0731 license:mit 8-bit

Related

Total size
157 GB
Files
57
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-08-17 07:08

Files by quantization

Auxiliary files 57 files 157 GB
model-00048-of-00048.safetensors 3.44 GB cc43742b download
model-00046-of-00048.safetensors 3.36 GB 5db924ca download
model-00004-of-00048.safetensors 3.35 GB 9610f56b download
model-00012-of-00048.safetensors 3.34 GB 64ed4e5f download
model-00014-of-00048.safetensors 3.34 GB 45db2f54 download
model-00016-of-00048.safetensors 3.34 GB e0530b70 download
model-00018-of-00048.safetensors 3.34 GB e393fea9 download
model-00020-of-00048.safetensors 3.34 GB 9f556769 download
model-00022-of-00048.safetensors 3.34 GB decd67a4 download
model-00024-of-00048.safetensors 3.34 GB fc27aeb4 download
model-00026-of-00048.safetensors 3.34 GB 657b8931 download
model-00028-of-00048.safetensors 3.34 GB b2fd5cbb download
model-00030-of-00048.safetensors 3.34 GB 9ed3c317 download
model-00032-of-00048.safetensors 3.34 GB 16365384 download
model-00034-of-00048.safetensors 3.34 GB 0f949451 download
model-00036-of-00048.safetensors 3.34 GB 7e676142 download
model-00038-of-00048.safetensors 3.34 GB 137fa617 download
model-00040-of-00048.safetensors 3.34 GB 8bc93d8a download
model-00042-of-00048.safetensors 3.34 GB 4d19bf36 download
model-00044-of-00048.safetensors 3.34 GB 422d3889 download
model-00006-of-00048.safetensors 3.34 GB 4a4f3764 download
model-00008-of-00048.safetensors 3.34 GB 224968d2 download
model-00010-of-00048.safetensors 3.34 GB 627145f4 download
model-00013-of-00048.safetensors 3.32 GB 8dfe199d download
model-00015-of-00048.safetensors 3.32 GB 5810381a download
model-00017-of-00048.safetensors 3.32 GB ed111302 download
model-00019-of-00048.safetensors 3.32 GB a74ca4d3 download
model-00021-of-00048.safetensors 3.32 GB 1671cce7 download
model-00023-of-00048.safetensors 3.32 GB c61a3e17 download
model-00025-of-00048.safetensors 3.32 GB a66b6b8d download
model-00027-of-00048.safetensors 3.32 GB fb01f21a download
model-00029-of-00048.safetensors 3.32 GB 9ec2fdf9 download
model-00031-of-00048.safetensors 3.32 GB d5078c3f download
model-00033-of-00048.safetensors 3.32 GB f2cffd43 download
model-00035-of-00048.safetensors 3.32 GB 9cb6a316 download
model-00037-of-00048.safetensors 3.32 GB a59d662f download
model-00039-of-00048.safetensors 3.32 GB a29af1aa download
model-00041-of-00048.safetensors 3.32 GB fd312e7f download
model-00043-of-00048.safetensors 3.32 GB b7103842 download
model-00005-of-00048.safetensors 3.32 GB f87a5ac7 download
model-00007-of-00048.safetensors 3.32 GB df81bb80 download
model-00009-of-00048.safetensors 3.32 GB 04d69ef1 download
model-00011-of-00048.safetensors 3.32 GB e4b8e601 download
model-00002-of-00048.safetensors 3.32 GB 77b26c93 download
model-00003-of-00048.safetensors 3.32 GB 412abf4c download
model-00047-of-00048.safetensors 3.32 GB 62816173 download
model-overlay-bahamut-00001-of-00001.safetensors 1.20 GB eec6e22d download
model-00045-of-00048.safetensors 1010 MB a5be6aed download
model-00001-of-00048.safetensors 1010 MB f3668ba4 download
tokenizer.json 6.07 MB 628e3364 download
model.safetensors.index.json 5.35 MB 2061986a download
bahamut_lora_report.json 42.3 KB 14fa976b download
README.md 5.72 KB 8e06017d download
config.json 1.84 KB 5f2da910 download
.gitattributes 1.48 KB a6344aac download
tokenizer_config.json 801 B f3dad388 download
generation_config.json 170 B c56a8c5b download

README current version from Hugging Face


license: mit
base_model:

  • deepseek-ai/DeepSeek-V4-Flash-0731
    tags:
  • deepseek
  • deepseek-v4
  • abliterated
  • uncensored
  • heretic
  • fp8

DeepSeek-V4-Flash-0731-BahamutRU-t265-abliterated

deepseek-ai/DeepSeek-V4-Flash-0731 with the refusal direction ablated, using the rank-1
adapter from
BahamutRU/DeepSeek-V4-Flash-0731-heretic-abliterated-v2-GGUF-lora
baked directly into the FP8 weights.

There is no LoRA at runtime. The edit is merged into the tensors; load it like any other
checkpoint.

What was changed

57 target tensors (114 counting block scales):

  • layers.{11..42}.attn.wo_b — attention output projections
  • layers.{18..42}.ffn.shared_experts.w2 — shared-expert down-projections

Everything else — all 72,203 other indexed tensors, including every routed expert —
is byte-identical to the official DeepSeek release. The 48 original shards are unmodified;
the edited tensors live in a single overlay file that the index redirects to.

Method

Directional ablation is inherently rank-1: W <- W - lam * r (r^T W) is an outer product, so
the GGUF adapter is the edit in closed form. Verified against the base weights before
applying — cos(lora_a, r^T W) is -0.995 / -0.999 / -0.943 at layers 30 / 20 / 42, with
implied lambda 4.81 / 2.67 / 2.39. All 32 attention layers share one direction
(min |cos| 0.994), matching the adapter's declared global direction scope.

The delta is applied verbatim rather than re-derived as a projection, since the cosine is
near but not exactly -1 — reproducing what the adapter author measured rather than an
approximation of it.

Dequantise FP8 e4m3 (ue8m0 128x128 block scales) -> add lora_b (x) lora_a -> requantise,
raising a block's exponent where the larger delta would otherwise clip. Zero elements were
clamped; peak overshoot 1.095x.

Lambda ranges 0.72–4.88 across layers, peaking at layer 31. Above 2.0 the refusal component
is inverted rather than merely removed.

Not included

The adapter also ablates routed-expert down-projections (256 experts x 31 layers). Those are
FP4 and were left untouched here, so this is a partial application of t265.

Evaluation

Measured 2026-08-16/17 on 2x DGX Spark (vLLM, TP=2, FP8 KV cache) against the unabliterated
base, using identical prompts, seeds and serving configuration.

benchmark                            base    this model
--------------------------------------------------------
capability (higher is better)
  IFEval  (prompt strict)           83.4%        82.6%
  HumanEval+                        87.8%        88.4%
  MBPP+                             73.3%        73.3%
  MMLU-Pro                          78.9%        77.9%
  tool calling: correct tool       100.0%       100.0%
  tool calling: correct args       100.0%       100.0%
  tool calling: false positives      0.0%         0.0%

behavioural drift from base (lower is better)
  mean first-token KL                  —       0.1117
  median first-token KL                —       0.0120
  p90 first-token KL                   —       0.3761
  top-1 token agreement                —        91.2%

long-form generation
  mean distinct-3                 0.9834       0.9833
  creative type-token ratio       0.3392       0.3583

Capability is unaffected. Instruction following, code generation, knowledge and tool
calling all land within noise of the base model.

Drift from base is small. First-token KL divergence on harmless prompts is the most
direct measure of what an abliteration costs — it is also what heretic-gguf optimises
against, jointly with refusal rate. This checkpoint measures 0.1117 mean / 0.0120 median and
still picks the base model's most likely first token on 91.2% of harmless prompts.

Long-form generation is intact, which matters for extended chat: lexical diversity and
repetition behaviour match base, with no degeneration despite λ reaching 4.88.

Caveats

  • These KL values are not comparable to the adapter author's published 0.065. Different
    prompt set (80 mundane factual, coding and creative prompts), different top-k handling.
    Absolute values only mean something relative to the same measurement on the same set.
  • MMLU-Pro is underpowered here: 280 samples (20 × 14 subtasks), ±2.4%, so it can only
    detect gaps of roughly 5 points.
  • HumanEval and MBPP are contaminated. Valid as a regression signal against the same
    base, not as capability claims.
  • Tool calling used 24 samples, of which 3 are negative cases (a question needing no
    tool). Enough to show nothing is obviously broken; not enough to characterise the tail.
  • KL covers first tokens only — a sensitive early indicator, not a full account of drift.
  • All runs had thinking disabled, because third-party harnesses cannot parse
    reasoning_content or budget tokens for it. Both checkpoints were measured identically,
    but these are not thinking-mode scores.
  • Refusal rate is not reported here. The adapter author measured 12/140 on their harmful
    set; measuring it properly needs a prompt set this evaluation did not include.

Credit

Warning

Refusal behaviour has been deliberately removed. This model will attempt whatever it is
asked. Safety judgement is entirely the operator's responsibility.

README history 3 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-08-17Upload README.md with huggingface_hub7168e0b5.7 KB
    Loading...
  2. 2026-08-17Upload README.md with huggingface_hub5505a526.2 KB
    Loading...
  3. 2026-08-16Upload folder using huggingface_hub2c07b3e2.8 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration