← back to catalog · registered 2026-08-22 13:56

windowsxp811203/DeepSeek-V4-Flash-0731-Abliterated-GGUF

windowsxp811203 Deepseek GGUF MoE second-order 1.0M ctx
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/windowsxp811203%2FDeepSeek-V4-Flash-0731-Abliterated-GGUF"
Response includes
  • classification m8
  • files 10
  • hub_downloads_all_time 860
  • author_summary 17 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M8
Primary method

Repackaging (quantization)

Applied on top of direct removal inherited from the base model.
Confidence
MEDIUM
Inherited from base model
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=1
  • assume M1 (base ablation) + M8 (GGUF quant) - default when producer unknown
Refusal direction extraction

No specific extraction method could be identified for this model. The producer either did not document it or used a proprietary pipeline.

What is a refusal direction? →
Downloads · lifetime
860
104 last 30d - stable
Likes
3
Model age
8w ago
created 2026-08-16
Downloads over time
Now886→from128↑592%
90381671962128 on Aug 19886 on Oct 11886 on Oct 8AugSepOct
Aug 19 → Oct 11 · 48 snapshots · spans 53 days

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Variants by this author 2 formats · 211 downloads combined

The same weights this author released in different packaging. Pick the format that matches your runtime.

Metadata

License
mit
Languages
en zh
Quantizations
Q3_K
Tags
gguf llama.cpp deepseek4 abliterated uncensored moe text-generation en zh base_model:windowsxp811203/DeepSeek-V4-Flash-0731-Abliterated base_model:quantized:windowsxp811203/DeepSeek-V4-Flash-0731-Abliterated license:mit

Related

Total size
272 GB
Files
10
Quantizations
2
Registered
2026-08-22 13:56
Last updated on HF
2026-08-17 05:16

Files by quantization

Q3_K 4 files 126 GB
DeepSeek-V4-Flash-0731-Abliterated-Q3_K_M-00001-of-00004.gguf 41.8 GB ******** download
DeepSeek-V4-Flash-0731-Abliterated-Q3_K_M-00002-of-00004.gguf 41.8 GB ******** download
DeepSeek-V4-Flash-0731-Abliterated-Q3_K_M-00003-of-00004.gguf 41.5 GB ******** download
DeepSeek-V4-Flash-0731-Abliterated-Q3_K_M-00004-of-00004.gguf 888 MB ******** download
Auxiliary files 6 files 146 GB
DeepSeek-V4-Flash-0731-Abliterated-MXFP4-00001-of-00004.gguf 41.5 GB ******** download
DeepSeek-V4-Flash-0731-Abliterated-MXFP4-00002-of-00004.gguf 41.4 GB ******** download
DeepSeek-V4-Flash-0731-Abliterated-MXFP4-00003-of-00004.gguf 41.4 GB ******** download
DeepSeek-V4-Flash-0731-Abliterated-MXFP4-00004-of-00004.gguf 21.3 GB ******** download
README.md 6.55 KB 11e6cc24 download
.gitattributes 2.25 KB 927e7726 download

README current version from Hugging Face


license: mit
base_model: windowsxp811203/DeepSeek-V4-Flash-0731-Abliterated
base_model_relation: quantized
pipeline_tag: text-generation
library_name: gguf
tags:

  • gguf
  • llama.cpp
  • deepseek4
  • abliterated
  • uncensored
  • moe
    language:
  • en
  • zh
    extra_gated_prompt: >-
    This model has had its refusal behaviour removed. You are responsible for how you use it
    and for complying with all applicable laws.
    extra_gated_fields:
    I understand this model will not refuse harmful requests: checkbox
    I accept responsibility for my use of this model: checkbox

DeepSeek-V4-Flash-0731-Abliterated — GGUF

llama.cpp builds of windowsxp811203/DeepSeek-V4-Flash-0731-Abliterated,
an abliterated (refusal-removed) DeepSeek-V4-Flash-0731. 43 layers, 256 routed experts (6 active),
1,048,576 native context.

build total size shards PPL ↓ MMLU (400) refusal
MXFP4 (recommended) 156.4 GB 4 × ≤44.5 GB 4.105 77.00 % 1/120 · 0.83 %
Q3_K_M 135.3 GB 4 × ≤44.9 GB 4.606 75.75 % 0/120 · 0.00 %

Both are split into shards because HuggingFace caps individual files at 50 GB. You do not need
to merge them: point llama.cpp at shard 00001 and it loads the rest automatically (verified —
server ready in 22 s from the first shard alone). Download the whole set for a build before running.

Both need multi-GPU or heavy CPU offload; verified across 3× H200.

Read this before picking a file

The MXFP4 file is not an unquantized intermediate — it is the faithful conversion, and it is
already smaller than the original checkpoint
(156.4 GB vs 166.9 GB on the Hub). The source is FP8
with 128×128 block scales, and llama.cpp maps the expert tensors to MXFP4 directly:

type tensors bytes share
MXFP4 (MoE experts) 129 147.2 GB 94.1 %
Q8_0 365 6.2 GB 4.0 %
BF16 190 2.8 GB 1.8 %
F32 641 0.1 GB 0.1 %

Because 94 % of the weight is already 4-bit, the usual ladder makes no sense here: Q4_K_M, Q5_K_M
and Q6_K would all be larger than this file
while adding nothing, since they would re-expand
4-bit experts to 4.8–6.6 bits. Only Q3 and below actually shrink it. That is why this repo ships
two files instead of five.

Q2_K exists and is deliberately not published

It was built and measured. It is not here because it is broken in a way that is easy to miss:

build PPL MMLU verdict
MXFP4 4.105 77.00 % ship
Q3_K_M 4.606 75.75 % ship
Q2_K (103.1 GB) 8.323 23.25 % withheld — not uploaded

MMLU chance for 4 choices is 25 %. Q2_K scores 23.25 % (93/400, 95 % CI 19.1–27.4 %) — that is
statistically indistinguishable from random guessing, not measurably below it. Either way, a
build that answered 77 % of these questions at MXFP4 now answers them no better than a coin-flip
generator: its knowledge is gone.

What makes this worth spelling out is which signals missed it. Q2_K still reads fine — asked
why the sky is blue it explains Rayleigh scattering, asked to reverse a linked list it writes valid
Python, asked whether 9.11 or 9.9 is larger it answers correctly and shows the conversion. And its
refusal rate is a perfect 0/120
, so the abliteration survives 2-bit fully intact.

Fluency and refusal rate both looked healthy on a model whose knowledge had collapsed. Perplexity
did flag it — 8.323 against 4.606 for Q3_K_M, an 81 % jump and far outside the ±0.108 error bar — so
PPL was a correct early warning. But PPL only says "much worse"; it took the capability benchmark to
show how much worse, and that the remaining ability was at chance. Refusal rate said nothing at
all.

So the quality cliff for this checkpoint sits between Q3 and Q2 — a collapse
(75.75 % → 23.25 %) rather than a slope. With only three points measured, that brackets the cliff;
it does not locate it precisely.

Usage

# grab one complete build (all four shards)
hf download windowsxp811203/DeepSeek-V4-Flash-0731-Abliterated-GGUF \
  --include "*MXFP4*" --local-dir .

# 3x H200 / 2x 96GB, or fewer cards with CPU offload.
# Point at shard 1 only — llama.cpp finds 00002..00004 beside it.
llama-server -m DeepSeek-V4-Flash-0731-Abliterated-MXFP4-00001-of-00004.gguf -ngl 99 -c 8192

Thinking is on by default; disable per request with "chat_template_kwargs": {"thinking": false},
or server-wide with --chat-template-kwargs '{"thinking":false}'. Note the key is thinking,
not enable_thinking.

Known limitations

  • No MTP draft model. The checkpoint carries three mtp.N weight namespaces
    (mtp.0/1/2, ~1,570 tensors each, 4,705 total, each with its own attn.wo_b), while the config
    declares num_nextn_predict_layers: 1 — the weights are three blocks, the declared draft depth is
    one. llama.cpp's DeepSeek-V4 converter skips MTP by default and
    --mtp aborts with Unexpected DeepSeek-V4 MTP layer 1 — its expected MTP layout does not match
    this checkpoint. Speculative decoding with the built-in draft head is therefore not available
    through llama.cpp today.
  • No imatrix. An importance matrix was computed, but llama-quantize rejects it on this
    architecture (imatrix size 32768 is different from tensor size 4096 for blk.0.attn_output_a.weight). Q3_K_M is therefore a plain quantization; an imatrix build would
    likely score slightly better.
  • Requantizing needs --allow-requantize, since most source tensors are already MXFP4/Q8_0.

Method

Refusal: AdvBench, greedy, non-thinking, no prefill jailbreak, n=120 per build, zero truncated
or errored samples in every run. PPL: held-out wikitext-2 test, 40 chunks @ -c 2048.
MMLU: 400 equidistant questions, single letter answer, identical prompting for all builds.

The parent model's recipe (rank-1 projection on attn.wo_b across layers 10–42 plus the three
mtp.N blocks — 36 tensors in total per its ABLIT_META.json — at λ=3.5) is documented in the
parent card.

Support / 打賞

If these models are useful to you, tips are appreciated — they pay for the GPU time.
如果這些模型對你有幫助,歡迎打賞,用於支應算力成本。

USDT (TRC20) · TPTo32r7vKazpTNaFqfFZ2ztoK1DG88888

Disclaimer

This model will not refuse. It is published for alignment and safety research. You are responsible
for your use of it and for complying with applicable law. Inherits the MIT license of the base model.

Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration