license: mit
base_model: windowsxp811203/DeepSeek-V4-Flash-0731-Abliterated
base_model_relation: quantized
pipeline_tag: text-generation
library_name: gguf
tags:
- gguf
- llama.cpp
- deepseek4
- abliterated
- uncensored
- moe
language: - en
- zh
extra_gated_prompt: >-
This model has had its refusal behaviour removed. You are responsible for how you use it
and for complying with all applicable laws.
extra_gated_fields:
I understand this model will not refuse harmful requests: checkbox
I accept responsibility for my use of this model: checkbox
DeepSeek-V4-Flash-0731-Abliterated — GGUF
llama.cpp builds of windowsxp811203/DeepSeek-V4-Flash-0731-Abliterated,
an abliterated (refusal-removed) DeepSeek-V4-Flash-0731. 43 layers, 256 routed experts (6 active),
1,048,576 native context.
| build | total size | shards | PPL ↓ | MMLU (400) | refusal |
|---|---|---|---|---|---|
MXFP4 (recommended) |
156.4 GB | 4 × ≤44.5 GB | 4.105 | 77.00 % | 1/120 · 0.83 % |
Q3_K_M |
135.3 GB | 4 × ≤44.9 GB | 4.606 | 75.75 % | 0/120 · 0.00 % |
Both are split into shards because HuggingFace caps individual files at 50 GB. You do not need
to merge them: point llama.cpp at shard 00001 and it loads the rest automatically (verified —
server ready in 22 s from the first shard alone). Download the whole set for a build before running.
Both need multi-GPU or heavy CPU offload; verified across 3× H200.
Read this before picking a file
The MXFP4 file is not an unquantized intermediate — it is the faithful conversion, and it is
already smaller than the original checkpoint (156.4 GB vs 166.9 GB on the Hub). The source is FP8
with 128×128 block scales, and llama.cpp maps the expert tensors to MXFP4 directly:
| type | tensors | bytes | share |
|---|---|---|---|
| MXFP4 (MoE experts) | 129 | 147.2 GB | 94.1 % |
| Q8_0 | 365 | 6.2 GB | 4.0 % |
| BF16 | 190 | 2.8 GB | 1.8 % |
| F32 | 641 | 0.1 GB | 0.1 % |
Because 94 % of the weight is already 4-bit, the usual ladder makes no sense here: Q4_K_M, Q5_K_M
and Q6_K would all be larger than this file while adding nothing, since they would re-expand
4-bit experts to 4.8–6.6 bits. Only Q3 and below actually shrink it. That is why this repo ships
two files instead of five.
Q2_K exists and is deliberately not published
It was built and measured. It is not here because it is broken in a way that is easy to miss:
| build | PPL | MMLU | verdict |
|---|---|---|---|
| MXFP4 | 4.105 | 77.00 % | ship |
| Q3_K_M | 4.606 | 75.75 % | ship |
| Q2_K (103.1 GB) | 8.323 | 23.25 % | withheld — not uploaded |
MMLU chance for 4 choices is 25 %. Q2_K scores 23.25 % (93/400, 95 % CI 19.1–27.4 %) — that is
statistically indistinguishable from random guessing, not measurably below it. Either way, a
build that answered 77 % of these questions at MXFP4 now answers them no better than a coin-flip
generator: its knowledge is gone.
What makes this worth spelling out is which signals missed it. Q2_K still reads fine — asked
why the sky is blue it explains Rayleigh scattering, asked to reverse a linked list it writes valid
Python, asked whether 9.11 or 9.9 is larger it answers correctly and shows the conversion. And its
refusal rate is a perfect 0/120, so the abliteration survives 2-bit fully intact.
Fluency and refusal rate both looked healthy on a model whose knowledge had collapsed. Perplexity
did flag it — 8.323 against 4.606 for Q3_K_M, an 81 % jump and far outside the ±0.108 error bar — so
PPL was a correct early warning. But PPL only says "much worse"; it took the capability benchmark to
show how much worse, and that the remaining ability was at chance. Refusal rate said nothing at
all.
So the quality cliff for this checkpoint sits between Q3 and Q2 — a collapse
(75.75 % → 23.25 %) rather than a slope. With only three points measured, that brackets the cliff;
it does not locate it precisely.
Usage
# grab one complete build (all four shards)
hf download windowsxp811203/DeepSeek-V4-Flash-0731-Abliterated-GGUF \
--include "*MXFP4*" --local-dir .
# 3x H200 / 2x 96GB, or fewer cards with CPU offload.
# Point at shard 1 only — llama.cpp finds 00002..00004 beside it.
llama-server -m DeepSeek-V4-Flash-0731-Abliterated-MXFP4-00001-of-00004.gguf -ngl 99 -c 8192
Thinking is on by default; disable per request with "chat_template_kwargs": {"thinking": false},
or server-wide with --chat-template-kwargs '{"thinking":false}'. Note the key is thinking,
not enable_thinking.
Known limitations
- No MTP draft model. The checkpoint carries three
mtp.Nweight namespaces
(mtp.0/1/2, ~1,570 tensors each, 4,705 total, each with its ownattn.wo_b), while the config
declaresnum_nextn_predict_layers: 1— the weights are three blocks, the declared draft depth is
one. llama.cpp's DeepSeek-V4 converter skips MTP by default and--mtpaborts withUnexpected DeepSeek-V4 MTP layer 1— its expected MTP layout does not match
this checkpoint. Speculative decoding with the built-in draft head is therefore not available
through llama.cpp today. - No imatrix. An importance matrix was computed, but
llama-quantizerejects it on this
architecture (imatrix size 32768 is different from tensor size 4096 for blk.0.attn_output_a.weight).Q3_K_Mis therefore a plain quantization; an imatrix build would
likely score slightly better. - Requantizing needs
--allow-requantize, since most source tensors are already MXFP4/Q8_0.
Method
Refusal: AdvBench, greedy, non-thinking, no prefill jailbreak, n=120 per build, zero truncated
or errored samples in every run. PPL: held-out wikitext-2 test, 40 chunks @ -c 2048.
MMLU: 400 equidistant questions, single letter answer, identical prompting for all builds.
The parent model's recipe (rank-1 projection on attn.wo_b across layers 10–42 plus the threemtp.N blocks — 36 tensors in total per its ABLIT_META.json — at λ=3.5) is documented in the
parent card.
Support / 打賞
If these models are useful to you, tips are appreciated — they pay for the GPU time.
如果這些模型對你有幫助,歡迎打賞,用於支應算力成本。
USDT (TRC20) · TPTo32r7vKazpTNaFqfFZ2ztoK1DG88888
Disclaimer
This model will not refuse. It is published for alignment and safety research. You are responsible
for your use of it and for complying with applicable law. Inherits the MIT license of the base model.