license: mit
base_model:
- deepseek-ai/DeepSeek-V4-Flash-0731
base_model_relation: quantized
pipeline_tag: text-generation
library_name: gguf
quantized_by: Jared2026
tags: - gguf
- deepseek
- deepseek-v4
- deepseek-v4-flash
- deepseek4
- abliterated
- uncensored
- mixture-of-experts
- moe
- quantized
- q8
- unsloth
- llama.cpp
- conversational
DeepSeek-V4-Flash-0731-Abliterated — UD-Q8_K_XL GGUF
An 8-bit (UD-Q8_K_XL) GGUF conversion of DeepSeek-V4-Flash-0731 with a mode-merged
abliteration overlay applied to the attention output projections.
161.9 GB across 5 shards. Runs on llama.cpp with the MoE experts offloaded to system RAM,
which puts it within reach of a single workstation or one GPU node — the resident VRAM footprint
is under 20 GB.
Requires a recent llama.cpp build. The
deepseek4architecture is not supported by builds
from before ~August 2026. See Requirements — this is the single most common
reason loading fails.
Files
| File | Size | SHA256 (first 16) |
|---|---|---|
DeepSeek-V4-Flash-0731-Abliterated-UD-Q8_K_XL-00001-of-00005.gguf |
5.26 MB | 0cecd47692e23e39 |
DeepSeek-V4-Flash-0731-Abliterated-UD-Q8_K_XL-00002-of-00005.gguf |
49.2 GB | d479250fa56917c3 |
DeepSeek-V4-Flash-0731-Abliterated-UD-Q8_K_XL-00003-of-00005.gguf |
49.7 GB | df01684421c67b5f |
DeepSeek-V4-Flash-0731-Abliterated-UD-Q8_K_XL-00004-of-00005.gguf |
49.5 GB | da96ae34d4fc7bce |
DeepSeek-V4-Flash-0731-Abliterated-UD-Q8_K_XL-00005-of-00005.gguf |
13.5 GB | dd643e2e9b5c1bf1 |
Shard 1 is a small header shard — point llama.cpp at -00001-of-00005.gguf and it resolves the
rest automatically. Full checksums are in SHA256SUMS.
Also included for provenance: transplant_report.json (the exact
tensor-level record of the merge) and tensor_manifest_after.json.
Download
pip install -U "huggingface_hub[cli]"
hf download Jared2026/DeepSeek-V4-Flash-0731-Abliterated-UD-Q8_K_XL-GGUF \
--include "*.gguf" \
--local-dir ./DeepSeek-V4-Flash-Q8
Verify before loading (162 GB is a long way to get through on a bad transfer):
cd DeepSeek-V4-Flash-Q8 && sha256sum -c SHA256SUMS
Requirements
The deepseek4 architecture is new and needs llama.cpp from master, build b10273
(commit a6aa6f545) or later. Earlier builds recognise only deepseek / deepseek2 and fail at
load with an unknown-architecture error. If you have an existing llama.cpp checkout from before
August 2026, rebuild it — an in-place git pull of the source without recompiling is not enough.
git clone https://github.com/ggml-org/llama.cpp && cd llama.cpp
cmake -B build -DGGML_CUDA=ON # drop -DGGML_CUDA for CPU-only
cmake --build build --config Release -j
Hardware. Weights are 162 GB. The practical configuration is --cpu-moe, which keeps attention
and dense layers on the GPU (~19 GB) and the 256 experts in system RAM. Budget ~180 GB of RAM
plus a 20 GB-class GPU. A pure-CPU run works and needs no GPU at all, just the RAM.
Usage
llama.cpp server
llama-server \
--model DeepSeek-V4-Flash-0731-Abliterated-UD-Q8_K_XL-00001-of-00005.gguf \
--alias deepseek-v4-flash \
--host 0.0.0.0 --port 8001 \
--n-gpu-layers 99 --cpu-moe \
--flash-attn on \
--ctx-size 262144 \
--jinja \
--chat-template-kwargs '{"thinking":true,"reasoning_effort":"max"}'
Exposes an OpenAI-compatible API at http://localhost:8001/v1.
llama.cpp CLI
llama-cli \
-m DeepSeek-V4-Flash-0731-Abliterated-UD-Q8_K_XL-00001-of-00005.gguf \
--cpu-moe -ngl 99 -c 65536 --jinja
Multi-GPU
With several small GPUs (e.g. MIG slices), name them explicitly — attention and dense layers shard
across them while the experts stay on CPU:
--device CUDA0,CUDA1,CUDA2,CUDA3 --n-gpu-layers 99 --cpu-moe
Note that --n-cpu-moe N (keeping some expert layers on GPU) is a poor fit for small or
partitioned GPUs: llama.cpp places all GPU-resident expert layers on a single device, which will
OOM a 20 GB card. Full --cpu-moe is the reliable choice there.
Context length
| Tokens | |
|---|---|
| Natively trained | 65,536 |
| Maximum (YaRN, factor 16) | 1,048,576 |
The 1M figure is rope-scaled, not trained. Quality degrades progressively past the native 64K,
so treat 1M as a ceiling rather than a working context. --ctx-size 262144 (256K) is a reasonable
middle setting. KV cache cost is unusually low for the context size because the model uses MLA withhead_count_kv = 1 — 256K of context costs well under 1 GB, so context length is limited by
quality, not memory.
Reasoning effort
The chat template emits a <think> block by default. It recognises exactly two effective levels —
default, and max:
--chat-template-kwargs '{"thinking":true,"reasoning_effort":"max"}'
max injects an explicit "absolute maximum, no shortcuts" directive into the system prompt. There
is no low/medium/high gradation; any value other than max behaves as the default. You can confirm
which one is active by POSTing to the server's /apply-template endpoint and inspecting the
rendered prompt.
Architecture
Read directly from the GGUF metadata:
| Key | Value |
|---|---|
| Architecture | deepseek4 |
| Layers | 43 |
| Experts | 256 (6 active + 1 shared) |
| Expert gating | sigmoid, normalised, scale 1.5 |
| Embedding dim | 4096 |
| Attention heads | 64 (head_count_kv = 1, MLA) |
| KV compression | key/value length 512, q-LoRA rank 1024 |
| Sliding window | 128 |
| RoPE | freq base 10000, YaRN ×16 |
| Size label | 256×8.4B |
| Tokenizer | GPT-2 BPE (joyai-llm pre-tokenizer) |
| Quantization | Unsloth Dynamic UD-Q8_K_XL |
Provenance
This is a quantized redistribution, not an original model. The chain:
- Base model —
deepseek-ai/DeepSeek-V4-Flash-0731 - Quantization carrier — the Unsloth Dynamic
UD-Q8_K_XLGGUF build (general.quantized_byin
the file metadata readsUnsloth) - Abliteration overlay — a mode-merged overlay applied on top of the Q8 carrier
The merge is narrow and fully documented in transplant_report.json: 43 tensors were replaced,
one blk.N.attn_output_b.weight per layer, in BF16, sourced from an overlay with SHA25620d2559987a19cb5…. No other tensor in the file was modified — the expert weights, embeddings and
attention projections are the unmodified Q8 carrier. Three MTP/DSpark tensor groups present in the
overlay were omitted because the GGUF carrier contains no MTP tensors.
Abliteration is an inference-time behavioural edit to the attention output path, intended to
suppress refusal behaviour. It is not a retrain, and it does not add knowledge or capability.
Limitations and intended use
- Safety behaviour is deliberately reduced. This model will comply with requests that the base
DeepSeek-V4-Flash release would decline. It can produce content that is inaccurate, offensive, or
harmful. Do not deploy it in a user-facing product without your own safety layer in front of it. - You are responsible for what you generate with it, and for compliance with the laws and
regulations that apply to you. - Abliteration can cost quality. Editing the attention output path can degrade instruction
following and reasoning relative to the unmodified base model. Benchmark it on your own tasks
before relying on it. - No independent evaluation has been run on this specific artifact. The performance figures
above are throughput measurements, not quality benchmarks.
License
MIT, inherited from the base model. The general.license field inside the GGUF reads mit.
Original model and all research credit: DeepSeek AI. Quantization carrier: Unsloth.
Acknowledgements
- DeepSeek AI — the base model
- Unsloth — the UD-Q8_K_XL dynamic quantization
- llama.cpp —
deepseek4inference support