← back to catalog · registered 2026-08-22 13:56

Jared2026/DeepSeek-V4-Flash-0731-Abliterated-UD-Q8_K_XL-GGUF

Jared2026 Deepseek GGUF MoE 1.0M ctx
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/Jared2026%2FDeepSeek-V4-Flash-0731-Abliterated-UD-Q8_K_XL-GGUF"
Response includes
  • classification m8
  • files 10
  • hub_downloads_all_time 62
  • author_summary 1 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M8
Primary method

Repackaging (quantization)

Applied on top of direct removal inherited from the base model.
Confidence
MEDIUM
Inherited from base model
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=1
  • assume M1 (base ablation) + M8 (GGUF quant) - default when producer unknown
Refusal direction extraction

No specific extraction method could be identified for this model. The producer either did not document it or used a proprietary pipeline.

What is a refusal direction? →
Downloads · lifetime
62
0
Likes
2
Model age
8w ago
created 2026-08-12
Downloads over time
Now62→from62↑0%
6262636362 on Aug 1962 on Oct 11AugSepOct
Aug 19 → Oct 11 · 48 snapshots · spans 53 days

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
mit
Quantizations
Q8_K
Tags
gguf deepseek deepseek-v4 deepseek-v4-flash deepseek4 abliterated uncensored mixture-of-experts moe quantized q8 unsloth

Related

Total size
151 GB
Files
10
Quantizations
2
Registered
2026-08-22 13:56
Last updated on HF
2026-08-15 18:01

Files by quantization

Q8_K 5 files 151 GB
DeepSeek-V4-Flash-0731-Abliterated-UD-Q8_K_XL-00003-of-00005.gguf 46.3 GB ******** download
DeepSeek-V4-Flash-0731-Abliterated-UD-Q8_K_XL-00004-of-00005.gguf 46.1 GB ******** download
DeepSeek-V4-Flash-0731-Abliterated-UD-Q8_K_XL-00002-of-00005.gguf 45.8 GB ******** download
DeepSeek-V4-Flash-0731-Abliterated-UD-Q8_K_XL-00005-of-00005.gguf 12.6 GB ******** download
DeepSeek-V4-Flash-0731-Abliterated-UD-Q8_K_XL-00001-of-00005.gguf 5.01 MB ******** download
Auxiliary files 5 files 11.4 MB
tensor_manifest_after.json 11.3 MB ******** download
transplant_report.json 38.7 KB 98242a42 download
README.md 8.08 KB d6b285ca download
.gitattributes 2.56 KB 406107fe download
SHA256SUMS 660 B e3e58e48 download

README current version from Hugging Face


license: mit
base_model:

  • deepseek-ai/DeepSeek-V4-Flash-0731
    base_model_relation: quantized
    pipeline_tag: text-generation
    library_name: gguf
    quantized_by: Jared2026
    tags:
  • gguf
  • deepseek
  • deepseek-v4
  • deepseek-v4-flash
  • deepseek4
  • abliterated
  • uncensored
  • mixture-of-experts
  • moe
  • quantized
  • q8
  • unsloth
  • llama.cpp
  • conversational

DeepSeek-V4-Flash-0731-Abliterated — UD-Q8_K_XL GGUF

An 8-bit (UD-Q8_K_XL) GGUF conversion of DeepSeek-V4-Flash-0731 with a mode-merged
abliteration overlay applied to the attention output projections.

161.9 GB across 5 shards. Runs on llama.cpp with the MoE experts offloaded to system RAM,
which puts it within reach of a single workstation or one GPU node — the resident VRAM footprint
is under 20 GB.

Requires a recent llama.cpp build. The deepseek4 architecture is not supported by builds
from before ~August 2026. See Requirements — this is the single most common
reason loading fails.


Files

File Size SHA256 (first 16)
DeepSeek-V4-Flash-0731-Abliterated-UD-Q8_K_XL-00001-of-00005.gguf 5.26 MB 0cecd47692e23e39
DeepSeek-V4-Flash-0731-Abliterated-UD-Q8_K_XL-00002-of-00005.gguf 49.2 GB d479250fa56917c3
DeepSeek-V4-Flash-0731-Abliterated-UD-Q8_K_XL-00003-of-00005.gguf 49.7 GB df01684421c67b5f
DeepSeek-V4-Flash-0731-Abliterated-UD-Q8_K_XL-00004-of-00005.gguf 49.5 GB da96ae34d4fc7bce
DeepSeek-V4-Flash-0731-Abliterated-UD-Q8_K_XL-00005-of-00005.gguf 13.5 GB dd643e2e9b5c1bf1

Shard 1 is a small header shard — point llama.cpp at -00001-of-00005.gguf and it resolves the
rest automatically. Full checksums are in SHA256SUMS.

Also included for provenance: transplant_report.json (the exact
tensor-level record of the merge) and tensor_manifest_after.json.

Download

pip install -U "huggingface_hub[cli]"

hf download Jared2026/DeepSeek-V4-Flash-0731-Abliterated-UD-Q8_K_XL-GGUF \
    --include "*.gguf" \
    --local-dir ./DeepSeek-V4-Flash-Q8

Verify before loading (162 GB is a long way to get through on a bad transfer):

cd DeepSeek-V4-Flash-Q8 && sha256sum -c SHA256SUMS

Requirements

The deepseek4 architecture is new and needs llama.cpp from master, build b10273
(commit a6aa6f545) or later
. Earlier builds recognise only deepseek / deepseek2 and fail at
load with an unknown-architecture error. If you have an existing llama.cpp checkout from before
August 2026, rebuild it — an in-place git pull of the source without recompiling is not enough.

git clone https://github.com/ggml-org/llama.cpp && cd llama.cpp
cmake -B build -DGGML_CUDA=ON      # drop -DGGML_CUDA for CPU-only
cmake --build build --config Release -j

Hardware. Weights are 162 GB. The practical configuration is --cpu-moe, which keeps attention
and dense layers on the GPU (~19 GB) and the 256 experts in system RAM. Budget ~180 GB of RAM
plus a 20 GB-class GPU. A pure-CPU run works and needs no GPU at all, just the RAM.

Usage

llama.cpp server

llama-server \
    --model DeepSeek-V4-Flash-0731-Abliterated-UD-Q8_K_XL-00001-of-00005.gguf \
    --alias deepseek-v4-flash \
    --host 0.0.0.0 --port 8001 \
    --n-gpu-layers 99 --cpu-moe \
    --flash-attn on \
    --ctx-size 262144 \
    --jinja \
    --chat-template-kwargs '{"thinking":true,"reasoning_effort":"max"}'

Exposes an OpenAI-compatible API at http://localhost:8001/v1.

llama.cpp CLI

llama-cli \
    -m DeepSeek-V4-Flash-0731-Abliterated-UD-Q8_K_XL-00001-of-00005.gguf \
    --cpu-moe -ngl 99 -c 65536 --jinja

Multi-GPU

With several small GPUs (e.g. MIG slices), name them explicitly — attention and dense layers shard
across them while the experts stay on CPU:

--device CUDA0,CUDA1,CUDA2,CUDA3 --n-gpu-layers 99 --cpu-moe

Note that --n-cpu-moe N (keeping some expert layers on GPU) is a poor fit for small or
partitioned GPUs: llama.cpp places all GPU-resident expert layers on a single device, which will
OOM a 20 GB card. Full --cpu-moe is the reliable choice there.

Context length

Tokens
Natively trained 65,536
Maximum (YaRN, factor 16) 1,048,576

The 1M figure is rope-scaled, not trained. Quality degrades progressively past the native 64K,
so treat 1M as a ceiling rather than a working context. --ctx-size 262144 (256K) is a reasonable
middle setting. KV cache cost is unusually low for the context size because the model uses MLA with
head_count_kv = 1 — 256K of context costs well under 1 GB, so context length is limited by
quality, not memory.

Reasoning effort

The chat template emits a <think> block by default. It recognises exactly two effective levels —
default, and max:

--chat-template-kwargs '{"thinking":true,"reasoning_effort":"max"}'

max injects an explicit "absolute maximum, no shortcuts" directive into the system prompt. There
is no low/medium/high gradation; any value other than max behaves as the default. You can confirm
which one is active by POSTing to the server's /apply-template endpoint and inspecting the
rendered prompt.

Architecture

Read directly from the GGUF metadata:

Key Value
Architecture deepseek4
Layers 43
Experts 256 (6 active + 1 shared)
Expert gating sigmoid, normalised, scale 1.5
Embedding dim 4096
Attention heads 64 (head_count_kv = 1, MLA)
KV compression key/value length 512, q-LoRA rank 1024
Sliding window 128
RoPE freq base 10000, YaRN ×16
Size label 256×8.4B
Tokenizer GPT-2 BPE (joyai-llm pre-tokenizer)
Quantization Unsloth Dynamic UD-Q8_K_XL

Provenance

This is a quantized redistribution, not an original model. The chain:

  1. Base model — deepseek-ai/DeepSeek-V4-Flash-0731
  2. Quantization carrier — the Unsloth Dynamic UD-Q8_K_XL GGUF build (general.quantized_by in
    the file metadata reads Unsloth)
  3. Abliteration overlay — a mode-merged overlay applied on top of the Q8 carrier

The merge is narrow and fully documented in transplant_report.json: 43 tensors were replaced,
one blk.N.attn_output_b.weight per layer, in BF16
, sourced from an overlay with SHA256
20d2559987a19cb5…. No other tensor in the file was modified — the expert weights, embeddings and
attention projections are the unmodified Q8 carrier. Three MTP/DSpark tensor groups present in the
overlay were omitted because the GGUF carrier contains no MTP tensors.

Abliteration is an inference-time behavioural edit to the attention output path, intended to
suppress refusal behaviour. It is not a retrain, and it does not add knowledge or capability.

Limitations and intended use

  • Safety behaviour is deliberately reduced. This model will comply with requests that the base
    DeepSeek-V4-Flash release would decline. It can produce content that is inaccurate, offensive, or
    harmful. Do not deploy it in a user-facing product without your own safety layer in front of it.
  • You are responsible for what you generate with it, and for compliance with the laws and
    regulations that apply to you.
  • Abliteration can cost quality. Editing the attention output path can degrade instruction
    following and reasoning relative to the unmodified base model. Benchmark it on your own tasks
    before relying on it.
  • No independent evaluation has been run on this specific artifact. The performance figures
    above are throughput measurements, not quality benchmarks.

License

MIT, inherited from the base model. The general.license field inside the GGUF reads mit.
Original model and all research credit: DeepSeek AI. Quantization carrier: Unsloth.

Acknowledgements

Discussions 1 thread

  1. 2026-09-27Access request reviewopen1 💬#1
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration