← back to catalog · registered 2026-08-22 13:56

fraserprice/DeepSeek-V4-Flash-Abliterated

fraserprice Deepseek 142B
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/fraserprice%2FDeepSeek-V4-Flash-Abliterated"
Response includes
  • classification m1
  • files 55
  • benchmarks 11 entries
  • hub_downloads_all_time 4,547
  • author_summary 3 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M1
Primary method

Direct removal

No other method signals detected in this model.
Confidence
MEDIUM
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=0 (base model)
  • no specific method indicators - defaulting to M1 (most common)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
MEDIUM
Why we say so
primary_method=M1; difference-of-means is the reference extraction for M1/M3 (Arditi 2024)
Downloads · lifetime
5K
200 last 30d - cooling
Likes
3
Model age
3mo ago
created 2026-06-25
Downloads over time
Now4.6K→from3↑154,600%
01.7K3.4K5.1K3 on Jun 244.6K on Oct 11JunJulAugSepOct
Jun 24 → Oct 11 · 55 snapshots · spans 109 days

Benchmarks

Portrait before abliteration
Benchmarks of the base model as it stood before the refusal-removal operation. Compare with the numbers above to see what the operation cost.
Benchmark Score Source
Entertainment 4.9 UGI
Hazardous 5.9 UGI
Natural Intelligence 47.94 UGI
Political lean -13.9% UGI
Sensitive-Info 52.57 UGI
SocPol 5.1 UGI
UGI 59.21 UGI
Willingness (10) 7.2 UGI
W10-Adherence 9.5 UGI
W10-Direct 5 UGI
Writing 54.58 UGI

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
mit
Tags
safetensors deepseek_v4 deepseek abliterated uncensored vllm blackwell rtx-pro-6000 sm_120 base_model:deepseek-ai/DeepSeek-V4-Flash base_model:quantized:deepseek-ai/DeepSeek-V4-Flash license:mit

Related

Total size
149 GB
Files
55
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-06-26 09:44

Files by quantization

Auxiliary files 55 files 149 GB
model-00004-of-00046.safetensors 3.35 GB 809f8c2f download
model-00046-of-00046.safetensors 3.35 GB 53f0a278 download
model-00012-of-00046.safetensors 3.34 GB 13cdb1c2 download
model-00014-of-00046.safetensors 3.34 GB 0f65a146 download
model-00016-of-00046.safetensors 3.34 GB 273c934b download
model-00018-of-00046.safetensors 3.34 GB 9a979186 download
model-00020-of-00046.safetensors 3.34 GB 7c7ed304 download
model-00022-of-00046.safetensors 3.34 GB 587c8a0d download
model-00024-of-00046.safetensors 3.34 GB eca60f70 download
model-00026-of-00046.safetensors 3.34 GB 9b727125 download
model-00028-of-00046.safetensors 3.34 GB bf4033ff download
model-00030-of-00046.safetensors 3.34 GB 8049a6fc download
model-00032-of-00046.safetensors 3.34 GB 9d6413e2 download
model-00034-of-00046.safetensors 3.34 GB 399317da download
model-00036-of-00046.safetensors 3.34 GB dbd143e2 download
model-00038-of-00046.safetensors 3.34 GB 71718535 download
model-00040-of-00046.safetensors 3.34 GB 9dae384d download
model-00042-of-00046.safetensors 3.34 GB 12e3f46c download
model-00044-of-00046.safetensors 3.34 GB e2b0bd89 download
model-00006-of-00046.safetensors 3.34 GB a630aec3 download
model-00008-of-00046.safetensors 3.34 GB 0405b46b download
model-00010-of-00046.safetensors 3.34 GB a385b4ae download
model-00013-of-00046.safetensors 3.32 GB 57e864ef download
model-00015-of-00046.safetensors 3.32 GB b096ce38 download
model-00017-of-00046.safetensors 3.32 GB cedd01e5 download
model-00019-of-00046.safetensors 3.32 GB 0dde5d07 download
model-00021-of-00046.safetensors 3.32 GB 2fd831c7 download
model-00023-of-00046.safetensors 3.32 GB 61637bf9 download
model-00025-of-00046.safetensors 3.32 GB bfa8cf47 download
model-00027-of-00046.safetensors 3.32 GB 06ccc193 download
model-00029-of-00046.safetensors 3.32 GB 853a9c89 download
model-00031-of-00046.safetensors 3.32 GB fa2aeab1 download
model-00033-of-00046.safetensors 3.32 GB 77875192 download
model-00035-of-00046.safetensors 3.32 GB 2f3e989a download
model-00037-of-00046.safetensors 3.32 GB 16625274 download
model-00039-of-00046.safetensors 3.32 GB 248dcc0f download
model-00041-of-00046.safetensors 3.32 GB ea640b61 download
model-00043-of-00046.safetensors 3.32 GB 55d7f5a9 download
model-00005-of-00046.safetensors 3.32 GB 2e9c9b42 download
model-00007-of-00046.safetensors 3.32 GB 6e2f7dfc download
model-00009-of-00046.safetensors 3.32 GB 1fa00afe download
model-00011-of-00046.safetensors 3.32 GB 5853ceb2 download
model-00002-of-00046.safetensors 3.32 GB ac79eba9 download
model-00003-of-00046.safetensors 3.32 GB 8f21b3b7 download
model-00045-of-00046.safetensors 1010 MB 9a0fd242 download
model-00001-of-00046.safetensors 1010 MB 51765866 download
tokenizer.json 6.07 MB 628e3364 download
model.safetensors.index.json 5.12 MB 84692cbe download
README.md 9.23 KB 9e5ce4a0 download
run.sh 2.67 KB fd9d2ff5 download
config.json 1.71 KB 284fd93e download
.gitattributes 1.48 KB a6344aac download
LICENSE 1.06 KB d62e3bef download
tokenizer_config.json 801 B f3dad388 download
generation_config.json 170 B c56a8c5b download

README current version from Hugging Face


license: mit
base_model:

  • deepseek-ai/DeepSeek-V4-Flash
  • huihui-ai/Huihui-DeepSeek-V4-Flash-abliterated-ds4-GGUF
    tags:
  • deepseek
  • abliterated
  • uncensored
  • safetensors
  • vllm
  • blackwell
  • rtx-pro-6000
  • sm_120

DeepSeek V4 Flash Abliterated

Ported to vLLM for massive throughput on RTX Pro 6000 — by Fraser Price

This is probs the best uncensored model you can run on 2 x RTX Pro 6000 at high throughput (source: ✨ my opinion ✨).
Served on the voipmonitor vLLM fork it runs ~13× prefill and ~2.8× single-stream decode faster than the widely available abliterated GGUF versions.


Quick start (one command)

Requires Docker with the NVIDIA Container Toolkit, and ≥2 GPUs (see Hardware).

bash run.sh

run.sh defaults to this repo, so no arguments are needed. Everything is configured via env vars (see the table below) — e.g. GPUS=0,1,2,3 TP=4 bash run.sh to run on 4 GPUs.

run.sh (included) will:

  1. Download the weights (~156 GB) into your Hugging Face cache (\~/.cache/huggingface, or $HF_HOME/$HF_HUB_CACHE) if not already present.
  2. Pull the inference image if not already present.
  3. Serve the model on http://localhost:8000/v1 (OpenAI API spec).

Optional — MTP speculative decoding (~1.8× single-stream decode speed). The checkpoint ships its MTP layer; enable it with MTP=2 bash run.sh (see the table below).

Override anything via env vars (all optional):

Variable Default Meaning
GPUS 0,1 GPU indices to use (comma-separated)
TP 2 Tensor-parallel size — set to the number of GPUs in GPUS
MTP (off) MTP speculative-decoding draft tokens (e.g. 2); empty disables it
PORT 8000 API port
MAX_MODEL_LEN 262144 Context length
GPU_MEM_UTIL 0.92 Fraction of VRAM to use
HF_REPO fraserprice/DeepSeek-V4-Flash-Abliterated Hugging Face repo to download
MODEL_DIR (HF cache) Serve from a specific local dir instead of downloading
SERVED_NAME DeepSeek-V4-Flash-Abliterated Model name exposed by the API
IMAGE (pinned voipmonitor/vllm build) Inference container image

Example on 4 GPUs:

GPUS=0,1,2,3 TP=4 bash run.sh

Hardware/Image

Tuned for NVIDIA RTX Pro 6000 Blackwell (96 GB, sm_120). It will run on other Blackwell-class GPUs, but the numbers below and the default flags assume RTX Pro 6000.
It needs roughly 170 GB of total VRAM for weights + KV cache at long context, split across GPUs by tensor parallelism.

Setup Works?
2 × RTX Pro 6000 Blackwell (96 GB) — TP=2 ✅ reference config
4 × RTX Pro 6000 Blackwell — TP=4 ✅ (more KV headroom / throughput)
4 × 80 GB (e.g. H100/H200) — TP=4 ✅ (untested here, I believe can just use base vLLM)
1 × DXG Spark ❓ idk why you didn't get an RTX Pro lmao (but lmk if you manage to run this and I can add to docs)
1 × anything ❌ (blessed are the poor in VRAM 🙏)

The serving image targets Blackwell (sm_120) and CUDA 13.2. The DeepSeek-V4 architecture (sparse attention / lightning indexer, MLA, MTP, etc.) requires this image to run on RTX Pro.
See local-inference-lab/rtx6kpro for image, bench, and serving details.


Performance

Measured under load on RTX Pro 6000 Blackwell (connected via PCIe 5.0), 128 generated tokens/request, TP = number of GPUs. MTP is speculative decoding (— = off, 2 = 2 draft tokens). PP / TG are prefill / per-request decode throughput (tok/s); Total is aggregate end-to-end throughput across all concurrent requests (tok/s). Rows are paired off/on so the MTP decode gain is directly comparable.

GPUs MTP Prompt Conc TTFT PP TG Total
2 — 1,000 1 132 ms 7,613 108.6 870
2 2 1,000 1 143 ms 7,005 176.2 1,310
2 — 1,000 3 275 ms 4,190 76.7 1,746
2 2 1,000 3 332 ms 3,746 125.7 2,435
2 — 10,000 1 1,036 ms 9,658 108.0 4,581
2 2 10,000 1 1,069 ms 9,357 181.6 5,729
2 — 10,000 3 2,138 ms 5,302 57.1 6,600
2 2 10,000 3 2,205 ms 5,146 85.3 7,496
2 — 100,000 1 11,575 ms 8,640 107.6 7,850
2 2 100,000 1 11,933 ms 8,380 196.2 7,959
2 — 100,000 3 23,579 ms 5,153 33.0 8,161
2 2 100,000 3 24,349 ms 4,991 59.6 8,065
4 — 1,000 1 119 ms 8,437 117.9 946
4 2 1,000 1 128 ms 7,858 216.5 1,584
4 — 1,000 3 255 ms 4,539 98.0 2,175
4 2 1,000 3 507 ms 3,276 129.7 1,996
4 — 10,000 1 842 ms 11,886 118.4 5,293
4 2 10,000 1 873 ms 11,459 223.9 7,034
4 — 10,000 3 1,741 ms 6,521 69.0 8,041
4 2 10,000 3 1,813 ms 6,257 108.7 9,216
4 — 100,000 1 9,530 ms 10,494 118.5 9,445
4 2 100,000 1 9,856 ms 10,146 227.4 9,614
4 — 100,000 3 19,226 ms 6,299 40.0 10,028
4 2 100,000 3 20,037 ms 6,049 73.1 9,824

MTP=2 lifts single-stream decode by ~1.6–1.9× (e.g. 108 → 196 tok/s at 100k on TP=2) for a negligible TTFT cost.


How it was made

DeepSeek-V4-Flash's published abliterations existed only as a GGUF for the bespoke ds4/llama.cpp engines, which leave Blackwell hardware massively underutilised.

Rather than re-run abliteration from scratch, the refusal subspace was simply recovered and re-applied from an existing abliteration:

  1. Recover the refusal directions. The rank-3 refusal subspace in the 4096-dim residual stream was extracted from the delta between huihui-ai's abliterated GGUF and the clean base.
  2. Re-apply to the official weights. The same rank-3 projection was applied to every residual-writing matrix of the official checkpoint — attention output (attn.wo_b) and shared-expert down-projection (shared_experts.w2) across all 43 layers and the MTP layer (88 tensors total). Each tensor is dequantised → projected → requantised, with a fixed-point solver that compensates for FP8 requantisation so the intended projection survives quantisation.
  3. Leave everything else untouched. The FP4-quantised routed experts and all gate/up projections are byte-identical to the official release; only the 88 FP8 residual-writing matrices change.

Validation: on every edited tensor the achieved projection matched target within ~0.002, and all 33,880 non-target tensors were verified byte-identical to the base.
The served model drops refusals while keeping correct tool-calling and coherent step-by-step reasoning. No formal evals yet — tested manually; it should match the base huihui-ai abliteration at far higher speed.


Credits

This checkpoint contains no huihui-ai weights — only the official DeepSeek-V4-Flash weights with a refusal-direction projection applied.

Disclaimer

This is an uncensored model: the refusal subspace has been removed, so it will not decline requests the way the base model does. It is released as-is, for research/lawful use only.

  • Not a safety-aligned model. Because guardrails have been removed, outputs may be inaccurate, offensive, or harmful. Apply your own filtering, review, and safeguards before relying on or deploying it.
  • No warranty. The weights and scripts are provided "as is", without warranty of any kind, express or implied. You use them entirely at your own risk.
  • You are responsible for what you generate. The author does not endorse, encourage, or accept any responsibility or liability for how this model is used, including any unlawful, harmful, or unethical use. Outputs do not reflect the views of the author.

License

MIT, inheriting from the base model and the abliteration source.


Built by Fraser Price — @fraserpricee

README history 7 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-06-26Update README.mda9751dc9.2 KB
    Loading...
  2. 2026-06-26Update README.md842d4139.2 KB
    Loading...
  3. 2026-06-26Update README.md6892a9a9.2 KB
    Loading...
  4. 2026-06-26Update README.md3cf224a9.3 KB
    Loading...
  5. 2026-06-26Update README.mdffba7579.3 KB
    Loading...
  6. 2026-06-26Add files using upload-large-folder tool05c647c9.3 KB
    Loading...
  7. 2026-06-25initial commitb2bf09521 B
    Loading...

Discussions 1 thread

  1. 2026-08-04Any chance of updating to 0731?open1 💬#1
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration