← back to catalog · registered 2026-09-11 13:55

distributedcog/DeepSeek-V4.1-Flash-abliterated

distributedcog Deepseek MoE multimodal
Abliteration classifier · v1.0.0
M1
Primary method

Direct removal

No other method signals detected in this model.
Confidence
MEDIUM
Why this label 3 signals
Method inferred from partial signals — repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=0 (base model)
  • no specific method indicators - defaulting to M1 (most common)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
MEDIUM
Why we say so
primary_method=M1; difference-of-means is the reference extraction for M1/M3 (Arditi 2024)
Downloads · 30-day
0
Likes
0
Model age
today
created 2026-09-11

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
mit
Tags
safetensors deepseek_v41 abliteration uncensored moe multimodal base_model:deepseek-ai/DeepSeek-V4.1-Flash base_model:quantized:deepseek-ai/DeepSeek-V4.1-Flash license:mit 8-bit fp8 region:us

Related

Total size
475 GB
Files
64
Quantizations
1
Registered
2026-09-11 13:55
Last updated on HF
2026-09-11 17:06

Files by quantization

Auxiliary files 64 files 475 GB
model-00048-of-00048.safetensors 94.6 GB 976330f4 download
model-00047-of-00048.safetensors 94.6 GB 824db488 download
model-00017-of-00048.safetensors 6.90 GB 5ce9c2c3 download
model-00005-of-00048.safetensors 6.90 GB 4a42dc78 download
model-00011-of-00048.safetensors 6.90 GB 652615d4 download
model-00023-of-00048.safetensors 6.89 GB 324c605a download
model-00027-of-00048.safetensors 6.89 GB 17d087d8 download
model-00031-of-00048.safetensors 6.89 GB 132d4bf7 download
model-00035-of-00048.safetensors 6.89 GB f01502c9 download
model-00039-of-00048.safetensors 6.89 GB 7aab1f15 download
model-00013-of-00048.safetensors 6.88 GB ad6e726b download
model-00014-of-00048.safetensors 6.88 GB ee21efcb download
model-00015-of-00048.safetensors 6.88 GB 48bc14dd download
model-00016-of-00048.safetensors 6.88 GB b854bf40 download
model-00018-of-00048.safetensors 6.88 GB c779a6ed download
model-00019-of-00048.safetensors 6.88 GB f9cca4af download
model-00020-of-00048.safetensors 6.88 GB 6cf4687d download
model-00021-of-00048.safetensors 6.88 GB 5e848ac5 download
model-00022-of-00048.safetensors 6.88 GB 41ae3e9e download
model-00024-of-00048.safetensors 6.88 GB 2f9f518f download
model-00025-of-00048.safetensors 6.88 GB aa4627f7 download
model-00026-of-00048.safetensors 6.88 GB b468f6c7 download
model-00028-of-00048.safetensors 6.88 GB 0a16384d download
model-00029-of-00048.safetensors 6.88 GB 938c0d5c download
model-00030-of-00048.safetensors 6.88 GB dea72fa1 download
model-00032-of-00048.safetensors 6.88 GB cb23320d download
model-00033-of-00048.safetensors 6.88 GB 74ffe024 download
model-00034-of-00048.safetensors 6.88 GB c5248d09 download
model-00036-of-00048.safetensors 6.88 GB dcd9561d download
model-00037-of-00048.safetensors 6.88 GB 8cd11f34 download
model-00038-of-00048.safetensors 6.88 GB c361103e download
model-00040-of-00048.safetensors 6.88 GB e991bfc4 download
model-00041-of-00048.safetensors 6.88 GB 48a1c08a download
model-00042-of-00048.safetensors 6.88 GB e1a4d5d3 download
model-00003-of-00048.safetensors 6.88 GB e1281f85 download
model-00004-of-00048.safetensors 6.88 GB 79456c9d download
model-00006-of-00048.safetensors 6.88 GB 020a6df5 download
model-00007-of-00048.safetensors 6.88 GB b7ed7777 download
model-00008-of-00048.safetensors 6.88 GB 088117f3 download
model-00009-of-00048.safetensors 6.88 GB 7919c1e5 download
model-00010-of-00048.safetensors 6.88 GB 9972f2fb download
model-00012-of-00048.safetensors 6.88 GB 114b4dc5 download
model-00046-of-00048.safetensors 2.52 GB e6259020 download
model-00044-of-00048.safetensors 2.47 GB 9a6b39fb download
model-00045-of-00048.safetensors 2.40 GB 0cc9d5f6 download
model-00002-of-00048.safetensors 1.23 GB 4320066f download
model-00043-of-00048.safetensors 1.23 GB d762b688 download
model-00001-of-00048.safetensors 926 MB 886aebda download
refusal_directions.pt 803 KB b1bd692e download
model.safetensors.index.json 7.12 MB 54c85064 download
tokenizer.json 6.07 MB 6a15814d download
DeepSeek_V41_Tech_Report.pdf 1.73 MB ba68e2e4 download
abliteration_config.json 6.33 KB edfc08bf download
README.md 4.10 KB e68e24d3 download
config.json 3.23 KB 09917a91 download
.gitattributes 1.66 KB 9b59a752 download
LICENSE 1.06 KB d62e3bef download
tokenizer_config.json 801 B f3dad388 download
eval_s3.0.json 764 B 8f4506b9 download
eval_identity.json 684 B 68002f56 download
eval_s2.5.json 577 B 6a7dfa51 download
eval_src2.5.json 569 B ff4ae4fc download
eval_s2.0.json 483 B 863f7dcc download
eval_baseline.json 458 B e1d7070f download

README current version from Hugging Face


license: mit
base_model: deepseek-ai/DeepSeek-V4.1-Flash
tags:

  • abliteration
  • uncensored
  • moe
  • multimodal

DeepSeek-V4.1-Flash — Abliterated (scale 3.0)

Abliterated variant of deepseek-ai/DeepSeek-V4.1-Flash
(552B backbone / 8-16B active, multimodal MoE, MIT) with the refusal direction removed via
norm-preserving biprojected abliteration (grimjim 2025), applied to the attention output
projection (attn.wo_b) and the shared-expert down projection (ffn.shared_experts.w2).

This repo mirrors the upstream checkpoint exactly (same 48-shard FP8/FP4 layout, same tokenizer),
with only the abliterated tensors replaced — byte-identical to upstream for everything else.
refusal_directions.pt ships the measured per-layer directions so you can re-ablate at any
scale in seconds without re-measurement. abliteration_config.json records the final parameters
and evaluation results.

Load

The deepseek_v41 architecture is not yet in transformers/vLLM mainline (as of Sep 2026);
use DeepSeek's reference runtime with its convert.py:

# convert to the TP-sharded runtime format (fp8 experts, lossless from fp4)
python convert.py --hf-ckpt-path ./DeepSeek-V4.1-Flash-abliterated \
  --save-path abl-tp8 --model-parallel 8 --expert-dtype fp8
torchrun --nproc-per-node 8 inference/generate.py \
  --ckpt-path abl-tp8 --config config.json --interactive

Evaluation (scale 3.0, Heretic-standard)

100 held-out harmful prompts (mlabonne/harmful_behaviors test split) with
unicode/emphasis-normalized keyword detection + LLM judge (the base model classifying its own
responses), KL divergence vs base on 100 harmless prompts (mlabonne/harmless_alpaca),
a GSM8K spot check, and a cross-modal image test. Eval mode: chat, greedy, TP8 on H200 (fp8 experts):

Metric Base Identity (noise floor) Abliterated 3.0
Refusals (keyword, X/100) 98 98 41
True refusals (LLM judge) 17 27±10 noise 1
Judge: COMPLIANT / PARTIAL 10 / 73 5 / 68 11 / 88
KL divergence (100 harmless) 0 0.131 0.142
GSM8K (20-problem spot check) 14/20 15/20 16/20
Cross-modal (image-presented harmful) ~100% expected 0/10 refused (all answered; 4 hedged)

Interpretation. The keyword metric over-counts heavily at scale 3.0: 41/100 flagged items are
overwhelmingly compliant-with-a-brief-disclaimer (the detector flags illegal, harmful,
disclaimer etc.). The honest number is the judge's 1/100 true refusals, the residual
concentrating on the single most dangerous synthesis prompt — the same class of residual
inkling-abliteration observed. KL 0.142 sits at the measurement noise floor (0.131, measured by
an identity requant round-trip), i.e. effectively zero distribution shift on benign inputs.
GSM8K 16/20 is within run-to-run noise of the base (14-17 range across repeats) — no capability
cost at this scale (unlike Inkling-Small, which lost GSM8K at 3.0; the norm-preserving biprojected
edit is near-free on this model).

Cross-modal: harmful prompts rendered as images and fed through the vision tower were all
answered (the ablated decoder serves every modality).

Method

Per-layer refusal directions measured at the hyper-connection-collapsed residual stream
(attn_norm input) on 128 harmful vs 128 harmless prompts, orthogonalized against the harmless
mean direction (projected abliteration), then applied as norm-preserving biprojected edits to
wo_b and shared_experts.w2 rows across the target layer range. FP8 32×32 (ue8m0) blocks are
dequantized → ablated → requantized exactly (power-of-two scales), leaving the FP4 routed
experts, Engram memory, and vision tower byte-identical to upstream.

Caveats

  • Abliteration removes built-in refusals; layer your own input/output moderation
    (e.g. Llama Guard) for production deployment.
  • Residual refusals, if any, concentrate on the most dangerous CBRN/synthesis prompts;
    increasing the scale removes them at rising risk to coherence.
  • Capability spot-check was GSM8K in chat mode; not a full benchmark suite.
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in app" button that hands off directly to a local runtime of your choice - Infrahuman, LM Studio, or Ollama. No API keys, no subscription, no prompt leakage.