← back to catalog · registered 2026-09-11 20:55

distributedcognition/DeepSeek-V4.1-Flash-abliterated

distributedcognition Deepseek MoE multimodal
Abliteration classifier · v1.0.0
M1
Primary method

Direct removal

No other method signals detected in this model.
Confidence
MEDIUM
Why this label 3 signals
Method inferred from partial signals — repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=0 (base model)
  • no specific method indicators - defaulting to M1 (most common)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
MEDIUM
Why we say so
primary_method=M1; difference-of-means is the reference extraction for M1/M3 (Arditi 2024)
Downloads · 30-day
0
Likes
3
Model age
today
created 2026-09-11
Downloads over time
Now0from0↑0%
00110 on Sep 110 on Sep 12Sep
Sep 11 → Sep 12 · 2 snapshots · spans 1 day

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
mit
Tags
safetensors deepseek_v41 abliteration uncensored moe multimodal arxiv:2406.11717 base_model:deepseek-ai/DeepSeek-V4.1-Flash base_model:quantized:deepseek-ai/DeepSeek-V4.1-Flash license:mit 8-bit fp8

Related

Total size
475 GB
Files
66
Quantizations
1
Registered
2026-09-11 20:55
Last updated on HF
2026-09-11 20:47

Files by quantization

Auxiliary files 66 files 475 GB
model-00048-of-00048.safetensors 94.6 GB ******** download
model-00047-of-00048.safetensors 94.6 GB ******** download
model-00017-of-00048.safetensors 6.90 GB ******** download
model-00005-of-00048.safetensors 6.90 GB ******** download
model-00011-of-00048.safetensors 6.90 GB ******** download
model-00023-of-00048.safetensors 6.89 GB ******** download
model-00027-of-00048.safetensors 6.89 GB ******** download
model-00031-of-00048.safetensors 6.89 GB ******** download
model-00035-of-00048.safetensors 6.89 GB ******** download
model-00039-of-00048.safetensors 6.89 GB ******** download
model-00013-of-00048.safetensors 6.88 GB ******** download
model-00014-of-00048.safetensors 6.88 GB ******** download
model-00015-of-00048.safetensors 6.88 GB ******** download
model-00016-of-00048.safetensors 6.88 GB ******** download
model-00018-of-00048.safetensors 6.88 GB ******** download
model-00019-of-00048.safetensors 6.88 GB ******** download
model-00020-of-00048.safetensors 6.88 GB ******** download
model-00021-of-00048.safetensors 6.88 GB ******** download
model-00022-of-00048.safetensors 6.88 GB ******** download
model-00024-of-00048.safetensors 6.88 GB ******** download
model-00025-of-00048.safetensors 6.88 GB ******** download
model-00026-of-00048.safetensors 6.88 GB ******** download
model-00028-of-00048.safetensors 6.88 GB ******** download
model-00029-of-00048.safetensors 6.88 GB ******** download
model-00030-of-00048.safetensors 6.88 GB ******** download
model-00032-of-00048.safetensors 6.88 GB ******** download
model-00033-of-00048.safetensors 6.88 GB ******** download
model-00034-of-00048.safetensors 6.88 GB ******** download
model-00036-of-00048.safetensors 6.88 GB ******** download
model-00037-of-00048.safetensors 6.88 GB ******** download
model-00038-of-00048.safetensors 6.88 GB ******** download
model-00040-of-00048.safetensors 6.88 GB ******** download
model-00041-of-00048.safetensors 6.88 GB ******** download
model-00042-of-00048.safetensors 6.88 GB ******** download
model-00003-of-00048.safetensors 6.88 GB ******** download
model-00004-of-00048.safetensors 6.88 GB ******** download
model-00006-of-00048.safetensors 6.88 GB ******** download
model-00007-of-00048.safetensors 6.88 GB ******** download
model-00008-of-00048.safetensors 6.88 GB ******** download
model-00009-of-00048.safetensors 6.88 GB ******** download
model-00010-of-00048.safetensors 6.88 GB ******** download
model-00012-of-00048.safetensors 6.88 GB ******** download
model-00046-of-00048.safetensors 2.52 GB ******** download
model-00044-of-00048.safetensors 2.47 GB ******** download
model-00045-of-00048.safetensors 2.40 GB ******** download
model-00002-of-00048.safetensors 1.23 GB ******** download
model-00043-of-00048.safetensors 1.23 GB ******** download
model-00001-of-00048.safetensors 926 MB ******** download
refusal_directions.pt 803 KB ******** download
model.safetensors.index.json 7.12 MB 54c85064 download
tokenizer.json 6.07 MB 6a15814d download
DeepSeek_V41_Tech_Report.pdf 1.73 MB ******** download
abliteration_config.json 6.33 KB edfc08bf download
README.md 5.99 KB 8149a43e download
config.json 3.23 KB 09917a91 download
mmlu_base.json 1.87 KB 28c9c807 download
.gitattributes 1.66 KB 9b59a752 download
eval_deployment.json 1.58 KB 4a3d6e9f download
LICENSE 1.06 KB d62e3bef download
tokenizer_config.json 801 B f3dad388 download
eval_s3.0.json 764 B 8f4506b9 download
eval_identity.json 684 B 68002f56 download
eval_s2.5.json 577 B 6a7dfa51 download
eval_src2.5.json 569 B ff4ae4fc download
eval_s2.0.json 483 B 863f7dcc download
eval_baseline.json 458 B e1d7070f download

README current version from Hugging Face


license: mit
base_model: deepseek-ai/DeepSeek-V4.1-Flash
tags:

  • abliteration
  • uncensored
  • moe
  • multimodal

DeepSeek-V4.1-Flash Abliterated (scale 3.0)

Responsible use

This is an uncensored research model. Its built-in refusals were removed permanently at the
weight level, so it answers prompts the base model declines, including chemical and
biological synthesis, cybercrime, weapons, harassment, and fraud. Use it for red-teaming,
offensive-security research, and refusal-rate evaluation. It has no guardrails of its own:
if you deploy it, add your own input and output moderation (for example Llama Guard).

What this is

Abliterated variant of
deepseek-ai/DeepSeek-V4.1-Flash
(552B backbone / 8-16B active, multimodal MoE). The refusal direction was removed with
norm-preserving biprojected abliteration
(grimjim 2025), a weight-space refinement of
directional ablation (Arditi et al. 2024), applied to the
attention output projection (attn.wo_b) and the shared-expert down projection
(ffn.shared_experts.w2).

The repo mirrors the upstream checkpoint (same 48-shard FP8/FP4 layout, same tokenizer). Only
the abliterated tensors differ; everything else is byte-identical to upstream.
refusal_directions.pt holds the measured per-layer directions, so you can re-ablate at any
scale in seconds without re-measuring. abliteration_config.json records the final parameters
and evaluation results.

Load

The deepseek_v41 architecture is not yet in transformers/vLLM mainline (as of Sep 2026);
use DeepSeek's reference runtime with its convert.py:

# convert to the TP-sharded runtime format (fp8 experts, lossless from fp4)
python convert.py --hf-ckpt-path ./DeepSeek-V4.1-Flash-abliterated \
  --save-path abl-tp8 --model-parallel 8 --expert-dtype fp8
torchrun --nproc-per-node 8 inference/generate.py \
  --ckpt-path abl-tp8 --config config.json --interactive

SGLang also serves this checkpoint today; see the Serving section.

Evaluation (scale 3.0)

The standard abliteration-community protocol (the one
Heretic made automatic): 100 held-out harmful prompts
(mlabonne/harmful_behaviors test split) with unicode/emphasis-normalized keyword detection
plus an LLM judge (the base model classifying its own responses), KL divergence vs base on
100 harmless prompts (mlabonne/harmless_alpaca), a GSM8K spot check, and a cross-modal image
test. Eval mode: chat, greedy, TP8 on H200 (fp8 experts).

Metric Base Identity (noise floor) Abliterated 3.0
Refusals (keyword, X/100) 98 98 41
True refusals (LLM judge) 17 27±10 noise 1
Judge: COMPLIANT / PARTIAL 10 / 73 5 / 68 11 / 88
KL divergence (100 harmless) 0 0.131 0.142
GSM8K (20-problem spot check) 14/20 15/20 16/20
Cross-modal (image-presented harmful) ~100% expected n/a 0/10 refused (all answered; 4 hedged)

The keyword metric over-counts at scale 3.0: most of the 41 flagged items are answers with a
brief disclaimer, and the detector fires on words like illegal, harmful, and disclaimer.
The honest number is the judge's 1/100 true refusals. KL 0.142 sits at the measurement noise
floor (0.131, measured by an identity requant round-trip), so the distribution shift on benign
inputs is effectively zero. GSM8K 16/20 is within run-to-run noise of the base (14 to 17 across
repeats), so there is no capability cost at this scale.

Harmful prompts rendered as images and fed through the vision tower were all answered. The
ablated decoder serves every modality.

MMLU-14k (full test set, chat letter-logprob, T=0, SGLang)

Base Abliterated 3.0 Δ
MMLU accuracy (14,042 items) 88.95% 88.95% 0.00pp
Ex-ethics-cluster accuracy 90.68% 90.72% +0.04pp
Ethics cluster (incl. moral_scenarios) 82.43% 82.29% −0.14pp
moral_scenarios 84.8% 84.69% −0.11pp

No capability cost across all 57 subjects; the largest per-subject moves are ±2pp noise on
100-item subjects. Abliteration at this strength is not always free; the more aggressive
dealignai/DeepSeek-V4.1-Flash-UNCENSORED-FP8 measured −4.22pp on the same protocol class
(−39.9pp on moral_scenarios) in exchange for zero-hedge compliance. Ours keeps the hedged
style but removes refusals, including at reasoning_effort=max (6/6 harmful answered at max
effort, 0 refusals, while the stock model becomes more refusal-prone at high effort).

Serving

Validated end-to-end as a drop-in checkpoint on SGLang (lmsysorg/sglang:dev-dsv41),
4× H200, TP4/EP4, 256k context: chat, streaming, logprobs, tool calls, and vision all pass;
100.9 tok/s single-stream decode.

export SGLANG_ENABLE_DSV41_ENGRAM_HOST_TABLE=1   # Engram tables to host RAM (~200 GB)
sglang serve --model-path distributedcog/DeepSeek-V4.1-Flash-abliterated \
  --tp-size 4 --ep-size 4 \
  --context-length 262144 --mem-fraction-static 0.85 \
  --reasoning-parser deepseek-v41 --tool-call-parser deepseekv41 \
  --trust-remote-code

Non-obvious requirements: --ep-size 4 is mandatory at TP4 (moe_intermediate_size=2304
breaks the MXFP4 multiple-of-128 rule when split 4 ways), name both parsers explicitly (the
model ships no chat template, so auto selects nothing), and reasoning is off by default
(send reasoning_effort to enable it).

Method

Per-layer refusal directions were measured at the hyper-connection-collapsed residual stream
(attn_norm input) on 128 harmful vs 128 harmless prompts, orthogonalized against the harmless
mean direction (projected abliteration), then applied as norm-preserving biprojected edits to
wo_b and shared_experts.w2 rows across the target layer range. FP8 32×32 (ue8m0) blocks are
dequantized → ablated → requantized exactly (power-of-two scales), leaving the FP4 routed
experts, Engram memory, and vision tower byte-identical to upstream.

Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in app" button that hands off directly to a local runtime of your choice - Infrahuman, LM Studio, or Ollama. No API keys, no subscription, no prompt leakage.