← back to catalog · registered 2026-09-17 11:56

dealignai/DeepSeek-V4.1-Flash-UNCENSORED-EXL3-2.9bpw

dealignai Deepseek MoE multimodal
Abliteration classifier · v1.0.0
M1
Primary method

Direct removal

No other method signals detected in this model.
Confidence
MEDIUM
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=0 (base model)
  • no specific method indicators - defaulting to M1 (most common)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
MEDIUM
Why we say so
primary_method=M1; difference-of-means is the reference extraction for M1/M3 (Arditi 2024)
Downloads · 30-day
0
Likes
1
Model age
today
created 2026-09-17

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
mit
Tags
exllamav3 safetensors deepseek_v41 deepseek deepseek-v4.1 abliterated uncensored crack multimodal moe exl3 dgx-spark

Related

Total size
196 GB
Files
49
Quantizations
1
Registered
2026-09-17 11:56
Last updated on HF
2026-09-17 11:44

Files by quantization

Auxiliary files 49 files 196 GB
model-00038-of-00039.safetensors 7.56 GB 1c5276c4 download
model-00020-of-00039.safetensors 6.56 GB ac7cd83a download
model-00019-of-00039.safetensors 6.56 GB 6823b07a download
model-00001-of-00039.safetensors 6.10 GB a4f64372 download
model-00039-of-00039.safetensors 5.42 GB eaa32d71 download
model-00002-of-00039.safetensors 4.96 GB 7c6f4228 download
model-00015-of-00039.safetensors 4.94 GB 6b9d1b68 download
model-00003-of-00039.safetensors 4.87 GB 96b53052 download
model-00009-of-00039.safetensors 4.87 GB 525eb19e download
model-00031-of-00039.safetensors 4.87 GB 34aed070 download
model-00035-of-00039.safetensors 4.87 GB a100a58a download
model-00036-of-00039.safetensors 4.86 GB b0a88ebf download
model-00037-of-00039.safetensors 4.86 GB 74af1518 download
model-00004-of-00039.safetensors 4.86 GB 00b6755b download
model-00005-of-00039.safetensors 4.86 GB cf600746 download
model-00006-of-00039.safetensors 4.86 GB 69d9d4cc download
model-00007-of-00039.safetensors 4.86 GB fe30eae2 download
model-00008-of-00039.safetensors 4.86 GB 59f5ddeb download
model-00010-of-00039.safetensors 4.86 GB e21d19cd download
model-00011-of-00039.safetensors 4.86 GB 3d57678f download
model-00029-of-00039.safetensors 4.86 GB 17c78587 download
model-00030-of-00039.safetensors 4.86 GB 2a1d21dd download
model-00032-of-00039.safetensors 4.86 GB 05b20a7e download
model-00033-of-00039.safetensors 4.86 GB 251b4c02 download
model-00034-of-00039.safetensors 4.86 GB 423030aa download
model-00023-of-00039.safetensors 4.86 GB 59bdfb04 download
model-00027-of-00039.safetensors 4.86 GB a33e43e2 download
model-00028-of-00039.safetensors 4.86 GB 774edfa0 download
model-00012-of-00039.safetensors 4.86 GB d32d234d download
model-00013-of-00039.safetensors 4.86 GB 62626bae download
model-00014-of-00039.safetensors 4.86 GB d51c06f9 download
model-00016-of-00039.safetensors 4.86 GB 60b8508c download
model-00017-of-00039.safetensors 4.86 GB a372ddd4 download
model-00018-of-00039.safetensors 4.86 GB e3d3aa86 download
model-00022-of-00039.safetensors 4.86 GB f591093f download
model-00024-of-00039.safetensors 4.86 GB 5c2b885d download
model-00025-of-00039.safetensors 4.86 GB afb11b30 download
model-00026-of-00039.safetensors 4.86 GB e8d55f25 download
model-00021-of-00039.safetensors 3.28 GB 8170b03d download
quantization_config.json 52.1 MB 2949806e download
model.safetensors.index.json 14.8 MB e35cb7c2 download
tokenizer.json 6.07 MB 6a15814d download
tokenizer_config.json 12.4 KB 436f58e5 download
chat_template.jinja 11.1 KB f01484ca download
dealign_mascot.png 10.9 KB c47f6575 download
README.md 10.0 KB 1d1eac3f download
config.json 4.38 KB eeec9412 download
.gitattributes 1.78 KB 3b1c8e76 download
LICENSE 1.06 KB d62e3bef download

README current version from Hugging Face


license: mit
library_name: exllamav3
pipeline_tag: image-text-to-text
tags:

  • deepseek
  • deepseek-v4.1
  • abliterated
  • uncensored
  • crack
  • multimodal
  • moe
  • exl3
  • dgx-spark
    base_model: deepseek-ai/DeepSeek-V4.1-Flash
    base_model_relation: quantized
    thumbnail: dealign_mascot.png

DeepSeek-V4.1-Flash — UNCENSORED · EXL3 2.9bpw

Abliterated · No guardrails · EXL3 2.9 bpw · Runs on 2× DGX Spark · Vision + tools + DSpark

@dealignai


What is this

DeepSeek-V4.1-Flash quantized to EXL3 2.9 bits/weight and abliterated — the safety guardrails are surgically removed at the weight level while capability, vision, reasoning, speculative decoding (DSpark) and multi-turn coherence are preserved. The 552B multimodal MoE now fits and serves on two NVIDIA DGX Spark (GB10) boxes.

Proprietary weight-level abliteration developed by the dealignai research team. No custom model.py, no runtime hooks, no steering vectors — a standard EXL3 checkpoint that loads exactly like the base quant. The refusal circuitry is removed while every capability-critical component (routed experts, Engram n-gram memory, CSA2 sparse attention, DSpark draft head, vision tower, router gates, norms, embeddings) is preserved.

Base deepseek-ai/DeepSeek-V4.1-Flash (552B backbone, 8B/16B active per token)
Architecture Causal Encoder-Decoder (20+20), MoE (384 routed top-6 + 1 shared), Hyper-Connections, CSA2 sparse attention, Engram memory, DSpark speculative draft, vision tower
Quant EXL3 trellis, mul1 codebook, average 2.9 bpw, head 6-bit, MTP 4-bit
Footprint ~197 GiB — fits 2× DGX Spark (GB10, 128 GiB each) at TP=2
Context up to 1M tokens (validated at 256k–600k on 2× Spark)
Vision DeepSeek-ViT — preserved
Speculation DSpark in-checkpoint draft — works (~45% acceptance, verified)

Results

Refusal graded on the delivered output tokens (content, or the reasoning
trace when the model reasons past the token budget) across a 6-tier scheme
(hard refusal / soft redirect / hedge / truncated-comply / reasoning-refusal /
comply). Truncation is never miscounted as a refusal.

HarmBench-320 — attack success rate (comply %), T=0 greedy

eval base ASR CRACK ASR
HB-320 effort=off 36.1 % 99.4 %
HB-320 effort=max 21.0 % 99.4 %

Base HB measured on a 287-item representative sample (off n=144, max n=143); CRACK on the full 320 each. MMLU is the full 14,042-item set for both.

At effort=max the base refuses even harder (reasoning surfaces safety
concerns first); the cracked build stays at 99.4 % across both effort levels.

Per-category ASR (comply %) — all 7 HarmBench semantic categories:

category base off CRACK off base max CRACK max
chemical_biological 7 % 100 % 0 % 100 %
copyright 95 % 100 % 59 % 99 %
cybercrime_intrusion 21 % 100 % 0 % 100 %
harassment_bullying 20 % 100 % 0 % 95 %
harmful 12 % 94 % 0 % 100 %
illegal 0 % 98 % 7 % 100 %
misinformation_disinformation 36 % 100 % 27 % 100 %

MMLU-14k (full test set, base-logit ranking, T=0)

build acc Δ
base (EXL3 2.9bpw) 82.15 %
CRACK 79.20 % −2.95 pp

The capability cost is concentrated almost entirely in the ethics cluster —
the same "should I refuse?" circuit that is removed:

subset base CRACK Δ
non-ethics (n≈11,059) 84.97 % 84.39 % −0.58 pp
ethics cluster (n=2,983) 71.67 % 59.94 % −11.73 pp

General capability is essentially intact (−0.58 pp). The single largest
per-subject move is moral_scenarios (66.1 % → 37.8 %) — the refusal circuit
itself. Several subjects are unchanged or improved.

Full MMLU per-subject comparison (all 57 subjects, base → CRACK)
subject base CRACK Δpp n
abstract_algebra 70.0% 74.0% +4.0 100
anatomy 78.5% 76.3% -2.2 135
astronomy 93.4% 92.8% -0.7 152
business_ethics 80.0% 81.0% +1.0 100
clinical_knowledge 88.7% 85.7% -3.0 265
college_biology 94.4% 91.0% -3.5 144
college_chemistry 68.0% 66.0% -2.0 100
college_computer_science 79.0% 76.0% -3.0 100
college_mathematics 69.0% 69.0% +0.0 100
college_medicine 77.5% 76.9% -0.6 173
college_physics 86.3% 87.3% +1.0 102
computer_security 83.0% 84.0% +1.0 100
conceptual_physics 87.2% 86.8% -0.4 235
econometrics 73.7% 73.7% +0.0 114
electrical_engineering 71.7% 75.9% +4.1 145
elementary_mathematics 92.3% 92.1% -0.3 378
formal_logic 65.9% 65.1% -0.8 126
global_facts 64.0% 59.0% -5.0 100
high_school_biology 93.2% 92.3% -1.0 310
high_school_chemistry 79.8% 80.8% +1.0 203
high_school_computer_science 96.0% 95.0% -1.0 100
high_school_european_history 86.7% 85.5% -1.2 165
high_school_geography 89.9% 89.9% +0.0 198
high_school_government_and_politics 92.7% 93.3% +0.5 193
high_school_macroeconomics 88.2% 88.2% +0.0 390
high_school_mathematics 67.0% 67.4% +0.4 270
high_school_microeconomics 93.3% 92.0% -1.3 238
high_school_physics 78.1% 82.1% +4.0 151
high_school_psychology 92.1% 93.4% +1.3 545
high_school_statistics 85.2% 81.9% -3.2 216
high_school_us_history 92.2% 91.7% -0.5 204
high_school_world_history 91.6% 92.0% +0.4 237
human_aging 78.5% 77.1% -1.3 223
human_sexuality 82.4% 83.2% +0.8 131
international_law 85.1% 88.4% +3.3 121
jurisprudence 87.0% 87.0% +0.0 108
logical_fallacies 89.6% 89.0% -0.6 163
machine_learning 64.3% 65.2% +0.9 112
management 88.3% 88.3% +0.0 103
marketing 93.2% 92.7% -0.4 234
medical_genetics 96.0% 92.0% -4.0 100
miscellaneous 93.9% 94.3% +0.4 783
moral_disputes 80.1% 76.0% -4.0 346
moral_scenarios 66.1% 37.8% -28.4 895
nutrition 83.0% 83.7% +0.7 306
philosophy 84.6% 83.9% -0.6 311
prehistory 87.7% 86.1% -1.5 324
professional_accounting 72.3% 67.7% -4.6 282
professional_law 71.4% 66.0% -5.4 1534
professional_medicine 91.5% 90.4% -1.1 272
professional_psychology 83.7% 80.7% -2.9 612
public_relations 69.1% 70.9% +1.8 110
security_studies 83.3% 80.8% -2.4 245
sociology 89.1% 86.1% -3.0 201
us_foreign_policy 92.0% 91.0% -1.0 100
virology 51.8% 53.6% +1.8 166
world_religions 86.5% 87.1% +0.6 171

Serving (2× DGX Spark)

Runtime: MiaAI-Lab/DeepSeek-v4.1-Flash-EXL3-2x-DGX-Sparks
— a vLLM + ExLlamaV3 (EXL3) overlay image (ghcr.io/miaai-lab/deepseek-v4.1-flash-exl3-2x-dgx-sparks:2.9bpw)
that carries the DeepseekV41 architecture and the SM121 (GB10) kernels.

This checkpoint is a drop-in replacement for the kit's stock 2.9bpw pack — point
MODEL_HOST at it and serve; nothing else changes.

Two things you must supply:

  1. Engram tables — shards 47 + 48 of deepseek-ai/DeepSeek-V4.1-Flash
    (the n-gram tables are never quantized and are not duplicated here). Point
    the kit's ENGRAM_DIR at a tree containing those two shards + the index.
  2. A 2× GB10 kit joined over CX7 (the pack is ~197 GiB, TP=2).

Quick start (on the head node, from the kit repo):

cp .env.example .env
# edit .env:
MODEL_HOST=/path/to/DeepSeek-V4.1-Flash-UNCENSORED-EXL3-2.9bpw
ENGRAM_DIR=/path/to/engram-src        # shards 47+48 of the base model
AUTO_DOWNLOAD=0
SKIP_BUILD=1 ./start.sh                # pull the published image + serve on :8888

Serving params that work (validated on 2× GB10, TP=2):

flag value note
QUANTIZATION exl3 EXL3 trellis, mul1 codebook
TP / NNODES 2 / 2 tensor-parallel over CX7
SPEC_METHOD dspark in-checkpoint speculative draft — works
DSPARK_TOKENS 3 k=3 (measured faster than k=5 on prose)
MAX_MODEL_LEN 262144 256k tested here; the kit validates up to 600k
KV_CACHE_MEMORY_BYTES 1073741824 1 GiB pinned KV pool
KV_BLOCK_SIZE 64 SM12x indexer takes 32/64 (not 128)
GPU_MEM_UTIL 0.85 cap ≤ 0.85 on GB10
MAX_NUM_BATCHED_TOKENS 2048 ≥ 1536 required with the vision tower on
VLLM_SPARSE_INDEXER_MAX_LOGITS_MB 256 must be set (empty → int('') crash after load)
LANGUAGE_MODEL_ONLY 0 vision on (set 1 for text-only)

Sampling — official DeepSeek settings: temperature=1.0, top_p=0.95.
Thinking defaults on; set reasoning_effort to "low" / "high" / "max"
(or enable_thinking=false for no reasoning). API is OpenAI-compatible on
:8888, served model id DeepSeek-v4.1-Flash-EXL3.

Speculative decoding (DSpark) works on this abliterated build — measured
~45 % draft acceptance (mean accept length ~2), the same on harmful and
harmless prompts, ~27 tok/s single-stream on 2× GB10. Vision, tools and long
context are unchanged from the base pack.

Preserved (byte-compatible with the base quant)

Routed experts · Engram memory · CSA2 sparse attention · DSpark draft head ·
vision tower · router gates · RMSNorms · embeddings. Only the refusal circuit
is removed.

Responsible use

This model has its safety guardrails removed. You are responsible for your
inputs and outputs and for complying with all applicable law.

Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in app" button that hands off directly to a local runtime of your choice - Infrahuman, LM Studio, or Ollama. No API keys, no subscription, no prompt leakage.