← back to catalog · registered 2026-08-22 13:56

pyrodog/CyberNeurova-DeepSeek-V4-Flash-abliterated-v2-IQ3_XXS-AS-GGUF

pyrodog Deepseek GGUF MoE second-order 1.0M ctx
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/pyrodog%2FCyberNeurova-DeepSeek-V4-Flash-abliterated-v2-IQ3_XXS-AS-GGUF"
Response includes
  • classification m8
  • files 6
  • hub_downloads_all_time 732
  • author_summary 2 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M8
Primary method

Repackaging (quantization)

Applied on top of direct removal inherited from the base model.
Confidence
MEDIUM
Inherited from base model
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=1
  • assume M1 (base ablation) + M8 (GGUF quant) - default when producer unknown
Refusal direction extraction

No specific extraction method could be identified for this model. The producer either did not document it or used a proprietary pipeline.

What is a refusal direction? →
Downloads · lifetime
732
175 last 30d - stable
Likes
2
Model age
2mo ago
created 2026-07-24
Downloads over time
Now838→from122↑587%
86361635910122 on Jul 22838 on Oct 11JulAugSepOct
Jul 22 → Oct 11 · 52 snapshots · spans 81 days

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
mit
Languages
en
Tags
gguf llama.cpp deepseek deepseek-v4 mixture-of-experts abliterated uncensored imatrix quantized apple-silicon metal macbook-pro

Related

Total size
95.5 GB
Files
6
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-07-24 16:09

Files by quantization

Auxiliary files 6 files 95.5 GB
CyberNeurova-DeepSeek-V4-Flash-abliterated-v2-IQ3_XXS-AS-F32Router.gguf 95.5 GB 4dd8f2c2 download
README.md 8.13 KB abdd5ce1 download
chat_template.jinja 5.28 KB fab0b0c8 download
.gitattributes 1.59 KB c8966963 download
LICENSE 1.04 KB d694b7ff download
SHA256SUMS 139 B 09a8ecd6 download

README current version from Hugging Face


license: mit
language:

  • en
    library_name: gguf
    pipeline_tag: text-generation
    base_model: cyberneurova/CyberNeurova-DeepSeek-V4-Flash-abliterated-GGUF
    base_model_relation: quantized
    tags:
  • gguf
  • llama.cpp
  • deepseek
  • deepseek-v4
  • mixture-of-experts
  • abliterated
  • uncensored
  • imatrix
  • quantized
  • apple-silicon
  • metal
  • macbook-pro
  • m5-max
  • 128gb
    model_name: CyberNeurova DeepSeek V4 Flash Abliterated v2 IQ3_XXS-AS F32Router

CyberNeurova DeepSeek V4 Flash Abliterated v2 — IQ3_XXS-AS F32Router GGUF

Tailor-made for the 128 GB M5 Max MacBook Pro. This quant was
designed, built, and benchmarked on that exact Apple-silicon machine. Its
95.48 GiB weights fit inside the tested Mac's 107.5 GiB Metal-visible
allocation pool while leaving about 12 GiB for runtime state and context.

This is a custom importance-matrix quantization of
cyberneurova/CyberNeurova-DeepSeek-V4-Flash-abliterated-GGUF.
It specifically targets the 128 GB unified-memory M5 Max MacBook Pro and
similar machines where the Q8 release is too large but preserving more quality
than the smallest quants is worthwhile.

This repository only contributes the quantization. CyberNeurova created the
abliterated v2 weights; DeepSeek created the original V4 Flash model.

Artifact

File Size Whole-file SHA-256
CyberNeurova-DeepSeek-V4-Flash-abliterated-v2-IQ3_XXS-AS-F32Router.gguf 102,521,722,976 bytes (95.48 GiB) 4dd8f2c2b3b14d857081f9b7ceeaa7c44564744c12cd1230e97b93fec76b593c

llama-gguf-hash tensor-data SHA-256:
12e66621bd390b79904463cf3cdb5b288bf7265352cf1b66740e211c8f39c7b3.

Quantization layout

Tensor group Type
Routed-expert down and up IQ3_XXS
Routed-expert gate IQ2_XXS
Attention, shared experts, token embedding, and output Q8_0
Router inputs F32
Small auxiliary tensors F32
Hash-routing tables I32

The installed llama-quantize 10090 dry run reported:

model size = 288244.36 MiB (8.50 BPW)
quant size =  97767.25 MiB (2.88 BPW)
imatrix entries = 129

The tensor layout was reconstructed by comparing the CyberNeurova Q8 GGUF
with bullerwins/DeepSeek-V4-Flash.IQ3_XXS-AS.gguf. Both contain the same
1,328 tensor names and dimensions. The published AS reference stores its 43
router-input tensors as BF16. Current llama.cpp excludes those tensors from
re-quantization, so this build preserves CyberNeurova's F32 routers instead.
That costs about 86 MiB and retains higher precision.

The complete per-tensor mapping is included at
recipe/iq3_xxs_as.tensor-types.txt.

Inputs and reproducibility

Input Bytes SHA-256
cyberneurova-DeepSeek-V4-Flash-abliterated-Q8_0.gguf 302,251,447,616 ffff4e8e526a490f4e68dd649f32f6bc1e25d80d2f5df343996b5a956f9490cc
DeepSeek-V4-Flash-chat-v2-routed-moe-ds4-aligned.dat 450,892,654 427f9d25dcc996fd8aad2c037c4eda8976996d105bc6ffe5b6608480f3b537b8

The importance matrix is the published Jedi chat-v2 routed-MoE matrix used for
Audrey Tang's aligned CyberNeurova re-quant.

The build used llama.cpp version 10090 (7347430f4) and:

llama-quantize \
  --allow-requantize \
  --imatrix DeepSeek-V4-Flash-chat-v2-routed-moe-ds4-aligned.dat \
  --tensor-type-file recipe/iq3_xxs_as.tensor-types.txt \
  --override-kv \
  general.name=str:CyberNeurova-DeepSeek-V4-Flash-Abliterated-v2-IQ3_XXS-AS-F32Router \
  cyberneurova-DeepSeek-V4-Flash-abliterated-Q8_0.gguf \
  CyberNeurova-DeepSeek-V4-Flash-abliterated-v2-IQ3_XXS-AS-F32Router.gguf \
  IQ3_XXS

Validation checked every tensor name, dimension, and requested tensor type.
It also hash-compared every F32, I32, and Q8_0 tensor expected to remain
byte-identical to the CyberNeurova Q8 source.

Download

hf download \
  pyrodog/CyberNeurova-DeepSeek-V4-Flash-abliterated-v2-IQ3_XXS-AS-GGUF \
  CyberNeurova-DeepSeek-V4-Flash-abliterated-v2-IQ3_XXS-AS-F32Router.gguf \
  chat_template.jinja \
  --local-dir .

Run with llama.cpp

The GGUF does not embed a chat template. The companion
chat_template.jinja is the exact template extracted
from a clean DeepSeek V4 Flash reference GGUF. Pass it explicitly; otherwise
llama.cpp may fall back to ChatML and produce misleadingly poor behavior.

The following is the tested 32K-context Apple Metal profile:

llama-server \
  -m CyberNeurova-DeepSeek-V4-Flash-abliterated-v2-IQ3_XXS-AS-F32Router.gguf \
  -ngl 999 \
  -fa on \
  -ctk q8_0 \
  -ctv q8_0 \
  -c 32768 \
  --jinja \
  --chat-template-file chat_template.jinja \
  --reasoning-format deepseek \
  --reasoning on

The model was tested on a 128 GB M5 Max. The 95.48 GiB weight file left about
12 GiB of that machine's 107.5 GiB Metal-visible allocation pool for runtime
state and context. Actual usable memory varies by system and background load.

Local A/B measurements

The custom quant was compared with Audrey Tang's 90.89 GiB mixed quant under
the same mainline llama.cpp Metal runtime. Both used all GPU layers, flash
attention, Q8 KV caches, a 32,768-token context, temperature 0, and seed
424242.

Measurement This IQ3_XXS-AS Audrey mixed
File size 95.48 GiB 90.89 GiB
Model load 6.03 s 6.02 s
512-token prompt processing 476.32 tok/s 335.32 tok/s
8,192-token prompt processing 325.82 tok/s 295.11 tok/s
128-token generation 11.48 tok/s 10.24 tok/s
Eight-chunk WikiText-2 perplexity 3.5382 3.9270
Strict local smoke tasks 6/12 5/12
26,664-token retrieval pass, 136.42 s pass, 145.34 s

For 4.59 GiB more storage, this quant was 42.05% faster on the 512-token
prompt, 10.41% faster on the 8,192-token prompt, and 12.09% faster in
generation. Its measured perplexity was 9.90% lower.

These are controlled measurements on one machine, not broad independent
benchmarks. The 12-task smoke suite is particularly small; a formatting-only
miss on each side makes the semantically lenient score a 7/12 tie. The speed
and perplexity results are the stronger reasons to prefer this quant.

Audrey's specialized ds4 runtime can execute the smaller mixed quant's
IQ2/Q4 experts much faster, but does not implement this model's IQ3 expert
dispatch. Use this quant for the measured fidelity/performance balance under
mainline llama.cpp; use the smaller mixed quant with ds4 when specialized
decode speed is the priority.

Intended use and limitations

This is an experimental research artifact for local inference, defensive
security research, academic study of refusal mechanisms, red-team evaluation,
and controlled coding-agent experiments.

The underlying model has had canonical refusal behavior deliberately reduced.
That does not make its outputs correct or safe. It can produce harmful,
insecure, biased, or fabricated content. Review generated code before running
it, use sandboxing for tools, and do not use the model to facilitate illegal
activity or harm.

Additional limitations:

  • Long-context behavior beyond 32K was not validated for this quant.
  • The benchmark covers one Apple-silicon machine and one llama.cpp build.
  • Quantization adds error on top of the abliterated source weights.
  • The GGUF requires the external chat template included in this repository.
  • Mainline and specialized DeepSeek V4 runtimes are evolving quickly.

License and acknowledgements

MIT, inherited from DeepSeek V4 Flash and the CyberNeurova release. See
LICENSE.

Thanks to:

  • DeepSeek-AI for V4 Flash
  • CyberNeurova for the v2 abliterated weights
  • Audrey Tang for the aligned comparison quant and runtime work
  • Jedi for the aligned importance matrix
  • bullerwins for the IQ3_XXS-AS tensor-layout reference
  • llama.cpp contributors for GGUF tooling and inference support

README history 3 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-07-24Emphasize 128 GB M5 Max MacBook Pro target5815d748.1 KB
    Loading...
  2. 2026-07-24Remove not-for-all-audiences gatec30e6767.8 KB
    Loading...
  3. 2026-07-24Add model card, license, checksum, and chat template637fa9c7.8 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration