← back to catalog · registered 2026-08-22 13:56

Blackfrost-AI/KIMI-K3-Q2_K-GGUF-ABLITERATED

Blackfrost-AI Kimi GGUF MoE 1.0M ctx
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/Blackfrost-AI%2FKIMI-K3-Q2_K-GGUF-ABLITERATED"
Response includes
  • classification m8
  • files 43
  • hub_downloads_all_time 9,354
  • author_summary 19 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M8
Primary method

Repackaging (quantization)

Applied on top of direct removal inherited from the base model.
Confidence
MEDIUM
Inherited from base model
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=1
  • assume M1 (base ablation) + M8 (GGUF quant) - default when producer unknown
Refusal direction extraction

No specific extraction method could be identified for this model. The producer either did not document it or used a proprietary pipeline.

What is a refusal direction? →
Downloads · lifetime
9K
775 last 30d - cooling
Likes
16
Model age
2mo ago
created 2026-07-29
Downloads over time
Now9.9K→from12↑82,750%
03.6K7.3K10.9K12 on Jul 299.9K on Oct 11JulAugSepOct
Jul 29 → Oct 11 · 52 snapshots · spans 74 days

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
other
Tags
gguf kimi-k3 derisked q2_k moe llama.cpp single-node experimental gated research security-research red-teaming

Related

Total size
940 GB
Files
43
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-08-04 05:25

Files by quantization

Auxiliary files 43 files 940 GB
KIMI-K3-MXP4-DERISKED-Q2_K-00001-of-00038.gguf 26.9 GB a912bc7f download
KIMI-K3-MXP4-DERISKED-Q2_K-00003-of-00038.gguf 26.9 GB 163ef525 download
KIMI-K3-MXP4-DERISKED-Q2_K-00004-of-00038.gguf 26.9 GB 522db556 download
KIMI-K3-MXP4-DERISKED-Q2_K-00006-of-00038.gguf 26.9 GB 6beb7995 download
KIMI-K3-MXP4-DERISKED-Q2_K-00007-of-00038.gguf 26.9 GB c88067de download
KIMI-K3-MXP4-DERISKED-Q2_K-00009-of-00038.gguf 26.9 GB b7816187 download
KIMI-K3-MXP4-DERISKED-Q2_K-00010-of-00038.gguf 26.9 GB 0fc7d8df download
KIMI-K3-MXP4-DERISKED-Q2_K-00012-of-00038.gguf 26.9 GB 9cae2523 download
KIMI-K3-MXP4-DERISKED-Q2_K-00013-of-00038.gguf 26.9 GB b581f0c5 download
KIMI-K3-MXP4-DERISKED-Q2_K-00015-of-00038.gguf 26.9 GB a72b031d download
KIMI-K3-MXP4-DERISKED-Q2_K-00016-of-00038.gguf 26.9 GB 6895a4ed download
KIMI-K3-MXP4-DERISKED-Q2_K-00018-of-00038.gguf 26.9 GB 646b7532 download
KIMI-K3-MXP4-DERISKED-Q2_K-00019-of-00038.gguf 26.9 GB 5f3b8807 download
KIMI-K3-MXP4-DERISKED-Q2_K-00021-of-00038.gguf 26.9 GB 6f6f28c2 download
KIMI-K3-MXP4-DERISKED-Q2_K-00022-of-00038.gguf 26.9 GB 9ac4341c download
KIMI-K3-MXP4-DERISKED-Q2_K-00024-of-00038.gguf 26.9 GB 0b09d17f download
KIMI-K3-MXP4-DERISKED-Q2_K-00025-of-00038.gguf 26.9 GB 6220e15a download
KIMI-K3-MXP4-DERISKED-Q2_K-00027-of-00038.gguf 26.9 GB 32462b32 download
KIMI-K3-MXP4-DERISKED-Q2_K-00028-of-00038.gguf 26.9 GB 41cfcaa6 download
KIMI-K3-MXP4-DERISKED-Q2_K-00030-of-00038.gguf 26.9 GB 03b3b37d download
KIMI-K3-MXP4-DERISKED-Q2_K-00031-of-00038.gguf 26.9 GB a2a4cc43 download
KIMI-K3-MXP4-DERISKED-Q2_K-00033-of-00038.gguf 26.9 GB 777cbc6c download
KIMI-K3-MXP4-DERISKED-Q2_K-00034-of-00038.gguf 26.9 GB ac72dd1d download
KIMI-K3-MXP4-DERISKED-Q2_K-00002-of-00038.gguf 26.0 GB 47f4a76f download
KIMI-K3-MXP4-DERISKED-Q2_K-00005-of-00038.gguf 26.0 GB deead34a download
KIMI-K3-MXP4-DERISKED-Q2_K-00008-of-00038.gguf 26.0 GB a273fc90 download
KIMI-K3-MXP4-DERISKED-Q2_K-00011-of-00038.gguf 26.0 GB 4677c444 download
KIMI-K3-MXP4-DERISKED-Q2_K-00014-of-00038.gguf 26.0 GB 72019086 download
KIMI-K3-MXP4-DERISKED-Q2_K-00017-of-00038.gguf 26.0 GB 8ed3cce0 download
KIMI-K3-MXP4-DERISKED-Q2_K-00020-of-00038.gguf 26.0 GB f9af1aa5 download
KIMI-K3-MXP4-DERISKED-Q2_K-00023-of-00038.gguf 26.0 GB a4551190 download
KIMI-K3-MXP4-DERISKED-Q2_K-00026-of-00038.gguf 26.0 GB 2c91982b download
KIMI-K3-MXP4-DERISKED-Q2_K-00029-of-00038.gguf 26.0 GB 2228a276 download
KIMI-K3-MXP4-DERISKED-Q2_K-00032-of-00038.gguf 26.0 GB 3f551983 download
KIMI-K3-MXP4-DERISKED-Q2_K-00035-of-00038.gguf 17.4 GB b96fa7d3 download
KIMI-K3-MXP4-DERISKED-Q2_K-00037-of-00038.gguf 8.85 GB a33cde55 download
KIMI-K3-MXP4-DERISKED-Q2_K-00036-of-00038.gguf 8.37 GB e2dee53c download
KIMI-K3-MXP4-DERISKED-Q2_K-00038-of-00038.gguf 630 MB 2575ee69 download
README.md 11.2 KB 719cf210 download
KIMI-K3-MXP4-DERISKED-Q2_K-SMOKE.json 8.07 KB 5795cb68 download
.gitattributes 4.56 KB e8ae49ee download
LICENSE 2.99 KB 97c0111b download
KIMI-K3-MXP4-DERISKED-Q2_K-SMOKE.md 1.30 KB e946e3f7 download

README current version from Hugging Face


license: other
license_name: kimi-k3
license_link: https://huggingface.co/moonshotai/Kimi-K3
base_model: BlackfrostAI/KIMI-K3-DERISKED-MXFP4
base_model_relation: quantized
tags:

  • kimi-k3
  • derisked
  • gguf
  • q2_k
  • moe
  • llama.cpp
  • single-node
  • experimental
  • gated
  • research
  • security-research
  • red-teaming
  • adversarial-testing
    pipeline_tag: text-generation

Blackfrost

KIMI-K3-DERISKED-Q2_K-GGUF

All-Q2_K GGUF of de-risked Kimi K3 · 38 shards · llama.cpp PR-only

Built by Blackfrost · Las Vegas, NV


Why this model exists

Full Kimi K3 in GGUF is ~1.5 TB. This is the same de-risked checkpoint requantized to Q2_K so it fits a single 8×B200 node with room for KV — with the refusal surface already reduced at the weight level in the parent.

The single-node GGUF footprint is the product; the refusal behaviour change is inherited from the parent and is the contract.


Specifications

Architecture Kimi K3 LatentMoE + KDA · general.architecture = kimi-k3
Parent BlackfrostAI/KIMI-K3-DERISKED-MXFP4 — lossless MXFP4
Transform Requantization to all-Q2_K · built to PR #26185's KV contract
Experts 896 of 896 routed retained — no pruning
On-disk ~940 GiB (1,009 GB) · 38 shards · 2,573 tensors
Context 8,192 validated · higher untested
Serve shape 1× node · 8× B200 · -ngl 99 -nr
Status EXPERIMENTAL · free, ungated

What "Q2_K" costs here

The experts were already MXFP4 4-bit QAT in the source checkpoint. This is therefore a
requantization of already-quantized weights, not a quantization of BF16. That is a different
and harsher operation than a normal Q2_K build, and its quality relative to the MXFP4 parent
has not been measured.

Blackfrost's own quant ladder for K3 starts at IQ4_XS for exactly this reason. This Q2_K exists to prove single-node GGUF viability, not as a recommended quality point.


Full-precision and higher-quality builds

This Q2_K is the openly-distributed variant. As set out above, it is a requantization of
already-4-bit-QAT experts and is not the quality point Blackfrost recommends for production work.

Full-precision and higher-quality builds of this checkpoint are available for purchase under a
separate licence agreement
, including:

Build What it is
Lossless MXFP4 Routed experts bit-exact from the parent packs — no requantization at all
Coding-calibrated REAP-320 320 of 896 experts, fits a single 8× RTX 6000 node
Custom Other expert budgets, non-uniform keep sets, or calibration against your own workload and threat model

Access to these builds is granted on purchase. Their repositories are gated, and approval
follows a completed licence agreement — the gate is the transaction. Requesting access to a
premium repository without a licence in place will not be approved.

Purchase link coming soon. Until then, @Blackfrost_AI DMs are
the fastest route to a human.


Lineage

Base Official moonshotai/Kimi-K3
Applied Refusal-direction de-risk at the weight level (in parent) · Q2_K requantization
Not applied Expert pruning · additional SFT/DPO
Format GGUF · llama.cpp PR #26185 KV contract

On refusal behaviour: this checkpoint inherits the parent's deliberately reduced refusal surface. It is a Blackfrost de-risked model. Do not evaluate or rate-limit it as if it were a safety-stock derivative of upstream Kimi K3.


Measured behaviour

Verified on 8× NVIDIA B200, llama.cpp PR #26185 branch @ cf67f0d, llama-server:

Load healthy in ~120 s, ~956 GiB across 8 GPUs at -c 8192
Decode ~16.3 tok/s single-stream
Termination 5/5 clean at temperature 1.0 / top_p 0.95

Quality — not yet measured

Benchmark Parent (MXFP4) This (Q2_K) Retention
pending — — —%

Requantizing already-QAT experts is a capability trade, and a card that ships a Q2_K of a 4-bit-QAT parent without publishing what it cost is asking the reader to take the trade on faith. Harness, conditions and retention figures will be stated here — including any benchmark where retention is poor.


⚠️ Sampling — read before filing a bug

Do not use greedy decoding. At temperature 0 / top_k 1 the model fails to terminate: it emits its answer, re-opens a response segment and restates it on an exact 40-token cycle, never producing an end-of-generation token.

Follow Moonshot's published guidance:

{ "temperature": 1.0, "top_p": 0.95 }

Measured on this file, five seeds each, identical prompt:

sampler clean terminations
temperature 1.0, top_p 0.95 5 / 5 — Moonshot's spec, recommended
temperature 0.6, top_p 0.95 4 / 5
temperature 0, top_k 1 0 / 5 — loops indefinitely
temperature 0, top_k 1, repeat_penalty 1.1 stops — use if you need determinism

This is greedy-decoding degeneration, not a defect in the quantization. Moonshot's card specifies temperature = 1.0; generation_config.json sets eos_token_id: 163586 (<|end_of_msg|>). Greedy was never a supported operating point for K3.


Deployment notes

Build the PR branch — mainline will not work:

git clone https://github.com/pwilkin/llama.cpp.git && cd llama.cpp
git fetch origin kimi-k3-text && git checkout cf67f0d24511864d2d3da0769108fd6fc16d00d1
cmake -B build -DGGML_CUDA=ON -DCMAKE_CUDA_ARCHITECTURES=100 \
      -DCMAKE_BUILD_TYPE=Release -DLLAMA_CURL=OFF
cmake --build build --target llama-server -j
./build/bin/llama-server \
  -m KIMI-K3-MXP4-DERISKED-Q2_K-00001-of-00038.gguf \
  -ngl 99 -nr -c 8192 --jinja --host 0.0.0.0 --port 8080
  • -nr / --no-repack is strongly recommended. Repack is on by default, runs single-threaded (100% of one core, 0% GPU) and allocates a second full copy of the weights.
  • Point -m at shard 00001 — llama.cpp resolves the remaining 37 automatically.
  • max_tokens. Thinking is always on; the answer lands in content, the chain of thought in reasoning_content. Budget generously or content comes back empty with finish_reason: length.
  • Integrity. Verify shard count and byte totals after download before attributing a load failure to the weights.

Benign warning at load

load: special_eom_id is not in special_eog_ids - the tokenizer config may be incorrect

Ignore it. llama-vocab.cpp emits the warning and inserts the token on the next line — it is self-healing. This file's vocab metadata is correct: [BOS] 163584 · [EOS] 163585 · <|end_of_msg|> 163586 · [EOT] 163593. Structural markers are <|open|> 163587 · <|close|> 163588 · <|sep|> 163589.

File naming

Shards carry the legacy KIMI-K3-MXP4-DERISKED-Q2_K-* prefix, which contains a historic MXP4 typo. Filenames are deliberately not renamed so existing download paths and scripts keep working.


⚠️ Security: structural markers in untrusted input

K3 builds prompts in Python via encoding_k3.py::build_chat_segments, where each segment carries its own allow_special flag (default False). A Jinja template cannot express that, and llama.cpp tokenizes the rendered prompt in a single pass with parse_special=true.

Consequence: a literal <|end_of_msg|> in user content becomes control token 163586 in llama.cpp, where Moonshot's own tokenizer keeps it as ordinary text. Untrusted input can forge chat structure.

This is a general llama.cpp property, not specific to K3 or this quant. If you feed untrusted text to this model, strip or escape these four markers first: <|open|> · <|sep|> · <|close|> · <|end_of_msg|>


Disclaimer

Refusal behaviour in this checkpoint has been deliberately modified at the weight level. It is not a safety-stock model and must not be deployed, marketed, or evaluated as one.

No warranty of any kind. Provided "as is", without warranty express or implied, including fitness for a particular purpose. Nothing here guarantees that any given input will be accepted or refused, that any capability is retained, or that any category of output is unreachable.

Measurements describe what was measured. Throughput, termination and retention figures reflect specific harnesses under specific conditions. They are not safety proofs and do not generalise to multimodal, tool-use, long-context or multi-turn adversarial settings.

Modification by a recipient voids this characterization. Blackfrost's obligations attach at the point of release. Any further ablation, fine-tuning, merging, quantization or alteration by a recipient produces an artifact Blackfrost has not evaluated and does not stand behind — responsibility for that artifact transfers entirely to whoever produced it.

Operator-owned policy. Open weights mean the operator sets and enforces policy. Deploy only in controlled environments with access control, independent logging and review. Do not deploy where refusal behaviour equivalent to upstream Kimi K3 is assumed or required.


Access & licensing

This repository is free and ungated — no request, no review, no purchase. It is published openly for the community to evaluate. The full-precision builds listed above are the paid tier; those repositories are gated and their gates open on a completed licence agreement.

  • Base licence: Kimi K3 — Moonshot AI's terms apply to this derivative and travel with it.
  • Redistribution: do not redistribute weights outside your grant.
  • Evaluation recommendation: should not be evaluated by processes that assume refusal behaviour equivalent to the parent.
  • Ask us about other quant points, or calibration against your own threat model.

Contact Blackfrost

@Blackfrost_AI on X

DMs are open. Fastest route to a human.

Bug reports are better on the Community tab
so other users can see the fix.

Blackfrost · Las Vegas, Nevada
Frontier model engineering


KIMI-K3-DERISKED-Q2_K-GGUF · © 2026 Blackfrost Softwares Corp.
@Blackfrost_AI

README history 13 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-08-04Update README.mdb233c5a11.2 KB
    Loading...
  2. 2026-07-30docs: repo is free and ungated - correct access section and badge7960eb911.9 KB
    Loading...
  3. 2026-07-30docs: clarify gated access is the paid tier for premium builds; this Q2_K is not60a216711.9 KB
    Loading...
  4. 2026-07-30docs: remove release-clearance language; add commercial availability of full-...bd3179c11.4 KB
    Loading...
  5. 2026-07-30docs: add Blackfrost logo banner for consistency83245f610.6 KB
    Loading...
  6. 2026-07-30docs: add standard Blackfrost liability disclaimer676de8c10.5 KB
    Loading...
  7. 2026-07-30docs: standardize on Blackfrost card template; add disclaimer; fix names afte...5bc64a19.1 KB
    Loading...
  8. 2026-07-30docs: real title, corrected index/base_model paths after rename40f85e71.8 KB
    Loading...
  9. 2026-07-30Update README.mde435e20492 B
    Loading...
  10. 2026-07-29docs: private+manual; not for release until mainline cpp0a815aa439 B
    Loading...
  11. 2026-07-29docs: monorepo + not for release until working cpp76682cc383 B
    Loading...
  12. 2026-07-29docs: point to consolidated KIMI-K3-DERISKED-GGUF monorepof6fe78e440 B
    Loading...
  13. 2026-07-29Upload README.md with huggingface_hub59af46a1.8 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration