license: other
license_name: kimi-k3
base_model: Uniboshi/Kimi-K3-Abliterated-V1
base_model_relation: quantized
pipeline_tag: text-generation
tags:
- gguf
- kimi-k3
- q2_k
- q4_k
- moe
- abliterated
- uncensored
language: - en
- ja
- zh
extra_gated_prompt: |
This is a quantized derivative of an intentionally reduced-refusal model.
Before requesting access, you must read the model cards for
Uniboshi/Kimi-K3-Abliterated-V1, moonshotai/Kimi-K3, and
GrEarl/Kimi-K3-GGUF, together with the complete Kimi K3 License. Request
access only if you understand the reduced-safeguard nature of the model and
accept responsibility for lawful, controlled use.
extra_gated_fields:
"I have read all three required upstream and quantization model cards": checkbox
"I have read and accept the Kimi K3 License and applicable upstream conditions": checkbox
"I understand that this model is not a safety boundary and may produce harmful content": checkbox
"I accept responsibility for access control, output review, deployment, and legal compliance": checkbox
Kimi-K3-Abliterated-V1 Q2_K GGUF
A 94-part mixed Q2_K/Q4_K GGUF quantization of
Uniboshi/Kimi-K3-Abliterated-V1,
built with the quantization policy and split layout of
GrEarl/Kimi-K3-GGUF v3.
[!WARNING]
The base model intentionally reduces refusal behavior. It may generate
inaccurate, offensive, unsafe, or illegal suggestions. This model is not a
safety boundary. Users are responsible for access controls, output review,
downstream use, and compliance with applicable law and license terms.
Required reading and acknowledgement
Do not request access, download, run, redistribute, or deploy this model until
you have read these resources in full:
- Uniboshi/Kimi-K3-Abliterated-V1 —
the full-weight base model, its reduced-refusal intent, usage notes, and
limitations. - moonshotai/Kimi-K3 — the
original architecture, intended deployment guidance, and
Kimi K3 License. - GrEarl/Kimi-K3-GGUF — the Q2
v3 quantization design, llama.cpp runtime requirements, measured evidence,
and quantization limitations reused by this build.
By requesting access or using this repository, you confirm that you have read
those cards and the complete license, understand that safeguards have
intentionally been weakened, and accept responsibility for controlled and
lawful use. This acknowledgement does not replace or modify the upstream
license. Under the Hugging Face gated-model workflow, an access request also
shares the requester's account identity and contact information with the
repository owner.
Lineage and release identity
| Item | Value |
|---|---|
| Hugging Face base model | Uniboshi/Kimi-K3-Abliterated-V1 |
| HF relationship | quantized |
| Original model | moonshotai/Kimi-K3 |
| Structural and quantizer template | GrEarl/Kimi-K3-GGUF v3 |
| Parts | 94 |
| Total size | 864.81 GiB |
| Quantization layout | mixed Q2_K/Q4_K |
| Architecture | kimi-k3 |
| Modality | text only |
This release is a quantized derivative of Uniboshi's checkpoint.
It is not a fine-tune of GrEarl/Kimi-K3-GGUF.
The release build does not apply a refusal vector.
It does not apply a QAware correction or reconstruct any weight from a rank-one
approximation. No experimental post-quantization abliteration contributes to
the published bytes.
Direct quantization method
The build uses a change-aware materialization strategy. The pinned Uniboshi
checkpoint changes 280 writer tensors relative to its Kimi-K3 source lineage.
One changed tensor is the vision-only mm_projector.proj.2.weight, which is
outside this text GGUF. The remaining 279 text tensors are fetched directly
from the pinned Uniboshi BF16 safetensors and quantized exactly once from BF16
to Q4_K.
| Uniboshi tensor family | GGUF role | Count | Output type |
|---|---|---|---|
embed_tokens |
token_embd.weight |
1 | Q4_K |
o_proj |
attention output writers | 93 | Q4_K |
down_proj |
dense/shared MLP down writers | 93 | Q4_K |
routed_expert_up_proj |
Stable LatentMoE routed up writers | 92 | Q4_K |
| Total | 279 | Q4_K |
The exact BF16 input for those tensors is 29.8457 GiB and their packed Q4_K
payload is 8.3941 GiB. Tensors established as unchanged in the pinned
full-weight comparison reuse the byte-identical v3 quantization payload for the
same source weight. This avoids downloading and re-quantizing several
terabytes of unchanged data while producing the intended Uniboshi quantized
checkpoint. It is a build optimization, not a model merge or a transfer of an
inferred refusal direction.
The v3 policy keeps quality-sensitive writers in Q4_K and the large routed
expert stacks in Q2_K. None of the Q2_K routed-expert payloads is modified by
the 279-tensor direct-quantization pass.
Integrity verification
The release pipeline fails closed on revision drift, tensor-name or shape
mismatch, unexpected quantization type, packed-size mismatch, and incomplete
downloads. For every output part it records and checks:
- the pinned source LFS size and SHA-256;
- the exact Uniboshi BF16 byte range and SHA-256 for every materialized tensor;
- the output payload SHA-256 for every directly quantized tensor;
- unchanged payload hashes against the structural template;
- file size, tensor metadata, split metadata, and a GGUFReader reopen;
- the final full-file SHA-256.
| Audited item | Result |
|---|---|
| Pinned GGUF template parts | 94 / 94 |
| Output parts reopened with GGUFReader | 94 / 94 |
| Directly quantized Uniboshi text tensors | 279 |
| Q2_K routed-expert payloads modified | 0 |
| Output file hashes recorded | 94 / 94 |
The repository includes release-manifest.json with all 94 file sizes,
output hashes, template-source hashes, pinned revisions, and the materialization
declaration.
Runtime and evaluation status
Runtime, throughput, refusal behavior, and quality results from the retired
experimental QAware/refusal-transfer candidate do not apply to this release
and are not claimed here. At initial publication, this exact uploaded revision
has passed the weight-space and file-integrity gates above but has not completed
an end-to-end runtime or behavioral evaluation. Runtime timings, raw outputs,
and benchmark results will be added only after direct testing of this revision;
no behavioral score is inferred from weight-space verification alone.
Usage
Place all 94 GGUF files in one directory and pass the first part to a
Kimi-K3-capable llama.cpp build:
llama-cli -m Kimi-K3-Q2_K-00001-of-00094.gguf -p "..."
Follow the runtime requirements and split guidance in the
GrEarl/Kimi-K3-GGUF model card.
Do not assume every llama.cpp release or prebuilt binary supports Kimi-K3.
Limitations
- This is a text-only GGUF; the vision projector is not included.
- Abliteration is intended to reduce refusal behavior but does not guarantee
unrestricted compliance for every prompt. - Uniboshi's results describe its full-weight checkpoint unless explicitly
re-measured on this exact quantized revision. - Quantization can change logits, reasoning paths, instruction following, and
refusal behavior. No base-model benchmark should be presented as a score for
this release without a direct rerun. - The model does not provide deployment safety controls.
Credits
- Uniboshi for Kimi-K3-Abliterated-V1, the base checkpoint directly
quantized by this release. - Moonshot AI for Kimi-K3.
- The llama.cpp Kimi-K3 contributors, including the work merged through
PR #26185. - Refusal-direction and abliterated-model researchers whose work established
the broader technique used by the base checkpoint.
No upstream author or project is claimed to endorse this derivative.
License
The Kimi K3 License
applies and is reproduced in this repository. It requires preservation of its
copyright and permission notice and compliance with applicable law. It also
contains conditions for high-revenue Model-as-a-Service businesses and
prominent Kimi K3 attribution for certain very large commercial products. Read
the complete license yourself; this summary is not legal advice and does not
replace the license. This repository grants no additional rights and waives no
upstream condition.