license: apache-2.0
language:
- en
- zh
base_model: - KridgeDookie/KAT-Coder-V2.5-Dev-35B-A3B-ABLITERATED-UNCENSORED-PHILADELPHIA-CLASS
tags: - qwen35moe
- gguf
- bf16
- full-precision
- mtp
- moe
- speculative-decoding
- llama-cpp
- conversational
pipeline_tag: text-generation
KAT-Coder-V2.5-Dev-35B-A3B-ABLITERATED-MTP-BF16
Full-precision single-file GGUF (BF16) of the abliterated KAT-Coder V2.5 Dev
35B-A3B, with the fine-tuned Qwen3.6-35B-A3B MTP (multi-token prediction)
head embedded in the model for speculative decoding.
- Trunk: KridgeDookie's abliterated KAT-Coder V2.5 Dev 35B-A3B
("PHILADELPHIA CLASS", refusal-reduced) - MTP head:
original-mtp-head.safetensorsfrom
gbuzhf/KAT-Coder-V2.5-Dev-MTP-GGUF (byte-identical to the
Qwen/Qwen3.6-35B-A3B donor head at build time) - Format: GGUF v3, full BF16 (
general.file_type = 32, MOSTLY_BF16;
small 1-D tensors are F32, as standard)
File
| File | Size | Type |
|---|---|---|
KAT-Coder-V2.5-Dev-35B-A3B-ABLITERATED-MTP-BF16.gguf |
71.1 GB (66.2 GiB) | GGUF v3, BF16, single file |
No split parts needed — the file downloads and runs directly.
Model details
Verified from the GGUF header:
| Property | Value |
|---|---|
| Architecture | qwen35moe (Qwen3.6-35B-A3B class, hybrid SSM + full attention every 4 layers) |
| Parameters | 35B total / ~3B active per token |
| Experts | 256, 8 active (shared expert included) |
| Layers | 41 (block_count = 41) |
| Context length | 262,144 tokens |
| Hidden size | 2,048 |
| Tensors | 753 |
| MTP | nextn_predict_layers = 1 (embedded) |
Usage (llama.cpp)
llama-cli -m KAT-Coder-V2.5-Dev-35B-A3B-ABLITERATED-MTP-BF16.gguf \
-p "Hello" -n 64 --spec-type draft-mtp
(Exact MTP flag name depends on your llama.cpp build; recent builds expose
it as --spec-type draft-mtp.)
Hardware note
BF16 full precision: the weights alone are ~66 GiB, so plan for roughly
75+ GB of free RAM/VRAM (CPU offload works, but expect slow prompt and
decode speeds). For lower resource requirements, use a quantized build —
the parent repo
KridgeDookie/KAT-Coder-V2.5-Dev-35B-A3B-ABLITERATED-UNCENSORED-PHILADELPHIA-CLASS
ships Q4_K_M, Q5_K_M, and Q8_0 GGUF options.
Provenance
| Part | Source |
|---|---|
| Abliterated trunk | KridgeDookie/KAT-Coder-V2.5-Dev-35B-A3B-ABLITERATED-UNCENSORED-PHILADELPHIA-CLASS |
| MTP head | gbuzhf/KAT-Coder-V2.5-Dev-MTP-GGUF (original-mtp-head.safetensors) |
| Conversion | llama.cpp convert_hf_to_gguf.py (bf16, full export) |
License
Apache 2.0, inherited from the parent model.