← back to catalog · registered 2026-08-22 13:56

moeshawky/KAT-Coder-V2.5-Dev-35B-A3B-ABLITERATED-MTP-BF16

moeshawky 35B GGUF MoE second-order 262K ctx
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/moeshawky%2FKAT-Coder-V2.5-Dev-35B-A3B-ABLITERATED-MTP-BF16"
Response includes
  • classification m8
  • files 3
  • hub_downloads_all_time 345
  • author_summary 2 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M8
Primary method

Repackaging (quantization)

Applied on top of direct removal inherited from the base model.
Confidence
MEDIUM
Inherited from base model
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=1
  • assume M1 (base ablation) + M8 (GGUF quant) - default when producer unknown
Refusal direction extraction

No specific extraction method could be identified for this model. The producer either did not document it or used a proprietary pipeline.

What is a refusal direction? →
Downloads · lifetime
345
91 last 30d - stable
Likes
1
Model age
2mo ago
created 2026-08-11
Downloads over time
Now373→from0↑0%
01372744100 on Aug 12373 on Oct 11373 on Oct 7AugSepOct
Aug 12 → Oct 11 · 49 snapshots · spans 60 days

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
apache-2.0
Languages
en zh
Quantizations
BF16
Tags
gguf qwen35moe bf16 full-precision mtp moe speculative-decoding llama-cpp conversational text-generation en zh

Related

Total size
66.2 GB
Files
3
Quantizations
2
Registered
2026-08-22 13:56
Last updated on HF
2026-08-11 20:04

Files by quantization

BF16 1 file 66.2 GB
KAT-Coder-V2.5-Dev-35B-A3B-ABLITERATED-MTP-BF16.gguf 66.2 GB 567f95a1 download
Auxiliary files 2 files 6.02 KB
.gitattributes 3.35 KB ac54488d download
README.md 2.67 KB 97e08003 download

README current version from Hugging Face


license: apache-2.0
language:

  • en
  • zh
    base_model:
  • KridgeDookie/KAT-Coder-V2.5-Dev-35B-A3B-ABLITERATED-UNCENSORED-PHILADELPHIA-CLASS
    tags:
  • qwen35moe
  • gguf
  • bf16
  • full-precision
  • mtp
  • moe
  • speculative-decoding
  • llama-cpp
  • conversational
    pipeline_tag: text-generation

KAT-Coder-V2.5-Dev-35B-A3B-ABLITERATED-MTP-BF16

Full-precision single-file GGUF (BF16) of the abliterated KAT-Coder V2.5 Dev
35B-A3B, with the fine-tuned Qwen3.6-35B-A3B MTP (multi-token prediction)
head embedded in the model for speculative decoding.

  • Trunk: KridgeDookie's abliterated KAT-Coder V2.5 Dev 35B-A3B
    ("PHILADELPHIA CLASS", refusal-reduced)
  • MTP head: original-mtp-head.safetensors from
    gbuzhf/KAT-Coder-V2.5-Dev-MTP-GGUF (byte-identical to the
    Qwen/Qwen3.6-35B-A3B donor head at build time)
  • Format: GGUF v3, full BF16 (general.file_type = 32, MOSTLY_BF16;
    small 1-D tensors are F32, as standard)

File

File Size Type
KAT-Coder-V2.5-Dev-35B-A3B-ABLITERATED-MTP-BF16.gguf 71.1 GB (66.2 GiB) GGUF v3, BF16, single file

No split parts needed — the file downloads and runs directly.

Model details

Verified from the GGUF header:

Property Value
Architecture qwen35moe (Qwen3.6-35B-A3B class, hybrid SSM + full attention every 4 layers)
Parameters 35B total / ~3B active per token
Experts 256, 8 active (shared expert included)
Layers 41 (block_count = 41)
Context length 262,144 tokens
Hidden size 2,048
Tensors 753
MTP nextn_predict_layers = 1 (embedded)

Usage (llama.cpp)

llama-cli -m KAT-Coder-V2.5-Dev-35B-A3B-ABLITERATED-MTP-BF16.gguf \
  -p "Hello" -n 64 --spec-type draft-mtp

(Exact MTP flag name depends on your llama.cpp build; recent builds expose
it as --spec-type draft-mtp.)

Hardware note

BF16 full precision: the weights alone are ~66 GiB, so plan for roughly
75+ GB of free RAM/VRAM (CPU offload works, but expect slow prompt and
decode speeds). For lower resource requirements, use a quantized build —
the parent repo
KridgeDookie/KAT-Coder-V2.5-Dev-35B-A3B-ABLITERATED-UNCENSORED-PHILADELPHIA-CLASS
ships Q4_K_M, Q5_K_M, and Q8_0 GGUF options.

Provenance

Part Source
Abliterated trunk KridgeDookie/KAT-Coder-V2.5-Dev-35B-A3B-ABLITERATED-UNCENSORED-PHILADELPHIA-CLASS
MTP head gbuzhf/KAT-Coder-V2.5-Dev-MTP-GGUF (original-mtp-head.safetensors)
Conversion llama.cpp convert_hf_to_gguf.py (bf16, full export)

License

Apache 2.0, inherited from the parent model.

README history 3 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-08-11Fix model card: single-file BF16 GGUF (remove stale split-file instructions, ...3d406f72.7 KB
    Loading...
  2. 2026-08-11Document 2-part splitdc4cd0c1.5 KB
    Loading...
  3. 2026-08-11Add model cardfeb0e101.2 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration