library_name: gguf
base_model: llmfan46/Qwen3.6-35B-A3B-uncensored-heretic-Native-MTP-Preserved-GGUF
tags:
- gguf
- qwen3
- moe
- mtp
- quantized
- llama.cpp
Qwen3.6-35B-A3B-uncensored-heretic-Native-MTP-Preserved — TQ3_4S GGUF
A TQ3_4S (~Q3) quantized GGUF of Qwen3.6-35B-A3B (MoE, native MTP preserved),
weighing in at ~14 GB. Built for running on a single 16 GB GPU with speculative
decoding via the native MTP draft head.
On an RTX 5060 Ti (16 GB) this runs at roughly 100–140 tokens/s (with MTP
speculative decoding enabled).
Files
| File | Size | Notes |
|---|---|---|
Qwen3.6-35B-A3B-uncensored-heretic-Native-MTP-Preserved-TQ3_4S.gguf |
~14 GB | Mixed K-quant + TQ3_4S layout |
A multimodal
mmprojis not included in this repo. If you need vision,
grab the matchingmmproj-BF16.gguffrom the base model repo below.
Sources / Attribution
This quant was built by referencing the following:
- Base weights (BF16 GGUF):
llmfan46/Qwen3.6-35B-A3B-uncensored-heretic-Native-MTP-Preserved-GGUF— the BF16 GGUF +mmprojwere the quantization input. - Quant recipe reference:
YTan2000/Qwen3.6-35B-A3B-MTP-TQ3_4S— the mixed-precision tensor layout below was matched to this model. - Tooling (TQ3_4S support):
turbo-tan/llama.cpp-tq3— llama.cpp fork that adds theTQ3_4Stype and--spec-type draft-mtpspeculative decoding.
Quantization recipe
Plain TQ3_4S would quantize all weights at 4 bpw (~18 GB). Instead this is a
mixed layout (base ftype Q3_K_M, with per-tensor overrides) matching the
reference model:
token_embd = Q5_K
output = Q4_K
attn_* / *_shexp / ssm_out = Q6_K
ffn_gate_exps / ffn_up_exps = Q2_K
ffn_down_exps = Q3_K (Q4_K on blocks 21, 28, 38)
ssm_alpha / ssm_beta / nextn.eh_proj = TQ3_4S
Produced with llama-quantize:
llama-quantize \
--token-embedding-type Q5_K \
--output-tensor-type Q4_K \
--tensor-type-file tensor_types.txt \
Qwen3.6-35B-A3B-...-BF16.gguf \
Qwen3.6-35B-A3B-...-TQ3_4S.gguf \
Q3_K_M
Running with llama-server
Requires the turbo-tan/llama.cpp-tq3
fork (for TQ3_4S + MTP speculative decoding). Tuned for a 16 GB GPU; adjust-ngl, --ctx-size, and --cache-ram to your hardware.
llama-server \
--host 0.0.0.0 --port 8080 \
--model Qwen3.6-35B-A3B-...-TQ3_4S.gguf \
--jinja \
--chat-template-file chat_template.jinja \
-ngl 55 \
-fa on \
-ctk q8_0 -ctv tq3_0 \
--batch-size 2048 \
--ubatch-size 512 \
--ctx-size 64000 \
--parallel 1 -np 1 \
--spec-type draft-mtp \
--spec-draft-ngl 99 \
--spec-draft-n-max 2 \
--spec-draft-n-min 1 \
--spec-draft-p-min 1.0 \
--spec-draft-type-k q4_0 \
--spec-draft-type-v tq3_0 \
--reasoning on \
--reasoning-format auto \
--warmup --perf \
--threads 4 --threads-batch 8 \
--cache-ram 16000 \
--ctx-checkpoints 32
Notes:
--ctx-size 64000is roughly the empirical max on 16 GB before OOM; lower it if you hit memory limits.-ctk q8_0 -ctv tq3_0quantizes the KV cache to fit more context.--spec-type draft-mtpuses the model's native MTP head as the speculative draft — no separate draft model needed.- For vision, add
--mmproj mmproj-BF16.gguf(runs on CPU). - To disable reasoning: replace
--reasoning onwith--reasoning off --reasoning-budget 0.