license: other
license_name: qwen-community-1.0
license_link: LICENSE
base_model:
- SC117/Qwen3.8-Flash-Next-GSQ-RCO-abliterated-GGUF
base_model_relation: quantized
pipeline_tag: image-text-to-text
tags: - gguf
- llama.cpp
- moe
- quantized
- abliterated
- vulkan
Qwen3.8-Flash-Next · GSQ-RCO abliterated · Hybrid
IQ3_XXS trunk and hot experts, Q2_0 cold experts: the perplexity of IQ3_XXS, 12-30% faster than IQ3_XXS on a 24 GB GPU
What it is
One GGUF assembled from two tiers of SC117's abliterated GSQ-RCO quants of Qwen3.8-Flash-Next:
- every non-expert tensor (attention, shared experts, embeddings, norms) is from the IQ3_XXS tier;
- every routed expert tensor (
ffn_{gate,up,down}_exps) is from the Q2_0 tier.
Nothing was re-quantized: the tensors are copied byte for byte (gguf_swap_experts.py). The Q2_0 tier has the smallest experts (1.44 MB per expert against 1.75 MB in IQ3_XXS), but it also cheapens the dense part: shared experts down to Q2_0, attention to Q3_K, ple_key from BF16 to Q2_0. That part is only 0.38 GB larger in IQ3_XXS, and it carries most of the Q2_0 tier's quality loss.
Running it
With the vulkan-moe-hot fork (recommended)
shefowl/llama.cpp, branch vulkan-moe-hot keeps copies of the most used ("hot") experts of every layer in VRAM and computes the rest ("cold") on the CPU from RAM. Made and tested on an RX 7900 XTX with RADV. With LLAMA_MOE_HOT_SRC the hot copies are read from another GGUF of the same model, which gives per-expert precision, something GGUF itself cannot express:
| precision | read from | |
|---|---|---|
| dense weights | IQ3_XXS | this file |
| hot experts, in VRAM | IQ3_XXS | SC117's IQ3_XXS shard 1, which you need too (47 GB) |
| cold experts, on the CPU | Q2_0 | this file |
LLAMA_MOE_HOT_SRC=Qwen3.8-Flash-Next-GSQ-RCO-abliterated-IQ3_XXS-00001-of-00002.gguf \
MODEL=Qwen3.8-Flash-Next-GSQ-RCO-abliterated-Hybrid-00001-of-00002.gguf \
HOT_LIST=hot-lists/hot-12.5GB-no-vision.txt \
MTP=mtp-Qwen3.8-Flash-Next-Q4_K_M.gguf \
tools/moe-hot/run-hot.sh
run-hot.sh is in the fork; the MTP head is from drluoto/Qwen3.8-Flash-Next-MTP-GGUF. The lists in hot-lists/ are for a 24 GB card that also drives a desktop (about 3 GB): 12.5 GB of hot experts without vision, 11 GB with the mmproj. For another budget, make one with tools/moe-hot/moe_hot_list.py and give it the IQ3_XXS file as --model, since the hot copies come from there. The fork's README has the details.
With stock llama.cpp
Load shard 1 as usual; every routed expert is then Q2_0. The file has the size of the Q2_0 tier (67 GB with shard 2) and a lower perplexity: +2.6% against IQ3_XXS, where the Q2_0 tier has +4.3%. Upstream llama.cpp has no SIMD code for Q2_0 on x86, so the dot product falls back to scalar code (48 cycles per 32 weights on Zen 4) and experts kept on the CPU (-ot exps=CPU) run slowly. The fork has AVX2 and AVX-512 VBMI versions (5.0 and 3.4 cycles).
Measurements
One machine: RX 7900 XTX 24 GB (Vulkan, RADV), Ryzen 7 7700X, 64 GB DDR5.
Speed with the fork
Decode t/s: greedy, 400 tokens, one prompt per language or topic, second pass; GPU power level high, cold experts locked in RAM, MTP with 3 draft tokens, 32k context, every model with about 1.9 GB of VRAM left free.
| t/s | hot experts | code | science | English prose | Cyrillic (Russian) | Chinese |
|---|---|---|---|---|---|---|
| this hybrid | 12.6 GB, 71.6% of calls | 54.5 | 40.1 | 25.5 | 27.2 | 23.1 |
| GSQ-RCO IQ3_XXS | 12.0 GB, 69.9% | 48.6 | 34.4 | 19.6 | 23.3 | 18.4 |
| GSQ-RCO Q2_0 tier | 12.9 GB, 80.1% | 65.8 | 47.5 | 29.8 | 30.5 | 27.0 |
| AD-4.27 (AtomicChat recipe, Navin-Models uncensored) | 11.5 GB, 63.4% | 43.8 | 28.6 | 17.7 | 20.2 | 17.1 |
Against IQ3_XXS the hybrid is 12-30% faster, against AD-4.27 24-44%. The Q2_0 tier is the fastest, at +4.3% perplexity (next table). At the same free VRAM the hybrid holds 0.6 GB more hot experts than IQ3_XXS: the prefill buffer holds the cold experts of one layer, and Q2_0 ones are smaller. Cold experts locked in RAM: 22.4 GiB for the hybrid, 28.9 for IQ3_XXS, 19.7 for Q2_0, 36.5 for AD-4.27. Run to run the numbers move by about 5%, between sessions by up to 10% (how much VRAM the desktop takes changes the hot list).
Quality against GSQ-RCO IQ3_XXS
Same model, so the differences come from the quantization alone. Perplexity on 20k tokens of mixed English documentation, C++ and Russian text (40 chunks of 512), paired with the IQ3_XXS logits; ± is the standard error.
| perplexity ratio | mean KLD | same top-1 token | |
|---|---|---|---|
| this hybrid, with the fork | 0.998 ± 0.008 | 0.219 | 84.0% |
| this file with stock llama.cpp | 1.026 ± 0.011 | 0.363 | 79.4% |
| GSQ-RCO Q2_0 tier | 1.043 ± 0.011 | 0.405 | 78.1% |
The hot/cold split by itself changes nothing: IQ3_XXS with another hot list gives a KLD of 0.000. The KLD of the hybrid comes from its cold Q2_0 experts, which move the distribution without moving the perplexity.
Tasks, reasoning_effort medium, greedy, the hybrid with the fork against IQ3_XXS:
| hybrid | IQ3_XXS | ||
|---|---|---|---|
| GSM-Plus, 100 tasks | 78 | 78 | the same answer on every task |
| CRUXEval-O, 100 tasks | 94 | 97 | 5 character-level slips by the hybrid (a count off by two, a dropped character, a letter case); not significant at n = 100 |
Vision works with the mmproj of the original GGUFs: 14 px text in a 1280x960 picture was read exactly.
Not measured: standard perplexity sets (wikitext), other hardware, refusal behaviour.
Files
| file | size | content |
|---|---|---|
Qwen3.8-Flash-Next-GSQ-RCO-abliterated-Hybrid-00001-of-00002.gguf |
38.4 GB | IQ3_XXS trunk, Q2_0 routed experts |
Qwen3.8-Flash-Next-GSQ-RCO-abliterated-Hybrid-00002-of-00002.gguf |
28.8 GB | per-layer n-gram embedding table, byte-identical to shard 2 of SC117's tiers |
hot-lists/ |
expert lists for the fork |
Credits and license
- Qwen: the Qwen3.8-Flash-Next base model, under the Qwen Community License 1.0 (see
LICENSE). The upstream GSQ-RCO repositories are tagged Apache-2.0; this repository follows the base model's license. - IST-DASLab: GSQ and RCO, and the quantized weights.
- orcarouter: the abliterated weights. SC117: the abliterated GSQ-RCO GGUFs this file is assembled from.
- drluoto: the MTP head GGUF.
- llama.cpp: the GGUF format and the runtime.
Not affiliated with any of them. The model is abliterated: its refusals were removed upstream, and you are responsible for how you use it.