license: other
license_name: qwen-community-1.0
license_link: LICENSE
base_model:
- shefowl/Qwen3.8-Flash-Next-GSQ-RCO-abliterated-Hybrid-GGUF
base_model_relation: quantized
pipeline_tag: text-generation
tags: - gguf
- strata
- moe
- pruned
- quantized
- abliterated
Qwen3.8-Flash-Next · GSQ-RCO abliterated · Hybrid · 352 experts
352 of 512 routed experts per layer: 10-11 tokens/s on a laptop with an RTX 4050 (6 GB) and 16 GB of RAM
What it is
A pruned copy of our Hybrid GGUF of Qwen3.8-Flash-Next (IQ3_XXS trunk, Q2_0 routed experts, assembled from SC117's abliterated GSQ-RCO quants). Every layer keeps 352 of its 512 routed experts; the router still picks 10 per token. Nothing was re-quantized or retrained: the kept experts and their router rows are copied byte for byte. kept_experts_352.json lists them.
The experts were chosen for general use, not for code alone: languages, code, math and facts were all in the calibration (see How it was made).
| this file | full hybrid | |
|---|---|---|
| routed experts per layer | 352 | 512 |
| shard 1 (everything the engine keeps in memory or reads per token) | 25.8 GiB | 35.8 GiB |
| routed experts alone | 21.8 GiB | 31.6 GiB |
| shard 2 (n-gram table, one row read per token) | 28.8 GB, identical | 28.8 GB |
Speed

Decode tokens/s on Strata. Five 2000-token greedy answers, one per topic, on a freshly started engine with Strata's own expert order ("first start").
6 GB of VRAM + 16 GB of RAM: a real laptop
Acer Nitro V 16 (ANV16-41): Ryzen 5 8645HS, RTX 4050 Laptop 6 GB, 2x8 GB DDR5-5600, Kingston OM8PGP4512Q OEM NVMe (1.4-1.6 GiB/s in the engine's read pattern). Nobara Linux, kernel 7.2, NVIDIA driver 595. Strata 0.1.42 built for CUDA, 8K context, no MTP draft layer; the settings are under Running it.
| t/s | code | science | English prose | Cyrillic (Russian) | Chinese | mean |
|---|---|---|---|---|---|---|
| this file | 9.4 | 11.7 | 11.5 | 11.9 | 9.7 | 10.8 |
| full hybrid | 8.0 | 9.8 | 9.6 | 9.4 | 8.3 | 9.0 |
Six more code tasks (React, C#, Go, Rust, C++, Polars), 1200 tokens each:
| t/s | first start | after the engine learned the workload |
|---|---|---|
| this file | 9.3 | 12.1 (with a 32K context, where the first start was 8.4) |
| full hybrid | 7.3 | 10.0 |
- "Learned the workload". Strata can save which experts a session used (
--expert-profile-save) and load them first next time. One session of six other code tasks gave +37-44% on code. The same profile cost 16-26% on other text, so keep one profile per kind of work. Two starters are inprofiles/. - 32K context on this laptop (
--kv k8v4 --kv-resident 20480): 10.2 t/s mean. A 6,000-token prompt was read in 61 s; a follow-up question in the same chat started after 4 s. - Windows. With identical settings (Strata's defaults), a 320-expert sibling of this file ran at 6.0-7.4 t/s under Windows and at 13.8-16.0 under Linux on this laptop: Windows left 6.2 GiB of RAM for experts, Linux 9.7. We did not run this file under Windows.
8 GB and more, with 16 GB of RAM: emulated
Our desktop (RX 7900 XTX, Ryzen 7 7700X) held to the budget: Strata limited to the given VRAM, engine and server in a cgroup with 12 GiB of RAM and no swap, 6 cores, model on a Samsung 980 PRO (2.45 GiB/s in the engine's read pattern). 32K context, MTP draft layer on.
| mean t/s | this file | full hybrid |
|---|---|---|
| 8 GB + 16 GB | 14.7 | 11.1 |
| 12 GB + 16 GB | 43.2 | 30.4 |
| 16 GB + 16 GB | not measured yet | 48.6 |
When the experts do not fit in VRAM plus RAM, the rest is read from the SSD for every token, and speed is roughly the SSD's read rate divided by the megabytes read per token. Pruning removes experts the router seldom picks, which the engine seldom read anyway: on the 6 GB budget this file reads 231 MiB per token and the full hybrid 273. That is why a 31% smaller expert set is only about 20% faster there.
What the speed costs

All rows ran on the same machine with the same prompts and graders. No thinking.
| facts: NQ-open, 1800 questions | code: HumanEval + MBPP, 591 tasks | loops in 100 code answers, greedy / sampled | loops in 64 long answers, 8 languages | |
|---|---|---|---|---|
| full hybrid, 512 experts | 33.1% | 83.6% | 6 / 4 | 1 |
| this file, 352 experts | 28.4% | 80.4% | 28 / 12 | 1 |
| ISTA Coder, 256 experts | 25.2% | 86.8% | 26 / 10 | 36 |
| dense Qwen3.8-27B, GSQ-RCO IQ3_XXS | 26.9% | 83.4% | not run | not run |
| Bonsai 27B, ternary | 22.6% | 78.3% | 14 / 4 | 8 |
- Facts: greedy; an answer is right when it contains one of the reference answers after normalization. This file keeps 86% of the full hybrid's score.
- Code: one Python function per task, 768 tokens, sampled (temperature 1.0, top-p 0.95, top-k 20), the tasks' own tests run in a sandbox.
- Loops: an answer counts as a loop when its last 200 characters occur earlier in it, or when more than 10% of its 20-word windows are repeats (25% for code, which repeats itself legitimately). Code answers are 1500 tokens. Long answers are 1000 greedy tokens on 8 topics in Chinese, Japanese, Korean, Cyrillic (Russian), Arabic, Hindi, Turkish and Portuguese.
- Use the model's sampling, not temperature 0. Greedy code answers loop four to five times as often as the full model's; with sampling the gap is 12 against 4.
- Differences were checked with an exact sign test on the paired answers. Against the full hybrid this file is lower on facts (p < 0.001) and loops more in greedy code (p < 0.001); the 3 points of code are borderline (p = 0.05). Against the dense 27B it ties on both facts and code.
Against the ISTA Coder
The Coder keeps 256 experts chosen for code and stores them at more bits, so its shard 1 is larger than this file's (27.6 against 25.8 GiB).
| this file | ISTA Coder | |
|---|---|---|
| code, 591 tasks | 80.4% | 86.8% (p < 0.001) |
| facts, 1800 questions | 28.4% | 25.2% (p = 0.001) |
| long answers that loop, of 64 | 1 | 36, in every language tested |
| 6 GB + 16 GB emulated, English prose, t/s | 13.0 | 6.8 |
| 12 GB + 16 GB emulated, English prose, t/s | 44.2 | 24.9 |
The Coder writes better code. This file stays usable outside code, and on small budgets it is faster: the Coder's experts are 1.95 MiB each against 1.32, so every expert that misses memory costs more to read. Its other speed answers looped or stopped early under greedy decoding, so prose is the one clean comparison; on 8 GB + 16 GB its engine ran out of the 12 GiB RAM limit in our setup. The Coder is not abliterated.
Running it
Tested with Strata only (0.1.42, commit 61b3fb5). We did not test stock llama.cpp or our Vulkan fork with this file.
Strata's releases carry Windows engines. For Linux we built the engine from source with CUDA 13 (-DSTRATA_ENABLE_CUDA=ON, as Strata's setup.py does).
# in the Strata folder
.venv/bin/python tools/iq_pack.py --gguf /path/to/Qwen3.8-Flash-Next-GSQ-RCO-abliterated-IQ3_XXS-Q2_0-352E-00001-of-00002.gguf \
--out packs/k352
cp /path/to/profiles/expert-profile-352E-general.bin packs/k352/
Strata's shipped data/expert-profile.bin numbers 512 experts and does not fit a pruned file; use the ones in profiles/.
strata-k352.json for 6 GB of VRAM and 16 GB of RAM:
{
"exe": "engine/strata",
"args": ["--pack", "packs/k352",
"--native", "/path/to/Qwen3.8-Flash-Next-GSQ-RCO-abliterated-IQ3_XXS-Q2_0-352E-00001-of-00002.gguf",
"--ple-gguf", "/path/to/Qwen3.8-Flash-Next-GSQ-RCO-abliterated-IQ3_XXS-Q2_0-352E-00002-of-00002.gguf",
"--expert-profile", "packs/k352/expert-profile-352E-general.bin",
"--expert-profile-save", "packs/k352/expert-profile-mine.bin", "--expert-profile-save-every", "0",
"--expert-cache", "auto", "--prefill", "256", "--spec", "4", "--spec-min-p", "0.5",
"--max-context", "8192", "--kv", "q4_0", "--mmap-experts", "--resident-budget-gib", "11"],
"cwd": ".", "tokenizer": "packs/k352/tokenizer", "model_name": "qwen3.8-flash-next-k352",
"log": "strata-k352.log", "host": "127.0.0.1", "port": 8080
}
STRATA_ROUTE_TAIL_SKIP=0 STRATA_RESIDENT_HEADROOM_GIB=2 STRATA_UNBUFFERED_LOAD=1 \
.venv/bin/python -m serve.server --engine strata --config strata-k352.json --port 8080
STRATA_ROUTE_TAIL_SKIP=0. On CUDA, Strata 0.1.42 skips the least likely experts of a missed token by default. That is about 20% faster, and on a 320-expert sibling of this file it made long Russian and Chinese answers loop. All numbers above are with it off.STRATA_RESIDENT_HEADROOM_GIB=2withSTRATA_UNBUFFERED_LOAD=1on 16 GB of RAM. With 1.5 GiB the laptop was left with 260 MiB free. With 3 GiB and buffered reads the engine moved its file reads into the page cache and fell to 3.5 t/s.- No MTP draft layer on 6 GB. The stock one does not fit next to a useful expert cache. A draft layer we cut down to fit made the laptop slower (9.4 against 10.6 t/s on a sibling file).
- 32K context on 6 GB:
"--max-context", "32768", "--kv", "k8v4", "--kv-resident", "20480". --expert-profile-savewrites the session's expert order when the server stops (stop it with TERM). Point--expert-profileat that file next time.- 8 GB and more:
"--max-context", "32768", "--kv", "int8", add"--mtp", "mtp/rt"as in the full hybrid's card, and leaveSTRATA_RESIDENT_HEADROOM_GIBat 2.
How it was made
The expert choice comes from RCO (IST-DASLab's search for the expert set that keeps the model's output distribution closest to the full model's), the method behind the ISTA Coder. We ran it on Qwen's original BF16 weights for 320 experts per layer with our own calibration: 128 sequences of 2048 tokens, the full model's answers in 16 domains at 4-7% each (text in English, other European languages, Chinese, Japanese, Korean, Cyrillic (Russian), Arabic, Hindi, Turkish and Portuguese; code, math, facts, tool calls, thinking and other). The resulting list was then applied to the abliterated GGUF. This file keeps those 320 and, in every layer, the 32 further experts the search ranked next. RCO's sets are nested in practice: the 320 set is 99.6% inside the 448 set from a separate search.
Counting how often the router picks an expert was a worse guide. Our earlier selection by routing counts kept more routing mass at 384 experts than RCO's choice (86% against 79%) and scored 5 points lower on facts (24.4% against 29.6%).
Routing profiles: expert-profile-352E-general.bin is Strata's shipped order restricted to the kept experts; expert-profile-352E-code.bin was saved by the engine on the laptop after six code tasks.
Files
| file | size | content |
|---|---|---|
Qwen3.8-Flash-Next-GSQ-RCO-abliterated-IQ3_XXS-Q2_0-352E-00001-of-00002.gguf |
27.7 GB | IQ3_XXS trunk, 352 Q2_0 routed experts per layer |
Qwen3.8-Flash-Next-GSQ-RCO-abliterated-IQ3_XXS-Q2_0-352E-00002-of-00002.gguf |
28.8 GB | per-layer n-gram embedding table, byte-identical to shard 2 of the full hybrid and of SC117's tiers |
profiles/ |
expert orders for Strata: general and code | |
kept_experts_352.json |
kept expert ids per layer, in the full model's numbering |
Not measured
Vision. Stock llama.cpp and our Vulkan fork. This file under Windows. Real 8, 12 and 16 GB cards (those rows are emulated on one 24 GB card). Thinking mode. Benchmarks beyond the four above. Each speed row is one session; repeated Strata runs of the full model varied by up to 8%.
Credits and license
- Qwen: the Qwen3.8-Flash-Next base model, under the Qwen Community License 1.0 (see
LICENSE). - IST-DASLab: GSQ and RCO, the quantized weights, and the Coder we compare with.
- orcarouter: the abliterated weights. SC117: the abliterated GSQ-RCO GGUFs this file is assembled from.
- Niko1221 and contributors: Strata (MIT).
Not affiliated with any of them. The model is abliterated: its refusals were removed upstream, and you are responsible for how you use it.