← back to catalog · registered 2026-10-11 19:59

shefowl/Qwen3.8-Flash-Next-GSQ-RCO-abliterated-Hybrid-352E-4050-16GB-Generalist-GGUF

shefowl GGUF MoE second-order
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/shefowl%2FQwen3.8-Flash-Next-GSQ-RCO-abliterated-Hybrid-352E-4050-16GB-Generalist-GGUF"
Response includes
  • classification unknown
  • files 6
  • author_summary 2 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
?
Primary method

Unclassified

No clear signals of an abliteration technique in this model.
Confidence
UNKNOWN
Why this label 1 signal
No classification signals present. This may not be an abliterated model at all - it could be a repackaging, a merge with unrelated goals, or unrelated content that mentions the term.
  • no classification signals present (no abliterated, uncensored, or known producer/method markers)
Refusal direction extraction

No specific extraction method could be identified for this model. The producer either did not document it or used a proprietary pipeline.

What is a refusal direction? →
Downloads · 30-day
0
Likes
0
Model age
today
created 2026-10-11

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
other
Tags
gguf strata moe pruned quantized abliterated text-generation base_model:shefowl/Qwen3.8-Flash-Next-GSQ-RCO-abliterated-Hybrid-GGUF base_model:quantized:shefowl/Qwen3.8-Flash-Next-GSQ-RCO-abliterated-Hybrid-GGUF license:other endpoints_compatible region:us

Related

Total size
52.7 GB
Files
6
Quantizations
1
Registered
2026-10-11 19:59
Last updated on HF
2026-10-11 20:05

Files by quantization

Auxiliary files 6 files 52.7 GB
Qwen3.8-Flash-Next-GSQ-RCO-abliterated-IQ3_XXS-Q2_0-352E-00002-of-00002.gguf 26.8 GB 316b46f3 download
Qwen3.8-Flash-Next-GSQ-RCO-abliterated-IQ3_XXS-Q2_0-352E-00001-of-00002.gguf 25.8 GB 195bf092 download
kept_experts_352.json 62.7 KB 0724e27d download
README.md 13.4 KB 9fd8463b download
LICENSE 3.16 KB 9557a896 download
.gitattributes 1.70 KB a3acfb09 download

README current version from Hugging Face


license: other
license_name: qwen-community-1.0
license_link: LICENSE
base_model:

  • shefowl/Qwen3.8-Flash-Next-GSQ-RCO-abliterated-Hybrid-GGUF
    base_model_relation: quantized
    pipeline_tag: text-generation
    tags:
  • gguf
  • strata
  • moe
  • pruned
  • quantized
  • abliterated

Qwen3.8-Flash-Next · GSQ-RCO abliterated · Hybrid · 352 experts

352 of 512 routed experts per layer: 10-11 tokens/s on a laptop with an RTX 4050 (6 GB) and 16 GB of RAM

runs on Strata 352 of 512 experts shard 1: 25.8 GiB

Read the speed table before you download. The full hybrid (all 512 experts) runs on the same laptop at 9.0 tokens/s, this file at 10.8. For that difference this file gives up about a seventh of the full model's factual knowledge and loops more often in greedy code. On cards with 12 GB or more the full model is fast enough that we would take it.

What it is

A pruned copy of our Hybrid GGUF of Qwen3.8-Flash-Next (IQ3_XXS trunk, Q2_0 routed experts, assembled from SC117's abliterated GSQ-RCO quants). Every layer keeps 352 of its 512 routed experts; the router still picks 10 per token. Nothing was re-quantized or retrained: the kept experts and their router rows are copied byte for byte. kept_experts_352.json lists them.

The experts were chosen for general use, not for code alone: languages, code, math and facts were all in the calibration (see How it was made).

this file full hybrid
routed experts per layer 352 512
shard 1 (everything the engine keeps in memory or reads per token) 25.8 GiB 35.8 GiB
routed experts alone 21.8 GiB 31.6 GiB
shard 2 (n-gram table, one row read per token) 28.8 GB, identical 28.8 GB

Speed

Decode speed by budget

Decode tokens/s on Strata. Five 2000-token greedy answers, one per topic, on a freshly started engine with Strata's own expert order ("first start").

6 GB of VRAM + 16 GB of RAM: a real laptop

Acer Nitro V 16 (ANV16-41): Ryzen 5 8645HS, RTX 4050 Laptop 6 GB, 2x8 GB DDR5-5600, Kingston OM8PGP4512Q OEM NVMe (1.4-1.6 GiB/s in the engine's read pattern). Nobara Linux, kernel 7.2, NVIDIA driver 595. Strata 0.1.42 built for CUDA, 8K context, no MTP draft layer; the settings are under Running it.

t/s code science English prose Cyrillic (Russian) Chinese mean
this file 9.4 11.7 11.5 11.9 9.7 10.8
full hybrid 8.0 9.8 9.6 9.4 8.3 9.0

Six more code tasks (React, C#, Go, Rust, C++, Polars), 1200 tokens each:

t/s first start after the engine learned the workload
this file 9.3 12.1 (with a 32K context, where the first start was 8.4)
full hybrid 7.3 10.0
  • "Learned the workload". Strata can save which experts a session used (--expert-profile-save) and load them first next time. One session of six other code tasks gave +37-44% on code. The same profile cost 16-26% on other text, so keep one profile per kind of work. Two starters are in profiles/.
  • 32K context on this laptop (--kv k8v4 --kv-resident 20480): 10.2 t/s mean. A 6,000-token prompt was read in 61 s; a follow-up question in the same chat started after 4 s.
  • Windows. With identical settings (Strata's defaults), a 320-expert sibling of this file ran at 6.0-7.4 t/s under Windows and at 13.8-16.0 under Linux on this laptop: Windows left 6.2 GiB of RAM for experts, Linux 9.7. We did not run this file under Windows.

8 GB and more, with 16 GB of RAM: emulated

Our desktop (RX 7900 XTX, Ryzen 7 7700X) held to the budget: Strata limited to the given VRAM, engine and server in a cgroup with 12 GiB of RAM and no swap, 6 cores, model on a Samsung 980 PRO (2.45 GiB/s in the engine's read pattern). 32K context, MTP draft layer on.

mean t/s this file full hybrid
8 GB + 16 GB 14.7 11.1
12 GB + 16 GB 43.2 30.4
16 GB + 16 GB not measured yet 48.6

When the experts do not fit in VRAM plus RAM, the rest is read from the SSD for every token, and speed is roughly the SSD's read rate divided by the megabytes read per token. Pruning removes experts the router seldom picks, which the engine seldom read anyway: on the 6 GB budget this file reads 231 MiB per token and the full hybrid 273. That is why a 31% smaller expert set is only about 20% faster there.

What the speed costs

Facts and code

All rows ran on the same machine with the same prompts and graders. No thinking.

facts: NQ-open, 1800 questions code: HumanEval + MBPP, 591 tasks loops in 100 code answers, greedy / sampled loops in 64 long answers, 8 languages
full hybrid, 512 experts 33.1% 83.6% 6 / 4 1
this file, 352 experts 28.4% 80.4% 28 / 12 1
ISTA Coder, 256 experts 25.2% 86.8% 26 / 10 36
dense Qwen3.8-27B, GSQ-RCO IQ3_XXS 26.9% 83.4% not run not run
Bonsai 27B, ternary 22.6% 78.3% 14 / 4 8
  • Facts: greedy; an answer is right when it contains one of the reference answers after normalization. This file keeps 86% of the full hybrid's score.
  • Code: one Python function per task, 768 tokens, sampled (temperature 1.0, top-p 0.95, top-k 20), the tasks' own tests run in a sandbox.
  • Loops: an answer counts as a loop when its last 200 characters occur earlier in it, or when more than 10% of its 20-word windows are repeats (25% for code, which repeats itself legitimately). Code answers are 1500 tokens. Long answers are 1000 greedy tokens on 8 topics in Chinese, Japanese, Korean, Cyrillic (Russian), Arabic, Hindi, Turkish and Portuguese.
  • Use the model's sampling, not temperature 0. Greedy code answers loop four to five times as often as the full model's; with sampling the gap is 12 against 4.
  • Differences were checked with an exact sign test on the paired answers. Against the full hybrid this file is lower on facts (p < 0.001) and loops more in greedy code (p < 0.001); the 3 points of code are borderline (p = 0.05). Against the dense 27B it ties on both facts and code.

Against the ISTA Coder

The Coder keeps 256 experts chosen for code and stores them at more bits, so its shard 1 is larger than this file's (27.6 against 25.8 GiB).

this file ISTA Coder
code, 591 tasks 80.4% 86.8% (p < 0.001)
facts, 1800 questions 28.4% 25.2% (p = 0.001)
long answers that loop, of 64 1 36, in every language tested
6 GB + 16 GB emulated, English prose, t/s 13.0 6.8
12 GB + 16 GB emulated, English prose, t/s 44.2 24.9

The Coder writes better code. This file stays usable outside code, and on small budgets it is faster: the Coder's experts are 1.95 MiB each against 1.32, so every expert that misses memory costs more to read. Its other speed answers looped or stopped early under greedy decoding, so prose is the one clean comparison; on 8 GB + 16 GB its engine ran out of the 12 GiB RAM limit in our setup. The Coder is not abliterated.

Running it

Tested with Strata only (0.1.42, commit 61b3fb5). We did not test stock llama.cpp or our Vulkan fork with this file.

Strata's releases carry Windows engines. For Linux we built the engine from source with CUDA 13 (-DSTRATA_ENABLE_CUDA=ON, as Strata's setup.py does).

# in the Strata folder
.venv/bin/python tools/iq_pack.py --gguf /path/to/Qwen3.8-Flash-Next-GSQ-RCO-abliterated-IQ3_XXS-Q2_0-352E-00001-of-00002.gguf \
  --out packs/k352
cp /path/to/profiles/expert-profile-352E-general.bin packs/k352/

Strata's shipped data/expert-profile.bin numbers 512 experts and does not fit a pruned file; use the ones in profiles/.

strata-k352.json for 6 GB of VRAM and 16 GB of RAM:

{
 "exe": "engine/strata",
 "args": ["--pack", "packs/k352",
  "--native", "/path/to/Qwen3.8-Flash-Next-GSQ-RCO-abliterated-IQ3_XXS-Q2_0-352E-00001-of-00002.gguf",
  "--ple-gguf", "/path/to/Qwen3.8-Flash-Next-GSQ-RCO-abliterated-IQ3_XXS-Q2_0-352E-00002-of-00002.gguf",
  "--expert-profile", "packs/k352/expert-profile-352E-general.bin",
  "--expert-profile-save", "packs/k352/expert-profile-mine.bin", "--expert-profile-save-every", "0",
  "--expert-cache", "auto", "--prefill", "256", "--spec", "4", "--spec-min-p", "0.5",
  "--max-context", "8192", "--kv", "q4_0", "--mmap-experts", "--resident-budget-gib", "11"],
 "cwd": ".", "tokenizer": "packs/k352/tokenizer", "model_name": "qwen3.8-flash-next-k352",
 "log": "strata-k352.log", "host": "127.0.0.1", "port": 8080
}
STRATA_ROUTE_TAIL_SKIP=0 STRATA_RESIDENT_HEADROOM_GIB=2 STRATA_UNBUFFERED_LOAD=1 \
  .venv/bin/python -m serve.server --engine strata --config strata-k352.json --port 8080
  • STRATA_ROUTE_TAIL_SKIP=0. On CUDA, Strata 0.1.42 skips the least likely experts of a missed token by default. That is about 20% faster, and on a 320-expert sibling of this file it made long Russian and Chinese answers loop. All numbers above are with it off.
  • STRATA_RESIDENT_HEADROOM_GIB=2 with STRATA_UNBUFFERED_LOAD=1 on 16 GB of RAM. With 1.5 GiB the laptop was left with 260 MiB free. With 3 GiB and buffered reads the engine moved its file reads into the page cache and fell to 3.5 t/s.
  • No MTP draft layer on 6 GB. The stock one does not fit next to a useful expert cache. A draft layer we cut down to fit made the laptop slower (9.4 against 10.6 t/s on a sibling file).
  • 32K context on 6 GB: "--max-context", "32768", "--kv", "k8v4", "--kv-resident", "20480".
  • --expert-profile-save writes the session's expert order when the server stops (stop it with TERM). Point --expert-profile at that file next time.
  • 8 GB and more: "--max-context", "32768", "--kv", "int8", add "--mtp", "mtp/rt" as in the full hybrid's card, and leave STRATA_RESIDENT_HEADROOM_GIB at 2.

How it was made

The expert choice comes from RCO (IST-DASLab's search for the expert set that keeps the model's output distribution closest to the full model's), the method behind the ISTA Coder. We ran it on Qwen's original BF16 weights for 320 experts per layer with our own calibration: 128 sequences of 2048 tokens, the full model's answers in 16 domains at 4-7% each (text in English, other European languages, Chinese, Japanese, Korean, Cyrillic (Russian), Arabic, Hindi, Turkish and Portuguese; code, math, facts, tool calls, thinking and other). The resulting list was then applied to the abliterated GGUF. This file keeps those 320 and, in every layer, the 32 further experts the search ranked next. RCO's sets are nested in practice: the 320 set is 99.6% inside the 448 set from a separate search.

Counting how often the router picks an expert was a worse guide. Our earlier selection by routing counts kept more routing mass at 384 experts than RCO's choice (86% against 79%) and scored 5 points lower on facts (24.4% against 29.6%).

Routing profiles: expert-profile-352E-general.bin is Strata's shipped order restricted to the kept experts; expert-profile-352E-code.bin was saved by the engine on the laptop after six code tasks.

Files

file size content
Qwen3.8-Flash-Next-GSQ-RCO-abliterated-IQ3_XXS-Q2_0-352E-00001-of-00002.gguf 27.7 GB IQ3_XXS trunk, 352 Q2_0 routed experts per layer
Qwen3.8-Flash-Next-GSQ-RCO-abliterated-IQ3_XXS-Q2_0-352E-00002-of-00002.gguf 28.8 GB per-layer n-gram embedding table, byte-identical to shard 2 of the full hybrid and of SC117's tiers
profiles/ expert orders for Strata: general and code
kept_experts_352.json kept expert ids per layer, in the full model's numbering

Not measured

Vision. Stock llama.cpp and our Vulkan fork. This file under Windows. Real 8, 12 and 16 GB cards (those rows are emulated on one 24 GB card). Thinking mode. Benchmarks beyond the four above. Each speed row is one session; repeated Strata runs of the full model varied by up to 8%.

Credits and license

  • Qwen: the Qwen3.8-Flash-Next base model, under the Qwen Community License 1.0 (see LICENSE).
  • IST-DASLab: GSQ and RCO, the quantized weights, and the Coder we compare with.
  • orcarouter: the abliterated weights. SC117: the abliterated GSQ-RCO GGUFs this file is assembled from.
  • Niko1221 and contributors: Strata (MIT).

Not affiliated with any of them. The model is abliterated: its refusals were removed upstream, and you are responsible for how you use it.

Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration