← back to catalog · registered 2026-10-03 09:58

shefowl/Qwen3.8-Flash-Next-GSQ-RCO-abliterated-Hybrid-GGUF

shefowl GGUF MoE multimodal second-order
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/shefowl%2FQwen3.8-Flash-Next-GSQ-RCO-abliterated-Hybrid-GGUF"
Response includes
  • classification unknown
  • files 5
  • author_summary 1 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
?
Primary method

Unclassified

No clear signals of an abliteration technique in this model.
Confidence
UNKNOWN
Why this label 1 signal
No classification signals present. This may not be an abliterated model at all - it could be a repackaging, a merge with unrelated goals, or unrelated content that mentions the term.
  • no classification signals present (no abliterated, uncensored, or known producer/method markers)
Refusal direction extraction

No specific extraction method could be identified for this model. The producer either did not document it or used a proprietary pipeline.

What is a refusal direction? →
Downloads · 30-day
0
Likes
0
Model age
today
created 2026-10-03

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
other
Tags
gguf llama.cpp moe quantized abliterated vulkan image-text-to-text base_model:SC117/Qwen3.8-Flash-Next-GSQ-RCO-abliterated-GGUF base_model:quantized:SC117/Qwen3.8-Flash-Next-GSQ-RCO-abliterated-GGUF license:other endpoints_compatible region:us
Total size
62.6 GB
Files
5
Quantizations
1
Registered
2026-10-03 09:58
Last updated on HF
2026-10-03 10:06

Files by quantization

Auxiliary files 5 files 62.6 GB
Qwen3.8-Flash-Next-GSQ-RCO-abliterated-Hybrid-00001-of-00002.gguf 35.8 GB 80bcf585 download
Qwen3.8-Flash-Next-GSQ-RCO-abliterated-Hybrid-00002-of-00002.gguf 26.8 GB 316b46f3 download
README.md 7.94 KB afd4c382 download
LICENSE 3.16 KB 9557a896 download
.gitattributes 1.68 KB 365c2acb download

README current version from Hugging Face


license: other
license_name: qwen-community-1.0
license_link: LICENSE
base_model:

  • SC117/Qwen3.8-Flash-Next-GSQ-RCO-abliterated-GGUF
    base_model_relation: quantized
    pipeline_tag: image-text-to-text
    tags:
  • gguf
  • llama.cpp
  • moe
  • quantized
  • abliterated
  • vulkan

Qwen3.8-Flash-Next · GSQ-RCO abliterated · Hybrid

IQ3_XXS trunk and hot experts, Q2_0 cold experts: the perplexity of IQ3_XXS, 12-30% faster than IQ3_XXS on a 24 GB GPU

vulkan-moe-hot fork 67 GB Qwen Community License 1.0

We recommend running it with our hot-experts fork: vulkan-moe-hot, a llama.cpp fork for AMD cards on Vulkan. It keeps the most used experts in VRAM, and only with it do you get the full speed and the IQ3_XXS precision of the hot experts. Stock llama.cpp runs the file too.

What it is

One GGUF assembled from two tiers of SC117's abliterated GSQ-RCO quants of Qwen3.8-Flash-Next:

  • every non-expert tensor (attention, shared experts, embeddings, norms) is from the IQ3_XXS tier;
  • every routed expert tensor (ffn_{gate,up,down}_exps) is from the Q2_0 tier.

Nothing was re-quantized: the tensors are copied byte for byte (gguf_swap_experts.py). The Q2_0 tier has the smallest experts (1.44 MB per expert against 1.75 MB in IQ3_XXS), but it also cheapens the dense part: shared experts down to Q2_0, attention to Q3_K, ple_key from BF16 to Q2_0. That part is only 0.38 GB larger in IQ3_XXS, and it carries most of the Q2_0 tier's quality loss.

Running it

With the vulkan-moe-hot fork (recommended)

shefowl/llama.cpp, branch vulkan-moe-hot keeps copies of the most used ("hot") experts of every layer in VRAM and computes the rest ("cold") on the CPU from RAM. Made and tested on an RX 7900 XTX with RADV. With LLAMA_MOE_HOT_SRC the hot copies are read from another GGUF of the same model, which gives per-expert precision, something GGUF itself cannot express:

precision read from
dense weights IQ3_XXS this file
hot experts, in VRAM IQ3_XXS SC117's IQ3_XXS shard 1, which you need too (47 GB)
cold experts, on the CPU Q2_0 this file
LLAMA_MOE_HOT_SRC=Qwen3.8-Flash-Next-GSQ-RCO-abliterated-IQ3_XXS-00001-of-00002.gguf \
MODEL=Qwen3.8-Flash-Next-GSQ-RCO-abliterated-Hybrid-00001-of-00002.gguf \
HOT_LIST=hot-lists/hot-12.5GB-no-vision.txt \
MTP=mtp-Qwen3.8-Flash-Next-Q4_K_M.gguf \
tools/moe-hot/run-hot.sh

run-hot.sh is in the fork; the MTP head is from drluoto/Qwen3.8-Flash-Next-MTP-GGUF. The lists in hot-lists/ are for a 24 GB card that also drives a desktop (about 3 GB): 12.5 GB of hot experts without vision, 11 GB with the mmproj. For another budget, make one with tools/moe-hot/moe_hot_list.py and give it the IQ3_XXS file as --model, since the hot copies come from there. The fork's README has the details.

With stock llama.cpp

Load shard 1 as usual; every routed expert is then Q2_0. The file has the size of the Q2_0 tier (67 GB with shard 2) and a lower perplexity: +2.6% against IQ3_XXS, where the Q2_0 tier has +4.3%. Upstream llama.cpp has no SIMD code for Q2_0 on x86, so the dot product falls back to scalar code (48 cycles per 32 weights on Zen 4) and experts kept on the CPU (-ot exps=CPU) run slowly. The fork has AVX2 and AVX-512 VBMI versions (5.0 and 3.4 cycles).

Measurements

One machine: RX 7900 XTX 24 GB (Vulkan, RADV), Ryzen 7 7700X, 64 GB DDR5.

Speed with the fork

Decode t/s: greedy, 400 tokens, one prompt per language or topic, second pass; GPU power level high, cold experts locked in RAM, MTP with 3 draft tokens, 32k context, every model with about 1.9 GB of VRAM left free.

t/s hot experts code science English prose Cyrillic (Russian) Chinese
this hybrid 12.6 GB, 71.6% of calls 54.5 40.1 25.5 27.2 23.1
GSQ-RCO IQ3_XXS 12.0 GB, 69.9% 48.6 34.4 19.6 23.3 18.4
GSQ-RCO Q2_0 tier 12.9 GB, 80.1% 65.8 47.5 29.8 30.5 27.0
AD-4.27 (AtomicChat recipe, Navin-Models uncensored) 11.5 GB, 63.4% 43.8 28.6 17.7 20.2 17.1

Against IQ3_XXS the hybrid is 12-30% faster, against AD-4.27 24-44%. The Q2_0 tier is the fastest, at +4.3% perplexity (next table). At the same free VRAM the hybrid holds 0.6 GB more hot experts than IQ3_XXS: the prefill buffer holds the cold experts of one layer, and Q2_0 ones are smaller. Cold experts locked in RAM: 22.4 GiB for the hybrid, 28.9 for IQ3_XXS, 19.7 for Q2_0, 36.5 for AD-4.27. Run to run the numbers move by about 5%, between sessions by up to 10% (how much VRAM the desktop takes changes the hot list).

Quality against GSQ-RCO IQ3_XXS

Same model, so the differences come from the quantization alone. Perplexity on 20k tokens of mixed English documentation, C++ and Russian text (40 chunks of 512), paired with the IQ3_XXS logits; ± is the standard error.

perplexity ratio mean KLD same top-1 token
this hybrid, with the fork 0.998 ± 0.008 0.219 84.0%
this file with stock llama.cpp 1.026 ± 0.011 0.363 79.4%
GSQ-RCO Q2_0 tier 1.043 ± 0.011 0.405 78.1%

The hot/cold split by itself changes nothing: IQ3_XXS with another hot list gives a KLD of 0.000. The KLD of the hybrid comes from its cold Q2_0 experts, which move the distribution without moving the perplexity.

Tasks, reasoning_effort medium, greedy, the hybrid with the fork against IQ3_XXS:

hybrid IQ3_XXS
GSM-Plus, 100 tasks 78 78 the same answer on every task
CRUXEval-O, 100 tasks 94 97 5 character-level slips by the hybrid (a count off by two, a dropped character, a letter case); not significant at n = 100

Vision works with the mmproj of the original GGUFs: 14 px text in a 1280x960 picture was read exactly.

Not measured: standard perplexity sets (wikitext), other hardware, refusal behaviour.

Files

file size content
Qwen3.8-Flash-Next-GSQ-RCO-abliterated-Hybrid-00001-of-00002.gguf 38.4 GB IQ3_XXS trunk, Q2_0 routed experts
Qwen3.8-Flash-Next-GSQ-RCO-abliterated-Hybrid-00002-of-00002.gguf 28.8 GB per-layer n-gram embedding table, byte-identical to shard 2 of SC117's tiers
hot-lists/ expert lists for the fork

Credits and license

  • Qwen: the Qwen3.8-Flash-Next base model, under the Qwen Community License 1.0 (see LICENSE). The upstream GSQ-RCO repositories are tagged Apache-2.0; this repository follows the base model's license.
  • IST-DASLab: GSQ and RCO, and the quantized weights.
  • orcarouter: the abliterated weights. SC117: the abliterated GSQ-RCO GGUFs this file is assembled from.
  • drluoto: the MTP head GGUF.
  • llama.cpp: the GGUF format and the runtime.

Not affiliated with any of them. The model is abliterated: its refusals were removed upstream, and you are responsible for how you use it.

Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration