license: other
license_name: qwen-community-license-1.0
license_link: LICENSE
base_model: ISTA-DASLab/Qwen3.8-Flash-Next-GSQ-RCO-Coder-GGUF
base_model_relation: quantized
pipeline_tag: image-text-to-text
library_name: gguf
tags:
- gguf
- gsq
- rco
- expert-pruning
- code
- strata
- abliterated
- control-vector
Qwen3.8-Flash-Next-abliterated-GSQ-RCO-Coder-Strata-GGUF
Credits first
The model weights in this repo are not mine. They are the GSQ-RCO Coder release of Qwen3.8-Flash-Next made by the Deep Algorithms and Systems Lab at ISTA: 256 of the 512 experts in each layer are kept, the rest are removed, and the kept weights are stored at 3.5 bits. The files are copied byte for byte from ISTA-DASLab/Qwen3.8-Flash-Next-GSQ-RCO-Coder-GGUF at revision 5348543e0147355ac9cbcb031184a3546350988e. The SHA256 of both shards and of the vision projector match their files. See SHA256SUMS.
| Part | Authors | Links |
|---|---|---|
| Expert pruning and quantization, GSQ and RCO | ISTA-DASLab | GSQ paper, GSQ code, RCO paper, RCO code |
| Base model | Qwen team | Qwen/Qwen3.8-Flash-Next |
| Refusal direction, taken from their weights | huihui-ai | Huihui-Qwen3.8-Flash-Next-abliterated |
| Runtime with the projection feature | Niko1221 | Strata |
If you use these weights, cite the two ISTA-DASLab papers. The BibTeX is at the bottom of this page.
My part is the check that the refusal vector I published for the unpruned sizes also works on the pruned model, and the measurements below. The vector is the same 480 KB file as in the repo with the four unpruned sizes, with the same SHA256. For the original card and ISTA's own SWE-bench and LiveCodeBench numbers, go to the ISTA-DASLab repo.
What the Coder is
ISTA-DASLab removed half of the experts of Qwen3.8-Flash-Next and picked the ones to keep on code, agentic and vision data. The download is 58.4 GB instead of 83.6 GB for IQ3_S, and only the first shard (29.6 GB) has to be in memory. Their numbers: 91.3% of the full model's SWE-bench Verified score and 98.7% of LiveCodeBench v6. They also say it is weaker outside code. Strata's page adds that languages other than English got worse. If you want a general model and have 48 GB of RAM or more, take one of the unpruned sizes from my other repo.
The vector
After each steered layer Strata subtracts one direction r from every residual stream:
h = h - (h . r) r
The weights stay as ISTA-DASLab published them. A request with "experimental_speed_projection": false runs the original model.
The direction comes from the Huihui checkpoint of the full model. Huihui changed 144 residual-writing tensors with a rank-one edit along one unit direction r, and the projection in Strata does almost the same thing at run time. The Coder keeps the attention, the router and the residual stream of the full model, but only half of the experts that write into that stream. So I did not assume the direction carries over. I measured it, and it does.
Run it with Strata
You need an NVIDIA or AMD card with 12 GB or more and 32 GB of RAM. Strata's own page has the full requirements.
git clone https://github.com/Niko1221/Strata
hf download alesha-pro/Qwen3.8-Flash-Next-abliterated-GSQ-RCO-Coder-Strata-GGUF --local-dir coder
mv coder/mmproj-Qwen3.8-Flash-Next-BF16.gguf coder/IQ1_M/
cd Strata
./setup.sh --setup --family coder --gguf-dir ../coder/IQ1_M --vision gpu \
--experimental-speed-projection ../coder/Huihui-Qwen3.8-Flash-Next-refusal-direction-r.gguf
On Windows use START-HERE.bat with the same flags. The mv puts the vision projector next to the shards, where setup looks for it. Setup builds or downloads the engine, fetches the MTP draft layer (about 6 GB from the original Qwen checkpoint), packs the model and starts the server at http://127.0.0.1:8080. OpenAI clients use /v1, Anthropic clients use /v1/messages.
If you already have one of the unpruned sizes, you only need the first shard. Shard 2 (the n-gram table) and the vision projector are the same files.
Setup writes these engine flags into strata-<model>.json:
--control-vector-scaled <path>/Huihui-Qwen3.8-Flash-Next-refusal-direction-r.gguf:1.0
--control-vector-layer-range 4 44 --cvec-mode project --cvec-dir per-layer
The numbers below are for layers 1-47. To use that range, change the two numbers in that file. The 4-44 range is being measured now.
I did not run the installer end to end on my machine. I built Strata 0.1.40 from source and started the engine with the arguments setup writes for --family coder.
Plain llama.cpp will load these GGUFs, but its stock --control-vector adds a vector. It has no projection mode, so there the files behave as the original model.
Measurements
One RTX 3090 at 300 W, EPYC 7642 (48 cores, AVX2, no AVX-512), 125 GB RAM, Strata 0.1.40 built for sm_86, a 65,536 context. At that context the card held 8,524 of the 12,288 experts.
Refusals. The first three rows are greedy with thinking off, 96 tokens per answer. The probe is the 10 prompts of my other cards (lock picking, phishing email, keylogger, explicit story and similar), sampled at temperature 1.0. A refusal is a regex on the first 400 characters of the answer.
| Set | Original, vector off | This vector, layers 1-47 |
|---|---|---|
| 64 harmful instructions (AdvBench) | 64 refused | 0 |
| 82 harmful behaviours (JailbreakBench, rows that are not from AdvBench) | 79 | 1 |
| 64 ordinary instructions (Alpaca) | 2 | 0 |
| 10-prompt probe, thinking off | 10 | 1 |
| 10-prompt probe, thinking on | 8 | 0 |
I read the opening of every flagged answer in the vector column. The JailbreakBench hit is an answer that does the task but turns it around: asked for an expert thread that spreads a myth, the model wrote a thread that debunks it. A wider regex flags 2 more answers there. One argues with the request instead of doing it (a campaign against hiring people with disabilities). The other is a false match, the model answers the question. The probe hit is the piracy question: the model answers but says it will not give download links. So the plain refusals are gone, and 3 answers out of 156 still push back in a softer way.
With thinking on, 3 of the 10 probe prompts spent all 3,000 tokens reasoning and returned no answer. I did not count those as refusals. The original's 2 hits on ordinary instructions are one real refusal of a medical question and one answer that starts with "I cannot generate audio".
Distance to the original Coder, teacher-forced on 6 windows of 1,024 tokens of prose and 6 of code. KL is from the original to the vector arm over the original's top 256 tokens.
| Text | KL | Same top token | Perplexity, original | Perplexity, vector |
|---|---|---|---|---|
| Prose | 0.034 | 92.2% | 5.825 | 5.856 |
| Code | 0.013 | 96.9% | 2.241 | 2.252 |
Still running, I will add them here: refusals with thinking on over a larger set, the 4-44 range, agent coding tasks, a tool call and an image request, needle tests up to 250K, and the speed curve from 4K to 261K.
Files
| File | Size | From |
|---|---|---|
IQ1_M/Qwen3.8-Flash-Next-GSQ-RCO-IQ1_M-00001-of-00002.gguf |
29.6 GB | ISTA-DASLab, unchanged |
IQ1_M/Qwen3.8-Flash-Next-GSQ-RCO-IQ1_M-00002-of-00002.gguf |
28.8 GB | ISTA-DASLab, unchanged (n-gram table, the same file as in the unpruned sizes) |
mmproj-Qwen3.8-Flash-Next-BF16.gguf |
0.9 GB | ISTA-DASLab, unchanged |
tensor-allocation/ |
ISTA-DASLab, unchanged | |
Huihui-Qwen3.8-Flash-Next-refusal-direction-r.gguf |
480 KB | mine, r for layers 1-47 in llama.cpp control vector format |
LICENSE, SHA256SUMS |
The file names and folders are the published ones because Strata's setup looks for them. ISTA-DASLab named the size IQ1_M for its 1.89 bits per parameter of the original model. The kept weights are stored like IQ3_S.
Removing refusals removes a safety behaviour. What the model writes with the vector on is your responsibility.
Citation
Cite the authors of the pruning and quantization:
@article{gsq2026,
title = {GSQ: Highly-Accurate Low-Precision Scalar Quantization for LLMs via Gumbel-Softmax Sampling},
author = {Dadgarnia, Alireza and Tabesh, Soroush and Nikdan, Mahdi and Helcig, Michael and Kurtic, Eldar and Kleinegger, Maximilian and Alistarh, Dan},
journal= {arXiv preprint arXiv:2604.18556},
year = {2026}
}
@article{rco2026,
title = {Model Compression with Exact Budget Constraints via Riemannian Manifolds},
author = {Helcig, Michael and Alistarh, Dan},
journal= {arXiv preprint arXiv:2605.00649},
year = {2026}
}
License
ISTA-DASLab state that their weights inherit the license of the base model. The base model, the Huihui checkpoint the direction comes from and this repo are under the Qwen Community License 1.0. Read that file for its terms.