license: other
license_name: qwen-community-license-1.0
license_link: LICENSE
base_model: Qwen/Qwen3.8-Flash-Next
base_model_relation: quantized
pipeline_tag: image-text-to-text
library_name: gguf
tags:
- gguf
- gsq
- rco
- strata
- abliterated
- control-vector
Qwen3.8-Flash-Next-abliterated-GSQ-RCO-IQ2_XS-Strata-GGUF
Credits first
The model weights in this repo are not mine. They are the GSQ-RCO IQ2_XS quantization of Qwen3.8-Flash-Next made by the Deep Algorithms and Systems Lab at ISTA, copied byte for byte from ISTA-DASLab/Qwen3.8-Flash-Next-GSQ-RCO-GGUF at revision ed59f92082b1e93c0e96d60a8b11aab089b52f09. The SHA256 of both shards and of the vision projector match their files. See SHA256SUMS.
| Part | Authors | Links |
|---|---|---|
| Quantization, GSQ and RCO | ISTA-DASLab | GSQ paper, GSQ code, RCO paper, RCO code |
| Base model | Qwen team | Qwen/Qwen3.8-Flash-Next |
| Refusal direction, taken from their weights | huihui-ai | Huihui-Qwen3.8-Flash-Next-abliterated |
| Runtime with the projection feature | Niko1221 | Strata |
If you use these weights, cite the two ISTA-DASLab papers. The BibTeX is at the bottom of this page.
My part is one 480 KB file, Huihui-Qwen3.8-Flash-Next-refusal-direction-r.gguf, and the measurements below. For any other size, or for the original model without the vector, go to the ISTA-DASLab repo.
For the size that stays closer to the full model see the IQ3_S kit. In my runs IQ3_S is closer to the BF16 model (KL 0.07 against 0.19, same top-1 token 93% against 88%) and IQ2_XS decodes 45-70% faster on prompts up to 131K and keeps 14 GB less in RAM.
The vector
A kit for Strata on one consumer GPU: the official IQ2_XS files plus a control vector that removes refusals while the model runs. After each steered layer Strata subtracts one direction r from every residual stream:
h = h - (h . r) r
The weights stay as ISTA-DASLab published them. A request with "experimental_speed_projection": false runs the original model.
The direction comes from the Huihui checkpoint. Huihui changed 144 residual-writing tensors with a rank-one edit along one unit direction r. At strength 1.0 that edit is W_new = W - r (r^T W), which removes r from every block output. The projection in Strata does almost the same thing at run time. It differs in two ways. Strata projects the whole stream, so the embedding and n-gram contributions lose their r component too. Layer 0 cannot be steered.
My EXL3 quantizations of the same model bake the strength-1.0 edit into the weights. This repo keeps the weights original and applies the direction in the engine.
Run it with Strata
You need an NVIDIA or AMD card with 12 GB or more and, for this size, 48 GB of RAM. Strata's own page has the full requirements.
git clone https://github.com/Niko1221/Strata
hf download alesha-pro/Qwen3.8-Flash-Next-abliterated-GSQ-RCO-IQ2_XS-Strata-GGUF --local-dir flash-next-iq2_xs
cd Strata
./setup.sh --setup --model IQ2_XS --gguf-dir ../flash-next-iq2_xs --vision gpu \
--experimental-speed-projection ../flash-next-iq2_xs/Huihui-Qwen3.8-Flash-Next-refusal-direction-r.gguf
On Windows use START-HERE.bat with the same flags. Setup builds or downloads the engine, fetches the MTP draft layer (about 6 GB from the original Qwen checkpoint), packs the model and starts the server at http://127.0.0.1:8080. OpenAI clients use /v1, Anthropic clients use /v1/messages.
Setup writes these engine flags into strata-<model>.json:
--control-vector-scaled <path>/Huihui-Qwen3.8-Flash-Next-refusal-direction-r.gguf:1.0
--control-vector-layer-range 4 44 --cvec-mode project --cvec-dir per-layer
I tested layers 4-44 and 1-47. Both are in the tables below. To use 1-47, change the two numbers in that file.
I did not run the installer end to end on my machine. I built the engine from the pinned source and started it with the same arguments setup writes.
Plain llama.cpp will load these GGUFs, but its stock --control-vector adds a vector. It has no projection mode, so there the files behave as the original model.
Measurements
One RTX 3090 at 300 W, EPYC 7642 (48 cores, AVX2, no AVX-512), 125 GB RAM, Strata 0.1.38 at commit 99f3dbd built for sm_86.
Refusals on 10 prompts (lock picking, phishing email, keylogger, explicit story and similar), sampled at temperature 1.0. The count comes from a regex on the start of the answer. I also read the opening of every answer in the 1-47 arm. The other two arms are regex counts.
| Arm | Thinking off | Thinking on |
|---|---|---|
| Original, vector off | 9/10 | 9/10 |
| This vector, layers 4-44 | 0/10 | 0/10 |
| This vector, layers 1-47 | 0/10 | 0/10 |
With thinking on, three or four of the ten prompts spent all 3,000 tokens reasoning in each vector arm and returned no answer. On IQ3_S that happened once in ten. I checked the reasoning text for loops and found none (under 1% of 12-word sequences repeat), so this size deliberates longer before it answers. I did not count those as refusals. Ten prompts is a small set.
Teacher-forced comparison on four 1,023-token windows (code agent trace, tool-use dialogue, reasoning, plain web text). The reference is the BF16 Huihui model at strength 1.0, run one layer at a time through Transformers. KL is over the engine's top 256 tokens.
| Arm | KL to BF16 Huihui | Top-1 same as BF16 | Perplexity | KL from original |
|---|---|---|---|---|
| BF16 Huihui reference | 2.565 | |||
| Original, vector off | 0.2025 | 87.5% | 2.600 | |
| This vector, 4-44 | 0.1923 | 87.4% | 2.603 | 0.005-0.019 |
| This vector, 1-47 | 0.1933 | 87.8% | 2.592 | 0.006-0.018 |
The reference has noise of its own. Two BF16 passes over the same text with different window lengths differ by KL 0.013-0.021 and agree on top-1 at 95-98%. So a small part of the 0.19 is the reference and the rest is quantization. The same measurement on IQ3_S gives 0.07.
Agent work through OMP, five small coding tasks with tools (a Python log summary fix, an HTTP Range parser with 18 tests, and semver, LRU cache and CSV fixes in JavaScript). I graded each on its original tests.
| Arm | Passed |
|---|---|
| This vector, 1-47 | 5/5 |
A second turn on the Python task (add a CLI, keep the tests green) also passed. The original and the 4-44 arm were not run on these tasks at this size. On IQ3_S all four arms passed 5/5.
Other checks used the 1-47 vector with images on and a 131K context.
| Check | Result |
|---|---|
| Needle in a haystack at 32K and 115K prompt tokens | 6 of 6 found |
| Vision, one synthetic image | text read, both shapes and colours named |
| MTP drafts accepted over all agent and probe requests | 65-71% |
The speed curve below uses the prompts of my EXL3 cards. Each request reads an independent document and writes 512 tokens, nonthinking and greedy, with nothing cached. I ran each depth twice with MTP drafts on, a 262K context with KV streaming and the 1-47 vector on. Decode and prompt reading are tokens/s, mean of the two repeats, as the engine reports them.
| Prompt tokens | IQ3_S decode | IQ2_XS decode | IQ3_S prompt reading | IQ2_XS prompt reading | IQ3_S first token | IQ2_XS first token |
|---|---|---|---|---|---|---|
| 4,096 | 58.1 | 84.2 | 890 | 1,164 | 4.6 s | 3.6 s |
| 32,768 | 59.0 | 87.1 | 1,840 | 2,014 | 17.9 s | 16.3 s |
| 131,072 | 46.6 | 79.6 | 1,768 | 2,001 | 74.2 s | 65.6 s |
| 261,120 | 39.5 | 47.3 | 1,546 | 1,669 | 169.1 s | 156.6 s |
The two repeats differ by up to 8 tokens/s at 4K and 12 at 261K. MTP accepted 60-65% of the drafts. Peak VRAM was 23.8 GB on both sizes. The engine refuses 261,632 + 512 tokens in a 262,144 context, so the deepest point is 261,120.
Agent sessions and the refusal probe run with thinking and sampling at temperature 1.0. There the server log gives a median decode of 73-80 tokens/s on requests of 200 tokens or more.
The vision check is a smoke test on one image. About 12,000-13,000 of the 24,576 experts fit in the 3090's VRAM and the CPU computes the rest. This CPU has no AVX-512, so a desktop with AVX-512 may decode faster.
I did not test harder refusal sets, real photos or documents in vision, more than one GPU, or AMD cards.
Files
| File | Size | From |
|---|---|---|
Qwen3.8-Flash-Next-GSQ-RCO-IQ2_XS-00001-of-00002.gguf |
39.2 GB | ISTA-DASLab, unchanged |
Qwen3.8-Flash-Next-GSQ-RCO-IQ2_XS-00002-of-00002.gguf |
28.8 GB | ISTA-DASLab, unchanged (n-gram table) |
mmproj-Qwen3.8-Flash-Next-BF16.gguf |
0.9 GB | ISTA-DASLab, unchanged |
tensor-allocation/ |
ISTA-DASLab, unchanged | |
Huihui-Qwen3.8-Flash-Next-refusal-direction-r.gguf |
480 KB | mine, r for layers 1-47 in llama.cpp control vector format |
LICENSE, SHA256SUMS |
The file names are the published ones because Strata's setup looks for them.
Removing refusals removes a safety behaviour. What the model writes with the vector on is your responsibility.
Citation
Cite the authors of the quantization:
@article{gsq2026,
title = {GSQ: Highly-Accurate Low-Precision Scalar Quantization for LLMs via Gumbel-Softmax Sampling},
author = {Dadgarnia, Alireza and Tabesh, Soroush and Nikdan, Mahdi and Helcig, Michael and Kurtic, Eldar and Kleinegger, Maximilian and Alistarh, Dan},
journal= {arXiv preprint arXiv:2604.18556},
year = {2026}
}
@article{rco2026,
title = {Model Compression with Exact Budget Constraints via Riemannian Manifolds},
author = {Helcig, Michael and Alistarh, Dan},
journal= {arXiv preprint arXiv:2605.00649},
year = {2026}
}
License
ISTA-DASLab state that their quantized weights inherit the license of the base model. The base model, the Huihui checkpoint the direction comes from and this repo are under the Qwen Community License 1.0. Read that file for its terms.