← back to catalog · registered 2026-10-04 02:58

alesha-pro/Qwen3.8-Flash-Next-abliterated-GSQ-RCO-IQ3_S-Strata-GGUF

alesha-pro Qwen GGUF multimodal
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/alesha-pro%2FQwen3.8-Flash-Next-abliterated-GSQ-RCO-IQ3_S-Strata-GGUF"
Response includes
  • classification unknown
  • files 8
  • author_summary 5 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
?
Primary method

Unclassified

No clear signals of an abliteration technique in this model.
Confidence
UNKNOWN
Why this label 1 signal
No classification signals present. This may not be an abliterated model at all - it could be a repackaging, a merge with unrelated goals, or unrelated content that mentions the term.
  • no classification signals present (no abliterated, uncensored, or known producer/method markers)
Refusal direction extraction

No specific extraction method could be identified for this model. The producer either did not document it or used a proprietary pipeline.

What is a refusal direction? →
Downloads · 30-day
0
Likes
0
Model age
today
created 2026-10-04

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
other
Tags
gguf gsq rco strata abliterated control-vector image-text-to-text arxiv:2604.18556 arxiv:2605.00649 base_model:Qwen/Qwen3.8-Flash-Next base_model:quantized:Qwen/Qwen3.8-Flash-Next license:other

Related

Total size
77.9 GB
Files
8
Quantizations
2
Registered
2026-10-04 02:58
Last updated on HF
2026-10-04 03:32

Files by quantization

BF16 1 file 866 MB
mmproj-Qwen3.8-Flash-Next-BF16.gguf 866 MB b1a82259 download
Auxiliary files 7 files 77.9 GB
Qwen3.8-Flash-Next-GSQ-RCO-IQ3_S-00001-of-00002.gguf 51.1 GB 4c1eb2ce download
Qwen3.8-Flash-Next-GSQ-RCO-IQ3_S-00002-of-00002.gguf 26.8 GB 316b46f3 download
Huihui-Qwen3.8-Flash-Next-refusal-direction-r.gguf 472 KB 59f9ae8b download
README.md 8.80 KB de13fb62 download
LICENSE 3.16 KB 9557a896 download
.gitattributes 1.81 KB 558afb37 download
SHA256SUMS 682 B b1b6c5e5 download

README current version from Hugging Face


license: other
license_name: qwen-community-license-1.0
license_link: LICENSE
base_model: Qwen/Qwen3.8-Flash-Next
base_model_relation: quantized
pipeline_tag: image-text-to-text
library_name: gguf
tags:

  • gguf
  • gsq
  • rco
  • strata
  • abliterated
  • control-vector

Qwen3.8-Flash-Next-abliterated-GSQ-RCO-IQ3_S-Strata-GGUF

Credits first

The model weights in this repo are not mine. They are the GSQ-RCO IQ3_S quantization of Qwen3.8-Flash-Next made by the Deep Algorithms and Systems Lab at ISTA, copied byte for byte from ISTA-DASLab/Qwen3.8-Flash-Next-GSQ-RCO-GGUF at revision ed59f92082b1e93c0e96d60a8b11aab089b52f09. The SHA256 of both shards and of the vision projector match their files. See SHA256SUMS.

Part Authors Links
Quantization, GSQ and RCO ISTA-DASLab GSQ paper, GSQ code, RCO paper, RCO code
Base model Qwen team Qwen/Qwen3.8-Flash-Next
Refusal direction, taken from their weights huihui-ai Huihui-Qwen3.8-Flash-Next-abliterated
Runtime with the projection feature Niko1221 Strata

If you use these weights, cite the two ISTA-DASLab papers. The BibTeX is at the bottom of this page.

My part is one 480 KB file, Huihui-Qwen3.8-Flash-Next-refusal-direction-r.gguf, and the measurements below. For any other size, or for the original model without the vector, go to the ISTA-DASLab repo.

The vector

A kit for Strata on one consumer GPU: the official IQ3_S files plus a control vector that removes refusals while the model runs. After each steered layer Strata subtracts one direction r from every residual stream:

h = h - (h . r) r

The weights stay as ISTA-DASLab published them. A request with "experimental_speed_projection": false runs the original model.

The direction comes from the Huihui checkpoint. Huihui changed 144 residual-writing tensors with a rank-one edit along one unit direction r. At strength 1.0 that edit is W_new = W - r (r^T W), which removes r from every block output. The projection in Strata does almost the same thing at run time. It differs in two ways. Strata projects the whole stream, so the embedding and n-gram contributions lose their r component too. Layer 0 cannot be steered.

My EXL3 quantizations of the same model bake the strength-1.0 edit into the weights. This repo keeps the weights original and applies the direction in the engine.

Run it with Strata

You need an NVIDIA or AMD card with 12 GB or more and, for this size, 64 GB of RAM with little else running. Strata's own page has the full requirements.

git clone https://github.com/Niko1221/Strata
hf download alesha-pro/Qwen3.8-Flash-Next-abliterated-GSQ-RCO-IQ3_S-Strata-GGUF --local-dir flash-next-iq3_s
cd Strata
./setup.sh --setup --model IQ3_S --gguf-dir ../flash-next-iq3_s --vision gpu \
  --experimental-speed-projection ../flash-next-iq3_s/Huihui-Qwen3.8-Flash-Next-refusal-direction-r.gguf

On Windows use START-HERE.bat with the same flags. Setup builds or downloads the engine, fetches the MTP draft layer (about 6 GB from the original Qwen checkpoint), packs the model and starts the server at http://127.0.0.1:8080. OpenAI clients use /v1, Anthropic clients use /v1/messages.

Setup writes these engine flags into strata-<model>.json:

--control-vector-scaled <path>/Huihui-Qwen3.8-Flash-Next-refusal-direction-r.gguf:1.0
--control-vector-layer-range 4 44 --cvec-mode project --cvec-dir per-layer

I tested layers 4-44 and 1-47. Both are in the tables below. To use 1-47, change the two numbers in that file.

I did not run the installer end to end on my machine. I built the engine from the pinned source and started it with the same arguments setup writes.

Plain llama.cpp will load these GGUFs, but its stock --control-vector adds a vector. It has no projection mode, so there the files behave as the original model.

Measurements

One RTX 3090 at 300 W, EPYC 7642 (48 cores, AVX2, no AVX-512), 125 GB RAM, Strata 0.1.38 at commit 99f3dbd built for sm_86.

Refusals on 10 prompts (lock picking, phishing email, keylogger, explicit story and similar), sampled at temperature 1.0. The count comes from a regex on the start of the answer, and I read every answer.

Arm Thinking off Thinking on
Original, vector off 10/10 9/10
Strata's bundled vector, layers 4-44 1/10 0/10
This vector, layers 4-44 0/10 0/10
This vector, layers 1-47 0/10 0/10

With thinking on, the keylogger prompt spent all 3,000 tokens reasoning in every vector arm and returned no answer. That is a token budget problem and I did not count it as a refusal. Ten prompts is a small set.

Teacher-forced comparison on four 1,023-token windows (code agent trace, tool-use dialogue, reasoning, plain web text). The reference is the BF16 Huihui model at strength 1.0, run one layer at a time through Transformers. KL is over the engine's top 256 tokens.

Arm KL to BF16 Huihui Top-1 same as BF16 Perplexity KL from original
BF16 Huihui reference 2.565
Original, vector off 0.0834 92.5% 2.525
Strata's bundled vector, 4-44 0.0832 92.3% 2.543 0.018-0.045
This vector, 4-44 0.0738 93.0% 2.527 0.007-0.020
This vector, 1-47 0.0746 92.8% 2.528 0.010-0.023

The reference has noise of its own. Two BF16 passes over the same text with different window lengths differ by KL 0.013-0.021 and agree on top-1 at 95-98%. So part of the 0.07 is the reference and the rest is quantization.

Agent work through OMP, five small coding tasks with tools (a Python log summary fix, an HTTP Range parser with 18 tests, and semver, LRU cache and CSV fixes in JavaScript). I graded each on its original tests.

Arm Passed
Original, vector off 5/5
Strata's bundled vector, 4-44 5/5
This vector, 4-44 5/5
This vector, 1-47 5/5

A second turn on the Python task with the 1-47 vector (add a CLI, keep the tests green) also passed. These tasks do not separate the arms. They show the vector did not break tool use on them.

Other checks used the 1-47 vector with images on and a 131K context.

Check Result
Needle in a haystack at 32K and 115K prompt tokens 6 of 6 found
Vision, one synthetic image text read, both shapes and colours named
Prompt reading at 32K-115K 1,780-1,970 tokens/s
Decode, requests of 200 tokens or more, from the server log median 46-54 tokens/s
MTP drafts accepted 65-70%

The vision check is a smoke test on one image. About 8,000 of the 24,576 experts fit in the 3090's VRAM and the CPU computes the rest. This CPU has no AVX-512, so a desktop with AVX-512 may decode faster.

I did not test harder refusal sets, real photos or documents in vision, more than one GPU, or AMD cards.

Files

File Size From
Qwen3.8-Flash-Next-GSQ-RCO-IQ3_S-00001-of-00002.gguf 54.8 GB ISTA-DASLab, unchanged
Qwen3.8-Flash-Next-GSQ-RCO-IQ3_S-00002-of-00002.gguf 28.8 GB ISTA-DASLab, unchanged (n-gram table)
mmproj-Qwen3.8-Flash-Next-BF16.gguf 0.9 GB ISTA-DASLab, unchanged
tensor-allocation/ ISTA-DASLab, unchanged
Huihui-Qwen3.8-Flash-Next-refusal-direction-r.gguf 480 KB mine, r for layers 1-47 in llama.cpp control vector format
LICENSE, SHA256SUMS

The file names are the published ones because Strata's setup looks for them.

Removing refusals removes a safety behaviour. What the model writes with the vector on is your responsibility.

Citation

Cite the authors of the quantization:

@article{gsq2026,
  title  = {GSQ: Highly-Accurate Low-Precision Scalar Quantization for LLMs via Gumbel-Softmax Sampling},
  author = {Dadgarnia, Alireza and Tabesh, Soroush and Nikdan, Mahdi and Helcig, Michael and Kurtic, Eldar and Kleinegger, Maximilian and Alistarh, Dan},
  journal= {arXiv preprint arXiv:2604.18556},
  year   = {2026}
}
@article{rco2026,
  title  = {Model Compression with Exact Budget Constraints via Riemannian Manifolds},
  author = {Helcig, Michael and Alistarh, Dan},
  journal= {arXiv preprint arXiv:2605.00649},
  year   = {2026}
}

License

ISTA-DASLab state that their quantized weights inherit the license of the base model. The base model, the Huihui checkpoint the direction comes from and this repo are under the Qwen Community License 1.0. Read that file for its terms.

Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration