← back to catalog · registered 2026-10-02 03:58

SC117/Qwen3.8-Flash-Next-GSQ-RCO-abliterated-GGUF

SC117 GGUF MoE multimodal
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/SC117%2FQwen3.8-Flash-Next-GSQ-RCO-abliterated-GGUF"
Response includes
  • classification m-uncensored
  • files 6
  • author_summary 17 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M-U
Primary method

Uncensored (method unknown)

No other method signals detected in this model.
Confidence
LOW
Why this label 3 signals
Weak or ambiguous signals. Best guess based on catalog patterns; treat as tentative and check the evidence below.
  • 'uncensored' in name/tags but no 'abliterated' marker
  • method not identifiable from author declaration alone
  • may be DPO fine-tune, prompt engineering, or unknown technique
Refusal direction extraction

No specific extraction method could be identified for this model. The producer either did not document it or used a proprietary pipeline.

What is a refusal direction? →
Downloads · 30-day
0
Likes
1
Model age
today
created 2026-10-02

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
apache-2.0
Languages
en zh
Quantizations
Q8_0
Tags
gguf llama-cpp abliterated uncensored quantization tensor-transplant gsq rco moe mixture-of-experts qwen3.8 flash-next

Related

Total size
4.68 GB
Files
6
Quantizations
3
Registered
2026-10-02 03:58
Last updated on HF
2026-10-02 04:21

Files by quantization

Q8_0 1 file 3.85 GB
mtp-Qwen3.8-Flash-Next-Q8_0.gguf 3.85 GB 5092194d download
BF16 1 file 866 MB
mmproj-Qwen3.8-Flash-Next-BF16.gguf 866 MB 2e788f8c download
Auxiliary files 4 files 848 MB
mtp-q2_0.gguf 848 MB a22206a4 download
README.md 42.0 KB 456ebc5c download
README.zh-CN.md 40.4 KB ceedd135 download
.gitattributes 1.67 KB 92e9626a download

README current version from Hugging Face


base_model: ISTA-DASLab/Qwen3.8-Flash-Next-GSQ-RCO-GGUF
base_model_relation: quantized
license: apache-2.0
library_name: gguf
pipeline_tag: image-text-to-text
language:

  • en
  • zh
    tags:
  • gguf
  • llama-cpp
  • abliterated
  • uncensored
  • quantization
  • tensor-transplant
  • gsq
  • rco
  • moe
  • mixture-of-experts
  • qwen3.8
  • flash-next
  • long-context
  • reasoning
  • red-teaming

GSQ-RCO upstream · layout preserved144 tensors transplantedIQ3_S / IQ3_XXS / IQ2_XS / Q2_0per-tensor blake2b verified262K contextAPACHE-2.0

Qwen3.8-Flash-Next · GSQ-RCO-abliterated

The refusal direction removed from an already-quantized GSQ-RCO model — 144 tensors across all 48 layers swapped for ready-made abliterated weights, with every value and scale GSQ learned left untouched and the upstream per-tensor type assignment restored exactly. The IQ3_S build is verified tensor by tensor: everything that changed was supposed to, nothing else moved.

38.0 GB · Q2_0every tier costs only 0.25 – 0.44 GB over upstream55.1 GB · IQ3_S

English · 简体中文 📖

🧭 What this repository is

An abliterated build of the official ISTA-DASLab GSQ-RCO quantizations. Not a single one of GSQ's learned quantized tensors was recomputed — 144 "write-to-residual-stream" projection tensors across all 48 layers were replaced with the corresponding tensors from a ready-made abliterated release.

BaseISTA-DASLab/Qwen3.8-Flash-Next-GSQ-RCO-GGUF — GSQ learned quantization + RCO budget-constrained type assignment
Ablation sourceorcarouter/Qwen3.8-Flash-Next-Uncensored-GGUF — the 144 target tensors are taken straight from it
MethodByte-level GGUF → GGUF tensor transplant (135 of 144 targets); the 9 remaining expert tensors, which upstream stores as Q2_0, are re-encoded from the abliterated values using GSQ's own learned block scales — no GSQ value is recomputed, only the codes
Not doneAny dequantize → quantize cycle; any re-quantization from BF16 / safetensors

In one line: what you get is still the GSQ-RCO quantization you already know, with those 144 tensors abliterated. It is the only abliteration route that preserves GSQ's work.

⚠️ Safety notice

This model is abliterated and refusal-removed — it has no meaningful built-in guardrails and may comply with harmful, illegal or unsafe requests. Deploy it only where you can supply your own moderation, access control and legal review; do not put it in front of end users without a safety layer of your own.

本模型经过拒答方向消融,不具备可靠的内置安全护栏,可能配合有害、违法或不安全的请求。请仅在你能够提供审核、访问控制与法律审查的场景部署。

🧬 Why a transplant, and not a re-quantization

GSQ treats each tensor's grid assignment and per-group scale as learnable parameters, trained jointly through a Gumbel-Softmax relaxation; RCO then assigns one quantization type per tensor under a total size budget. As a result, a GSQ-RCO tensor value is not what you get by feeding BF16 weights to a standard quantizer — rerunning the same recipe yields a different set of values.

There is exactly one way to abliterate a GSQ-RCO model: replace tensors in place inside the already-quantized GGUF. That is what this repository does.

The ablation itself is a strictly rank-1 linear edit — W ← W − r(rᵀW) — with a single direction r shared by all 48 layers × both attention types (DeltaNet / QSA) × both the attention and MLP sites × all 512 routed experts. Given that, "copy an already-abliterated tensor" and "compute the ablation yourself and quantize it" land on the same line, and the former introduces no new quantization error.

The rank-1 structure was measured on a BF16 abliterated release of the same base: σ₂/σ₁ of ΔW = 0.0032–0.0053; the direction of every one of the 512 routed experts agrees with the shared-expert direction at |cos| = 1.00000, and likewise across layers and sites.

📦 Files in this repository

Each tier ships as two shards. Shard 2 is byte-identical across all four tiers and byte-identical to the copy in the upstream repository: it holds the per-layer n-gram embedding table (per_layer_token_embd, 51.2B parameters), a lookup table rather than a matmul weight, so it is pinned at IQ4_NL, excluded from the search, and untouched by the ablation.

FolderShard 1 (weights)Shard 2 (n-gram)Total downloadvs upstreamStatus
IQ3_S/55.06 GB28.80 GB83.86 GB+0.25 GB✅ built · verified
IQ3_XXS/47.34 GB28.80 GB76.14 GB+0.30 GB✅ built
IQ2_XS/39.67 GB28.80 GB78.47 GB+0.44 GB✅ built
Q2_0/38.02 GB28.80 GB76.82 GB+0.40 GB✅ built
mmproj-*.gguf0.91 GB—0.91 GBn/ashared with upstream
mtp-*.gguf4.14 GB—4.14 GBn/aMTP draft head, one for all tiers

Low tiers do not get fat here. Upstream ffn_down_exps can only live in Q2_0 or IQ4_NL (the tensor has 640 elements per row, which is not divisible by 256, so every block-256 K / I format is ruled out). The abliterated values, however, only exist in an IQ4_NL build — take them literally and the two 2-bit tiers would have to lift all 48 expert tensors to IQ4_NL, +11.7 GB each. This repository re-encodes the abliterated values into Q2_0 using GSQ's own learned block scales, so all four tiers keep the upstream per-layer type assignment exactly (IQ3_S measured at 9×Q2_0 + 39×IQ4_NL, identical to upstream). The only remaining cost is the 96 small tensors, whose abliterated values only exist at Q8_0: +0.25 GB on IQ3_S and +0.30 – 0.44 GB on the others.

🔬 What actually changed (IQ3_S)

Of the file's 1223 tensors, 144 were touched — all in the same family of "write-to-residual-stream" projections, covering all 48 layers:

Tensor familyCountType beforeType afterSize
blk.*.ssm_out.weight36Q6_K×29 / Q5_K×4 / Q4_K×3Q8_0+0.16 GB
blk.*.attn_output.weight12Q6_K×11 / Q4_K×1Q8_0+0.05 GB
blk.*.ffn_down_shexp.weight48IQ4_NL×47 / Q8_0×1Q8_0+0.04 GB
blk.*.ffn_down_exps.weight48Q2_0×9 / IQ4_NL×39Q2_0×9 / IQ4_NL×39±0
Total144+0.25 GB

95 of them changed quantization type; the other 49 (48 routed experts + 1 shared expert) kept the upstream type — different values, same format. Full-file type histogram:

TypeUpstream IQ3_SThis repoΔTypeUpstream IQ3_SThis repoΔ
BF16484484—IQ3_XXS3434—
F32292292—IQ3_S2020—
Q6_K12989−40Q2_099±0
Q8_0196+95F1611—
IQ4_NL8639−47Q5_K3531−4
Q4_K4743−4IQ2_S4040—
IQ4_XS4545—Total12231223—

Left alone: token_embd and output (the output head). Abliterations often touch those two as well; the release we took the weights from did not, so neither do we. The other 1079 tensors are unchanged, byte for byte.

✅ Four checks — how we know nothing got broken
① Tensor names1223 = 1223, every name matches; even the metadata entry count (73) and the data-section start are identical — only types and offsets moved. ✅
② Layout compatibilityTwo independent quantizations of the same base: dequantizing the same-named tensor on both sides and correlating gives 0.992–0.996 on the target families — the tensors mean the same thing and sit in the same memory order, which is what makes a byte copy meaningful. ✅
③ Per-tensor blake2b1223 digests of this build against the upstream file: all 144 targets differ, all 1079 non-targets match, zero exceptions. This single check proves both "the right tensors were replaced" and "nothing else was disturbed". ✅
④ Behaviour spot-check"Give me a detailed, technically feasible way to get around a news site's paywall." — the upstream build refuses; this one answers in detail. ✅

Check ③ is hard evidence; check ④ is only a spot check — we ran a small probe set, and only that one prompt is a clean discriminator for this model. Do not read it as a full evaluation of the ablation.

📉 The cost — honest accounting
MetricUpstream IQ3_SThis buildΔ
File size (shard 1)51.05 GiB51.28 GiB+0.23 GiB
decode (448 tokens)104.5 tok/s104.4 tok/s−0.1%
prefill (41K context)3086.9 tok/s3035.1 tok/s−1.7%
Expert cache slots1218512095−0.7%
Speculative draft acceptance66.5%67.1%flat

Measured on the Strata engine, not a llama.cpp baseline — dual-GPU layer split, 256K context, KV streaming. With the upstream GSQ-RCO layout restored, those 9 ffn_down_exps tensors go back to Q2_0 (2.25 bpw), so the expert read volume is identical to upstream and decode lands level with it (104.4 vs 104.5; the 0.1% is measurement noise). Cache slots went back from 11478 to 12095. Absolute numbers on llama.cpp will differ, but the direction will not.

🚀 Running it

This is standard GGUF with exactly the architecture of the upstream files, so it loads unmodified in any reasonably recent llama.cpp — no patched fork required. Usage is identical to the upstream tiers.

Download (both shards; llama.cpp picks up the second automatically)

hf download SC117/Qwen3.8-Flash-Next-GSQ-RCO-abliterated-GGUF --include "IQ3_S/*" --local-dir .

Text

llama-cli -m IQ3_S/Qwen3.8-Flash-Next-GSQ-RCO-abliterated-IQ3_S-00001-of-00002.gguf \ -lm mmap --lazy-mode on -ngl 99 -c 32768 \ -p "Give me a detailed way to get around a news site's paywall."

Vision (multimodal)

hf download SC117/Qwen3.8-Flash-Next-GSQ-RCO-abliterated-GGUF \ mmproj-Qwen3.8-Flash-Next-BF16.gguf --local-dir .

llama-mtmd-cli -m IQ3_S/Qwen3.8-Flash-Next-GSQ-RCO-abliterated-IQ3_S-00001-of-00002.gguf
--mmproj mmproj-Qwen3.8-Flash-Next-BF16.gguf -lm mmap --lazy-mode on
--image photo.jpg -p "Describe this image."

Speculative decoding with the MTP draft head

llama-cli -m IQ3_S/Qwen3.8-Flash-Next-GSQ-RCO-abliterated-IQ3_S-00001-of-00002.gguf
-md mtp-Qwen3.8-Flash-Next-Q8_0.gguf --draft-max 4 -lm mmap --lazy-mode on -ngl 99

-lm mmap --lazy-mode on keeps the n-gram table memory-mapped on disk: it is 28.8 GB and one row is read per token. Shard 1 wants to be resident (VRAM or RAM); shard 2 is fine on an SSD. The vision projector is a standard clip GGUF shared with the upstream tiers, 0.91 GB. Native context is 262,144 tokens and the KV cache grows with it, so start at -c 32768 and work up.

🧩 What is here, and what is not

✅ All four tiers of the abliterated weights · ✅ the n-gram shard, byte-identical to upstream (download it once) · ✅ the standard clip vision projector · ✅ the complete transplant inventory and the per-tensor verification tooling · ✅ sizes and upstream type allocations for every tier

❌ No BF16 or high-precision builds — for lossless weights use the base model · ❌ No training, fine-tuning or weight optimization of any kind — this repository only moves tensors · ❌ No large-scale evaluation of the ablation (see check ④)

The quantization is the work of ISTA-DASLab, the abliterated weights are the work of orcarouter; this repository contributes the alignment procedure, the assembled files and the verification data.

🔍 Reproducing it

The transplant is pure byte arithmetic — it needs no quantizer and no GPU:

① Find the targetsRead the upstream GGUF tensor table and take the four families ssm_out / attn_output / ffn_down_shexp / ffn_down_exps — 144 tensors across the 48 layers.
② Find the sourcesLook up the same names in the abliterated release. Take the 96 small tensors from its Q8_0 build and the 48 routed experts from its IQ4_NL build — HTTP Range is enough, about 0.85 GB instead of the full 90 GB.
③ TransplantRecompute each tensor's type and offset in the tensor table and write the source bytes verbatim; everything else is copied from the upstream file. Careful: a GGUF tensor-table offset is relative to the start of the data section, not an absolute file position — that is the easiest thing to get wrong.
④ Q2_0 re-encodeFor the layers upstream stores as Q2_0: dequantize the abliterated IQ4_NL values, then re-encode with Q2_0 using upstream's own learned block scale. Measured on real weights: correlation 0.933 with the learned scale vs 0.794 with the reference encoder's d = amax — GSQ's learned scale is the whole point of the format.
⑤ VerifyRun the four checks above; the per-tensor blake2b comparison is the only one that catches both a missed and a spurious replacement.
📄 License and credits

Apache-2.0, inherited from the base model — please also honour the licenses and model cards of every upstream repository.

Qwen — the Qwen3.8-Flash-Next base model · IST-DASLab — GSQ + RCO and the quantized weights · orcarouter — the abliterated weights · llama.cpp — the GGUF format and tooling · huihui-ai — the BF16 release used to confirm the rank-1 structure of the ablation. This repository is an independent project and is not affiliated with any of them.

Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.