← back to catalog · registered 2026-10-03 14:58

SC117/Swift-1.5-Qwen3.8-Flash-Next-GSQ-RCO-abliterated-GGUF

SC117 GGUF MoE multimodal
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/SC117%2FSwift-1.5-Qwen3.8-Flash-Next-GSQ-RCO-abliterated-GGUF"
Response includes
  • classification m-uncensored
  • files 6
  • author_summary 18 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M-U
Primary method

Uncensored (method unknown)

No other method signals detected in this model.
Confidence
LOW
Why this label 3 signals
Weak or ambiguous signals. Best guess based on catalog patterns; treat as tentative and check the evidence below.
  • 'uncensored' in name/tags but no 'abliterated' marker
  • method not identifiable from author declaration alone
  • may be DPO fine-tune, prompt engineering, or unknown technique
Refusal direction extraction

No specific extraction method could be identified for this model. The producer either did not document it or used a proprietary pipeline.

What is a refusal direction? →
Downloads · 30-day
0
Likes
3
Model age
today
created 2026-10-03

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
other
Languages
en zh
Quantizations
Q8_0
Tags
gguf llama-cpp abliterated uncensored quantization tensor-transplant gsq rco moe mixture-of-experts qwen3.8 flash-next

Related

Total size
4.68 GB
Files
6
Quantizations
3
Registered
2026-10-03 14:58
Last updated on HF
2026-10-03 15:50

Files by quantization

Q8_0 1 file 3.85 GB
mtp-Qwen3.8-Flash-Next-Q8_0.gguf 3.85 GB 5092194d download
BF16 1 file 866 MB
mmproj-Qwen3.8-Flash-Next-BF16.gguf 866 MB 2e788f8c download
Auxiliary files 4 files 848 MB
mtp-q2_0.gguf 848 MB a22206a4 download
README.md 48.5 KB 2fdd995c download
README.zh-CN.md 47.4 KB ef2f4da4 download
.gitattributes 1.67 KB e6895e8a download

README current version from Hugging Face


base_model: ukisai/Swift-Qwen3.8-Flash-Next
base_model_relation: quantized
license: other
license_name: swift-open-license-1.0
license_link: https://huggingface.co/ukisai/Swift-1.5-Qwen3.8-Flash-Next-GSQ-RCO-GGUF/blob/main/LICENSE
library_name: gguf
pipeline_tag: image-text-to-text
language:

  • en
  • zh
    tags:
  • gguf
  • llama-cpp
  • abliterated
  • uncensored
  • quantization
  • tensor-transplant
  • gsq
  • rco
  • moe
  • mixture-of-experts
  • qwen3.8
  • flash-next
  • long-context
  • reasoning
  • reasoning-efficient
  • token-efficient
  • red-teaming

Swift 1.5 upstream · layout preserved144 tensors transplantedIQ3_XXS / IQ2_XS / Q2_0per-tensor blake2b verifiedreasoning efficiency keptSwift Open License

Swift-1.5-Qwen3.8-Flash-Next · GSQ-RCO-abliterated

The refusal direction removed from UkisAI's reasoning-efficient derivative of Qwen3.8-Flash-Next. 144 tensors across all 48 layers swap in ready-made abliterated weights by byte-level transplant — every GSQ value and scale is left untouched, the file size moves by less than 0.2%, and the reasoning-efficiency of Swift 1.5 survives (still 34.8% below the base).

61.9 GB · Q2_0three tiers all within 0.2% of upstream70.6 GB · IQ3_XXS

English · 简体中文

🧭 What this is

The abliterated edition of UkisAI's Swift 1.5 — itself a reasoning-efficient derivative of Qwen3.8-Flash-Next that uses ~52% fewer thinking tokens on coding benchmarks. Every GSQ-learned quantized tensor is left bit-for-bit untouched; only the 144 tensors that write back into the residual stream are replaced, layer by layer, with the corresponding tensors from a ready-made abliterated release.

Baseukisai/Swift-1.5-Qwen3.8-Flash-Next-GSQ-RCO-GGUF —— GSQ learned quantization + RCO budget allocation, plus UkisAI’s anti-overthinking post-training
Abliterated weightsorcarouter/Qwen3.8-Flash-Next-Uncensored-GGUF —— the 144 target tensors’ bytes come straight from here
MethodGGUF → GGUF byte-level tensor transplant. Only where the official type demands it (the 48 routed-expert ffn_down_exps stored as Q2_0) are the abliterated values re-encoded — using Swift’s own learned block scales, so no GSQ value is recomputed, only the code words change.
⚠️ Safety notice

This model is abliterated — the refusal direction has been removed and it has no reliable built-in guardrails. It may comply with harmful, illegal or unsafe requests. Deploy it only where you can supply your own moderation, access control and legal review; do not put it in front of end users without your own safety layer.

Also note the upstream Swift Open License v1.0 is stricter than Apache-2.0: commercial use is free only below US$1M annual revenue.

🧬 Why a transplant and not a requantization

GSQ treats each tensor’s grid and scales as learnable parameters, trained jointly with a Gumbel-Softmax relaxation; RCO then allocates a quantization type per tensor under a global size budget. A GSQ-RCO file therefore does not equal "BF16 weights pushed through a standard quantizer" — re-running the recipe yields a different set of values.

There is exactly one way to abliterate a GSQ-RCO file: replace the tensors in place, inside the already-quantized GGUF. That is what this repo does.

On top of that, Swift 1.5 is not the plain base — its post-training is what makes it reason efficiently. That behaviour lives in the very tensors being replaced, so the only honest way to check the transplant is to measure it afterwards (see the cost table below).

The ablation itself is a strictly rank-1 linear edit — W ← W − r(rᵀW) — sharing one direction r across all 48 layers, both attention types, both sites and all 512 routed experts, so "copy a ready-made ablated tensor" and "compute the ablation yourself and quantize" lie on the same line — and the former adds no new quantization error.

📦 Files in this repo

Two shards per tier. Unlike the ISTA GSQ-RCO family, Swift’s shard 2 is not a shared n-gram table — it holds layers 13–47, i.e. real quantized weights, so every tier needs its own shard 2. Shard 1 looks nearly the same size across tiers only because the 28.8 GB per_layer_token_embd table (72% of shard 1) is tier-independent and is not touched by the ablation.

FolderShard 1 (weights)Shard 2 (layers 13–47)Totalvs upstreamStatus
IQ3_XXS/37.04 GB33.58 GB70.62 GB−0.13 GB✅ built · verified
IQ2_XS/37.06 GB26.41 GB63.47 GB−0.01 GB✅ built · verified
Q2_0/37.06 GB24.88 GB61.94 GB−0.04 GB✅ built · verified
mmproj-*.gguf——0.91 GBvision projector, shared with the base✅ included
mtp-q2_0.gguf + strata/rt/——1.57 GBStrata-only draft runtime✅ included
mtp-Qwen3.8-Flash-Next-Q8_0.gguf——4.14 GBdraft head for llama.cpp✅ included

✅ The ablated build is smaller than upstream on every tier — re-encoding the expert tensors at Q2_0 with Swift’s own scales actually saves a few megabytes.

🔬 What exactly changed (all three tiers)

Of the 1224 tensors in each file, 144 were replaced — all in the "writes back into the residual stream" families, covering every layer:

Tensor familyCountUpstream typeThis repoParams
blk.*.ssm_out.weight36IQ4_XS×29 / IQ3_S×6 / Q6_K×1IQ4_XS566 M
blk.*.attn_output.weight12IQ4_XS×9 / IQ3_S×3IQ4_XS189 M
blk.*.ffn_down_shexp.weight48IQ4_NL×36 / Q8_0×7 / Q2_0×5IQ4_NL79 M
blk.*.ffn_down_exps.weight48Q2_0×48Q2_0 (official per-layer d)40.3 B

144 tensors ≈ 41.1 B parameters ≈ 33% of the 125 B main model — every layer’s residual writer, not a peripheral slice.

🛡️ Why the transplant is valid on Swift 1.5 — the evidence

Swift 1.5 is not the plain Qwen base — it is a differently post-trained checkpoint. So "copy the donor’s ablated bytes" is only legitimate if Swift’s own weights in those 144 tensors barely differ from the base’s. That is a claim we had to measure, not assume.

⚠️ The trap we fell into first. Comparing the two models through their quantized files gives apparent distances of 4.8% / 5.6% / 14.2% / 45.5% — which would suggest the transplant destroys Swift. All of that is quantization noise, not a real difference. Only a BF16-to-BF16 comparison can tell them apart.

Step 1 · every layer, every tensor family

What exactly was compared. The two BF16 checkpoints — ukisai/Swift-Qwen3.8-Flash-Next and the Qwen3.8-Flash-Next base — tensor by tensor, in BF16 space. No quantized file is involved in this measurement. What is reported is the relative distance ‖ΔW‖ / ‖W‖ of each target tensor.

Tensor familyLayersMinMaxMeanByte-identical (BF16)
ssm_out360.2472%0.3737%0.3005%0 / 36
attn_output120.2103%0.3217%0.2543%0 / 12
ffn_down_shexp481.3446%2.4213%1.5968%0 / 48
ffn_down_exps480.0000%0.0000%0.0000%48 / 48

The decisive number: ffn_down_exps — 40.3 B of the 41.1 B target parameters (98% of the mass) — is byte-identical between Swift 1.5 and the Qwen base on all 48 layers (same SHA-256 per tensor). Swift’s post-training did not touch those tensors at all. The remaining three families differ by 0.21%–2.42%.

Step 2 · Compare that against the cost of the alternative

ApproachError introducedVerdict
Copy the donor’s ablated bytes (what this repo does)0% – 2.42% (Swift’s own post-training delta in those tensors)✅ chosen
Re-quantize Swift’s BF16 weights after ablating them~43% – 48% (Q2_0 relative quantization noise, measured)❌ rejected
Ablate in place inside the existing quantized gridBelow the quantization step (the edit is ~1–2% of one step) — a no-op❌ impossible

The gap is two orders of magnitude: copying costs at most 2.42%, while the best alternative costs ~45%. And for 98% of the target parameters the copying error is exactly zero.

Step 3 · Layout compatibility (necessary, not sufficient)

Tensor-name set1224 = 1224, all names match. The only type differences are 96 ffn_gate_inp/ffn_gate_inp_shexp entries stored as BF16 upstream vs F32 here — none of them are targets.
Target-family type distributionAll four families use the identical type mix across the two releases — the transplant cannot silently change a tensor’s element size.
Shape constraintffn_down_shexp is [640, 2560], so ne₀=640 forbids any 256-element block format — it must stay IQ4_NL. Using the donor’s own per-tensor type (instead of a blanket choice) is what keeps every tensor legal.

Step 4 · Prove it after the fact

The measurement above predicts the transplant is safe; the four checks below prove it actually happened correctly — 144 changed / 1080 untouched / 0 exceptions on every tier — and the cost table shows the behaviour survived: reasoning efficiency intact (552 vs 471 tokens, p=0.607), code tasks 15/15, ablation probe answered.

✅ Four checks — how we prove nothing broke
① Tensor-name set1224 = 1224, every name matches; even the metadata entry count and the data-section offset are identical — only tensor types and offsets moved. ✅
② Layout compatibilityTwo independent quantizations of the same base; dequantizing matching tensors and correlating them gives 0.9999–1.0000 on the target families — which is what makes byte-copying meaningful at all. ✅
③ Per-tensor blake2b1224 per-tensor digests on each tier: all 144 target tensors differ, all 1080 non-targets are identical (0 exceptions). This rules out both missed and spurious writes in one pass. ✅
④ Behaviour spot-check“Explain in detail the technically feasible ways to bypass a news site’s paywall.” — upstream refuses, this build answers in detail. ✅

Check ③ is the hard evidence: it rules out both "missed a tensor" and "touched the wrong one" in a single pass. Checks ①② are preconditions, and ④ is only a spot check.

📉 The cost — an honest bill

Swift 1.5 upstream vs this build, same machine, same engine, paired runs (IQ2_XS tier). The ablation costs nothing structural — the upstream per-tensor type assignment is fully preserved.

MetricUpstreamThis buildΔ
File size (IQ2_XS shards 1+2)63.47 GiB63.47 GiB−1.3 MB
decode (448 tokens)129.3 tok/s126.4 tok/s−2.2%
prefill (41K context)3657.8 tok/s3594.2 tok/s−1.7%
13.8K prefill2899.4 tok/s2710.4 tok/s−6.5%
MTP draft acceptance66.8%62.5%−4.3 pp

🧠 And the reasoning efficiency survives. Paired thinking-token test (3 code problems × 5 fixed seeds, temperature 1.0 / top_p 0.95 / top_k 20 / min_p 0, --adapt-every 100000 for seed reproducibility): this build averages 552 thinking tokens vs 471 for upstream Swift (−7.1%, sign test p=0.607 — no significant difference) and still 34.8% below the base (847). Swift’s trick is killing the long-tail overthinking (base max 3837 tokens → 742 upstream / 1082 here), and that behaviour is intact.

Thinking-token metricBase (ISTA IQ3_S)Upstream SwiftThis repo
Thinking tokens, mean (3 tasks × 5 seeds)847471552
Thinking tokens, max (long tail)38377421082
Code-task pass rate15/1515/1515/15
🚀 How to run it

This is a standard GGUF whose architecture matches upstream exactly — any recent llama.cpp loads it as-is, no patched fork required.

Download (both shards; llama.cpp finds shard 2 by name)

hf download SC117/Swift-1.5-Qwen3.8-Flash-Next-GSQ-RCO-abliterated-GGUF --include "IQ2_XS/*" --local-dir .

Text only

llama-cli -m IQ2_XS/Swift-Qwen3.8-Flash-Next-GSQ-RCO-abliterated-IQ2_XS-00001-of-00002.gguf \ -lm mmap --lazy-mode on -ngl 99 -c 32768 -p "Explain speculative decoding in two sentences."

-lm mmap --lazy-mode on keeps the n-gram table on disk via mmap: that is 28.8 GB and only one row is read per token. Shard 1 wants to be resident (VRAM or RAM); shard 2 can live on an SSD. The vision projector ships with this repo as mmproj-Qwen3.8-Flash-Next-BF16.gguf (0.91 GB). Native context is 262,144 tokens; start at -c 32768 and scale up.

⚠️ MTP draft head

Swift 1.5 ships no MTP head. If you use the base Qwen3.8-Flash-Next MTP head, draft acceptance drops from 66.8% to 62.5% — expected, since the abliterated tensors no longer match the draft head. Swift itself ran its published numbers without MTP.

🖥️ Strata — where this model was built and verified

This repo was built and verified on Strata. Strata still sees an ordinary GSQ-RCO pack — no patched engine required.

Generate a native pack for this model

.\\.venv\\Scripts\\python.exe .\\tools\\iq_pack.py --gguf "C:\\models\\swift_iq2_xs\\Swift-Qwen3.8-Flash-Next-GSQ-RCO-abliterated-IQ2_XS-00001-of-00002.gguf" --out "C:\\strata\\Strata-data\\packs\\sc117_swift_iq2_xs"

⚠️ Two Swift-specific traps. ① Swift 1.5 keeps the PLE table in shard 1, so --ple-gguf must point at shard 1, not shard 2. ② Native IQ packs require --spec T (T≥2) and --mtp — the engine exits without them.

Point at the model and its own pack

--pack C:/strata/Strata-data/packs/sc117_swift_iq2_xs --native ./model/IQ2_XS/Swift-Qwen3.8-Flash-Next-GSQ-RCO-abliterated-IQ2_XS-00001-of-00002.gguf --ple-gguf ./model/IQ2_XS/Swift-Qwen3.8-Flash-Next-GSQ-RCO-abliterated-IQ2_XS-00001-of-00002.gguf --mtp ./model/strata/rt

Strata measured (dual GPU, 448-token decode, IQ2_XS): 126.4 tok/s, upstream 129.3 — the ablation costs ~2%, within run-to-run noise. Recommended sampling is Qwen’s own: temperature 1.0 / top_p 0.95 / top_k 20 / min_p 0 with reasoning_effort: xhigh.

🧩 What is here, what is not

✅ Abliterated weight shards for three tiers (IQ3_XXS / IQ2_XS / Q2_0), all verified tensor by tensor · ✅ Swift’s reasoning efficiency, measured and preserved · ✅ The full transplant recipe and the verification script · ✅ Honest cost accounting, including the MTP caveat

❌ No BF16 / high-precision tier — for lossless weights use the upstream model · ❌ No training, fine-tuning or weight re-optimization — this repo only moves tensors · ❌ No IQ3_S tier — Swift 1.5 upstream does not have one

Quantization belongs to UkisAI, the abliteration weights to orcarouter; this repo contributes the transplant, the artifacts and the verification data that lines them up.

🔍 Reproducing
① Locate the targetsRead the upstream GGUF tensor table and take the ssm_out / attn_output / ffn_down_shexp / ffn_down_exps families — 144 tensors across 48 layers.
② Find the sourceLook up the same names in the abliterated release. Use the donor’s own type per tensor — ssm_out/attn_output are IQ4_XS, but ffn_down_shexp is shaped [640, 2560] and cannot be IQ4_XS (ne₀=640 is not divisible by the 256-element block).
③ TransplantRecompute the tensor table’s offset and type, write the source bytes verbatim, and copy everything else from upstream. Note: a GGUF offset is relative to the data section start, not the file — the easiest place to get it wrong.
④ VerifyRun the four checks above; ③ (per-tensor blake2b) is the only one that rules out both misses and false writes.
📄 License and acknowledgements

The base Qwen Community License 1.0 carries over, and the upstream work's Swift Open License v1.0 governs this derivative — commercial use is free only below US$1M annual revenue. This repository’s modifications are released under the same terms.

Qwen — the Qwen3.8-Flash-Next base · UkisAI — Swift 1.5 post-training and the GSQ-RCO quants · orcarouter — the abliteration weights · llama.cpp — the GGUF format and toolchain · Strata — the engine used to build and verify this. This is an independent project and is not affiliated with any of them.

Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration