← back to catalog · registered 2026-10-08 01:58

sanasol2008/Qwen3.8-Flash-Next-Abliterated-Sushi-2bpw

sanasol2008 MoE multimodal second-order
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/sanasol2008%2FQwen3.8-Flash-Next-Abliterated-Sushi-2bpw"
Response includes
  • classification unknown
  • files 66
  • author_summary 1 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Refusal direction extraction

No specific extraction method could be identified for this model. The producer either did not document it or used a proprietary pipeline.

What is a refusal direction? →
Downloads · 30-day
0
Likes
0
Model age
today
created 2026-10-08

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
other
Tags
mlx safetensors qwen4_exp sushi exl3 2bit abliterated qwen moe image-text-to-text conversational base_model:orcarouter/Qwen3.8-Flash-Next-Uncensored

Related

Total size
64.8 GB
Files
66
Quantizations
1
Registered
2026-10-08 01:58
Last updated on HF
2026-10-08 01:35

Files by quantization

Auxiliary files 66 files 64.9 GB
ngram_table.bin 29.8 GB 53e12156 download
trunk.safetensors 4.96 GB dd5f3b87 download
vision.safetensors 856 MB d9208f7b download
experts-10.safetensors 609 MB 6a3c6deb download
experts-11.safetensors 609 MB f77695f8 download
experts-12.safetensors 609 MB e91eba58 download
experts-13.safetensors 609 MB 166194ec download
experts-14.safetensors 609 MB dbbb0b94 download
experts-15.safetensors 609 MB 1e01db91 download
experts-16.safetensors 609 MB 74e6d6ef download
experts-17.safetensors 609 MB 320163d3 download
experts-18.safetensors 609 MB 2a34fe21 download
experts-19.safetensors 609 MB a677f5c7 download
experts-20.safetensors 609 MB 4a65ab5c download
experts-21.safetensors 609 MB 61c2ccb6 download
experts-22.safetensors 609 MB 7ededf47 download
experts-23.safetensors 609 MB 05aa99ad download
experts-24.safetensors 609 MB 096f5894 download
experts-25.safetensors 609 MB 7e21e82c download
experts-26.safetensors 609 MB cdf76b7f download
experts-27.safetensors 609 MB e578626b download
experts-28.safetensors 609 MB 67c5468b download
experts-29.safetensors 609 MB 88472655 download
experts-30.safetensors 609 MB bf1e46f0 download
experts-31.safetensors 609 MB d0a2d979 download
experts-32.safetensors 609 MB a0553c6b download
experts-33.safetensors 609 MB c1587181 download
experts-34.safetensors 609 MB de8098d0 download
experts-35.safetensors 609 MB b016ac26 download
experts-36.safetensors 609 MB b4d4c296 download
experts-37.safetensors 609 MB c9e5af45 download
experts-38.safetensors 609 MB e459110b download
experts-39.safetensors 609 MB 0c70aaac download
experts-40.safetensors 609 MB 2c57479e download
experts-41.safetensors 609 MB 2273b2f7 download
experts-42.safetensors 609 MB 86ed94e2 download
experts-43.safetensors 609 MB 18697fc5 download
experts-44.safetensors 609 MB 956735bd download
experts-45.safetensors 609 MB 0f6907b7 download
experts-46.safetensors 609 MB cee79300 download
experts-47.safetensors 609 MB 15a6eb8d download
experts-0.safetensors 609 MB 02bb42d4 download
experts-1.safetensors 609 MB f16d0015 download
experts-2.safetensors 609 MB afcb2600 download
experts-3.safetensors 609 MB 2f20ae69 download
experts-4.safetensors 609 MB 954f9d96 download
experts-5.safetensors 609 MB 0730986e download
experts-6.safetensors 609 MB 899042c4 download
experts-7.safetensors 609 MB ac9f53a2 download
experts-8.safetensors 609 MB 4f46cdec download
experts-9.safetensors 609 MB 0707bf29 download
mtp-experts.safetensors 609 MB b1053487 download
mtp-trunk.safetensors 93.1 MB 157eac53 download
tokenizer.json 12.2 MB 0997f410 download
vocab.json 6.41 MB 0aa0ce06 download
merges.txt 3.20 MB a494e019 download
model.safetensors.index.json 274 KB 6e20395d download
tokenizer_config.json 17.5 KB 5de744b3 download
chat_template.jinja 8.74 KB c0c686f9 download
README.md 4.34 KB 8b490b43 download
config.json 4.20 KB 501352e0 download
LICENSE 3.16 KB 9557a896 download
.gitattributes 1.53 KB 52373fe2 download
preprocessor_config.json 390 B 2ea84a43 download
video_preprocessor_config.json 385 B 3ba673a5 download
generation_config.json 202 B 023756cf download

README current version from Hugging Face


license: other
license_name: qwen-community-license-1.0
license_link: LICENSE
base_model: orcarouter/Qwen3.8-Flash-Next-Uncensored
base_model_relation: quantized
pipeline_tag: image-text-to-text
tags:

  • sushi
  • mlx
  • exl3
  • 2bit
  • abliterated
  • qwen
  • moe

Qwen3.8-Flash-Next-Abliterated-Sushi-2bpw

Experimental Sushi-format quantization made directly from the BF16 weights of
orcarouter/Qwen3.8-Flash-Next-Uncensored,
revision e096800036ec20da7e2442dcd4044a004d4e99fa.
The abliteration comes from that upstream checkpoint; this conversion adds no
further abliteration or fine-tuning. Refusal behavior has not been independently
characterized here.

Format and conversion

  • Routed experts: uniform EXL3 K2, MCG codebook, window15, including MTP experts.
  • Dense trunk: affine8 where the reference Sushi pack uses it; remaining floating
    tensors retain their reference dtype. Norm folding/layout changes follow Sushi.
  • N-gram sidecar: affine4, group32, stored in ngram_table.bin.
  • Vision tower is retained. MTP weights are retained; MTP inference is not validated.
  • Calibration:64 rows x1024 tokens. Converter based on ExLlamaV3
    151539c77abc7ab7425d30da7a4e8e3c5c154e7b, with an isolated MCG-window15 encoder
    adaptation and parallel dispatch for quantized experts alongside BF16 linears.

This is a Sushi pack, not a general ExLlamaV3 or Transformers checkpoint. 2bpw
refers to the routed-expert code rate, not every tensor or the total file-size
average. No lossless/BF16-equivalent quality claim is made. The private Sashimi
quantizer and its calibration recipe are not reproduced.

Running and validation

The complete download is about 64.87 GiB (including the 29.8 GiB disk-backed
n-gram sidecar). Keep every shard, tokenizer file and ngram_table.bin together.

Download with the Hugging Face CLI:

hf download sanasol2008/Qwen3.8-Flash-Next-Abliterated-Sushi-2bpw \
  --local-dir ~/.sushi/models/Qwen3.8-Flash-Next-Abliterated-Sushi-2bpw

Tested on Apple M3 Max / 48 GiB unified memory with Sushi 1.1.1, MLX 0.32.3,
and a local vision-streaming patch on Sushi commit
4d32cb0bb16802df1038d3fbc518f3d36d57c92e.
Vision under expert streaming was tested with that patched engine; this does not
establish support in an unmodified Sushi release.
The retained vision weights
alone do not remove an engine's streaming restrictions.

The local serving configuration used:

sushi serve --model ~/.sushi/models/Qwen3.8-Flash-Next-Abliterated-Sushi-2bpw \
  --host 127.0.0.1 --port 12345 --ssd-budget-gb 22 --no-mtp \
  --kv-quant 8 --ctx-size 262144 --prefill-chunk 2048 --metrics \
  --max-tokens 8192 --prefix-cache-entries 2 --prefix-cache-mem 1GB \
  --prefix-cache-disk 10GB --temp 1

The 256k context is a configured ceiling, not a measured full-context memory or
quality guarantee. Memory usage grows with context and image inputs. MTP is
disabled in this streaming setup; MTP and video inference were not validated.

Controlled text comparison used sequential fresh servers, 22 GiB streaming,
KV8, 16k context / prefill chunk512, prefix-cache entries0, temperature0,
thinking off, the same 40-token prompt, and 256 generated tokens. No downloads
or checksum jobs ran during measurement.

Pack Cold decode Warm decode, mean of runs2–3
Original Sushi 2bpw 22.89 tok/s 31.60 tok/s
This BF16-derived pack 16.78 tok/s 26.48 tok/s

These are short-prompt expert-cache measurements, not representative long-context
agent throughput. On the final serving settings, arithmetic and Russian text
checks passed; an image OCR check correctly transcribed three fields including
all numbers (915 prompt tokens, 9.91s total request time). A tool-call check
correctly returned get_weather with city=Belgrade.

These smoke checks establish basic functionality, not model-wide quality,
BF16 equivalence, or a measured refusal rate.

License and attribution

The upstream LICENSE is included verbatim: Qwen Community License1.0. The upstream
model card labels Apache2.0, but its actual bundled license differs; this repository
preserves the bundled license rather than asserting a new Apache license.
See LICENSE for the complete terms. Credit to Qwen for the base model, OrcaRouter
for the BF16 derivative, turboderp-org for ExLlamaV3, and beamivalice for Sushi.

Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration