license: other
license_name: qwen-community-license-1.0
license_link: LICENSE
base_model: orcarouter/Qwen3.8-Flash-Next-Uncensored
base_model_relation: quantized
pipeline_tag: image-text-to-text
tags:
- sushi
- mlx
- exl3
- 2bit
- abliterated
- qwen
- moe
Qwen3.8-Flash-Next-Abliterated-Sushi-2bpw
Experimental Sushi-format quantization made directly from the BF16 weights of
orcarouter/Qwen3.8-Flash-Next-Uncensored,
revision e096800036ec20da7e2442dcd4044a004d4e99fa.
The abliteration comes from that upstream checkpoint; this conversion adds no
further abliteration or fine-tuning. Refusal behavior has not been independently
characterized here.
Format and conversion
- Routed experts: uniform EXL3 K2, MCG codebook, window15, including MTP experts.
- Dense trunk: affine8 where the reference Sushi pack uses it; remaining floating
tensors retain their reference dtype. Norm folding/layout changes follow Sushi. - N-gram sidecar: affine4, group32, stored in
ngram_table.bin. - Vision tower is retained. MTP weights are retained; MTP inference is not validated.
- Calibration:64 rows x1024 tokens. Converter based on ExLlamaV3
151539c77abc7ab7425d30da7a4e8e3c5c154e7b, with an isolated MCG-window15 encoder
adaptation and parallel dispatch for quantized experts alongside BF16 linears.
This is a Sushi pack, not a general ExLlamaV3 or Transformers checkpoint. 2bpw
refers to the routed-expert code rate, not every tensor or the total file-size
average. No lossless/BF16-equivalent quality claim is made. The private Sashimi
quantizer and its calibration recipe are not reproduced.
Running and validation
The complete download is about 64.87 GiB (including the 29.8 GiB disk-backed
n-gram sidecar). Keep every shard, tokenizer file and ngram_table.bin together.
Download with the Hugging Face CLI:
hf download sanasol2008/Qwen3.8-Flash-Next-Abliterated-Sushi-2bpw \
--local-dir ~/.sushi/models/Qwen3.8-Flash-Next-Abliterated-Sushi-2bpw
Tested on Apple M3 Max / 48 GiB unified memory with Sushi 1.1.1, MLX 0.32.3,
and a local vision-streaming patch on Sushi commit4d32cb0bb16802df1038d3fbc518f3d36d57c92e.
Vision under expert streaming was tested with that patched engine; this does not
establish support in an unmodified Sushi release. The retained vision weights
alone do not remove an engine's streaming restrictions.
The local serving configuration used:
sushi serve --model ~/.sushi/models/Qwen3.8-Flash-Next-Abliterated-Sushi-2bpw \
--host 127.0.0.1 --port 12345 --ssd-budget-gb 22 --no-mtp \
--kv-quant 8 --ctx-size 262144 --prefill-chunk 2048 --metrics \
--max-tokens 8192 --prefix-cache-entries 2 --prefix-cache-mem 1GB \
--prefix-cache-disk 10GB --temp 1
The 256k context is a configured ceiling, not a measured full-context memory or
quality guarantee. Memory usage grows with context and image inputs. MTP is
disabled in this streaming setup; MTP and video inference were not validated.
Controlled text comparison used sequential fresh servers, 22 GiB streaming,
KV8, 16k context / prefill chunk512, prefix-cache entries0, temperature0,
thinking off, the same 40-token prompt, and 256 generated tokens. No downloads
or checksum jobs ran during measurement.
| Pack | Cold decode | Warm decode, mean of runs2–3 |
|---|---|---|
| Original Sushi 2bpw | 22.89 tok/s | 31.60 tok/s |
| This BF16-derived pack | 16.78 tok/s | 26.48 tok/s |
These are short-prompt expert-cache measurements, not representative long-context
agent throughput. On the final serving settings, arithmetic and Russian text
checks passed; an image OCR check correctly transcribed three fields including
all numbers (915 prompt tokens, 9.91s total request time). A tool-call check
correctly returned get_weather with city=Belgrade.
These smoke checks establish basic functionality, not model-wide quality,
BF16 equivalence, or a measured refusal rate.
License and attribution
The upstream LICENSE is included verbatim: Qwen Community License1.0. The upstream
model card labels Apache2.0, but its actual bundled license differs; this repository
preserves the bundled license rather than asserting a new Apache license.
See LICENSE for the complete terms. Credit to Qwen for the base model, OrcaRouter
for the BF16 derivative, turboderp-org for ExLlamaV3, and beamivalice for Sushi.