license: other
license_name: qwen-community-1.0
license_link: LICENSE
base_model: orcarouter/Qwen3.8-Flash-Next-Uncensored
base_model_relation: quantized
pipeline_tag: text-generation
tags:
- mlx
- sushi
- qwen
- affine
- mixed-precision
Qwen3.8-Flash-Next-Uncensored — ORIGINAL mixed 3/4/8 affine, BF16 SSD Ngram
UPLOAD IN PROGRESS — INCOMPLETE MODEL. Only metadata and bounded pilot shards are being uploaded. The required BF16 Ngram and most weight shards are absent. Do not attempt to load or run this repository. Completion and full-model reproducibility have not been verified.
Approved new public destination: yomie4343/Qwen3.8-Flash-Next-Uncensored-MLX-Sushi-mixed-3-4-8bit-BF16-SSDNgram. Existing public 3/8 and 4/8 repositories are separate artifacts and are not replaced.
This is the existing ORIGINAL mixed 3/4/8 affine group-size-64 campaign candidate, prior to the separate weighted-affine allocation experiments. “ORIGINAL” identifies this campaign artifact; it does not mean the original full-precision Qwen weights. There is no new training in this publication preparation. The direct raw source is the refusal-modified orcarouter checkpoint, revision 8336e613ea508b13c2159bd0f68965d97a606b95. Its card identifies Qwen/Qwen3.8-Flash-Next as its base. The exact upstream training ancestor commit has not been independently established.
Pinned conversion lineage
The inherited MLX 4/8 intermediate is yomie4343/Qwen3.8-Flash-Next-Uncensored-MLX-Serve-mixed-4-8bit at 94af1f9fad6e422b8dda2fd9efe187208e388ceb. All 57 unchanged current shards match that fixed revision's remote LFS SHA-256 values. The other 44 current generated shards match the saved generation receipts, and all44 recorded raw input file hashes match the direct BF16 source at 8336e613ea508b13c2159bd0f68965d97a606b95. This verifies the pinned file links; the complete quantization procedure has not been independently rerun.
The required Ngram was assembled by copying 128 original BF16 tensor ranges from33 source containers, in numeric shard order. The saved build records verify source/written SHA-256 per range and full row coverage. Later Sushi adapters alter only the header and preserve the recorded payload hash. It is not dequantized from the older intermediate's 4-bit Ngram. These records plus the fresh current full-file hash establish the retained artifact; the remote128 ranges have not been independently downloaded and reconstructed again during publication preparation.
The unverified links are the exact official-Qwen training ancestor of the refusal-modified source, a new independent raw-to-intermediate/generated-shard conversion, and a new independent reconstruction of the original Ngram ranges. They are provenance proof limits, not evidence of a conflicting license. The intermediate's actual Qwen Community License1.0 is byte-identical to the direct source and separately inspected base LICENSE. NOTICE preserves this intermediate attribution as well.
Payload and format
| Component | Size and format |
|---|---|
| Weight containers | 101 safetensors shards, 66,075,781,474 bytes; indexed tensor payload 66,075,324,408 bytes across 3167 tensors |
| Affine projection inventory | Header-derived: 88 packed 3-bit, 60 packed 4-bit, 646 packed 8-bit projections; group size64 |
| Required SSD Ngram sidecar | ngram_table.bin, 102,400,491,728 bytes; safetensors-format raw BF16 [320001536,160] merged table |
| Model data plus serving metadata | 168,499,510,943 bytes before this card, license, notice and manifests; about156.93GiB |
MODEL-MANIFEST.json records fresh SHA-256 for every original shard, the entire BF16 Ngram file and serving metadata. All101 weight headers match the retained ORIGINAL inventory; 44 known generated-shard hashes also match. These hashes establish file identity, not model quality or complete training provenance. No weighted allocation variants are included.
The Ngram binary header says bits16 and BF16. It is not a 4-bit Ngram table. An old config field and generation receipt described a smaller 4-bit sidecar; the proposed config.json corrects only ngram_table.bits from4 to16. The header retains legacy group_size32 metadata, but the BF16 Sushi loader uses no quantization groups. No original weight or Ngram bytes are changed. The proposed config overlay has not yet been tested in a fresh clean-package server.
Engine compatibility
Local text serving was tested with isolated Sushi pinned at 711572e91c491b50b973e57045b8d4a7f5137764, MLX d73eb752ef2e6288fd95b032c0bff0a15a4a9e93 and MLX-C 56b2d39fc831f2c0eb5bb94d82ef7191f7b31fa6, with retained opt-in M3 Ultra patches documented in the separate engine source proposal. The personal engine fork destination and final published revision are still pending. No engine binary or runtime dependency is bundled with this model.
This is a Sushi native affine checkpoint plus a required SSD-read Ngram sidecar. Generic Transformers, GGUF, EXL3, arbitrary mlx-lm versions and other engines are not qualified. Upstream architecture names in config do not establish portable framework support. Vision has not been qualified here; text testing used --no-vision. Keep ngram_table.bin alongside the index/config and all101 shards on fast local SSD. Storage size is not an all-resident RAM requirement: the Ngram is SSD-backed.
The tested machine was M3 Ultra96GB with KV8 and context32768. This is an observed configuration, not a minimum-memory guarantee or qualification of other hardware/context sizes. Use an isolated server and the exact dependencies; retain limits and process isolation in the engine proposal's reproduction instructions.
Performance scope
The latest corrected exact-MTP controls produced256 completion tokens from an8130-token prompt at 76.9759 tokens/s (256 divided by mean complete post-prefill awake wall). Cold prefill was926.419 /926.174 tokens/s. Separate earlier ordinary-decode 32K controls reached972.709 /975.208 cold-prefill tokens/s. These are distinct experiments. Neither cached-prefix rates nor legacy tick-only MTP rates are advertised here. The1000-prefill /80-decode target is not robustly achieved. A corrected full-suite repeat, independent public-corpus quality study and fresh full-model/GPU qualification remain pending. A clean combined engine source build and 13 targeted native CPU tests have passed locally; this is not a new serving/performance repeat.
These serving measurements are not a general intelligence, refusal, safety or task-quality benchmark. Private prompts, calibration logs, outputs and request-state traces are excluded from publication. The source checkpoint describes refusal removal; it may produce inappropriate or inaccurate text and no new quality/safety qualification is claimed.
License and modifications
The actual fixed source LICENSE is Qwen Community License1.0, copyright2026 Qwen. It is reproduced unchanged in LICENSE. Its card's Apache-2.0 claim conflicts with that file. The independently inspected Qwen base license has identical bytes; that inspected revision is not claimed to be the proven training ancestor. This proposal therefore uses license: other with the actual Qwen license, not an Apache-only or unrestricted-commercial-use label.
The license contains conditions on notices, certain large commercial services, and commercial Model-as-a-Service/AI Work Assistant businesses. Read the full actual license before use or redistribution. No separate Qwen commercial license has been obtained or accepted by this preparation. The card mismatch is corrected by following the actual supplied license and retaining its notices; it is not treated as a permanent distribution prohibition by itself. Exact training ancestry and a complete independent raw-to-final conversion reconstruction remain disclosed provenance limits. The actual license expressly permits redistribution and derivative works with its retained notices and conditions. The requested individual weight publication is not itself a new commercial inference or assistant service. No conflicting model license was identified in the inspected fixed sources; this does not establish unrestricted commercial-use rights or third-party-rights clearance.
NOTICE names the original Qwen model, the orcarouter refusal-modified source, the fixed MLX intermediate, the mixed affine quantization/packaging changes and the BF16 Ngram requirement. Full private generation/calibration receipts are excluded. File hashing does not reverify the complete raw BF16-to-final conversion chain. Distribution must preserve the license and notices.