← back to catalog · registered 2026-10-05 05:58

yomie4343/Qwen3.8-Flash-Next-Uncensored-MLX-Sushi-mixed-3-4-8bit-BF16-SSDNgram

yomie4343 second-order
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/yomie4343%2FQwen3.8-Flash-Next-Uncensored-MLX-Sushi-mixed-3-4-8bit-BF16-SSDNgram"
Response includes
  • classification m-uncensored
  • files 14
  • author_summary 3 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M-U
Primary method

Uncensored (method unknown)

No other method signals detected in this model.
Confidence
LOW
Why this label 3 signals
Weak or ambiguous signals. Best guess based on catalog patterns; treat as tentative and check the evidence below.
  • 'uncensored' in name/tags but no 'abliterated' marker
  • method not identifiable from author declaration alone
  • may be DPO fine-tune, prompt engineering, or unknown technique
Refusal direction extraction

No specific extraction method could be identified for this model. The producer either did not document it or used a proprietary pipeline.

What is a refusal direction? →
Downloads · 30-day
0
Likes
1
Model age
today
created 2026-10-05

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
other
Tags
mlx safetensors qwen4_exp sushi qwen affine mixed-precision text-generation conversational base_model:orcarouter/Qwen3.8-Flash-Next-Uncensored base_model:quantized:orcarouter/Qwen3.8-Flash-Next-Uncensored license:other

Related

Total size
0 B
Files
14
Quantizations
1
Registered
2026-10-05 05:58
Last updated on HF
2026-10-05 06:35

Files by quantization

Auxiliary files 14 files 22.2 MB
tokenizer.json 12.2 MB 0997f410 download
vocab.json 6.41 MB 0aa0ce06 download
merges.txt 3.20 MB a494e019 download
model.safetensors.index.json 298 KB 7af446d4 download
MODEL-MANIFEST.json 40.6 KB 73611053 download
config.json 19.8 KB 4fe3ce94 download
tokenizer_config.json 17.5 KB 5de744b3 download
README.md 9.24 KB be1b6487 download
chat_template.jinja 8.74 KB c0c686f9 download
LICENSE 3.16 KB 9557a896 download
NOTICE 2.20 KB 7ad49d86 download
MANIFEST.json 1.91 KB 32c38ed0 download
.gitattributes 1.53 KB 52373fe2 download
generation_config.json 202 B 023756cf download

README current version from Hugging Face


license: other
license_name: qwen-community-1.0
license_link: LICENSE
base_model: orcarouter/Qwen3.8-Flash-Next-Uncensored
base_model_relation: quantized
pipeline_tag: text-generation
tags:

  • mlx
  • sushi
  • qwen
  • affine
  • mixed-precision

Qwen3.8-Flash-Next-Uncensored — ORIGINAL mixed 3/4/8 affine, BF16 SSD Ngram

UPLOAD IN PROGRESS — INCOMPLETE MODEL. Only metadata and bounded pilot shards are being uploaded. The required BF16 Ngram and most weight shards are absent. Do not attempt to load or run this repository. Completion and full-model reproducibility have not been verified.

Approved new public destination: yomie4343/Qwen3.8-Flash-Next-Uncensored-MLX-Sushi-mixed-3-4-8bit-BF16-SSDNgram. Existing public 3/8 and 4/8 repositories are separate artifacts and are not replaced.

This is the existing ORIGINAL mixed 3/4/8 affine group-size-64 campaign candidate, prior to the separate weighted-affine allocation experiments. “ORIGINAL” identifies this campaign artifact; it does not mean the original full-precision Qwen weights. There is no new training in this publication preparation. The direct raw source is the refusal-modified orcarouter checkpoint, revision 8336e613ea508b13c2159bd0f68965d97a606b95. Its card identifies Qwen/Qwen3.8-Flash-Next as its base. The exact upstream training ancestor commit has not been independently established.

Pinned conversion lineage

The inherited MLX 4/8 intermediate is yomie4343/Qwen3.8-Flash-Next-Uncensored-MLX-Serve-mixed-4-8bit at 94af1f9fad6e422b8dda2fd9efe187208e388ceb. All 57 unchanged current shards match that fixed revision's remote LFS SHA-256 values. The other 44 current generated shards match the saved generation receipts, and all44 recorded raw input file hashes match the direct BF16 source at 8336e613ea508b13c2159bd0f68965d97a606b95. This verifies the pinned file links; the complete quantization procedure has not been independently rerun.

The required Ngram was assembled by copying 128 original BF16 tensor ranges from33 source containers, in numeric shard order. The saved build records verify source/written SHA-256 per range and full row coverage. Later Sushi adapters alter only the header and preserve the recorded payload hash. It is not dequantized from the older intermediate's 4-bit Ngram. These records plus the fresh current full-file hash establish the retained artifact; the remote128 ranges have not been independently downloaded and reconstructed again during publication preparation.

The unverified links are the exact official-Qwen training ancestor of the refusal-modified source, a new independent raw-to-intermediate/generated-shard conversion, and a new independent reconstruction of the original Ngram ranges. They are provenance proof limits, not evidence of a conflicting license. The intermediate's actual Qwen Community License1.0 is byte-identical to the direct source and separately inspected base LICENSE. NOTICE preserves this intermediate attribution as well.

Payload and format

Component Size and format
Weight containers 101 safetensors shards, 66,075,781,474 bytes; indexed tensor payload 66,075,324,408 bytes across 3167 tensors
Affine projection inventory Header-derived: 88 packed 3-bit, 60 packed 4-bit, 646 packed 8-bit projections; group size64
Required SSD Ngram sidecar ngram_table.bin, 102,400,491,728 bytes; safetensors-format raw BF16 [320001536,160] merged table
Model data plus serving metadata 168,499,510,943 bytes before this card, license, notice and manifests; about156.93GiB

MODEL-MANIFEST.json records fresh SHA-256 for every original shard, the entire BF16 Ngram file and serving metadata. All101 weight headers match the retained ORIGINAL inventory; 44 known generated-shard hashes also match. These hashes establish file identity, not model quality or complete training provenance. No weighted allocation variants are included.

The Ngram binary header says bits16 and BF16. It is not a 4-bit Ngram table. An old config field and generation receipt described a smaller 4-bit sidecar; the proposed config.json corrects only ngram_table.bits from4 to16. The header retains legacy group_size32 metadata, but the BF16 Sushi loader uses no quantization groups. No original weight or Ngram bytes are changed. The proposed config overlay has not yet been tested in a fresh clean-package server.

Engine compatibility

Local text serving was tested with isolated Sushi pinned at 711572e91c491b50b973e57045b8d4a7f5137764, MLX d73eb752ef2e6288fd95b032c0bff0a15a4a9e93 and MLX-C 56b2d39fc831f2c0eb5bb94d82ef7191f7b31fa6, with retained opt-in M3 Ultra patches documented in the separate engine source proposal. The personal engine fork destination and final published revision are still pending. No engine binary or runtime dependency is bundled with this model.

This is a Sushi native affine checkpoint plus a required SSD-read Ngram sidecar. Generic Transformers, GGUF, EXL3, arbitrary mlx-lm versions and other engines are not qualified. Upstream architecture names in config do not establish portable framework support. Vision has not been qualified here; text testing used --no-vision. Keep ngram_table.bin alongside the index/config and all101 shards on fast local SSD. Storage size is not an all-resident RAM requirement: the Ngram is SSD-backed.

The tested machine was M3 Ultra96GB with KV8 and context32768. This is an observed configuration, not a minimum-memory guarantee or qualification of other hardware/context sizes. Use an isolated server and the exact dependencies; retain limits and process isolation in the engine proposal's reproduction instructions.

Performance scope

The latest corrected exact-MTP controls produced256 completion tokens from an8130-token prompt at 76.9759 tokens/s (256 divided by mean complete post-prefill awake wall). Cold prefill was926.419 /926.174 tokens/s. Separate earlier ordinary-decode 32K controls reached972.709 /975.208 cold-prefill tokens/s. These are distinct experiments. Neither cached-prefix rates nor legacy tick-only MTP rates are advertised here. The1000-prefill /80-decode target is not robustly achieved. A corrected full-suite repeat, independent public-corpus quality study and fresh full-model/GPU qualification remain pending. A clean combined engine source build and 13 targeted native CPU tests have passed locally; this is not a new serving/performance repeat.

These serving measurements are not a general intelligence, refusal, safety or task-quality benchmark. Private prompts, calibration logs, outputs and request-state traces are excluded from publication. The source checkpoint describes refusal removal; it may produce inappropriate or inaccurate text and no new quality/safety qualification is claimed.

License and modifications

The actual fixed source LICENSE is Qwen Community License1.0, copyright2026 Qwen. It is reproduced unchanged in LICENSE. Its card's Apache-2.0 claim conflicts with that file. The independently inspected Qwen base license has identical bytes; that inspected revision is not claimed to be the proven training ancestor. This proposal therefore uses license: other with the actual Qwen license, not an Apache-only or unrestricted-commercial-use label.

The license contains conditions on notices, certain large commercial services, and commercial Model-as-a-Service/AI Work Assistant businesses. Read the full actual license before use or redistribution. No separate Qwen commercial license has been obtained or accepted by this preparation. The card mismatch is corrected by following the actual supplied license and retaining its notices; it is not treated as a permanent distribution prohibition by itself. Exact training ancestry and a complete independent raw-to-final conversion reconstruction remain disclosed provenance limits. The actual license expressly permits redistribution and derivative works with its retained notices and conditions. The requested individual weight publication is not itself a new commercial inference or assistant service. No conflicting model license was identified in the inspected fixed sources; this does not establish unrestricted commercial-use rights or third-party-rights clearance.

NOTICE names the original Qwen model, the orcarouter refusal-modified source, the fixed MLX intermediate, the mixed affine quantization/packaging changes and the BF16 Ngram requirement. Full private generation/calibration receipts are excluded. File hashing does not reverify the complete raw BF16-to-final conversion chain. Distribution must preserve the license and notices.

README history 1 version

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-10-05Initialize incomplete upload with model license, notices and pinned provenancef366a5a9.2 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration