← back to catalog · registered 2026-09-17 12:56

ddark-il/Qwen3.8-Flash-Next-Uncensored

ddark-il MoE multimodal second-order
Abliteration classifier · v1.0.0
M-U
Primary method

Uncensored (method unknown)

No other method signals detected in this model.
Confidence
LOW
Why this label 3 signals
Weak or ambiguous signals. Best guess based on catalog patterns; treat as tentative and check the evidence below.
  • 'uncensored' in name/tags but no 'abliterated' marker
  • method not identifiable from author declaration alone
  • may be DPO fine-tune, prompt engineering, or unknown technique
Refusal direction extraction

No specific extraction method could be identified for this model. The producer either did not document it or used a proprietary pipeline.

What is a refusal direction? →
Downloads · 30-day
2K
Likes
3
Model age
2w ago
created 2026-08-31

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
apache-2.0
Tags
mlx safetensors qwen4_exp omlx quantized mixed-precision affine moe mtp vision-language apple-silicon m5

Related

Total size
102 GB
Files
35
Quantizations
1
Registered
2026-09-17 12:56
Last updated on HF
2026-09-17 12:41

Files by quantization

Auxiliary files 35 files 103 GB
model-00008-of-00021.safetensors 4.95 GB 4816f3fa download
model-00013-of-00021.safetensors 4.83 GB c4673e8f download
model-00016-of-00021.safetensors 4.83 GB bf57deb8 download
model-00020-of-00021.safetensors 4.83 GB 2aa92d77 download
model-00014-of-00021.safetensors 4.83 GB 28d2fc0d download
model-00015-of-00021.safetensors 4.83 GB 9c606289 download
model-00017-of-00021.safetensors 4.83 GB 7038cd85 download
model-00012-of-00021.safetensors 4.83 GB 142d29cd download
model-00018-of-00021.safetensors 4.83 GB 136cad40 download
model-00011-of-00021.safetensors 4.83 GB ed884ea2 download
model-00019-of-00021.safetensors 4.83 GB bff6463b download
model-00009-of-00021.safetensors 4.83 GB d7c95a1a download
model-00010-of-00021.safetensors 4.83 GB 88eec400 download
model-00001-of-00021.safetensors 4.75 GB e7bc9c6b download
model-00007-of-00021.safetensors 4.68 GB 0bac88c9 download
model-00006-of-00021.safetensors 4.66 GB d4e082d0 download
model-00003-of-00021.safetensors 4.66 GB a3cd36bb download
model-00005-of-00021.safetensors 4.66 GB 6158c244 download
model-00004-of-00021.safetensors 4.66 GB 44f4257d download
model-00002-of-00021.safetensors 4.66 GB 71cdf499 download
model-00021-of-00021.safetensors 3.52 GB 9c43daee download
model-mtp.safetensors 2.59 GB 39171918 download
model-mtp.safetensors.bak-mlp4bit 1.42 GB dcfefb57 download
tokenizer.json 12.2 MB 0997f410 download
vocab.json 6.41 MB 0aa0ce06 download
merges.txt 3.20 MB a494e019 download
model.safetensors.index.json 325 KB b04cf713 download
config.json 163 KB 0e85ad59 download
tokenizer_config.json 17.5 KB 5de744b3 download
README.md 13.2 KB 96a2064b download
.DS_Store 10.0 KB f47b8038 download
chat_template.jinja 8.74 KB c0c686f9 download
.gitattributes 1.60 KB 9eba0816 download
preprocessor_config.json 390 B 2ea84a43 download
generation_config.json 202 B 023756cf download

README current version from Hugging Face


license: apache-2.0
base_model: orcarouter/Qwen3.8-Flash-Next-Uncensored
base_model_relation: quantized
library_name: mlx
pipeline_tag: image-text-to-text
tags:

  • mlx
  • omlx
  • quantized
  • mixed-precision
  • affine
  • qwen4_exp
  • moe
  • mtp
  • vision-language
  • apple-silicon
  • m5
  • nax

Qwen3.8-Flash-Next-Uncensored — MLX mixed-precision 4.86 bpw (v2, range-search + 8-bit o_proj)

Mixed-precision affine quantization of
orcarouter/Qwen3.8-Flash-Next-Uncensored
(qwen4_exp, 180B) sized for an M5 Max 128 GB at up to 262k context.

109.29 GB · 4.857 bpw · 179.99B params · 3.29× vs bf16 (360 GB)

This is the second revision. It keeps the same bit map as the first release except for one
module, and changes how per-group quantization ranges are chosen. Both changes are
size-neutral: +47 MB (+0.04 %) over v1, and every tensor measured equal or better.

What changed vs. v1

  1. Per-group range search instead of min/max rounding. MLX's mx.quantize derives each
    group's affine range from the group's min/max, so one outlier stretches the grid and wastes
    levels for the other weights. Ranges are now chosen by a multi-step search over uniform
    grids, scored in bf16 — the format that is actually stored, which matters because MLX
    keeps scales/biases in bf16 and a grid that wins in float32 can lose once stored. In this
    build the search drives the two 4-bit tables (routed experts, n-gram PLE), which hold 91 % of
    the bytes; the 6- and 8-bit modules carry the runtime's own edge-factor clipping search.
  2. o_proj 6-bit → 8-bit gs64. +44 MB for a 3.1× smaller error on that module.

Nothing else changed: same modules, same group sizes, same bits everywhere else, and the MTP
head is byte-identical to v1 (so speculative acceptance is unaffected by construction).

Measured headroom not in this build

Both numbers below were produced by the same tooling described in Verification,
on this checkpoint's own source tensors. Neither is applied in this artifact, and the first one
costs nothing.

  • Same bits, same size, better (scale, bias) at ≥ 5 bits: −20 % error on the 8-bit modules.
    The runtime scores its own clipping candidates in float32 and then casts parameters to bf16, so
    at 256 levels per 64-weight group it optimises a quantity that is not stored. Scoring after the
    cast, o_proj 0.0073 → 0.0058, lm_head 0.0081 → 0.0062, embed_tokens 0.0083 → 0.0065, vision
    linear_fc1 0.0075 → 0.0062. Identical bytes, identical group sizes, identical kernels.
  • More bytes would buy a lot; the honest cost is cache, not RAM. Sweeping the whole
    NAX-eligible (bits × group_size) lattice: routed experts at 5-bit gs128 measure 0.0496 against
    0.0888 at the shipped 4-bit gs64 — a 31.6 % cut in element-weighted error across the two
    4-bit tables — for +11.3 GB of file, which lands as +11.3 GB resident (76.9 → 88.2 GB,
    since only ~76.9 GB of this 109.29 GB file is resident while the n-gram table pages through
    mmap). Taking it means evicting roughly that much n-gram page cache, so the price is paid in
    table cache hits and SSD reads on the table's hot subset, plus 1.33 → 1.55 GB of expert
    weights streamed per token. 5-bit on the table itself is a further −20 % on the same aggregate
    but grows the table to 38.4 GB (at gs32, metadata is a full 1.0 bit per weight, so 4→5 bit
    costs +6.4 GB). Both are measured as possible, not yet as free: they need a real session —
    tok/s plus SSD read volume — before they belong in a recipe. Worth knowing regardless of the
    verdict here: the shipped 4-bit gs64 and 8-bit gs64 points are exactly on the
    error-per-byte frontier, no lattice point is cheaper and not worse, so quality at this size is
    bought and never simply found.
  • The floating-point formats are not worth it. At the same nominal bits, mxfp4 measures
    −26 %, nvfp4 −22 % and mxfp8 −250 % against affine, because a power-of-two or fp8 shared
    scale with no per-group offset places levels worse than a free (scale, bias) pair does. They
    are all NAX-eligible; eligible is not the same as good.

Files

21 safetensors shards (model-00001…00021-of-00021) plus the MTP sidecar
(model-mtp.safetensors, 58 tensors), config.json, generation_config.json,
chat_template.jinja, tokenizer.json / tokenizer_config.json / vocab.json /
merges.txt, and preprocessor_config.json. The MTP sidecar is byte-identical to v1's.

Measured accuracy

All numbers below are reproduced by the tooling described in
Verification — they are measurements, not expectations. The reference is the
bf16 weights of the same source checkpoint for tensor-level error, and a hosted bf16
endpoint for output-level agreement (see the caveat in Limitations).

metric v1 v2 (this model)
Weighted mean rel-L2 vs bf16 source 0.01890 0.01768 (−6.5 %)
Tensors improved / worse 23 / 0
o_proj rel-L2 0.0242 0.0073 (−69.8 %)
Routed experts rel-L2 (4-bit) 0.0969 0.0913 (−5.8 %)
n-gram table rel-L2 (4-bit gs32) 0.0808 0.0733 (−9.3 %)
Top-1 agreement with bf16 oracle 81.5 % 90.7 %
Distribution distance (mean JSD) 0.0736 0.0575
p99 JSD 0.1984 0.1868

Exact-width tensors (norms, routers, hyper-connections, the sparse indexer, vision
linear_fc2) remain bit-identical to the source, and the MTP head is bit-identical to v1.

Quantization map

Base format: MLX affine — packed U32 weights + bf16 scales + bf16 biases, i.e.
2 × 16 / group_size bits of overhead per weight on top of the nominal width. 599 explicit
per-module entries plus the config default (4-bit gs64) applied to the routed experts; 743
quantized modules in total.

Module Bits Group Params Size Rationale
Routed experts MLP (48 × 512) 4 64 120.8B 67.9 GB dominant weight mass; 4-bit gs64 is the NAX tensor-op path
Shared experts MLP (48 × 3) 8 128 0.24B 0.24 GB fires on every token, tiny
GDN / linear attention (36 × 5 projections) 8 64 2.09B 2.22 GB recurrent state accumulates error over sequence length
Full-attn q/k/v proj 8 64 0.41B 0.44 GB attention-sensitive, small tensors
Full-attn o_proj 8 64 0.18B 0.19 GB residual-writing path; was 6-bit in v1
Sparse-attn indexer bf16 0.02B 0.04 GB decides which blocks sparse attention reads — exact
embed_tokens / lm_head 8 64 1.27B 1.35 GB output quality
N-gram PLE table (128 shards) 4 32 51.2B 32.0 GB lookup; gs32 is structural — rows are 160 wide, 64 does not divide 160
MTP head (experts / attn / FC / indexer) 8/8/6/6 64/64/128/64 2.6B 2.79 GB draft head: trunk verifies every token, so precision moves acceptance rate only
Vision tower 8 128 0.31B 0.32 GB inline in the main shards
Vision linear_fc2, patch_embed bf16 0.14B 0.27 GB input dim 4304 admits no supported group size
Routers, hyper-connections, norms, convs, SSM state bf16 ~0.7B 1.3 GB small; hyper-connections measured ~30 % slower decode at 8-bit with MTP on

Byte accounting as measured from the checkpoint: 107.46 GB quantized + 1.83 GB exact bf16.

NAX (M5 tensor unit) eligibility

100 % of quantized bytes are eligible. Verified against the kernel grid that ships inside
MLX rather than from documentation: affine NAX kernels exist for
group_size ∈ {32, 64, 128} × bits ∈ {2, 3, 4, 5, 6, 8}, in both dense (affine_qmm_{n,t}_nax)
and MoE gather (affine_gather_qmm_rhs_nax) flavours. Every format used here is in that grid,
including o_proj at gs64/8-bit.

Two caveats: kernel existence is eligibility, not engagement (the dispatch heuristic lives
in compiled host code), and worth knowing for planning — at compute-bound (prefill) sizes the
quantized NAX kernels are ~1.5× slower than bf16 NAX, because they run smaller tiles
(bm64/bn64/bk64 vs bm128/bn128/bk512). Quantization pays at decode, where streamed bytes
dominate, which is where this model spends its time.

How it was made

Streaming, tensor-by-tensor quantization with ~2 GB host RAM for a 360 GB source. The bit map
is read from the previous release's own config.json rather than from a rule list, so the
map is guaranteed identical apart from the explicit o_proj override; the source-to-artifact
tensor mapping was pre-flighted against all 1658 source tensors (0 unmatched, 0 mismatches),
and 2134 predicate decisions were logged with 0 fallbacks.

  • MTP head: grafted byte-identically from v1 (copy model-mtp.safetensors, add its 58 keys to
    the index, restore the 13 mtp.* quantization entries and the mtp config fields). The stock
    library helper does not accept a qwen4_exp recipient, so the graft is explicit.
  • The range search is safe by construction: each group's candidates are scored against
    mx.dequantize output — the runtime's own dequantizer on the stored bf16 parameters — and
    the built-in min/max result is always among the candidates, so no group can come out worse.
  • Tokenizer, chat template, preprocessor and generation configs are the source's own files.
  • No training, fine-tuning, distillation or weight editing: this is quantization only.

Requirements and serving

Needs oMLX ≥ 0.6.3rc3 (model_type: qwen4_exp, VL engine, pre-quantization sanitize).
Two settings are engine-level model settings, not request parameters: mtp_enabled
(Lightning MTP decode — the "enable_mtp" request field is ignored) and PLE mode.

  • Set PLE to mode="mmap". With the n-gram table resident (~107.6 GB hot) macOS killed the
    server under memory pressure on a 128 GB desktop; mmap loads in ~11 s and resides at ~76.9 GB.
    Because the table is a raw mmap with no row cache, keep RAM pressure low so its hot pages stay
    in page cache (oMLX hot_cache_max_size lowered to 6 GB for this reason).
  • Raise the Metal wired limit: sudo sysctl iogpu.wired_limit_mb=120000.
  • Capping served context at 128k halves KV (6.4 → 3.2 GB); only 12 of 48 layers cache KV
    (2-head GQA, 24 KB/token), and GDN state is context-independent.
  • Speculative-prefill token pruning does not apply: GDN state needs every token and pruning
    breaks n-gram hash adjacency.

Decode throughput was measured on v1's recipe, which this model matches within 0.04 % of bytes:
~30.7 tok/s aggregate on M5 Max 128 GB / oMLX 0.6.4 (25,505 tokens over 30 agentic
tool-calling turns, 12.5k → 79k context, 0 errors), 91.7 % prompt-cache hits. It has not been
re-measured on v2; the only weight change is o_proj at a higher precision, so expect parity
rather than an improvement.

Verification

Everything above is reproducible with the tooling in
qwen3.8-next-flash (skills/omlx-quant-accuracy/):

rung tool result
L0 structure audit_mlx_quant.py 743 quantized + 761 exact modules, complete weight/scales/biases triplets, no indivisible group sizes, index ↔ disk parity
L1 weights stream_compare.py weighted mean rel-L2 0.01768 vs the bf16 source, 23 improved / 0 worse
L2 outputs capture_logits.py, compare_captures.py, oracle.py full-vocab logprobs captured from the runtime; A/B and oracle agreement as tabled above
NAX nax_compat.py 100 % of quantized bytes eligible for the M5 tensor unit

Limitations

  • The output-level reference is not this checkpoint's own bf16 weights. Agreement/JSD were
    measured against a hosted bf16 endpoint for Qwen's official Qwen3.8-Flash-Next, while this
    model derives from orcarouter/…-Uncensored. Those figures therefore include whatever the
    Uncensored variant changed, not quantization error alone. The relative comparison between
    v1 and v2 is sound because both were measured the same way against the same reference; the
    absolute scores are not a quantization-error measurement, and the distribution gate should
    not be read as a pass/fail on precision.
  • The evaluation corpus is small (21 cases) and three of its checks are known to be
    uninformative — two expectations are provably wrong and one cannot be satisfied by any
    continuation. Treat the pass count (15/20, unchanged from v1) as indicative only.
  • Long-context behaviour (16k/32k/128k needle) and MTP acceptance rate were not re-measured on
    v2. MTP is byte-identical to v1, so acceptance should be unchanged.
  • A per-module range search optimizes per-group weight error, not end-to-end task accuracy.
    The oracle-agreement improvement is evidence that it helped on this corpus, not a guarantee
    on every workload.
  • This is an uncensored derivative of the upstream model, inherited from the base
    checkpoint; see the base model card for its behaviour and safety considerations.

Attribution

Quantization recipe and measurement tooling by the qwen3.8-next-flash project. Base weights:
orcarouter/Qwen3.8-Flash-Next-Uncensored. Architecture and the original model:
Qwen. Quantizer/runtime: oMLX and MLX. Licensed Apache-2.0, following the
base model.

Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in app" button that hands off directly to a local runtime of your choice - Infrahuman, LM Studio, or Ollama. No API keys, no subscription, no prompt leakage.