← back to catalog · registered 2026-09-18 08:56

yomie4343/Qwen3.8-Flash-Next-Uncensored-MLX-Serve-mixed-3-8bit

yomie4343 multimodal second-order
Abliteration classifier · v1.0.0
M-U
Primary method

Uncensored (method unknown)

No other method signals detected in this model.
Confidence
LOW
Why this label 3 signals
Weak or ambiguous signals. Best guess based on catalog patterns; treat as tentative and check the evidence below.
  • 'uncensored' in name/tags but no 'abliterated' marker
  • method not identifiable from author declaration alone
  • may be DPO fine-tune, prompt engineering, or unknown technique
Refusal direction extraction

No specific extraction method could be identified for this model. The producer either did not document it or used a proprietary pipeline.

What is a refusal direction? →
Downloads · 30-day
0
Likes
0
Model age
today
created 2026-09-18

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
other
Tags
mlx safetensors qwen4_exp mlx-serve apple-silicon mixed-precision 3-bit 4-bit 8-bit uncensored mtp image-text-to-text

Related

Total size
85.6 GB
Files
113
Quantizations
1
Registered
2026-09-18 08:56
Last updated on HF
2026-09-18 07:32

Files by quantization

Auxiliary files 113 files 85.6 GB
ngram_table.bin 29.8 GB c8ab74bc download
model-vision.safetensors 856 MB bba0b8dd download
model-00031.safetensors 779 MB 34e19f87 download
model-00029.safetensors 759 MB 4a791bcf download
model-00099.safetensors 704 MB a6d50a78 download
model-00007.safetensors 700 MB 00770f1d download
model-00009.safetensors 700 MB 05f3692c download
model-00011.safetensors 700 MB 4c84d69a download
model-00013.safetensors 700 MB 847d7d74 download
model-00015.safetensors 700 MB d1014987 download
model-00017.safetensors 700 MB cf679902 download
model-00019.safetensors 700 MB 7956a83d download
model-00021.safetensors 700 MB f91eb477 download
model-00023.safetensors 700 MB 2c5b91b7 download
model-00025.safetensors 700 MB aebace04 download
model-00030.safetensors 700 MB 902d4099 download
model-00032.safetensors 700 MB ebef3278 download
model-00034.safetensors 700 MB c34cb850 download
model-00036.safetensors 700 MB 4f867bdd download
model-00038.safetensors 700 MB 5344a6fa download
model-00040.safetensors 700 MB 1bc4d9a9 download
model-00042.safetensors 700 MB 3d66f49e download
model-00044.safetensors 700 MB 4d625073 download
model-00046.safetensors 700 MB 8badbe59 download
model-00048.safetensors 700 MB 305cf8b0 download
model-00052.safetensors 700 MB 23b4e2d0 download
model-00054.safetensors 700 MB 52b66abf download
model-00056.safetensors 700 MB 17864caa download
model-00058.safetensors 700 MB e40038d4 download
model-00060.safetensors 700 MB 75354c10 download
model-00062.safetensors 700 MB 7c4d077d download
model-00064.safetensors 700 MB e1739d87 download
model-00066.safetensors 700 MB 275d4ea1 download
model-00068.safetensors 700 MB 236db864 download
model-00070.safetensors 700 MB f0eceb41 download
model-00074.safetensors 700 MB 004e265d download
model-00076.safetensors 700 MB a95c480c download
model-00078.safetensors 700 MB 18853339 download
model-00080.safetensors 700 MB 5e314ae1 download
model-00082.safetensors 700 MB e6d3e2c4 download
model-00084.safetensors 700 MB 32abef63 download
model-00086.safetensors 700 MB a6c0c2a7 download
model-00088.safetensors 700 MB 08c2d535 download
model-00002.safetensors 700 MB 49ac6d1b download
model-00004.safetensors 700 MB f6dca2f9 download
model-00027.safetensors 700 MB 20a7e0b8 download
model-00050.safetensors 700 MB 81515c3b download
model-00072.safetensors 700 MB 382a0a4f download
model-00090.safetensors 700 MB 54960b71 download
model-00092.safetensors 700 MB eb03b012 download
model-00094.safetensors 700 MB f21aea08 download
model-00096.safetensors 700 MB abc2ef96 download
model-00098.safetensors 700 MB af28fee4 download
model-00100.safetensors 644 MB 6518523a download
model-00089.safetensors 487 MB 19732e27 download
model-00055.safetensors 484 MB a34ed73d download
model-00010.safetensors 484 MB 99d638f3 download
model-00026.safetensors 484 MB 35f8c057 download
model-00045.safetensors 482 MB 4115aa1e download
model-00018.safetensors 481 MB 819592d2 download
model-00051.safetensors 481 MB fbecc7c0 download
model-00081.safetensors 481 MB 84087971 download
model-00071.safetensors 481 MB 69bc2d62 download
model-00063.safetensors 481 MB 2e26090c download
model-00037.safetensors 481 MB cbcb5cdb download
model-00095.safetensors 481 MB 67f937ea download
model-00077.safetensors 463 MB c61d3e32 download
model-00075.safetensors 445 MB ec9710d7 download
model-00041.safetensors 432 MB baa07d81 download
model-00022.safetensors 432 MB 8b7a7d5b download
model-00057.safetensors 432 MB 3340a59a download
model-00012.safetensors 432 MB 566f4231 download
model-00059.safetensors 430 MB acc91ca9 download
model-00047.safetensors 429 MB c4e888ed download
model-00039.safetensors 429 MB 0e4f6800 download
model-00083.safetensors 429 MB 201ead1a download
model-00014.safetensors 429 MB 573efc4f download
model-00020.safetensors 429 MB 0ae748d4 download
model-00033.safetensors 429 MB 3433a2cc download
model-00065.safetensors 429 MB b200e5d8 download
model-00067.safetensors 429 MB 06fc5a53 download
model-00085.safetensors 429 MB a129a082 download
model-00073.safetensors 429 MB 5dafbe16 download
model-00003.safetensors 429 MB 0a611918 download
model-00091.safetensors 429 MB 8829cda1 download
model-00097.safetensors 429 MB 81e4ca0c download
model-00005.safetensors 390 MB 6ea0b00b download
model-00093.safetensors 376 MB 053c335a download
model-00035.safetensors 371 MB a7d72430 download
model-00016.safetensors 370 MB 08849e2c download
model-00087.safetensors 370 MB 37bf936c download
model-00069.safetensors 370 MB a72fea4b download
model-00028.safetensors 370 MB 2b634017 download
model-00008.safetensors 370 MB d29145d1 download
model-00024.safetensors 370 MB 6479097a download
model-00043.safetensors 370 MB a1f8467e download
model-00053.safetensors 370 MB 711f439f download
model-00061.safetensors 370 MB 6f92e294 download
model-00079.safetensors 370 MB 08f321a8 download
model-00049.safetensors 370 MB 167028aa download
model-00006.safetensors 72.2 MB bde08726 download
model-00001.safetensors 72.2 MB e308f597 download
tokenizer.json 12.2 MB 0997f410 download
vocab.json 6.41 MB 0aa0ce06 download
merges.txt 3.20 MB a494e019 download
model.safetensors.index.json 298 KB c3290539 download
tokenizer_config.json 17.5 KB 5de744b3 download
chat_template.jinja 8.74 KB c0c686f9 download
README.md 5.76 KB 331fca12 download
config.json 4.19 KB 3325acfc download
LICENSE 3.16 KB 9557a896 download
.gitattributes 1.53 KB 52373fe2 download
generation_config.json 202 B 023756cf download

README current version from Hugging Face


license: other
license_name: qwen-community-1.0
license_link: LICENSE
library_name: mlx
pipeline_tag: image-text-to-text
base_model: orcarouter/Qwen3.8-Flash-Next-Uncensored
base_model_relation: quantized
tags:

  • mlx
  • mlx-serve
  • qwen4_exp
  • apple-silicon
  • mixed-precision
  • 3-bit
  • 4-bit
  • 8-bit
  • uncensored
  • mtp

Qwen3.8-Flash-Next-Uncensored — MLX-Serve mixed 3/8-bit

An experimental MLX-Serve mixed-precision pack derived from
orcarouter/Qwen3.8-Flash-Next-Uncensored.
Routed MoE expert weights use 3-bit affine quantization (group size 64); the
remaining eligible projections use 8-bit quantization. The token embeddings
and the separate ngram_table.bin remain 4-bit. The MTP head and vision files
are included.

This pack is for the patched mlx-serve runtime used in the validation. It is
not a generic Transformers, mlx-lm, or mlx-vlm checkpoint.

The matching V2 runtime patch and validation manifests are in
engine/README.md.

Which variant should I use?

Two MLX-Serve-specific mixed-precision variants are available:

Variant Positioning Routed MoE experts Other eligible projections Token embeddings N-gram table Approx. pack size
4/8-bit Recommended default / quality-oriented 4-bit 8-bit 4-bit 4-bit ~107 GB
3/8-bit Lower-memory experimental variant 3-bit 8-bit 4-bit 4-bit ~86 GB

Choose the 4/8-bit variant as the default starting point when memory
capacity is sufficient. It uses less aggressive quantization for the routed
MoE expert weights.

Choose the 3/8-bit variant when reducing model memory or storage
requirements is more important. Only the routed MoE expert weights are
reduced to 3-bit; the token embeddings and N-gram table remain 4-bit, and the
other eligible projections remain 8-bit.

These labels describe the intended trade-off between the two packs. They are
not a claim that the 4/8-bit variant is universally more accurate, or that
the 3/8-bit variant has universally equivalent quality or speed
. Measured
quality and performance are workload-, context-, runtime-, and
hardware-dependent.

Runtime compatibility

These are specialized checkpoints for the patched mlx-serve runtime used
by this repository. They are not generic Transformers, mlx-lm, or
mlx-vlm checkpoints
. Do not rely on Hugging Face's automatically
generated mlx-lm / mlx-vlm usage examples for these packs; follow this
repository's engine/README.md and runtime instructions instead.

Validation notes

Results reported for each variant should be interpreted only under the
conditions documented on that variant's model card. Scores obtained with
different task sets, runtime revisions, generation limits, or evaluation
procedures should not be compared directly.

The 3/8-bit variant has additionally been validated on a Mac Studio M3 Ultra
with 96 GB unified memory using its matching patched MLX-Serve V2 runtime. In
that workload, it reduced measured MLX peak memory relative to the tested
4/8-bit reference. This is a workload-specific observation, not a general
memory or quality guarantee.

Download

hf download yomie4343/Qwen3.8-Flash-Next-Uncensored-MLX-Serve-mixed-3-8bit \
  --local-dir ./qwen-flash-next-mixed-3-8bit

Requirements

Use an Apple Silicon Mac with the matching patched mlx-serve V2 runtime and
keep the complete pack on a fast local SSD. The model files alone are not a
drop-in installation of the runtime patch.

Runtime

Use the V2 engine with the 3-bit kernel and these environment variables:

export MLX_SERVE_Q3_PACK=8
export MLX_SERVE_Q3_BYTES=1
export MLX_SERVE_NGRAM_WARM=0
export MLX_SERVE_CACHE_LIMIT=2147483648
./zig-out/bin/mlx-serve \
  --model /absolute/path/to/Qwen3.8-Flash-Next-Uncensored-MLX-Serve-mixed-3-8bit \
  --serve --host 127.0.0.1 --port 11234 --ctx-size 204800 \
  --kv-quant 8 --prefill-chunk 1024 --no-pld

The model pack is about 86 GB on disk, including the approximately 32 GB
4-bit n-gram table. Keep the complete directory on a fast local SSD.

Validation on Mac Studio M3 Ultra, 96 GB, 2026-09-18 JST

The patched engine reached practical speed parity with the existing 4/8-bit
pack: across approximately 3K, 12K, and 47K input tokens, median prefill and
decode differences stayed below 0.4% in the repeated comparison. This is a
workload-bounded comparison, not a claim of universal statistical equivalence.

Actual 127,744–200,094-token inputs passed the fixed JSON and tool-retrieval
checks in all 6/6 runs for this pack. Its cumulative MLX peak in that phase
was 63.774 GB, about 14.7 GB below the 4/8-bit reference.

The original 25-task Hermes/TypeScript comparison was independently graded at
72/75 (three repeated runs of the same 25 tasks). The three failures were the
same binary-search-tree implementation type error (missing return values).
A separate fixed-8192-output supplement for the two tasks affected by Hermes'
adaptive output cap passed 6/6 across this pack; that supplement is reported
separately and is not pooled with the 72/75 score.

Limitations

This is an experimental quantization and runtime combination. The validation
sets are fixed and do not guarantee general coding quality, reasoning quality,
or performance on other hardware. The source model is uncensored and can
produce unsafe, incorrect, or offensive output. Use least-privilege tools and
application-level safeguards.

The Qwen Community License 1.0 is retained; see LICENSE and review
its terms before redistribution or commercial use.

Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in app" button that hands off directly to a local runtime of your choice - Infrahuman, LM Studio, or Ollama. No API keys, no subscription, no prompt leakage.