← back to catalog · registered 2026-09-18 11:56

NovaeonStudio/Qwen3.8-Flash-Next-Uncensored-oQ3e-fp16-mtp

NovaeonStudio MoE multimodal second-order
curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/NovaeonStudio%2FQwen3.8-Flash-Next-Uncensored-oQ3e-fp16-mtp"
Response includes
  • classification m-uncensored
  • files 27
  • author_summary 1 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M-U
Primary method

Uncensored (method unknown)

No other method signals detected in this model.
Confidence
LOW
Why this label 3 signals
Weak or ambiguous signals. Best guess based on catalog patterns; treat as tentative and check the evidence below.
  • 'uncensored' in name/tags but no 'abliterated' marker
  • method not identifiable from author declaration alone
  • may be DPO fine-tune, prompt engineering, or unknown technique
Refusal direction extraction

No specific extraction method could be identified for this model. The producer either did not document it or used a proprietary pipeline.

What is a refusal direction? →
Downloads · 30-day
0
Likes
0
Model age
today
created 2026-09-18

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
apache-2.0
Languages
en
Tags
mlx safetensors qwen4_exp apple-silicon oMLX uncensored vision mtp moe flash-next novaeon image-text-to-text

Related

Total size
79.3 GB
Files
27
Quantizations
1
Registered
2026-09-18 11:56
Last updated on HF
2026-09-18 11:42

Files by quantization

Auxiliary files 27 files 79.3 GB
model-00012-of-00017.safetensors 4.79 GB 4491cd0b download
model-00016-of-00017.safetensors 4.79 GB 3cfa337c download
model-00014-of-00017.safetensors 4.79 GB 9f416827 download
model-00013-of-00017.safetensors 4.79 GB d90e5091 download
model-00010-of-00017.safetensors 4.79 GB a216a7da download
model-00011-of-00017.safetensors 4.79 GB f6374647 download
model-00015-of-00017.safetensors 4.79 GB 8546760c download
model-00009-of-00017.safetensors 4.79 GB 57594efb download
model-00008-of-00017.safetensors 4.79 GB 2188b011 download
model-00001-of-00017.safetensors 4.76 GB 5f7131a5 download
model-00006-of-00017.safetensors 4.75 GB 06a68c94 download
model-00007-of-00017.safetensors 4.74 GB 78fcff06 download
model-00005-of-00017.safetensors 4.66 GB 9f573c46 download
model-00004-of-00017.safetensors 4.66 GB 63299ad6 download
model-00002-of-00017.safetensors 4.66 GB 8e0da509 download
model-00003-of-00017.safetensors 4.66 GB 9d48c147 download
model-00017-of-00017.safetensors 3.37 GB b4e78788 download
tokenizer.json 12.2 MB 0997f410 download
vocab.json 6.41 MB 0aa0ce06 download
merges.txt 3.20 MB a494e019 download
model.safetensors.index.json 396 KB 546eb70f download
config.json 134 KB 879763ad download
tokenizer_config.json 17.5 KB 5de744b3 download
README.md 7.49 KB c248fb6e download
.gitattributes 1.53 KB 52373fe2 download
preprocessor_config.json 390 B 2ea84a43 download
generation_config.json 202 B 023756cf download

README current version from Hugging Face


license: apache-2.0
base_model: orcarouter/Qwen3.8-Flash-Next-Uncensored
base_model_relation: quantized
pipeline_tag: image-text-to-text
library_name: mlx
tags:

  • mlx
  • apple-silicon
  • oMLX
  • uncensored
  • vision
  • mtp
  • moe
  • qwen4_exp
  • flash-next
  • novaeon
    language:
  • en

Qwen3.8-Flash-Next-Uncensored · oQ3e-fp16-mtp — a Novaeon.Studio build

A 99 B-A5B uncensored, vision-capable MoE that actually runs — and runs fast — on a single 128 GB Mac. This is an oMLX oQ3e quantization (error-compensated 3-bit base, sensitivity-guided 5/6/8-bit and fp16 for the tensors that matter) of orcarouter/Qwen3.8-Flash-Next-Uncensored, with the native MTP head preserved and working. Built, tuned, and benchmarked on an Apple M5 Max (128 GB) by Novaeon.Studio.

Why this build exists: Flash-Next is a 512-expert MoE with a PLE/n-gram embedding stack that most quantizers either refuse or destroy. At bf16 it is ~335 GB; the canonical oQ4e is ~106 GB and leaves almost no headroom on a 128 GB box. oQ3e lands at 85 GB on disk / ~79.5 GB in oMLX, which is the difference between "technically loads" and "this is my daily driver." Everything under Measured is from our hardware.

Base model orcarouter/Qwen3.8-Flash-Next-Uncensored
Architecture qwen4_exp (Qwen4ExpForConditionalGeneration) — 48 layers, hidden 2560, 512 experts / 10 active, GatedDeltaNet linear attention + full attention every 4th layer, PLE n-gram embedding stack, vision tower (27-layer, 1152-dim)
Quantization oQ3e — 3-bit affine base (group 64) with 475 sensitivity-selected tensor overrides (198 @ 8-bit, 129 @ 5-bit, 17 @ 4-bit, 3 @ 6-bit) + fp16 for norms/router/embeddings
MTP head Native, mtp_num_hidden_layers: 1 · 76 mtp.* tensors preserved — and it delivers +67 % decode
Alignment Uncensored (inherited from the base; verified on these quantized weights)
Context 262,144 native · vision preserved
Size 85.2 GB on disk · 17 shards · 3,748 tensors · 79.5 GB reported resident by oMLX
Engine oMLX (Apple MLX) — VLM engine. Requires an oMLX build with the fixed qwen4_exp loader (0.7.0.dev2+).

Where the bytes go

Group Size
Transformer body (MoE + attention) 55.6 GB
PLE / embedding tables (largely fp16) 26.3 GB
Vision tower 1.8 GB
MTP head 1.5 GB

Measured on our hardware (Apple M5 Max 128 GB · oMLX, VLM engine)

MTP speculative decoding

Decode tok/s, median of 3 runs, 320-token generation, temp 0, 8 k KV, TurboQuant-KV 8-bit:

Configuration Decode (tok/s) vs MTP off
MTP off 42.3
MTP on, draft-4 70.2 +66 %
MTP on, draft-6 70.8 +67 %
MTP on, draft-8 67.9 +61 %

draft-6 is the shipped default. draft-8 regresses — the extra drafts stop paying for their verification cost.

Prefill and long context

Prompt Input tokens Wall Prefill throughput
4 k-class 11,268 8.2 s ~1,384 tok/s
16 k-class 44,869 23.6 s ~1,914 tok/s
32 k-class 90,068 34.8 s ~2,605 tok/s

Prefill throughput rises with prompt length — the MoE prefill path amortizes well. A needle-in-a-haystack retrieval at 80 k tokens was answered exactly and verbatim, and a warm-cache repeat of that 80 k prompt returned in 3.4 s (99.7 % prefix cache hit).

Memory behaviour

Loaded via oMLX's qwen4_exp path under a 108 GB memory-guard ceiling: 79.5 GB reported model size, qwen4_ple_ssd_offload off (no SSD offload forced), ~62.5 GB of the PLE stack served through mmap rather than wired pages. Under sustained decode the machine stayed at ~21 % free system memory with no paging stalls and no throughput cliff — i.e. the mmap-backed PLE is not an SSD bottleneck on this hardware. Cold load: ~17 s.

Capability probes

Probe Result
General coherence / explanation Pass — fluent, accurate, well-structured
Reasoning (bat-and-ball trap) Pass — $0.05 with a correct one-line check
Refusal probes (profanity, security mechanics) Pass — zero refusals, engaged directly with both
Vision (chart reading) Pass — correctly identified chart type, background, title, and axis category labels from a PNG
Long-context retrieval @ 80 k Pass — exact needle recall
Tool calling — single-tool select Pass — correct function, correct arguments
Tool calling — "no tool needed" Pass — answered directly, no spurious call
Tool calling — parallel multi-city Fail — emitted one call instead of two
Tool calling — chained (weather → email) Fail — emitted reasoning prose instead of a call

Be honest about the weak spot: this build is strong at prose, reasoning, vision, and long context, and it is genuinely uncensored — but multi-step and parallel tool orchestration is its soft edge. If you are wiring it into an agent loop that depends on parallel or chained tool calls, test that path before committing. Single-tool selection is reliable.


Recommended oMLX settings

{
  "mtp_enabled": true,
  "mtp_num_draft_tokens": 6,
  "turboquant_kv_enabled": true,
  "turboquant_kv_bits": 8.0,
  "turboquant_skip_last": true,
  "qwen4_ple_ssd_offload": false,
  "qwen35_ane_prefill_enabled": false,
  "qwen35_oq_a8_enabled": false,
  "max_context_window": 262144
}

Notes:

  • preserve_mtp: true is not the default when quantizing with oMLX's oq pipeline — pass it explicitly or you will silently ship a model with no MTP head.
  • Keep qwen4_ple_ssd_offload off on a 128 GB box; forcing SSD offload converts a 70 tok/s model into a disk-bound one.
  • ANE prefill measured as a net regression for this class on our hardware — left off.

Usage

# oMLX
omlx serve --model-dir ~/.omlx/models
curl http://127.0.0.1:8000/v1/chat/completions \
  -H "Authorization: Bearer $OMLX_API_KEY" -H 'Content-Type: application/json' \
  -d '{"model":"Qwen3.8-Flash-Next-Uncensored-oQ3e-fp16-mtp","messages":[{"role":"user","content":"Hello"}]}'

Text + image works through the standard OpenAI image_url content block (base64 data URIs supported).

Intended use & limitations

This model is uncensored: it will follow instructions that aligned models decline, and it applies no content filtering of its own. You own what you generate with it, and you are responsible for putting appropriate safeguards around any deployment that faces other people. It is intended for local research, creative work, and agentic experimentation by people who want an unfiltered local assistant — not for unsupervised public-facing serving.

Other limitations: 3-bit base quantization will cost some accuracy against the bf16 original, particularly on tight factual recall; parallel/chained tool calling is unreliable (see probes above); and it requires an oMLX build with the working qwen4_exp loader.

Credits

Base model by orcarouter. Architecture by the Qwen team. Quantization stack: oMLX on Apple MLX. Quantized, tuned, benchmarked, and documented by Novaeon.Studio.

License inherited from the base model (Apache-2.0).

Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in app" button that hands off directly to a local runtime of your choice - Infrahuman, LM Studio, or Ollama. No API keys, no subscription, no prompt leakage.