← back to catalog · registered 2026-09-18 13:56

LMLiutenant/Qwen3.8-Flash-Next-Uncensored-oQ5e-mtp

LMLiutenant multimodal second-order
curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/LMLiutenant%2FQwen3.8-Flash-Next-Uncensored-oQ5e-mtp"
Response includes
  • classification m1
  • files 39
  • author_summary 2 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M1
Primary method

Direct removal

No other method signals detected in this model.
Confidence
MEDIUM
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=0 (base model)
  • no specific method indicators - defaulting to M1 (most common)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
MEDIUM
Why we say so
primary_method=M1; difference-of-means is the reference extraction for M1/M3 (Arditi 2024)
Downloads · 30-day
0
Likes
0
Model age
today
created 2026-09-18

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
other
Tags
mlx safetensors qwen4_exp omlx qwen4 mtp vision tool-calling uncensored abliterated image-text-to-text conversational

Related

Total size
120 GB
Files
39
Quantizations
1
Registered
2026-09-18 13:56
Last updated on HF
2026-09-18 13:05

Files by quantization

Auxiliary files 39 files 120 GB
model-00001-of-00025.safetensors 4.89 GB cef36e92 download
model-00009-of-00025.safetensors 4.86 GB 53fa1590 download
model-00017-of-00025.safetensors 4.83 GB 28057456 download
model-00019-of-00025.safetensors 4.83 GB 3c525536 download
model-00013-of-00025.safetensors 4.83 GB 66a8ee61 download
model-00015-of-00025.safetensors 4.83 GB 2d169d62 download
model-00024-of-00025.safetensors 4.83 GB 9d3fc684 download
model-00021-of-00025.safetensors 4.83 GB 84fabefe download
model-00016-of-00025.safetensors 4.83 GB 08602413 download
model-00018-of-00025.safetensors 4.83 GB 26293a09 download
model-00020-of-00025.safetensors 4.83 GB 828d8883 download
model-00022-of-00025.safetensors 4.83 GB c169715b download
model-00023-of-00025.safetensors 4.83 GB d5238c3c download
model-00014-of-00025.safetensors 4.83 GB 53c95bdb download
model-00012-of-00025.safetensors 4.83 GB 3b89d6e5 download
model-00010-of-00025.safetensors 4.83 GB 08e05139 download
model-00011-of-00025.safetensors 4.83 GB 060f118e download
model-00007-of-00025.safetensors 4.75 GB 3935d81d download
model-00003-of-00025.safetensors 4.75 GB e5f7ec63 download
model-00006-of-00025.safetensors 4.75 GB 83edceda download
model-00002-of-00025.safetensors 4.75 GB 63a045b9 download
model-00005-of-00025.safetensors 4.75 GB 4fbd3e5c download
model-00004-of-00025.safetensors 4.75 GB 809f4085 download
model-00008-of-00025.safetensors 4.66 GB 17537b08 download
model-00025-of-00025.safetensors 4.30 GB c897973e download
tokenizer.json 12.2 MB 0997f410 download
vocab.json 6.41 MB 0aa0ce06 download
merges.txt 3.20 MB a494e019 download
model.safetensors.index.json 396 KB 3936649e download
config.json 248 KB 6117ca80 download
oq_imatrix_report.json 65.8 KB f843552b download
tokenizer_config.json 17.5 KB 5de744b3 download
chat_template.jinja 8.74 KB c0c686f9 download
README.md 5.26 KB 8fc8e9c2 download
LICENSE 3.16 KB 9557a896 download
SHA256SUMS 2.42 KB d3c709e3 download
.gitattributes 1.53 KB 52373fe2 download
preprocessor_config.json 390 B 2ea84a43 download
generation_config.json 202 B 023756cf download

README current version from Hugging Face


base_model: orcarouter/Qwen3.8-Flash-Next-Uncensored
base_model_relation: quantized
library_name: mlx
pipeline_tag: image-text-to-text
license: other
license_name: qwen-community-1.0
license_link: LICENSE
tags:

  • omlx
  • mlx
  • qwen4
  • mtp
  • vision
  • tool-calling
  • uncensored
  • abliterated

Qwen3.8-Flash-Next-Uncensored-oQ5e-mtp

An oMLX-native oQ5e quantization of
orcarouter/Qwen3.8-Flash-Next-Uncensored,
the BF16 abliterated build of Qwen3.8-Flash-Next. Targets a 128 GB Apple
Silicon Mac and preserves the vision tower, the MTP head, and the checkpoint's
native 262,144-token architecture setting.

For the abliteration method, refusal-direction details, and MTP-consistency
notes, see the source model card linked above. This repository only documents
the quantization.

Companion build: LMLiutenant/Qwen3.8-Flash-Next-Uncensored-oQ6e-mtp
(higher precision, ~6.9 bpw).

Nature of this model

This is an uncensored (abliterated) model: the source weights have had their
refusal direction removed, so it will attempt most requests without the
safety refusals present in the official Qwen release. It has no added
guardrails. You are responsible for how you use it and for complying with the
license and applicable law. Do not expose it to untrusted input in an agentic
setup without your own safeguards; with no refusal behaviour, it will not push
back on injected instructions.

Quantization

Property Value
Base model orcarouter/Qwen3.8-Flash-Next-Uncensored
Quantization oQ5e
Quantizer oMLX 0.6.4
Enhanced / imatrix mode Yes
Nominal group size 64
Non-quantized dtype BF16
MTP preserved Yes
Vision preserved Yes
Sensitivity model jedisct1/Qwen3.8-Flash-Next-Uncensored-oQ4e-100K-MTP

The oQ5e weights were quantized directly from the abliterated BF16 source. The
sensitivity model was used only to guide mixed-precision allocation; its
quantized weights are not the source of this model. It was chosen for matched
abliterated lineage with the source. Nominal group size was 64; oQ mixed-precision
may use different effective settings for selected tensors.

Hardware and memory

Qwen3.8-Flash-Next contains a very large N-gram/PLE component. With SSD N-gram
Offload it does not all stay resident in unified memory, which is what makes a
model this size practical on a 128 GB machine. Long-context and KV-cache use
raise memory further with context length.

Tested on a MacBook Pro, Apple M4 Max, 128 GB, oMLX, with SSD N-gram Offload
enabled. Peak MLX allocation (weights + KV + activations), measured in oMLX with SSD
N-gram Offload on: ~88 GB at 4K context, ~92 GB at 128K. These are
process-level allocator peaks, not total-system memory.

Recommended oMLX settings

For a 128 GB Apple Silicon Mac:

  • SSD N-gram Offload: ON
  • Lightning MTP: ON for chat/reasoning; consider OFF for coding agents

Serve via oMLX's OpenAI-compatible API. The weights are standard mlx-lm
compatible safetensors and should also load in mlx-lm and other MLX apps (untested).

Note: an unrelated oMLX engine bug on 0.7.0.dev1/dev2 can silently drop tool
calls to unregistered function names on the streaming path
(jundot/omlx#3660, open as of
2026-09; 0.6.4 unaffected). It does not affect the weights.

Benchmarks

Practical local tests, limited by compute time, not a standardized academic
suite, performed on oMLX 0.7.0.dev2.

Thinking OFF

Benchmark Samples oQ5e [this repo] oQ6e [this repo] oQ4e ¹ oQ5e ²
MMLU 2000 88.0% 88.2% 87.2% 88.1%
MMLU-Pro 1000 70.4% 73.4% 64.4% 69.2%

¹ jedisct1/Qwen3.8-Flash-Next-Uncensored-oQ4e-100K-MTP (abliterated lineage)
² GBP-DE/Qwen3.8-Flash-Next-oQ5e-mtp (base, non-abliterated)

At 2000 samples MMLU is flat across all four builds (~88%), within sampling
noise: the abliterated quants are at parity with the base build on broad
knowledge. On MMLU-Pro, accuracy rises with bit width across the
oQ4e/oQ5e/oQ6e series.

Smaller (n=100) GSM8K and HumanEval runs were near ceiling (90-98%) for every
build and are omitted as non-discriminating.

Model architecture

Quantized from Qwen3.8-Flash-Next (sparse Mixture-of-Experts): ~125B LM
parameters, ~6B activated, ~51B N-gram embedding parameters, ~4B MTP
parameters, 48 layers, 512 experts, 10 routed + 1 shared activated. See the
official Qwen model card for full architecture and context details.

Credits

License

Qwen Community License 1.0, inherited from the source repository and included as
LICENSE. Note the source card labels itself Apache-2.0, which does not match
the license file it ships; review before use or redistribution.

Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in app" button that hands off directly to a local runtime of your choice - Infrahuman, LM Studio, or Ollama. No API keys, no subscription, no prompt leakage.