← back to catalog · registered 2026-08-22 13:56

pyros-vault/Qwen3.8-27B-Uncensored-oQ4e-mtp

pyros-vault Qwen 27B multimodal second-order
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/pyros-vault%2FQwen3.8-27B-Uncensored-oQ4e-mtp"
Response includes
  • classification m-uncensored
  • files 16
  • hub_downloads_all_time 7,172
  • author_summary 5 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M-U
Primary method

Uncensored (method unknown)

No other method signals detected in this model.
Confidence
LOW
Why this label 3 signals
Weak or ambiguous signals. Best guess based on catalog patterns; treat as tentative and check the evidence below.
  • 'uncensored' in name/tags but no 'abliterated' marker
  • method not identifiable from author declaration alone
  • may be DPO fine-tune, prompt engineering, or unknown technique
Refusal direction extraction

No specific extraction method could be identified for this model. The producer either did not document it or used a proprietary pipeline.

What is a refusal direction? →
Downloads · lifetime
7K
3K last 30d - stable
Likes
16
Model age
7w ago
created 2026-08-18
Downloads over time
Now8K→from1.2K↑579%
1303K5.9K8.7K1.2K on Aug 198K on Oct 11AugSepOct
Aug 19 → Oct 11 · 49 snapshots · spans 53 days

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
apache-2.0
Tags
mlx safetensors qwen3_5 omlx oq quantized qwen3.8 uncensored mtp multimodal conversational image-text-to-text

Related

Total size
15.8 GB
Files
16
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-08-21 15:33

Files by quantization

Auxiliary files 16 files 15.8 GB
model-00002-of-00004.safetensors 4.67 GB d1b7a7aa download
model-00001-of-00004.safetensors 4.67 GB 43f60bf5 download
model-00003-of-00004.safetensors 4.66 GB cab35461 download
model-00004-of-00004.safetensors 1.80 GB 9dc42929 download
tokenizer.json 12.2 MB 0997f410 download
vocab.json 6.41 MB 0aa0ce06 download
merges.txt 3.20 MB a494e019 download
model.safetensors.index.json 207 KB 9798d031 download
config.json 46.2 KB 8eaffa5a download
oq_imatrix_report.json 30.2 KB cac88e4e download
tokenizer_config.json 17.5 KB 5de744b3 download
chat_template.jinja 8.74 KB c0c686f9 download
README.md 4.93 KB 1e8ae661 download
.gitattributes 1.53 KB 52373fe2 download
preprocessor_config.json 390 B 2ea84a43 download
generation_config.json 202 B 023756cf download

README current version from Hugging Face


library_name: mlx
pipeline_tag: image-text-to-text
inference: false
license: apache-2.0
base_model: orcarouter/Qwen3.8-27B-Uncensored
base_model_relation: quantized
tags:

  • mlx
  • omlx
  • oq
  • quantized
  • qwen3.8
  • qwen3_5
  • uncensored
  • mtp
  • multimodal
  • conversational

Qwen3.8-27B-Uncensored-oQ4e-mtp

This repository is a complete Apple MLX deployment of orcarouter/Qwen3.8-27B-Uncensored, converted with oMLX v0.6.1 using importance-matrix-enhanced oQ mixed-precision quantization.

Download the whole repository: the Safetensors shards require the included index, model config, tokenizer, chat template, and image/video processor files. This is not a GGUF, Transformers, or NInfer artifact.

Quick facts

Item Value
Model type qwen3_5
Quantization layout Affine Q4/G64 by default, with 166 Q5/G64 tensor overrides.
Tensor payload 16,971,681,558 bytes / 15.81 GiB
Safetensors shards 4
Conversion runtime oMLX 0.6.1
Calibration oqe_code_multilingual, 128 samples × 512 tokens
Included model features Vision resources and one MTP layer
Intended runtime oMLX on Apple Silicon/macOS

Choose a variant

Variant Nominal tier Tensor payload Shards
oQ4e + FP16 MTP auxiliaries 4-bit 17,893,140,142 bytes / 16.66 GiB 4
oQ4e (this repo) 4-bit 16,971,681,558 bytes / 15.81 GiB 4
oQ6e 6-bit 23,716,288,460 bytes / 22.09 GiB 5
oQ8e 8-bit 30,001,641,934 bytes / 27.94 GiB 6

These tiers differ in storage and quantization layout. No same-Mac quality, memory, TTFT, or throughput comparison is published here, so the table should not be read as a benchmark.

Download

Install the Hugging Face CLI, then place the complete repository below oMLX's model directory:

mkdir -p "$HOME/.omlx/models/pyros-vault"
hf download pyros-vault/Qwen3.8-27B-Uncensored-oQ4e-mtp \
  --local-dir "$HOME/.omlx/models/pyros-vault/Qwen3.8-27B-Uncensored-oQ4e-mtp"

Serve with oMLX

Install the current oMLX runtime and start its OpenAI-compatible server:

brew tap jundot/omlx https://github.com/jundot/omlx
brew install jundot/omlx/omlx
omlx serve --model-dir "$HOME/.omlx/models"

Discover the exact model ID exposed by your installed oMLX version:

curl http://127.0.0.1:8000/v1/models

Use that returned ID with the OpenAI-compatible endpoint. MTP files being present does not automatically enable speculative decoding: Lightning MTP is opt-in through oMLX model settings, and behavior can vary by runtime version and Apple chip.

Quantization and verification

The bundled oq_imatrix_report.json records calibration with oqe_code_multilingual over 128 sequences of 512 tokens. The included report records 504 importance entries, 503 applied modules, two missing names, and no shape mismatches.

The report and tensor metadata establish how the artifact was built; they are not an end-to-end quality benchmark. Repository structure, configs, shard counts, payload sizes, and quantization metadata were audited for this card. Inference was not rerun on a Mac, so no local speed, memory, MTP-acceptance, Vision-quality, or long-context claim is made.

Provenance

This is a deployment conversion of orcarouter/Qwen3.8-27B-Uncensored. The Uncensored label and all behavior or training claims are inherited from that source and were not independently verified here.
The repository retains the source model's Vision resources and one MTP layer. It does not contain NInfer DFlash weights. The label does not guarantee unrestricted, safe, correct, or policy-compliant output.

Limitations

  • MLX/oMLX targets Apple Silicon and macOS; this repository is not runnable through CUDA on Windows.
  • Hugging Face hosted inference does not serve this custom oMLX layout.
  • The config advertises a 262,144-token maximum context. That value is model metadata, not a claim that this full context was tested or will fit your machine.
  • Vision preprocessing, tool use, MTP acceptance, memory use, and throughput depend on the oMLX version, client, prompt, and Apple hardware.
  • Quantization can change output quality. Evaluate this exact variant on your workload.

License and credits

The direct upstream declares Apache-2.0. Review its gated model card and repository files for the full attribution and usage terms.

Quantized and packaged by pyros-vault with oMLX/oQ.

README history 4 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-08-21docs: document oMLX artifact and usage06c10b84.9 KB
    Loading...
  2. 2026-08-18Update README.md13ec629730 B
    Loading...
  3. 2026-08-18Update README.md6e4218c497 B
    Loading...
  4. 2026-08-18Upload Qwen3.8-27B-Uncensored-oQ4e-mtp via oMLX54fe072320 B
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration