← back to catalog · registered 2026-08-22 13:56

intheblue/Qwen3.8-27B-AEON-Ultimate-Uncensored-Multimodal-MLX-6bit

intheblue Qwen 27B multimodal second-order
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/intheblue%2FQwen3.8-27B-AEON-Ultimate-Uncensored-Multimodal-MLX-6bit"
Response includes
  • classification m1
  • files 16
  • author_summary 2 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M1
Primary method

Direct removal

No other method signals detected in this model.
Confidence
MEDIUM
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=0 (base model)
  • no specific method indicators - defaulting to M1 (most common)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
MEDIUM
Why we say so
primary_method=M1; difference-of-means is the reference extraction for M1/M3 (Arditi 2024)
Downloads · 30-day
2K
↑ 60% in 90 days
Likes
1
Model age
7w ago
created 2026-08-17
Downloads over time
Now1.9K→from1.2K↑60%
1.1K1.4K1.7K1.9K1.2K on Aug 191.9K on Sep 31.9K on Sep 2AugSep
Aug 19 → Sep 3 · 10 snapshots · spans 15 days

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
apache-2.0
Languages
en zh multilingual
Tags
mlx safetensors qwen3_5 mlx-vlm qwen3.8 apple-silicon quantized 6-bit multimodal vision vision-language gated-deltanet

Related

Total size
21.2 GB
Files
16
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-08-17 18:51

Files by quantization

Auxiliary files 16 files 21.2 GB
model-00001-of-00005.safetensors 4.99 GB f32d3483 download
model-00002-of-00005.safetensors 4.98 GB ff2ad5d8 download
model-00004-of-00005.safetensors 4.97 GB 17661ece download
model-00003-of-00005.safetensors 4.97 GB b5aaab4e download
model-00005-of-00005.safetensors 1.31 GB 5af47818 download
tokenizer.json 19.1 MB 6f32ce20 download
model.safetensors.index.json 213 KB 8b98b99f download
chat_template.jinja 8.74 KB c0c686f9 download
config.json 4.91 KB 5d160c1c download
README.md 4.83 KB 3b3d0a98 download
.gitattributes 1.53 KB 52373fe2 download
tokenizer_config.json 1.24 KB a9eacca6 download
processor_config.json 991 B 8f29fe38 download
preprocessor_config.json 390 B 2ea84a43 download
video_preprocessor_config.json 385 B 3ba673a5 download
generation_config.json 214 B 3f25ead4 download

README current version from Hugging Face


license: apache-2.0
base_model: AEON-7/Qwen3.8-27B-AEON-ULTIMATE-UNCENSORED-BF16
base_model_relation: quantized
pipeline_tag: image-text-to-text
library_name: mlx
language:

  • en
  • zh
  • multilingual
    tags:
  • mlx
  • mlx-vlm
  • qwen3_5
  • qwen3.8
  • apple-silicon
  • quantized
  • 6-bit
  • multimodal
  • vision
  • vision-language
  • gated-deltanet
  • hybrid-attention
  • mtp
  • speculative-decoding
  • abliterated
  • uncensored
  • conversational

Qwen3.8-27B-AEON-Ultimate-Uncensored — Multimodal MLX 6-bit

A 6-bit MLX quantization of AEON-7/Qwen3.8-27B-AEON-ULTIMATE-UNCENSORED-BF16 (revision 8f76e82), AEON's abliterated release of Qwen/Qwen3.8-27B, for local inference on Apple Silicon.

The vision tower is fully preserved (333/333 tensors — AEON's release keeps it hash-identical to stock Qwen3.8), so image and video understanding work through mlx-vlm. The model's native MTP head is published separately as a drafter for lossless speculative decoding: intheblue/Qwen3.8-27B-AEON-Ultimate-Uncensored-MLX-MTP-Drafter — pairing it typically speeds decode 1.4–1.9× at identical output quality.

Requirements

  • Apple Silicon Mac with MLX support; pip install mlx-vlm (converted with mlx-vlm 0.6.13 / mlx 0.32.0).
  • Unified memory: ~21 GB for weights, with an observed runtime peak of ~25 GiB at short context and ~30 GiB at 13k-token context. Comfortable on 48 GB+ machines, workable on 36 GB; not recommended below that.
  • Because decode is memory-bandwidth-bound, tokens/sec scales roughly linearly with the chip's memory bandwidth (see measured numbers below).

Usage

# text / vision
python -m mlx_vlm generate \
  --model intheblue/Qwen3.8-27B-AEON-Ultimate-Uncensored-Multimodal-MLX-6bit \
  --prompt "Describe this image." --image photo.jpg

# faster decode with the MTP drafter (lossless speculative decoding)
python -m mlx_vlm generate \
  --model intheblue/Qwen3.8-27B-AEON-Ultimate-Uncensored-Multimodal-MLX-6bit \
  --draft-model intheblue/Qwen3.8-27B-AEON-Ultimate-Uncensored-MLX-MTP-Drafter \
  --draft-kind mtp --draft-block-size 3 \
  --prompt "Write a short story about a lighthouse keeper."

Tips:

  • Draft block size 3 is the all-round sweet spot; 4 edges ahead on code; ≥5 regresses.
  • When serving (mlx_vlm.server), --prefill-step-size 512 cuts peak prefill memory by ~6 GB at no measured speed cost.
  • Recommended sampling (from the Qwen3.8 card): thinking temp=1.0, top_p=0.95, top_k=20; non-thinking temp=0.7, top_p=0.8, presence_penalty=1.5.

Conversion recipe

python -m mlx_vlm convert \
  --hf-path AEON-7/Qwen3.8-27B-AEON-ULTIMATE-UNCENSORED-BF16 \
  -q --q-bits 6 --q-group-size 64 \
  --mlx-path Qwen3.8-27B-AEON-Ultimate-Uncensored-mlx-6bit

Affine mode, group size 64, no calibration (RTN is near-lossless at 6-bit). Language model and vision tower are both quantized at the global setting; mtp.* tensors are excluded by design — mlx-vlm loads the MTP drafter as a separate model (see the drafter repo for the split recipe).

Measured performance

Test machine: Mac mini M4 Pro, 48 GB unified memory (~273 GB/s). Higher-bandwidth chips (M-series Max/Ultra) scale decode proportionally — serial decode on this quantization measures at ~100% of the device's streaming-bandwidth roofline.

Workload Config Decode
Coding (temp 0.2) MTP block 4 21.5 tok/s
Document QA @ 13k ctx (temp 0.7) MTP block 3 16.9 tok/s
Creative prose (temp 0.7) MTP block 3 15.9 tok/s
Serial (no drafter) — 11.4 tok/s

Prefill measured ~105–110 tok/s, flat with context length up to the 13k tested. Draft acceptance is workload-dependent: ~46% on open-ended prose, substantially higher on code and grounded QA.

Validation

  • Text, vision, and long-context (13k) generation smoke-tested after conversion.
  • Two-phase coding gate (spec research → implementation, executed against 28 hidden edge-case asserts across an SSE-parser task and a stack-VM task): 28/28.
  • Speculative decoding verified lossless (drafter rejections fall back to the target model's own tokens; the sampling distribution is unchanged by construction).

Provenance & responsibility

Qwen/Qwen3.8-27B → AEON-7 SSM-conv1d repair + abliterix abliteration (BF16, vision and MTP untouched — see their card for methodology and KL evidence) → this repo (6-bit MLX quantization, nothing else changed).

This is an abliterated, refusal-removed model. As the upstream card puts it: the model does not decide whether to comply — you do. Outputs are the responsibility of the operator; use within the law of your jurisdiction. Apache-2.0, inherited from base.

README history 6 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-10-02Update links after rename; drop uncensored tag/title62e86017.4 KB
    Loading...
  2. 2026-09-10Model card: status section (2026-09-10) with measured comparison, faster serv...b3709057.5 KB
    Loading...
  3. 2026-09-05Update repo ids to VisualInference (account renamed)494e28a4.9 KB
    Loading...
  4. 2026-08-17Update README.mdecb89f14.8 KB
    Loading...
  5. 2026-08-17Model card: community-facing revision (device-neutral copy; test hardware dis...8d18e505.2 KB
    Loading...
  6. 2026-08-17Add files using upload-large-folder tool98f0baa4.5 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration