← back to catalog · registered 2026-08-22 13:56

tejones36/Qwen-3.8-27B-Uncensored-mlx-mixed_3_6

tejones36 Qwen 27B multimodal second-order
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/tejones36%2FQwen-3.8-27B-Uncensored-mlx-mixed_3_6"
Response includes
  • classification m1
  • files 16
  • hub_downloads_all_time 6,267
  • author_summary 1 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M1
Primary method

Direct removal

No other method signals detected in this model.
Confidence
MEDIUM
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=0 (base model)
  • no specific method indicators - defaulting to M1 (most common)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
MEDIUM
Why we say so
primary_method=M1; difference-of-means is the reference extraction for M1/M3 (Arditi 2024)
Downloads · lifetime
6K
843 last 30d - stable
Likes
5
Model age
8w ago
created 2026-08-15
Downloads over time
Now6.6K→from594↑1,005%
2962.6K4.9K7.2K594 on Aug 186.6K on Oct 11AugSepOct
Aug 18 → Oct 11 · 49 snapshots · spans 54 days

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
apache-2.0
Tags
mlx safetensors qwen3_5 qwen image-text-to-text mlx-vlm mixed-precision uncensored abliterated zerofuse vision multimodal

Related

Total size
13.7 GB
Files
16
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-08-15 23:27

Files by quantization

Auxiliary files 16 files 13.7 GB
model-00001-of-00003.safetensors 5.00 GB 236b0059 download
model-00002-of-00003.safetensors 4.98 GB 6c9fbed5 download
model-00003-of-00003.safetensors 3.74 GB 764170b8 download
tokenizer.json 19.1 MB 06b95093 download
vocab.json 6.41 MB 0aa0ce06 download
model.safetensors.index.json 213 KB 4ccb06f2 download
config.json 126 KB 9280f7ff download
chat_template.jinja 8.74 KB c0c686f9 download
README.md 3.48 KB 2d8d4209 download
zerofuse_run.json 2.40 KB 6f783c1c download
.gitattributes 1.53 KB 52373fe2 download
tokenizer_config.json 1.14 KB 1d134cd2 download
processor_config.json 991 B 8f29fe38 download
preprocessor_config.json 390 B 2ea84a43 download
video_preprocessor_config.json 385 B 3ba673a5 download
generation_config.json 202 B 023756cf download

README current version from Hugging Face


license: apache-2.0
library_name: mlx
pipeline_tag: image-text-to-text
tags:

  • qwen
  • image-text-to-text
  • mlx
  • mlx-vlm
  • mixed-precision
  • uncensored
  • abliterated
  • zerofuse
  • vision
  • multimodal
    base_model: junafinity/Qwen-3.8-27B-Uncensored

Qwen-3.8-27B-Uncensored (MLX Mixed 3/6-bit)

This repository contains Apple Silicon-optimized weights for junafinity/Qwen-3.8-27B-Uncensored converted to MLX format using mlx-vlm.

Quantization Specifications

  • Converter: mlx-vlm
  • Base Model: junafinity/Qwen-3.8-27B-Uncensored
  • Recipe: --quant-predicate mixed_3_6
  • Group Size: 64
  • Format: MLX Safetensors
  • Modality: Multimodal (Vision & Text)

Precision Breakdown (mixed_3_6)

The mixed_3_6 quantization recipe dynamically allocates bit-width across model layers:

  • Attention & Critical Projection Layers: Quantized to 6-bit to preserve attention fidelity, reasoning capability, and overall coherence.
  • Feed-Forward / MLP Layers: Compressed to 3-bit to significantly reduce memory footprint and memory bandwidth pressure.
  • Effective Precision: ~3.6 bits per weight, allowing the 27B parameter model to run comfortably on Apple Silicon Macs with 18 GB – 24 GB+ Unified Memory.

Installation & Requirements

Ensure you have mlx-vlm installed:

pip install -U mlx-vlm

Usage

1. Command Line Interface (CLI)

Text Generation:

mlx_vlm.generate \
  --model tejones36/Qwen-3.8-27B-Uncensored-mlx-mixed_3_6 \
  --prompt "Explain the concept of speculative decoding in MLX." \
  --max-tokens 512 \
  --verbose

Multimodal (Image Analysis):

mlx_vlm.generate \
  --model tejones36/Qwen-3.8-27B-Uncensored-mlx-mixed_3_6 \
  --image "/path/to/image.png" \
  --prompt "Describe the visual details and composition of this image." \
  --max-tokens 512

2. Local OpenAI-Compatible API Server

Launch the server (e.g., on port 2077):

mlx_vlm.server \
  --model tejones36/Qwen-3.8-27B-Uncensored-mlx-mixed_3_6 \
  --port 2077

Test via curl:

curl http://localhost:2077/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{
    "model": "tejones36/Qwen-3.8-27B-Uncensored-mlx-mixed_3_6",
    "messages": [
      {"role": "user", "content": "Server check: confirm operational status."}
    ],
    "max_tokens": 128,
    "temperature": 0.3
  }'

3. Python API

from mlx_vlm import load, generate
from mlx_vlm.prompt_utils import apply_chat_template
from mlx_vlm.utils import load_config

model_path = "tejones36/Qwen-3.8-27B-Uncensored-mlx-mixed_3_6"

# 1. Load quantized model and multimodal processor
model, processor = load(model_path)
config = load_config(model_path)

# 2. Format prompt using model chat template
prompt = "Write a concise technical summary of mixed-bit quantization advantages."
formatted_prompt = apply_chat_template(processor, config, prompt)

# 3. Generate response
output = generate(
    model,
    processor,
    prompt=formatted_prompt,
    max_tokens=512,
    verbose=True
)

print(output)

Acknowledgements

README history 1 version

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-08-15Upload folder using huggingface_hub47a5d633.5 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration