license: apache-2.0
library_name: mlx
pipeline_tag: image-text-to-text
tags:
- qwen
- image-text-to-text
- mlx
- mlx-vlm
- mixed-precision
- uncensored
- abliterated
- zerofuse
- vision
- multimodal
base_model: junafinity/Qwen-3.8-27B-Uncensored
Qwen-3.8-27B-Uncensored (MLX Mixed 3/6-bit)
This repository contains Apple Silicon-optimized weights for junafinity/Qwen-3.8-27B-Uncensored converted to MLX format using mlx-vlm.
Quantization Specifications
- Converter:
mlx-vlm - Base Model: junafinity/Qwen-3.8-27B-Uncensored
- Recipe:
--quant-predicate mixed_3_6 - Group Size: 64
- Format: MLX Safetensors
- Modality: Multimodal (Vision & Text)
Precision Breakdown (mixed_3_6)
The mixed_3_6 quantization recipe dynamically allocates bit-width across model layers:
- Attention & Critical Projection Layers: Quantized to 6-bit to preserve attention fidelity, reasoning capability, and overall coherence.
- Feed-Forward / MLP Layers: Compressed to 3-bit to significantly reduce memory footprint and memory bandwidth pressure.
- Effective Precision: ~3.6 bits per weight, allowing the 27B parameter model to run comfortably on Apple Silicon Macs with 18 GB – 24 GB+ Unified Memory.
Installation & Requirements
Ensure you have mlx-vlm installed:
pip install -U mlx-vlm
Usage
1. Command Line Interface (CLI)
Text Generation:
mlx_vlm.generate \
--model tejones36/Qwen-3.8-27B-Uncensored-mlx-mixed_3_6 \
--prompt "Explain the concept of speculative decoding in MLX." \
--max-tokens 512 \
--verbose
Multimodal (Image Analysis):
mlx_vlm.generate \
--model tejones36/Qwen-3.8-27B-Uncensored-mlx-mixed_3_6 \
--image "/path/to/image.png" \
--prompt "Describe the visual details and composition of this image." \
--max-tokens 512
2. Local OpenAI-Compatible API Server
Launch the server (e.g., on port 2077):
mlx_vlm.server \
--model tejones36/Qwen-3.8-27B-Uncensored-mlx-mixed_3_6 \
--port 2077
Test via curl:
curl http://localhost:2077/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{
"model": "tejones36/Qwen-3.8-27B-Uncensored-mlx-mixed_3_6",
"messages": [
{"role": "user", "content": "Server check: confirm operational status."}
],
"max_tokens": 128,
"temperature": 0.3
}'
3. Python API
from mlx_vlm import load, generate
from mlx_vlm.prompt_utils import apply_chat_template
from mlx_vlm.utils import load_config
model_path = "tejones36/Qwen-3.8-27B-Uncensored-mlx-mixed_3_6"
# 1. Load quantized model and multimodal processor
model, processor = load(model_path)
config = load_config(model_path)
# 2. Format prompt using model chat template
prompt = "Write a concise technical summary of mixed-bit quantization advantages."
formatted_prompt = apply_chat_template(processor, config, prompt)
# 3. Generate response
output = generate(
model,
processor,
prompt=formatted_prompt,
max_tokens=512,
verbose=True
)
print(output)
Acknowledgements
- Base model weights provided by junafinity.
- Converted and powered by Apple's MLX and mlx-vlm.