license: other
base_model:
- llmfan46/Qwen3.6-27B-uncensored-heretic-v2-Native-MTP-Preserved
tags: - mlx
- omlx
- oq
- oq8
- qwen3.6
- mtp
pipeline_tag: text-generation
library_name: mlx
Qwen3.6-27B-uncensored-heretic-v2-Text-Only-oQ8-MLX
This repository contains an oMLX oQ8 mixed-precision MLX quantization ofllmfan46/Qwen3.6-27B-uncensored-heretic-v2-Native-MTP-Preserved.
This build tracks the BF16 artifact lineage from llmfan46/Qwen3.6-27B-uncensored-heretic-v2-Native-MTP-Preserved-GGUF. oMLX oQ
quantization operates on MLX/safetensors checkpoints rather than GGUF files, so
this build uses the corresponding BF16 safetensors checkpoint fromllmfan46/Qwen3.6-27B-uncensored-heretic-v2-Native-MTP-Preserved and excludes the existing GGUF quantizations.
Variant
- Quantization:
oQ8 - Variant: Text Only
- MTP tensors: stripped
- Text-only:
true - Approximate target density: 8.63 bpw.
- The vision/audio-side weights are removed from the output config and shards.
Usage
omlx serve dawncr0w/Qwen3.6-27B-uncensored-heretic-v2-Text-Only-oQ8-MLX
or with mlx-lm:
from mlx_lm import generate, load
model, tokenizer = load("dawncr0w/Qwen3.6-27B-uncensored-heretic-v2-Text-Only-oQ8-MLX")
print(generate(model, tokenizer, "Hello", max_tokens=64))
Validation
Local validation completed with the bundled oMLX runtime:
loader: mlx_lm.load
generation smoke test: passed
prompt: Hello
max tokens: 4
peak memory: 26.798 GB
Source
- BF16 source checkpoint:
llmfan46/Qwen3.6-27B-uncensored-heretic-v2-Native-MTP-Preserved - BF16 GGUF reference:
llmfan46/Qwen3.6-27B-uncensored-heretic-v2-Native-MTP-Preserved-GGUF - Quantization tool: oMLX oQ
Upstream model card license tag: apache-2.0.