license: mit
base_model: huihui-ai/Huihui-Ornith-1.5-9B-abliterated
tags: [mlx, apple-silicon, quantized, uncensored]
library_name: mlx-vlm
Ornith-1.5-9B-Abliterated — MLX MXFP4
MLX MXFP4 quantisation of huihui-ai/Huihui-Ornith-1.5-9B-abliterated.
Changes: weights quantised to MXFP4 from the BF16 source withmlx_vlm.convert. No fine-tuning, no merging, no re-alignment.
Measured
Converted and measured on one machine — Apple M3 Ultra, 96 GB unified memory,
macOS 27 — as part of a full ladder. Every rung in this family came from the
same BF16 source with the same group size, so bit width is the only
variable between them.
| Size on disk | 6.18 GB |
| Perplexity | 5.783 |
| Relative to best rung in family | 1.09× |
| Throughput (1 req / 8 concurrent) | 67.6 / 164.4 tok/s |
Perplexity measured on allenai/tulu-3-sft-mixture, 192 samples of 512 tokens,
seed 123 — identical for every rung.
Perplexity is only comparable within this family. Tokenizers differ
between model families, so a number here should never be compared against a
different base model's. The×column above is the meaningful one.
Usage
pip install mlx-vlm
mlx_vlm.generate --model shoemoney/Ornith-1.5-9B-Abliterated-MLX-mxfp4 --prompt "Hello" --max-tokens 256
Load with mlx-vlm, not mlx-lm — this architecture is registered in mlx-vlm.
Provenance
mlx_vlm.convert --hf-path huihui-ai/Huihui-Ornith-1.5-9B-abliterated \
--mlx-path Ornith-1.5-9B-Abliterated-mxfp4 -q --q-mode mxfp4
License
mit, inherited from the base model. Attribution above.