license: apache-2.0
library_name: mlx
pipeline_tag: text-generation
language:
- en
- ja
base_model: - prism-ml/Ternary-Bonsai-2-27B-gguf
- BoldingBuilds/Ternary-Bonsai-2-27B-Abliterated-v2-PQ2_0-MTP-GGUF
tags: - mlx
- omlx
- mtp
- speculative-decoding
- ternary
- 2-bit
- pq2_0
- qwen3_5
- bonsai-2
- abliterated
- uncensored
- reasoning
- apple-silicon
Ternary-Bonsai-2-27B-v2-Abliterated-MLX-PQ2_0-MTP
An MLX-native conversion of BoldingBuilds' Ternary-Bonsai-2-27B-Abliterated-v2-PQ2_0-MTP, preserving exact ternary weight representations and enabling self-speculative decoding (Multi-Token Prediction) on Apple Silicon.
Key Technical Facts
- Lossless Quant-Native Weight Preservation: Converted directly from the official PQ2_0 GGUF without re-quantization. All 866 tensors (27.3 billion elements) are mapped bitwise-exact to MLX affine packed representation.
- Thinking Mode Fix (v2 Update): Incorporates BoldingBuilds' v2 output-projection adjustment, preventing runaway reasoning loops and ensuring proper closure (
</think>) before emitting the final answer. - MTP Self-Speculative Decoding: Integrates the grafted Qwen3.8 MTP draft head. Correctly resolves the Sylvester-Walsh-Hadamard orthogonal rotation coordinates across both primary and MTP embedding projections, achieving an empirical ~87.5% draft acceptance rate in oMLX.
- Low Memory Footprint: Runs within ~8.6 - 9.4 GB unified memory (VRAM), making 27B parameter reasoning accessible on 16GB Apple Silicon machines (MacBook Air / Pro / Mac mini).
Architecture & Specifications
| Parameter | Specification |
|---|---|
| Base Architecture | Qwen3.5 / Bonsai 2 (Hybrid GDN + Full Attention) |
| Parameters | 27B total |
| Modality | Text-only (Vision / Multimodal not supported) |
| Effective Bit-Width | 2.13 bpw (Ternary weights with block scales) |
| Context Length | Up to 131,072 tokens |
| Vocabulary Size | 248,320 tokens (ByteLevel BPE, full tail tokens preserved) |
| Speculative Engine | Multi-Token Prediction (MTP) depth=1 |
| Runtime Target | oMLX (native Metal kernel & MTP pipeline support) |
Quickstart (oMLX)
For optimal performance with speculative MTP decoding on Apple Silicon, run using oMLX (free, open-source local LLM runner).
This model requires custom architecture code (model.py) to handle Hadamard orthogonal transformations and MTP draft grafting.
1. Python Inference via oMLX Runtime
from omlx.model_settings import ModelSettings
from omlx.utils.model_loading import lm_load_compat, maybe_apply_pre_load_patches
import mlx_lm
model_path = "path/to/Ternary-Bonsai-2-27B-v2-Abliterated-MLX-PQ2_0-MTP"
# Enable MTP speculative decoding (depth=1)
settings = ModelSettings(mtp_enabled=True, mtp_fixed_depth=1)
maybe_apply_pre_load_patches(model_path, model_settings=settings)
# Load model (strict=True, trust_remote_code=True)
model, tokenizer = lm_load_compat(model_path, trust_remote_code=True, lazy=False)
# Generate
prompt = "<|im_start|>user\nExplain quantum entanglement in simple terms.<|im_end|>\n<|im_start|>assistant\n"
response = mlx_lm.generate(model, tokenizer, prompt=prompt, max_tokens=512, verbose=True)
print(response)
2. Standalone MLX-LM CLI
When using standard mlx_lm, run with --trust-remote-code:
mlx_lm.generate \
--model path/to/Ternary-Bonsai-2-27B-v2-Abliterated-MLX-PQ2_0-MTP \
--prompt "Hello!" \
--trust-remote-code
(Note: Speculative MTP acceleration requires the oMLX patch layer or an MTP-enabled mlx-lm runtime branch).
Provenance & Attribution
- Base Weights & Architecture: Prism ML (Sylvester-Walsh-Hadamard transform + Qwen3.5-based hybrid).
- Abliteration & Output Tuning: BoldingBuilds (quant-native ternary bit flip, thinking-closure fix).
- MTP Architecture: Qwen Team, Alibaba Cloud.
- License: Apache-2.0.