← back to catalog · registered 2026-09-29 13:57

2001Y/Ternary-Bonsai-2-27B-Abliterated-v2-PQ2_0-MTP-MLX

2001Y 27B GGUF
curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/2001Y%2FTernary-Bonsai-2-27B-Abliterated-v2-PQ2_0-MTP-MLX"
Response includes
  • classification m-uncensored
  • files 26
  • author_summary 2 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M-U
Primary method

Uncensored (method unknown)

No other method signals detected in this model.
Confidence
LOW
Why this label 3 signals
Weak or ambiguous signals. Best guess based on catalog patterns; treat as tentative and check the evidence below.
  • 'uncensored' in name/tags but no 'abliterated' marker
  • method not identifiable from author declaration alone
  • may be DPO fine-tune, prompt engineering, or unknown technique
Refusal direction extraction

No specific extraction method could be identified for this model. The producer either did not document it or used a proprietary pipeline.

What is a refusal direction? →
Downloads · 30-day
0
Likes
1
Model age
today
created 2026-09-29

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
apache-2.0
Languages
en ja
Tags
mlx safetensors qwen3_5 omlx mtp speculative-decoding ternary 2-bit pq2_0 bonsai-2 abliterated uncensored

Related

Total size
8.37 GB
Files
26
Quantizations
1
Registered
2026-09-29 13:57
Last updated on HF
2026-09-29 13:07

Files by quantization

Auxiliary files 26 files 8.38 GB
model-00001.safetensors 1019 MB e5751668 download
model-00006.safetensors 1019 MB de17cca7 download
model-00005.safetensors 1013 MB 054a51c8 download
model-00004.safetensors 1002 MB 9d4d0940 download
model-00002.safetensors 1001 MB 60cb93b8 download
model-00003.safetensors 998 MB b0134713 download
model-00007.safetensors 957 MB f6b48c85 download
model-00008.safetensors 804 MB 0df21670 download
model-00009.safetensors 758 MB 045a87d4 download
hadamard-signs.safetensors 112 KB 285bad6d download
tokenizer.json 12.2 MB 0dcda57a download
conversion-report.json 920 KB a7273af2 download
tensor-map.json 633 KB 67983320 download
model-manifest.json 313 KB 96e66c6a download
model.safetensors.index.json 147 KB c186dcbc download
model.py 39.3 KB d0f3d8d0 download
tokenizer_config.json 16.1 KB c74de7e0 download
README.md 4.06 KB 14b493b0 download
config.json 3.97 KB e8cd9931 download
tokenizer-gate-report.json 3.86 KB 9021b8f9 download
build-result.json 2.66 KB b440c677 download
tokenization_bolding_mtp.py 2.17 KB f0a7175f download
PROVENANCE.md 1.79 KB 8794527a download
.gitattributes 1.53 KB 52373fe2 download
LICENSE-SOURCE.md 475 B c71a563d download
generation_config.json 202 B 023756cf download

README current version from Hugging Face


license: apache-2.0
library_name: mlx
pipeline_tag: text-generation
language:

  • en
  • ja
    base_model:
  • prism-ml/Ternary-Bonsai-2-27B-gguf
  • BoldingBuilds/Ternary-Bonsai-2-27B-Abliterated-v2-PQ2_0-MTP-GGUF
    tags:
  • mlx
  • omlx
  • mtp
  • speculative-decoding
  • ternary
  • 2-bit
  • pq2_0
  • qwen3_5
  • bonsai-2
  • abliterated
  • uncensored
  • reasoning
  • apple-silicon

Ternary-Bonsai-2-27B-Abliterated-v2-PQ2_0-MTP-MLX

An MLX-native conversion of BoldingBuilds' Ternary-Bonsai-2-27B-Abliterated-v2-PQ2_0-MTP, preserving exact ternary weight representations and enabling self-speculative decoding (Multi-Token Prediction) on Apple Silicon.

Key Technical Facts

  • Lossless Quant-Native Weight Preservation: Converted directly from the official PQ2_0 GGUF without re-quantization. All 866 tensors (27.3 billion elements) are mapped bitwise-exact to MLX affine packed representation.
  • Thinking Mode Fix (v2 Update): Incorporates BoldingBuilds' v2 output-projection adjustment, preventing runaway reasoning loops and ensuring proper closure (</think>) before emitting the final answer.
  • MTP Self-Speculative Decoding: Integrates the grafted Qwen3.8 MTP draft head. Correctly resolves the Sylvester-Walsh-Hadamard orthogonal rotation coordinates across both primary and MTP embedding projections, achieving an empirical ~87.5% draft acceptance rate in oMLX.
  • Low Memory Footprint: Runs within ~8.6 - 9.4 GB unified memory (VRAM), making 27B parameter reasoning accessible on 16GB Apple Silicon machines (MacBook Air / Pro / Mac mini).

Architecture & Specifications

Parameter Specification
Base Architecture Qwen3.5 / Bonsai 2 (Hybrid GDN + Full Attention)
Parameters 27B total
Effective Bit-Width 2.13 bpw (Ternary weights with block scales)
Context Length Up to 131,072 tokens
Vocabulary Size 248,320 tokens (ByteLevel BPE, full tail tokens preserved)
Speculative Engine Multi-Token Prediction (MTP) depth=1
Runtime Target oMLX (native Metal kernel & MTP pipeline support)

Quickstart (oMLX)

For optimal performance with speculative MTP decoding on Apple Silicon, run using oMLX (free, open-source macOS-native LLM runner with smart caching and native MTP acceleration).

This model requires custom architecture code (model.py) to handle Hadamard orthogonal transformations and MTP draft grafting.

1. Python Inference via oMLX Runtime

from omlx.model_settings import ModelSettings
from omlx.utils.model_loading import lm_load_compat, maybe_apply_pre_load_patches
import mlx_lm

model_path = "path/to/Ternary-Bonsai-2-27B-Abliterated-v2-PQ2_0-MTP-MLX"

# Enable MTP speculative decoding (depth=1)
settings = ModelSettings(mtp_enabled=True, mtp_fixed_depth=1)
maybe_apply_pre_load_patches(model_path, model_settings=settings)

# Load model (strict=True, trust_remote_code=True)
model, tokenizer = lm_load_compat(model_path, trust_remote_code=True, lazy=False)

# Generate
prompt = "<|im_start|>user\nExplain quantum entanglement in simple terms.<|im_end|>\n<|im_start|>assistant\n"
response = mlx_lm.generate(model, tokenizer, prompt=prompt, max_tokens=512, verbose=True)
print(response)

2. Standalone MLX-LM CLI

When using standard mlx_lm, run with --trust-remote-code:

mlx_lm.generate \
  --model path/to/Ternary-Bonsai-2-27B-Abliterated-v2-PQ2_0-MTP-MLX \
  --prompt "Hello!" \
  --trust-remote-code

(Note: Speculative MTP acceleration requires the oMLX patch layer or an MTP-enabled mlx-lm runtime branch).

Provenance & Attribution

  • Base Weights & Architecture: Prism ML (Sylvester-Walsh-Hadamard transform + Qwen3.5-based hybrid).
  • Abliteration & Output Tuning: BoldingBuilds (quant-native ternary bit flip, thinking-closure fix).
  • MTP Architecture: Qwen Team, Alibaba Cloud.
  • License: Apache-2.0.
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.