license: apache-2.0
library_name: mlx
pipeline_tag: image-text-to-text
base_model:
- OS-Software/Ternary-Bonsai-2-27B-Uncensored-Heretic-v2-GGUF
tags: - mlx
- mlx-vlm
- ternary
- 2-bit
- bonsai
- heretic
- uncensored
- conversational
- image-text-to-text
- tool-calling
- apple-silicon
Ternary Bonsai 2 27B Uncensored Heretic v2: MLX 2-bit
Community MLX conversion of OS-Software/Ternary-Bonsai-2-27B-Uncensored-Heretic-v2-GGUF, with the unchanged vision tower from Prism ML's Bonsai 2 MLX pack. The Safetensors weights occupy 8,595,478,392 bytes, approximately 8.60 GB (8.01 GiB), including vision.
The published language weights were repacked from Heretic v2 PQ2_0 into MLX affine 2-bit/group-128 storage without additional quantization or training. This repository also documents conversion from the source model's denser PTQ1_0 packing. Both conversion paths produce native MLX 2-bit/group-128 storage; converting PTQ1_0 does not create a 1.75-bit MLX format. The original developers are credited below.
Model and format
| Item | This pack |
|---|---|
| Variant | Ternary Bonsai 2 27B Uncensored Heretic v2 |
| Runtime | Apple Silicon macOS, MLX and native Hadamard-aware mlx-vlm |
| Model type | prism_hadamard_qwen35, schema version 2 |
| Language storage | Affine 2-bit, groups of 128 |
| Vision | Unchanged FP16 vision tower reused from the Prism MLX base |
| Configured context limit | 262,144 tokens; full-length operation was not tested for this conversion |
| License | Apache 2.0; see LICENSE and retained attribution in NOTICE.txt |
“2-bit” describes the MLX storage format. The underlying low-bit language weights remain ternary. Metadata, group scales, biases, higher-precision state tensors, and the FP16 vision tower also take space.
Runtime requirement
Use mlx-vlm 0.7.2 with native prism_hadamard_qwen35 support. This model stores its language weights in a rotated basis. The loader must apply the matching Hadamard activation transforms and embedding lookup. Generic MLX or Transformers loaders that omit these operations can produce incorrect output. Keep config.json, hadamard.json, tokenizer and processor files together with the weights.
The model files are intended for Apple Silicon Macs (M1 or later) with sufficient unified memory and a macOS/Python build supported by the pinned runtime. They are not specific to M1 hardware. This conversion was created and tested on an Apple M1 Max with 64 GiB unified memory; other Mac models have not been tested for this release. Merely being able to run MLX is insufficient: the loader must support this model's Hadamard-aware architecture. See the MLX installation requirements and the pinned versions below. Memory use depends on context length, image resolution and workload; the weight-file size alone is not a RAM requirement.
The tested environment used Python 3.11 on Apple Silicon with:
mlx==0.32.2
mlx-vlm==0.7.2
transformers==5.14.1
Newer versions may work, but were not validated for this release. The older runtime/ helper bundled with the original Prism pack is not required by this schema 2 native pack.
Run locally
Download this repository into a directory named bonsai-heretic-mlx, including the weights and all configuration, tokenizer, processor, and Hadamard files. Change into that folder and create a Python environment:
cd bonsai-heretic-mlx
python3.11 -m venv .venv-heretic
source .venv-heretic/bin/activate
python -m pip install -r requirements.txt
python -m mlx_vlm.generate \
--model . \
--prompt "What is the capital of France? Answer with only the city name." \
--max-tokens 64 \
--temperature 0
That short command uses the native CLI's default of thinking disabled. For serving, use the included serve.py helper:
python serve.py \
--model . \
--model-id Ternary-Bonsai-2-27B-Uncensored-Heretic-v2-mlx-2bit \
--host 127.0.0.1 \
--enable-thinking
The helper exposes an OpenAI-compatible API at http://127.0.0.1:8080/v1, with the fixed public model ID shown above. Read /v1/models to obtain that ID. The native server default is port 8080; use --port to choose another free port and update your client URL accordingly. The helper uses the native mlx-vlm loader and generation engine behind an HTTP boundary that translates model identity and limits model selection. Its public ID does not depend on your directory name.
For a short text request, run this in another terminal with Python available:
import json
from urllib.request import Request, urlopen
base = "http://127.0.0.1:8080/v1"
with urlopen(base + "/models") as response:
models = json.load(response)["data"]
model_id = models[0]["id"]
payload = {
"model": model_id,
"messages": [{"role": "user", "content": "Explain ternary model weights briefly."}],
"max_tokens": 256,
"temperature": 0.7,
"top_p": 0.8,
"top_k": 20,
"enable_thinking": False,
"stream": False,
}
request = Request(
base + "/chat/completions",
data=json.dumps(payload).encode(),
headers={"Content-Type": "application/json"},
)
with urlopen(request) as response:
result = json.load(response)
print(result["choices"][0]["message"]["content"])
For image input, use an OpenAI image_url content part; data URLs are supported. For tool use, send an OpenAI tools array, execute the returned function in your application, and send its result back as a tool message with the matching tool_call_id. Thinking can be selected per request with enable_thinking; the sample disables it for a short answer. Model loading and image processing require additional memory beyond the weight-file size.
Conversion and provenance
Both source packings are from the same pinned source revision:
Repository: OS-Software/Ternary-Bonsai-2-27B-Uncensored-Heretic-v2-GGUF
Revision: 6db27547a8f25667cdb43a28d15cc5d33937b7dc
| Source file | Size in bytes | SHA256 |
|---|---|---|
Ternary-Bonsai-2-27B-Uncensored-Heretic-v2-PQ2_0.gguf |
7,206,168,928 | 9ba04345402584cf7f581ff72c00ea8925a35f3d78ccd34ca3eb195fe552393c |
Ternary-Bonsai-2-27B-Uncensored-Heretic-v2-PTQ1_0.gguf |
5,946,648,928 | c30f3862e4809db58366aab0b97a50e0a144dd010061d931a61c000adaf50683 |
PQ2_0 uses two-bit slots in GGUF, while PTQ1_0 packs ternary values more densely. The converter decodes each source packing and preserves its ternary codes and group scales when repacking into MLX affine storage. MLX codes 0, 1, and 2 represent -s, 0, and +s through scale=s and bias=-s. Hadamard signs and transforms are preserved. No additional quantization or training is applied by either conversion path.
A complete conversion from the real PTQ1_0 file reproduced all 2,057 language tensors bit for bit against the PQ2_0-derived pack. The resulting entire 2,390-tensor Safetensors file also has the same size and SHA256: e56ba66424066818c7c89f0f1ba74111705b9062367905228736013a48aaf08d. The loading configurations matched, and a native CLI check of the separately converted PTQ1_0 pack returned Paris. Both source routes therefore produce the same published MLX weights; one shared model.safetensors serves both. See the full PTQ1_0 comparison and the independent codec checks.
The SSM A_log tensors are reconstructed as FP32 log(-stored GGUF SSM A), following the Prism runtime convention. Thus “lossless repacking” refers to the ternary codes and scales; it does not claim that this logarithmic reconstruction is a bit-exact inverse of an unavailable pre-GGUF checkpoint.
The output contains 2,390 tensors, of which 2,057 are language tensors. Tokenizer IDs and BPE merge ranks matched the source. The converter was validated against an already installed original Bonsai 2 GGUF/MLX pair: all 2,057 reconstructed language tensors matched that official MLX pack bit for bit.
The Heretic and original GGUF projectors had identical SHA256 hashes, permitting reuse of the original MLX vision tower and processor files unchanged. The Prism base MLX template revision is fcba37d2117a7077eac6b613b2668d14d9779edd. Its Safetensors SHA256 was 130de5925082c168b7866b2e91b52e44abbafc99017e3ca352b77b5b55a269ed; the corresponding GGUF projector SHA256 was 6807ede61d570bb86ba34b756a0fa109edc33668604de867c6ea6d8f1d631903.
See conversion.json for the original PQ2_0 provenance, conversion-PTQ1_0.json for the PTQ1_0 conversion, reproduce.md for conversion instructions, and SHA256SUMS for publication-file checksums. The original-pair validation is preserved in evidence/mlx-base-conversion-validation.json.
Portable metadata and local paths
The published model configuration and conversion metadata contain no personal filesystem paths or credentials. Provenance records public source and template repository IDs, pinned revisions, file basenames, and content hashes. The configuration omits _name_or_path; your local checkout location is not embedded as model identity.
The native loader still needs filesystem paths internally on your machine. The included serve.py keeps those paths local at the documented HTTP boundary: it publishes one fixed model ID, replaces known model, home, and working-directory paths in runtime metadata, response headers, and error responses, and disables arbitrary local model selection and discovery. See evidence/privacy-runtime-verification.json for the checked API behavior.
This handling applies to model identity and runtime metadata. User messages, generated answers, and tool arguments remain unchanged. It does not remove private content that an application or user supplies in a prompt or tool result. Running the native server directly can expose local model paths through its IDs and discovery responses; the serve.py command above is the recommended serving entry point for this repository.
Verification
Functional smoke checks completed on 2026-10-09, on an Apple M1 Max with 64 GiB unified memory, using the tested stack above and the Bonsai demo's model-ID adapter around the native server:
| Check | Result |
|---|---|
| Text: capital of France | Paris |
| Vision: solid red 128 × 128 image | Red |
| Native function call | get_weather with city set to Vienna, and finish_reason: "tool_calls" |
| Tool round trip | Used the supplied mock result: sunny, 17°C |
The weather response was supplied by the test harness; the model did not fetch live weather. These tests used temperature 0, thinking disabled, and a maximum of 192 output tokens. They verify basic text, image, and tool plumbing. They do not constitute a benchmark or measure OCR, broad vision accuracy, long-context reliability, reasoning quality, or refusal rates. Results published for the original Bonsai model or the source GGUF are not presented as measurements of this MLX conversion.
The original requests, responses, and checks are retained in evidence/mlx-verification.json. The revised fixed-ID server passed the same four checks, plus streaming and private-path boundary checks; see evidence/privacy-runtime-verification.json and evidence/privacy-boundary-verification.json.
Intended use and limitations
OS-Software describes Heretic v2 as a decensored derivative with an OT-Ridge LoRA merged into the ternary language weights. The source publisher recommends research and experimentation, including alignment evaluation and red-teaming, and advises against public end-user service deployment.
This conversion carries the source model's reduced safety alignment. Outputs can be inaccurate, biased, offensive, or harmful and should be checked before use. The label “uncensored” does not establish universal refusal-free behavior. The publisher reports that vision was not evaluated for its release; this conversion adds only the simple image smoke check documented above. A comprehensive independent evaluation of the converted Heretic model has not been performed.
Credits and license
- OS-Software: Heretic v2 language-weight modification and source GGUF publication.
- Prism ML: Bonsai 2 ternary model, Hadamard-aware representation, original MLX vision pack, and compatible runtime work. Created using Bonsai by Prism ML.
- Qwen / Alibaba Cloud: the underlying Qwen model lineage, identified by the retained upstream notice as Qwen3.8-27B.
- p-e-w: Heretic, acknowledged by the source publisher.
- Bonsai demo: conversion and runtime reference project.
The model is distributed under Apache 2.0. The original LICENSE and attribution notice are retained. Changes made for this distribution are the GGUF-to-MLX repack, native schema 2 packaging, reconstructed SSM A_log tensors, publication documentation and provenance metadata, and the public-model-ID server helper. The source-model use guidance above is attributed to its publisher; the governing redistribution terms are in LICENSE.