license: apache-2.0
library_name: mlx
pipeline_tag: image-text-to-text
base_model:
- 3hs4n/Ternary-Bonsai-2-27B-Uncensored-Heretic-v2-mlx-2bit
tags: - mlx
- mlx-vlm
- ternary
- 2-bit
- bonsai
- heretic
- uncensored
- conversational
- reasoning
- thinking
- vision
- image-text-to-text
- tool-calling
- apple-silicon
Ternary Bonsai 2 27B Uncensored Heretic v2: Compact MLX 2-bit
A smaller native MLX variant of the original Heretic v2 MLX pack. All 2,057 language tensors are copied byte for byte. The size reduction comes from quantizing the 27 vision MLP expansion layers (linear_fc1) from FP16 to native MLX affine 2-bit/group-128 storage. Vision attention, MLP contraction layers, the merger, positional embeddings and all other vision tensors remain unchanged.
The compact weights occupy 8,365,392,904 bytes, approximately 8.37 GB (7.79 GiB). This saves 230,085,488 bytes, approximately 230 MB or 2.68%, compared with the original weight file. Configuration and tokenizer files add a small amount to the complete download.
Vision quantization is lossy. Image understanding, screenshots and OCR can change. Use the original pack with its FP16 vision tower when preserving that vision representation matters. The language checkpoint, thinking template and tool-call format are retained.
What changes
| Item | Original MLX pack | Compact MLX pack |
|---|---|---|
| Weight file | 8.60 GB | 8.37 GB |
| Language tensors | Native affine 2-bit/group-128 | Identical tensors |
| Vision linear layers | FP16 | 27 MLP expansion layers use native affine 2-bit/group-128 |
| Remaining vision tensors | Original precision | Identical tensors |
| Architecture | prism_hadamard_qwen35, schema 2 |
Same architecture |
| Thinking and tool calling | Supported by the language model and template | Retained |
| Configured context limit | 262,144 tokens | Same limit; full-length operation was not tested |
This is a real change to native inference weights, rather than a compressed download archive. A compatible runtime keeps the quantized vision weights packed, reducing the weight allocation as well as disk usage. Total RAM still includes the context cache, image activations, temporary buffers and the runtime. The weight-file size is not a minimum system-memory requirement, and the file-size saving is not a promise of an identical reduction in peak process memory.
This is not a PTQ1_0 MLX format. PQ2_0 and PTQ1_0 are GGUF packing formats. Both source conversions produced the same original MLX language weights. This compact variant gains space by changing vision precision, while retaining native MLX 2-bit language storage.
LM Studio and thinking
Open the published LM Studio catalog, click Use Model in LM Studio, then choose Compact or the original pack in Download Options. Both options were verified in LM Studio 0.4.26+4. The catalog declares Thinking, Vision and Tool Use and provides a working Enable Thinking on/off control. The portable catalog definition is included for reference.
With the LM Studio CLI installed, you can also select a variant with:
lms get ehs4n/ternary-bonsai-2-27b-uncensored-heretic-v2-mlx --select
LM Studio 0.4.26 global search combines Hugging Face results with Staff Picks and does not index every personal Hub entry. Use the catalog link or CLI command above to reach this entry. A raw Hugging Face repository still offers one native MLX download and displays only Vision and Tool Use badges in that view; the missing Thinking badge is a display limitation. The catalog is under LM Studio account ehs4n, while both weight repositories remain under Hugging Face account 3hs4n. Keep the virtual catalog definition separate from the concrete weight folder.
Use LM Studio on Apple Silicon with an MLX runtime that supports native prism_hadamard_qwen35 schema 2 and its Hadamard transforms. A generic MLX loader that omits the transforms can produce incorrect output.
This remains a thinking model. The language tensors and thinking chat template are unchanged. A model can generate thinking tokens even when a model browser does not display a Thinking badge. LM Studio's capability badges and download-option grouping depend on its own model catalog and runtime metadata; the Hugging Face thinking and reasoning tags do not guarantee either UI behavior.
For Python serving, thinking can be enabled with the command below and selected per request with enable_thinking. The short CLI example deliberately uses the native CLI default of thinking disabled. Reasoning quality has not been benchmarked for this compact release.
The final Compact pack loaded successfully in LM Studio 0.4.26+4 with its installed native MLX backend. Catalog thinking on/off, an OpenAI native tool call, a mock-result tool round trip, and OCR of the public 32-pixel-font fixture all passed. See LM Studio verification.
Run with Python
Use native mlx-vlm 0.7.2 with MLX 0.32.2. Keep the weight, configuration, tokenizer, processor and hadamard.json files together. The repository is intended for Apple Silicon Macs with sufficient unified memory and a macOS/Python build compatible with the runtime. It does not require a particular Mac model or machine-specific port.
Create an environment, install the included dependencies and download the complete repository:
python3.11 -m venv .venv-heretic
source .venv-heretic/bin/activate
python -m pip install huggingface-hub==1.33.0
hf download 3hs4n/Ternary-Bonsai-2-27B-Uncensored-Heretic-v2-mlx-2bit-compact \
--local-dir bonsai-heretic-mlx-compact
cd bonsai-heretic-mlx-compact
python -m pip install -r requirements.txt
python -m mlx_vlm.generate \
--model . \
--prompt "What is the capital of France? Answer with only the city name." \
--max-tokens 64 \
--temperature 0
The pinned dependencies include:
mlx==0.32.2
mlx-vlm==0.7.2
transformers==5.14.1
For an OpenAI-compatible API with a fixed public model ID:
python serve.py \
--model . \
--model-id Ternary-Bonsai-2-27B-Uncensored-Heretic-v2-mlx-2bit-compact \
--host 127.0.0.1 \
--enable-thinking
The default API address is http://127.0.0.1:8080/v1. Use --port to choose another free port if needed. Read /v1/models to obtain the public model ID. The helper uses the native mlx-vlm inference engine; LM Studio uses its own engine and does not need this Python server.
For a short request, run this in another terminal:
import json
from urllib.request import Request, urlopen
base = "http://127.0.0.1:8080/v1"
with urlopen(base + "/models") as response:
model_id = json.load(response)["data"][0]["id"]
payload = {
"model": model_id,
"messages": [{"role": "user", "content": "Explain ternary model weights briefly."}],
"max_tokens": 256,
"temperature": 0.7,
"enable_thinking": False,
"stream": False,
}
request = Request(
base + "/chat/completions",
data=json.dumps(payload).encode(),
headers={"Content-Type": "application/json"},
)
with urlopen(request) as response:
result = json.load(response)
print(result["choices"][0]["message"]["content"])
Image requests use OpenAI image_url content parts. Tool requests use an OpenAI tools array; your application executes a returned function and sends its result as a tool message with the matching tool_call_id.
Conversion and verification
The source is the original Heretic v2 MLX release at revision 866836d7b57012fa964aee462e60bf053e67084d. Its language weights originate from OS-Software's Heretic v2 GGUF release, and its vision tower originates from Prism ML's Bonsai 2 MLX pack.
| File | SHA256 |
|---|---|
Source model.safetensors |
e56ba66424066818c7c89f0f1ba74111705b9062367905228736013a48aaf08d |
Compact model.safetensors |
4d118dbab09ecf968b43088b9607cbeb3a0e646fb959ec17362b47544d55675b |
The compact file contains 2,444 tensors. Quantized vision layers add scale and bias tensors, accounting for the increase from the original 2,390 tensors. All 2,057 language tensors matched the source byte for byte, and 2,363 tensors in total were copied unchanged. See reproduction instructions, the included conversion/evaluation scripts, and compact-conversion.json for the pinned source, quantization method and per-layer error measurements.
On 2026-10-09, the original and conservative Compact packs each passed all 20 targeted synthetic checks: text, two bounded thinking cases, a native weather tool call and mock-result round trip, colors, spatial questions, counting, five OCR font sizes, and a bar chart. Their generated answers were identical on these fixtures. The loaded parameter allocation decreased from 8,595,174,880 to 8,365,083,040 bytes, saving 230,091,840 bytes (2.68%). MLX active allocations after clearing the cache decreased from 8,602,139,112 to 8,372,107,752 bytes. These counters exclude Python/process RSS and cached buffers; they are not total RAM requirements. See the comparison, baseline evaluation, Compact evaluation and independent tensor equivalence.
A more aggressive 83-layer vision prototype regressed on one OCR fixture and was not selected for publication. The published variant quantizes only the 27 MLP expansion layers.
The conversion checks establish unchanged language tensors and native quantization structure. Quantization error statistics do not establish image or OCR accuracy. Broader vision accuracy, long-context reliability and reasoning quality have not been benchmarked for this compact release. Functional evidence for the original pack is available in its model card; those results should not be treated as measurements of this quantized vision tower.
Portable metadata
Published configuration and conversion metadata use public repository IDs, pinned revisions, file basenames and hashes. Personal model paths, credentials, machine logs and local caches are excluded. No machine-specific launch tuning is prescribed.
The included serve.py exposes a fixed model ID and limits model discovery and selection. It replaces known model, home and working-directory paths in runtime metadata, response headers and error responses. Local filesystem paths are still required internally by the loader. This handling does not remove private content supplied in user prompts, generated answers or tool results, and it does not control LM Studio's own logs or API behavior.
Intended use, credits and license
OS-Software describes Heretic v2 as a decensored derivative with an OT-Ridge LoRA merged into the ternary language weights. Its publisher recommends research and experimentation and advises against public end-user service deployment. The compact conversion does not change the source language weights or restore safety alignment. Outputs can be inaccurate, biased, offensive or harmful; the word "uncensored" does not establish universal refusal-free behavior.
- OS-Software: Heretic v2 language-weight modification and source GGUF publication.
- Prism ML: Bonsai 2 ternary model, Hadamard-aware representation, original MLX vision pack and compatible runtime work. Created using Bonsai by Prism ML.
- Qwen / Alibaba Cloud: underlying Qwen lineage, identified by the retained notice as Qwen3.8-27B.
- p-e-w: Heretic, acknowledged by the source publisher.
- Bonsai demo: conversion and runtime reference project.
- 3hs4n: community native MLX packaging and this compact vision quantization.
Distributed under Apache 2.0. The original LICENSE and attribution in NOTICE.txt are retained. Changes for this variant are quantization of the 27 vision MLP expansion layers, matching native quantization configuration, and compact-release documentation and provenance. The original MLX release's packaging and server-helper changes are retained. Source-publisher use guidance is attributed above; redistribution terms are governed by the license.