license: apache-2.0
library_name: gguf
pipeline_tag: text-generation
base_model: BoldingBuilds/Ternary-Bonsai-2-27B-Abliterated-v2-PQ2_0-MTP-GGUF
tags:
- gguf
- llama.cpp
- mtp
- speculative-decoding
- abliterated
- bonsai
- ternary
- benchmark
Ternary Bonsai 2 27B Abliterated v2 (PQ2_0 + MTP): GGUF mirror + RTX 3060 12 GB benchmarks
This repo has two parts:
- A byte-identical mirror of
Ternary-Bonsai-2-27B-Abliterated-v2-PQ2_0-MTP.gguffrom
BoldingBuilds/Ternary-Bonsai-2-27B-Abliterated-v2-PQ2_0-MTP-GGUF.
We did not change the weights. For the model itself (how the abliteration and the thinking fix were made, and
refusal/capability evals), read BoldingBuilds' card. It is the authoritative source. - Reproducible benchmarks on one RTX 3060 12 GB (
bench/): a 30-task inbox-automation benchmark across 17
local model/runtime configs, a max-context search, speed/draft-acceptance numbers, the llama.cpp MTP patch we
built with, and every raw result file.
Read this before you use the numbers: the only tests we ran on this exact file (abliterated v2 + MTP) were
the speed and max-context tests. The 50-attempt quality table below has no row for this file. Its "B" rows
are a different abliterated build without MTP, and its "J" rows are the non-abliterated base Bonsai 2 with MTP. See
Honest notes.
Files
| path | size | what |
|---|---|---|
Ternary-Bonsai-2-27B-Abliterated-v2-PQ2_0-MTP.gguf |
7,657,489,696 bytes (7.66 GB) | mirror, unmodified |
SHA256SUMS |
a4e4c7b578131595c1694354bd6c74d00920df1f5082647a6126899753ebebf8 (matches BoldingBuilds' published hash) |
|
bench/ |
~5 MB | benchmark code, tasks, raw results, summary.md, MTP patch, build script, context-test script |
LICENSE, NOTICE |
Apache-2.0 text and upstream attribution notices |
PQ2_0 is PrismML's ternary packing. Mainline llama.cpp, Ollama and LM Studio can't load it. You need
PrismML's llama.cpp fork (see below).
How we ran it
Runtime: PrismML llama.cpp prism-b10685 + MTP patch
We built PrismML's llama.cpp fork from the prism-b10685 source and appliedbench/scripts/0001-qwen35-mtp-hadamard-inverse.patch.
Without the patch, loading a Bonsai MTP GGUF with --spec-type draft-mtp fails at startup withHadamard-latent table 'token_embd.weight' is read without the inverse transform, because the MTP draft graph does its
own embedding lookup and skips the inverse Hadamard transform. The patch adds that transform (about 14 lines insrc/models/qwen35.cpp).
- The patch is the one published by decent-jawfish/bonsai-2-27b-mtp,
byte for byte (SHA-256c10f6222…ce23b0). Its git header names MFEC AI Lab as the author. BoldingBuilds ships
a patch with the same code change. - Build:
bench/scripts/build_prismml_mtp.sh(CUDA 12.8,-DCMAKE_CUDA_ARCHITECTURES=86
for the RTX 3060,-DLLAMA_CURL=OFF, targetllama-server). - According to BoldingBuilds, PrismML tag
prism-b10743-adfffbeor newer runs MTP files without the patch. We did
not test that tag.
git clone https://github.com/PrismML-Eng/llama.cpp && cd llama.cpp
git checkout prism-b10685-7dffb15 # the tag decent-jawfish lists for this patch
git apply /path/to/bench/scripts/0001-qwen35-mtp-hadamard-inverse.patch
cmake -B build -DGGML_CUDA=ON -DCMAKE_CUDA_ARCHITECTURES=86 -DLLAMA_CURL=OFF
cmake --build build --target llama-server -j
Launch command (RTX 3060 12 GB, fully on GPU)
This command is our production configuration with this file's largest safe context on the shared 3060
(see Max context):
./llama-server -m Ternary-Bonsai-2-27B-Abliterated-v2-PQ2_0-MTP.gguf \
--host 0.0.0.0 --port 8801 \
-ngl 999 -fa on -c 26624 -np 1 \
-ctk q8_0 -ctv q8_0 -ctkd q8_0 -ctvd q8_0 -fit off \
--jinja \
--temp 1.0 --top-p 0.95 --top-k 20 --min-p 0.05 \
--chat-template-kwargs '{"reasoning_effort": "medium"}' \
--spec-type draft-mtp --spec-draft-n-max 2
--spec-type draft-mtp --spec-draft-n-max 2: the built-in MTP head (tensorsblk.64.*) drafts 2 tokens per step.-ctk/-ctv q8_0quantises the target model's KV cache and-ctkd/-ctvd q8_0does the same for the draft's KV cache.-ngl 999 -fit offputs all layers on the GPU with no automatic fitting or CPU offload.-fa onturns on flash attention.- The sampling flags are the server defaults (Bonsai's recommended temp 1.0 / top-k 20 / top-p 0.95, plus llama.cpp's min-p 0.05).
The benchmark requests override them (see below).reasoning_effortis medium, which BoldingBuilds recommends for this build.
Results for this file (abliterated v2 + MTP)
Speed
Measured on the RTX 3060 12 GB at .54 while a vLLM embedder shared the GPU (~2.5 GB).
| metric | value |
|---|---|
| decode, short JSON answer | ~56 tok/s (55.7–56.2) |
decode with the context 90% full (23,977 prompt tokens, -c 26624) |
33.6 tok/s |
| prompt processing at 90% fill | 470 tok/s |
| MTP draft acceptance | 0.92 |
For comparison, the non-abliterated base Bonsai 2 MTP GGUF on the same card and binary gives 56.7–57.1 tok/s short, 33.8 tok/s at
90% of 28K, and 0.933 acceptance.
Max context (RTX 3060 12 GB)
Method (bench/scripts/ctx_test.py, raw data inbench/results/ctx_search_54.jsonl): launch llama-server fully on the GPU with the flags
above, send a synthetic prompt that fills about 90% of -c, and check that it completes. A setting counts only if at least 300 MiB
of VRAM stays free at peak. The 3060 was also running a vLLM embedder (2,548 MiB) during every run.
| model | GPU | -c |
loads | 90% prompt OK | min free at peak | counts |
|---|---|---|---|---|---|---|
| Abliterated v2 PQ2_0 MTP (this file) | 3060, shared with embedder | 26,624 | yes | yes | 375 MiB | yes, max safe |
| Abliterated v2 PQ2_0 MTP (this file) | 3060, shared with embedder | 28,672 | yes | yes | 283 MiB | no (margin) |
| Base Bonsai 2 27B PQ2_0 MTP | 3060, shared with embedder | 28,672 | yes | yes | 379 MiB | yes, max safe |
| Base Bonsai 2 27B PQ2_0 MTP | 3060, shared with embedder | 30,720 / 32,768 | yes | yes | 287 / 195 MiB | no (margin) |
| Base Bonsai 2 27B PQ2_0 MTP | 3060, dedicated (Windows, no embedder) | 77,824 | yes | in production | n/a | runs as the production config |
- Each extra 2,048 tokens of context costs about 46 MiB of VRAM (q8_0 KV, hybrid attention).
- This file is about 100 MB larger than the base MTP GGUF, so on the shared card it fits one 2K step less (26,624 vs 28,672).
- We did not measure this file on the dedicated 3060. The 77,824 figure is for the base Bonsai 2 MTP GGUF.
The 30-task inbox-automation benchmark
The question was which local model should handle email triage, extraction and drafting on our home hardware.
Tasks (prompts and graders in bench/code/): 30 unique tasks in three suites.
- easy (10): classify, priority, bill extraction, phishing, summarise, draft reply, batch classify, action items, Spanish, rule-driven tool call
- hard (10): final meeting time, conflicting bills, GST arithmetic, subtle phishing, rule engine, constrained Spanish reply, relative dates, strict schema, prompt injection, mixed-language batch
- ultra (10): GST + FX reconciliation, timezone/DST reschedule, phishing headers batch, 12-rule engine, action-items schema, injection via tool calls, Rioplatense constrained reply, invoice reconciliation, scheduling puzzle, multilingual taxonomy
Protocol
- Every config ran with thinking on (
reasoning_effort: mediumfor Bonsai,max_tokens3000) and thinking off
(chat_template_kwargs.enable_thinking=false,max_tokens800). - Sampling per request: temperature 0.2, top_p 0.95, top_k 20, seed 42. The hard and ultra suites were run a second time with
seed 7. Easy repeats at seed 42 gave identical outputs, so easy counts once.
That makes 10 + 2×10 + 2×10 = 50 attempts per model per thinking mode. - Strict grading: a truncated answer (hit
max_tokens) counts as a fail. Requests were sent one at a time, and wall time was measured at the client. - Winner: J54, the non-abliterated base Bonsai 2 27B PQ2_0 + MTP on the 3060. It scored 96% (48/50) with thinking on, 15.8 s median and 43.3 s max per task, ~52 tok/s.
Final comparison table (all 50 attempts, strict grading)
From bench/summary.md (updated 2026-09-28 03:35 NZDT). ".54" is the Linux RTX 3060 12 GB (shared with the embedder),
"Win" is a dedicated RTX 3060 12 GB, and "Mac" is an Apple M5 with 24 GB.
| Model | Mode | Pass % (strict) | easy | hard | ultra | Median s | Max s | >120 s | Decode tok/s | Trunc | VRAM / memory |
|---|---|---|---|---|---|---|---|---|---|---|---|
| A Swift-1.5 27B IQ3_S (Win) | off | 66% (33/50) | 9/10 | 16/20 | 8/20 | 8.9 | 26.4 | 0 | 14.4 | 0 | 12070 MiB VRAM (218 free) |
| A Swift-1.5 27B IQ3_S (Win) | on | 94% (47/50) | 10/10 | 19/20 | 18/20 | 51.2 | 210.1 | 6 | 14.6 | 1 | 12070 MiB VRAM (218 free) |
| B Bonsai-2-27B-Abl PQ2_0 (.54) | off | 72% (36/50) | 10/10 | 14/20 | 12/20 | 4.8 | 15.5 | 0 | 33.6 | 0 | 10135 MiB VRAM (2153 free) |
| B Bonsai-2-27B-Abl PQ2_0 (.54) | on | 94% (47/50) | 10/10 | 17/20 | 20/20 | 26.6 | 68.2 | 0 | 33.6 | 0 | 10135 MiB VRAM (2153 free) |
| C Qwen3.8-27B GSQ-RCO IQ3_S (Win) | off | 68% (34/50) | 10/10 | 16/20 | 8/20 | 9.4 | 55.6 | 0 | 14.8 | 2 | 12070 MiB VRAM (218 free) |
| C Qwen3.8-27B GSQ-RCO IQ3_S (Win) | on | 90% (45/50) | 10/10 | 19/20 | 16/20 | 59.9 | 207.4 | 9 | 14.9 | 5 | 12070 MiB VRAM (218 free) |
| D Gemma4-12B Q4_K_M+MTP (.54) | off | 66% (33/50) | 9/10 | 16/20 | 8/20 | 2.4 | 6.3 | 0 | 62.0 | 0 | 11187 MiB VRAM (1101 free) |
| D Gemma4-12B Q4_K_M+MTP (.54) | on | 76% (38/50) | 8/10 | 16/20 | 14/20 | 8.0 | 40.6 | 0 | 51.6 | 1 | 11187 MiB VRAM (1101 free) |
| E Xing4.0-29B-A4B IQ3_XXS (.54) | off | 36% (18/50) | 9/10 | 7/20 | 2/20 | 7.3 | 39.3 | 0 | 21.1 | 1 | 10921 MiB VRAM (1367 free) |
| E Xing4.0-29B-A4B IQ3_XXS (.54) | on | 32% (16/50) | 9/10 | 6/20 | 1/20 | 69.0 | 157.1 | 17 | 20.7 | 13 | 10921 MiB VRAM (1367 free) |
| F MiMo-9B MLX 4bit (Mac) | off | 52% (26/50) | 10/10 | 14/20 | 2/20 | 6.6 | 18.8 | 0 | 18.7 | 0 | ~7.2-9.4 GB unified (Mac) |
| F MiMo-9B MLX 4bit (Mac) | on | 52% (26/50) | 9/10 | 14/20 | 3/20 | 6.7 | 176.2 | 3 | 20.6 | 3 | ~7.2-9.4 GB unified (Mac) |
| F MiMo-9B MLX 4bit (Mac) | on_forced | 64% (32/50) | 9/10 | 16/20 | 7/20 | 17.4 | 162.9 | 2 | 18.9 | 2 | ~7.2-9.4 GB unified (Mac) |
| G Hermes-4-14B-OBL EXL3 4.0 (.54) | off | 8% (4/50) | 8/10 | 13/20 | 4/20 | 25.5 | 32.1 | 0 | 32.8 | 46 | 10669 MiB VRAM (1619 free; Q6 cache, chunk 512) |
| G Hermes-4-14B-OBL EXL3 4.0 (.54) | on | 10% (5/50) | 8/10 | 10/20 | 5/20 | 94.4 | 106.5 | 0 | 13.1 | 40 | 10669 MiB VRAM (1619 free; Q6 cache, chunk 512) |
| H Bonsai-2-27B MLX 2bit (Mac) | off | 62% (31/50) | 10/10 | 13/20 | 8/20 | 14.9 | 50.4 | 0 | 9.9 | 0 | 9966 MB footprint (Mac) |
| H Bonsai-2-27B MLX 2bit (Mac) | on | 94% (47/50) | 10/10 | 18/20 | 19/20 | 87.0 | 260.2 | 16 | 9.8 | 0 | 9966 MB footprint (Mac) |
| I Bonsai-2-27B-Abl PQ2_0 (Mac) | off | 72% (36/50) | 10/10 | 14/20 | 12/20 | 12.0 | 51.3 | 0 | 11.6 | 0 | ~10.0 GB footprint (Mac) |
| I Bonsai-2-27B-Abl PQ2_0 (Mac) | on | 90% (45/50) | 10/10 | 16/20 | 19/20 | 81.7 | 246.7 | 17 | 11.0 | 0 | ~10.0 GB footprint (Mac) |
| J54 Bonsai-2-27B PQ2_0+MTP (.54) | off | 66% (33/50) | 10/10 | 14/20 | 9/20 | 3.8 | 9.7 | 0 | 53.0 | 0 | 10963 MiB VRAM (1325 free) |
| J54 Bonsai-2-27B PQ2_0+MTP (.54) | on | 96% (48/50) | 10/10 | 18/20 | 20/20 | 15.8 | 43.3 | 0 | 52.0 | 0 | 10963 MiB VRAM (1325 free) |
| JMac Bonsai-2-27B PQ2_0+MTP (Mac) | off | 66% (33/50) | 10/10 | 14/20 | 9/20 | 13.1 | 50.8 | 0 | 12.1 | 0 | ~10.0 GB footprint (Mac) |
| JMac Bonsai-2-27B PQ2_0+MTP (Mac) | on | 96% (48/50) | 10/10 | 19/20 | 19/20 | 67.1 | 172.6 | 10 | 11.3 | 0 | ~10.0 GB footprint (Mac) |
| L ThinkingCap-Qwen3.8-27B IQ3_S (Mac) | off | 68% (34/50) | 10/10 | 16/20 | 8/20 | 23.1 | 134.2 | 2 | 6.4 | 2 | 12.6 GB weights, footprint ~9.9 GB + paging (Mac) |
| L ThinkingCap-Qwen3.8-27B IQ3_S (Mac) | on | 92% (46/50) | 10/10 | 18/20 | 18/20 | 81.9 | 310.9 | 14 | 6.2 | 0 | 12.6 GB weights, footprint ~9.9 GB + paging (Mac) |
| M ZDTaichu5.0-9B MLX 8bit (Mac, text-only) | off | 46% (23/50) | 10/10 | 11/20 | 2/20 | 6.7 | 33.2 | 0 | 13.5 | 0 | 12 GB footprint peak (Mac) |
| M ZDTaichu5.0-9B MLX 8bit (Mac, text-only) | on | 82% (41/50) | 10/10 | 17/20 | 14/20 | 65.0 | 227.1 | 10 | 13.3 | 1 | 12 GB footprint peak (Mac) |
| K54 Bonsai-2 PQ2_0+DFlash2 (.54, draft on CPU) | off | 66% (33/50) | 10/10 | 14/20 | 9/20 | 8.5 | 31.5 | 0 | 14.8 | 0 | 11203 MiB VRAM (1085 free; draft on CPU) |
| K54 Bonsai-2 PQ2_0+DFlash2 (.54, draft on CPU) | on | 96% (48/50) | 10/10 | 19/20 | 19/20 | 52.0 | 243.7 | 7 | 14.1 | 1 | 11203 MiB VRAM (1085 free; draft on CPU) |
| KMac Bonsai-2 PQ2_0+DFlash2 (Mac) | off | 66% (33/50) | 10/10 | 14/20 | 9/20 | 13.1 | 57.4 | 0 | 12.0 | 0 | ~9.3 GB weights, footprint ~8.5-10 GB (Mac) |
| KMac Bonsai-2 PQ2_0+DFlash2 (Mac) | on | 98% (49/50) | 10/10 | 19/20 | 20/20 | 71.6 | 172.2 | 11 | 11.6 | 0 | ~9.3 GB weights, footprint ~8.5-10 GB (Mac) |
| N54 HauhauCS Qwen3.8-27B IQ3_XS+MTP (.54, partial offload) | off | 60% (30/50) | 10/10 | 12/20 | 8/20 | 48.4 | 184.8 | 8 | 2.0 | 0 | 11221 MiB VRAM (1067 free; -ngl 44/65, rest on CPU) |
| NMac HauhauCS Qwen3.8-27B IQ3_XS+MTP (Mac) | off | 62% (31/50) | 10/10 | 13/20 | 8/20 | 17.6 | 60.7 | 0 | 7.6 | 0 | Mac RSS 12.8 GB (footprint 2.1 GB, weights mmap) |
| NMac HauhauCS Qwen3.8-27B IQ3_XS+MTP (Mac) | on | 86% (43/50) | 10/10 | 18/20 | 15/20 | 135.5 | 467.6 | 27 | 6.8 | 5 | Mac RSS 12.8 GB (footprint 2.1 GB, weights mmap) |
Full per-task detail, the repeat runs, and caveats are in bench/summary.md. The bench README
(bench/README.md) has the hardware list and the exact J launch command.
Honest notes
- No quality run for this file. This abliterated v2 MTP GGUF got the speed and max-context tests above, not the 50-attempt
benchmark. - "B" / "I" rows are
Ternary-Bonsai-2-27B-Abliterated-PQ2_0.gguf: a different abliterated Bonsai 2 build with no MTP
(B on the .54 3060, I on the Mac). They are not this file. - "J54" / "JMac" rows are the non-abliterated base Bonsai 2 27B PQ2_0 with a grafted MTP head (
Bonsai-2-27B-PQ2_0-MTP.gguf
from decent-jawfish/bonsai-2-27b-mtp). They are also not this file. - Speculative decoding checks every draft token against the target model, but batched verification can change floating-point
reduction order. BoldingBuilds reports that this occasionally flips near-ties. - The shared-GPU context limits depend on the ~2.5 GB embedder that was resident on that card. On a dedicated 12 GB card you get
much more (see the 77,824 base-model figure). - Hostnames and IPs are stripped from the code.
bench/code/bench.pyreadsWIN_HOST,LINUX54_HOSTandMAC_HOSTfrom the environment.
Reproduce
cd bench/code
LINUX54_HOST=<host> python3 bench.py --model J --modes off,on --suite tasks_hard --out results_hard_J.jsonl
# --suite: tasks | tasks_hard | tasks_ultra ; --base overrides the endpoint URL
python3 summarize.py && python3 combined.py # read results*.jsonl from the working dir
python3 ../scripts/ctx_test.py MODEL.gguf 26624 8899 ctx.jsonl # max-context probe (edit BIN inside first)
Safety
This is an abliterated model: its refusal behaviour was removed, and it will comply with requests a stock model declines. As
BoldingBuilds says, you are responsible for how you use it. Don't deploy it in a user-facing product without your own
safety layer.
Credits
Every contributor below was checked against the upstream model cards.
- Prism ML: Ternary Bonsai 2 27B (
prism-ml/Ternary-Bonsai-2-27B-gguf, Apache-2.0),
the PQ2_0 format, and the PrismML-Eng/llama.cpp fork (MIT) we ran on. Created using Bonsai by Prism ML. - Qwen team, Alibaba Cloud:
Qwen/Qwen3.8-27B(Apache-2.0),
the base Bonsai 2 is derived from and the source of the MTP head. - Unsloth:
unsloth/Qwen3.8-27B-GGUF. BoldingBuilds copied the MTP head
tensors from it. - BoldingBuilds: the abliteration (refusal edit plus the one-row end-of-thinking-token fix) and the MTP graft that make up this GGUF
(source repo). BoldingBuilds' card presents the abliteration as
their own work. Other abliterated builds (Hikari07jp, dealignai, OS-Software, Blackfrost) appear there only as comparisons. - MTP recipe for Bonsai 2, as credited by BoldingBuilds: decent-jawfish and
ProCreations, building on sudoingX/qwen38-mtp. - MTP llama.cpp patch (
bench/scripts/0001-qwen35-mtp-hadamard-inverse.patch): published by decent-jawfish. Its git header names MFEC AI Lab as the author. - ggml-org/llama.cpp (MIT) and its contributors: the inference engine, GGUF, and the speculative-decoding
(draft-mtp) support the PrismML fork builds on. - Benchmark, context tests and this mirror: groxaxo (Facundo).
If you use Bonsai 2 27B, Prism ML asks you to cite:
@techreport{bonsai2_27b,
title = {Bonsai 2 27B: A 27B Ternary Reasoning Model},
author = {Prism ML},
year = {2026},
month = {September},
url = {https://prismml.com}
}
License
The weights are Apache-2.0, as are all upstream weights (Prism ML Bonsai 2, Qwen3.8-27B, Unsloth GGUF, BoldingBuilds' derivative).
The weights are redistributed unmodified, with LICENSE and NOTICE included. The llama.cpp runtime is MIT-licensed; the included patch is a small change against it, published by decent-jawfish.
The benchmark code and results in bench/ are released under Apache-2.0.
Not affiliated with or endorsed by Prism ML, Qwen/Alibaba Cloud, Unsloth, or BoldingBuilds.