← back to catalog · registered 2026-09-27 22:57

groxaxo/Ternary-Bonsai-2-27B-Abliterated-v2-PQ2_0-MTP-GGUF-3060-bench

groxaxo 27B GGUF second-order
curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/groxaxo%2FTernary-Bonsai-2-27B-Abliterated-v2-PQ2_0-MTP-GGUF-3060-bench"
Response includes
  • classification unknown
  • files 6
  • author_summary 24 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
?
Primary method

Unclassified

No clear signals of an abliteration technique in this model.
Confidence
UNKNOWN
Why this label 1 signal
No classification signals present. This may not be an abliterated model at all - it could be a repackaging, a merge with unrelated goals, or unrelated content that mentions the term.
  • no classification signals present (no abliterated, uncensored, or known producer/method markers)
Refusal direction extraction

No specific extraction method could be identified for this model. The producer either did not document it or used a proprietary pipeline.

What is a refusal direction? →
Downloads · 30-day
0
Likes
0
Model age
today
created 2026-09-27
Downloads over time
Now0→from0↑0%
00110 on Sep 270 on Sep 28Sep
Sep 27 → Sep 28 · 2 snapshots · spans 1 day

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
apache-2.0
Tags
gguf llama.cpp mtp speculative-decoding abliterated bonsai ternary benchmark text-generation base_model:BoldingBuilds/Ternary-Bonsai-2-27B-Abliterated-v2-PQ2_0-MTP-GGUF base_model:quantized:BoldingBuilds/Ternary-Bonsai-2-27B-Abliterated-v2-PQ2_0-MTP-GGUF license:apache-2.0

Related

Total size
7.13 GB
Files
6
Quantizations
1
Registered
2026-09-27 22:57
Last updated on HF
2026-09-27 22:13

Files by quantization

Auxiliary files 6 files 7.13 GB
Ternary-Bonsai-2-27B-Abliterated-v2-PQ2_0-MTP.gguf 7.13 GB a4e4c7b5 download
README.md 18.5 KB 56e27522 download
LICENSE 9.94 KB 66a27ec5 download
.gitattributes 1.57 KB 82f5e642 download
NOTICE 977 B b40e1ac6 download
SHA256SUMS 117 B 4d93f82f download

README current version from Hugging Face


license: apache-2.0
library_name: gguf
pipeline_tag: text-generation
base_model: BoldingBuilds/Ternary-Bonsai-2-27B-Abliterated-v2-PQ2_0-MTP-GGUF
tags:

  • gguf
  • llama.cpp
  • mtp
  • speculative-decoding
  • abliterated
  • bonsai
  • ternary
  • benchmark

Ternary Bonsai 2 27B Abliterated v2 (PQ2_0 + MTP): GGUF mirror + RTX 3060 12 GB benchmarks

This repo has two parts:

  1. A byte-identical mirror of Ternary-Bonsai-2-27B-Abliterated-v2-PQ2_0-MTP.gguf from
    BoldingBuilds/Ternary-Bonsai-2-27B-Abliterated-v2-PQ2_0-MTP-GGUF.
    We did not change the weights. For the model itself (how the abliteration and the thinking fix were made, and
    refusal/capability evals), read BoldingBuilds' card. It is the authoritative source.
  2. Reproducible benchmarks on one RTX 3060 12 GB (bench/): a 30-task inbox-automation benchmark across 17
    local model/runtime configs, a max-context search, speed/draft-acceptance numbers, the llama.cpp MTP patch we
    built with, and every raw result file.

Read this before you use the numbers: the only tests we ran on this exact file (abliterated v2 + MTP) were
the speed and max-context tests. The 50-attempt quality table below has no row for this file. Its "B" rows
are a different abliterated build without MTP, and its "J" rows are the non-abliterated base Bonsai 2 with MTP. See
Honest notes.

Files

path size what
Ternary-Bonsai-2-27B-Abliterated-v2-PQ2_0-MTP.gguf 7,657,489,696 bytes (7.66 GB) mirror, unmodified
SHA256SUMS a4e4c7b578131595c1694354bd6c74d00920df1f5082647a6126899753ebebf8 (matches BoldingBuilds' published hash)
bench/ ~5 MB benchmark code, tasks, raw results, summary.md, MTP patch, build script, context-test script
LICENSE, NOTICE Apache-2.0 text and upstream attribution notices

PQ2_0 is PrismML's ternary packing. Mainline llama.cpp, Ollama and LM Studio can't load it. You need
PrismML's llama.cpp fork (see below).

How we ran it

Runtime: PrismML llama.cpp prism-b10685 + MTP patch

We built PrismML's llama.cpp fork from the prism-b10685 source and applied
bench/scripts/0001-qwen35-mtp-hadamard-inverse.patch.
Without the patch, loading a Bonsai MTP GGUF with --spec-type draft-mtp fails at startup with
Hadamard-latent table 'token_embd.weight' is read without the inverse transform, because the MTP draft graph does its
own embedding lookup and skips the inverse Hadamard transform. The patch adds that transform (about 14 lines in
src/models/qwen35.cpp).

  • The patch is the one published by decent-jawfish/bonsai-2-27b-mtp,
    byte for byte (SHA-256 c10f6222…ce23b0). Its git header names MFEC AI Lab as the author. BoldingBuilds ships
    a patch with the same code change.
  • Build: bench/scripts/build_prismml_mtp.sh (CUDA 12.8, -DCMAKE_CUDA_ARCHITECTURES=86
    for the RTX 3060, -DLLAMA_CURL=OFF, target llama-server).
  • According to BoldingBuilds, PrismML tag prism-b10743-adfffbe or newer runs MTP files without the patch. We did
    not test that tag.
git clone https://github.com/PrismML-Eng/llama.cpp && cd llama.cpp
git checkout prism-b10685-7dffb15        # the tag decent-jawfish lists for this patch
git apply /path/to/bench/scripts/0001-qwen35-mtp-hadamard-inverse.patch
cmake -B build -DGGML_CUDA=ON -DCMAKE_CUDA_ARCHITECTURES=86 -DLLAMA_CURL=OFF
cmake --build build --target llama-server -j

Launch command (RTX 3060 12 GB, fully on GPU)

This command is our production configuration with this file's largest safe context on the shared 3060
(see Max context):

./llama-server -m Ternary-Bonsai-2-27B-Abliterated-v2-PQ2_0-MTP.gguf \
  --host 0.0.0.0 --port 8801 \
  -ngl 999 -fa on -c 26624 -np 1 \
  -ctk q8_0 -ctv q8_0 -ctkd q8_0 -ctvd q8_0 -fit off \
  --jinja \
  --temp 1.0 --top-p 0.95 --top-k 20 --min-p 0.05 \
  --chat-template-kwargs '{"reasoning_effort": "medium"}' \
  --spec-type draft-mtp --spec-draft-n-max 2
  • --spec-type draft-mtp --spec-draft-n-max 2: the built-in MTP head (tensors blk.64.*) drafts 2 tokens per step.
  • -ctk/-ctv q8_0 quantises the target model's KV cache and -ctkd/-ctvd q8_0 does the same for the draft's KV cache.
  • -ngl 999 -fit off puts all layers on the GPU with no automatic fitting or CPU offload. -fa on turns on flash attention.
  • The sampling flags are the server defaults (Bonsai's recommended temp 1.0 / top-k 20 / top-p 0.95, plus llama.cpp's min-p 0.05).
    The benchmark requests override them (see below). reasoning_effort is medium, which BoldingBuilds recommends for this build.

Results for this file (abliterated v2 + MTP)

Speed

Measured on the RTX 3060 12 GB at .54 while a vLLM embedder shared the GPU (~2.5 GB).

metric value
decode, short JSON answer ~56 tok/s (55.7–56.2)
decode with the context 90% full (23,977 prompt tokens, -c 26624) 33.6 tok/s
prompt processing at 90% fill 470 tok/s
MTP draft acceptance 0.92

For comparison, the non-abliterated base Bonsai 2 MTP GGUF on the same card and binary gives 56.7–57.1 tok/s short, 33.8 tok/s at
90% of 28K, and 0.933 acceptance.

Max context (RTX 3060 12 GB)

Method (bench/scripts/ctx_test.py, raw data in
bench/results/ctx_search_54.jsonl): launch llama-server fully on the GPU with the flags
above, send a synthetic prompt that fills about 90% of -c, and check that it completes. A setting counts only if at least 300 MiB
of VRAM stays free at peak
. The 3060 was also running a vLLM embedder (2,548 MiB) during every run.

model GPU -c loads 90% prompt OK min free at peak counts
Abliterated v2 PQ2_0 MTP (this file) 3060, shared with embedder 26,624 yes yes 375 MiB yes, max safe
Abliterated v2 PQ2_0 MTP (this file) 3060, shared with embedder 28,672 yes yes 283 MiB no (margin)
Base Bonsai 2 27B PQ2_0 MTP 3060, shared with embedder 28,672 yes yes 379 MiB yes, max safe
Base Bonsai 2 27B PQ2_0 MTP 3060, shared with embedder 30,720 / 32,768 yes yes 287 / 195 MiB no (margin)
Base Bonsai 2 27B PQ2_0 MTP 3060, dedicated (Windows, no embedder) 77,824 yes in production n/a runs as the production config
  • Each extra 2,048 tokens of context costs about 46 MiB of VRAM (q8_0 KV, hybrid attention).
  • This file is about 100 MB larger than the base MTP GGUF, so on the shared card it fits one 2K step less (26,624 vs 28,672).
  • We did not measure this file on the dedicated 3060. The 77,824 figure is for the base Bonsai 2 MTP GGUF.

The 30-task inbox-automation benchmark

The question was which local model should handle email triage, extraction and drafting on our home hardware.

Tasks (prompts and graders in bench/code/): 30 unique tasks in three suites.

  • easy (10): classify, priority, bill extraction, phishing, summarise, draft reply, batch classify, action items, Spanish, rule-driven tool call
  • hard (10): final meeting time, conflicting bills, GST arithmetic, subtle phishing, rule engine, constrained Spanish reply, relative dates, strict schema, prompt injection, mixed-language batch
  • ultra (10): GST + FX reconciliation, timezone/DST reschedule, phishing headers batch, 12-rule engine, action-items schema, injection via tool calls, Rioplatense constrained reply, invoice reconciliation, scheduling puzzle, multilingual taxonomy

Protocol

  • Every config ran with thinking on (reasoning_effort: medium for Bonsai, max_tokens 3000) and thinking off
    (chat_template_kwargs.enable_thinking=false, max_tokens 800).
  • Sampling per request: temperature 0.2, top_p 0.95, top_k 20, seed 42. The hard and ultra suites were run a second time with
    seed 7. Easy repeats at seed 42 gave identical outputs, so easy counts once.
    That makes 10 + 2×10 + 2×10 = 50 attempts per model per thinking mode.
  • Strict grading: a truncated answer (hit max_tokens) counts as a fail. Requests were sent one at a time, and wall time was measured at the client.
  • Winner: J54, the non-abliterated base Bonsai 2 27B PQ2_0 + MTP on the 3060. It scored 96% (48/50) with thinking on, 15.8 s median and 43.3 s max per task, ~52 tok/s.

Final comparison table (all 50 attempts, strict grading)

From bench/summary.md (updated 2026-09-28 03:35 NZDT). ".54" is the Linux RTX 3060 12 GB (shared with the embedder),
"Win" is a dedicated RTX 3060 12 GB, and "Mac" is an Apple M5 with 24 GB.

Model Mode Pass % (strict) easy hard ultra Median s Max s >120 s Decode tok/s Trunc VRAM / memory
A Swift-1.5 27B IQ3_S (Win) off 66% (33/50) 9/10 16/20 8/20 8.9 26.4 0 14.4 0 12070 MiB VRAM (218 free)
A Swift-1.5 27B IQ3_S (Win) on 94% (47/50) 10/10 19/20 18/20 51.2 210.1 6 14.6 1 12070 MiB VRAM (218 free)
B Bonsai-2-27B-Abl PQ2_0 (.54) off 72% (36/50) 10/10 14/20 12/20 4.8 15.5 0 33.6 0 10135 MiB VRAM (2153 free)
B Bonsai-2-27B-Abl PQ2_0 (.54) on 94% (47/50) 10/10 17/20 20/20 26.6 68.2 0 33.6 0 10135 MiB VRAM (2153 free)
C Qwen3.8-27B GSQ-RCO IQ3_S (Win) off 68% (34/50) 10/10 16/20 8/20 9.4 55.6 0 14.8 2 12070 MiB VRAM (218 free)
C Qwen3.8-27B GSQ-RCO IQ3_S (Win) on 90% (45/50) 10/10 19/20 16/20 59.9 207.4 9 14.9 5 12070 MiB VRAM (218 free)
D Gemma4-12B Q4_K_M+MTP (.54) off 66% (33/50) 9/10 16/20 8/20 2.4 6.3 0 62.0 0 11187 MiB VRAM (1101 free)
D Gemma4-12B Q4_K_M+MTP (.54) on 76% (38/50) 8/10 16/20 14/20 8.0 40.6 0 51.6 1 11187 MiB VRAM (1101 free)
E Xing4.0-29B-A4B IQ3_XXS (.54) off 36% (18/50) 9/10 7/20 2/20 7.3 39.3 0 21.1 1 10921 MiB VRAM (1367 free)
E Xing4.0-29B-A4B IQ3_XXS (.54) on 32% (16/50) 9/10 6/20 1/20 69.0 157.1 17 20.7 13 10921 MiB VRAM (1367 free)
F MiMo-9B MLX 4bit (Mac) off 52% (26/50) 10/10 14/20 2/20 6.6 18.8 0 18.7 0 ~7.2-9.4 GB unified (Mac)
F MiMo-9B MLX 4bit (Mac) on 52% (26/50) 9/10 14/20 3/20 6.7 176.2 3 20.6 3 ~7.2-9.4 GB unified (Mac)
F MiMo-9B MLX 4bit (Mac) on_forced 64% (32/50) 9/10 16/20 7/20 17.4 162.9 2 18.9 2 ~7.2-9.4 GB unified (Mac)
G Hermes-4-14B-OBL EXL3 4.0 (.54) off 8% (4/50) 8/10 13/20 4/20 25.5 32.1 0 32.8 46 10669 MiB VRAM (1619 free; Q6 cache, chunk 512)
G Hermes-4-14B-OBL EXL3 4.0 (.54) on 10% (5/50) 8/10 10/20 5/20 94.4 106.5 0 13.1 40 10669 MiB VRAM (1619 free; Q6 cache, chunk 512)
H Bonsai-2-27B MLX 2bit (Mac) off 62% (31/50) 10/10 13/20 8/20 14.9 50.4 0 9.9 0 9966 MB footprint (Mac)
H Bonsai-2-27B MLX 2bit (Mac) on 94% (47/50) 10/10 18/20 19/20 87.0 260.2 16 9.8 0 9966 MB footprint (Mac)
I Bonsai-2-27B-Abl PQ2_0 (Mac) off 72% (36/50) 10/10 14/20 12/20 12.0 51.3 0 11.6 0 ~10.0 GB footprint (Mac)
I Bonsai-2-27B-Abl PQ2_0 (Mac) on 90% (45/50) 10/10 16/20 19/20 81.7 246.7 17 11.0 0 ~10.0 GB footprint (Mac)
J54 Bonsai-2-27B PQ2_0+MTP (.54) off 66% (33/50) 10/10 14/20 9/20 3.8 9.7 0 53.0 0 10963 MiB VRAM (1325 free)
J54 Bonsai-2-27B PQ2_0+MTP (.54) on 96% (48/50) 10/10 18/20 20/20 15.8 43.3 0 52.0 0 10963 MiB VRAM (1325 free)
JMac Bonsai-2-27B PQ2_0+MTP (Mac) off 66% (33/50) 10/10 14/20 9/20 13.1 50.8 0 12.1 0 ~10.0 GB footprint (Mac)
JMac Bonsai-2-27B PQ2_0+MTP (Mac) on 96% (48/50) 10/10 19/20 19/20 67.1 172.6 10 11.3 0 ~10.0 GB footprint (Mac)
L ThinkingCap-Qwen3.8-27B IQ3_S (Mac) off 68% (34/50) 10/10 16/20 8/20 23.1 134.2 2 6.4 2 12.6 GB weights, footprint ~9.9 GB + paging (Mac)
L ThinkingCap-Qwen3.8-27B IQ3_S (Mac) on 92% (46/50) 10/10 18/20 18/20 81.9 310.9 14 6.2 0 12.6 GB weights, footprint ~9.9 GB + paging (Mac)
M ZDTaichu5.0-9B MLX 8bit (Mac, text-only) off 46% (23/50) 10/10 11/20 2/20 6.7 33.2 0 13.5 0 12 GB footprint peak (Mac)
M ZDTaichu5.0-9B MLX 8bit (Mac, text-only) on 82% (41/50) 10/10 17/20 14/20 65.0 227.1 10 13.3 1 12 GB footprint peak (Mac)
K54 Bonsai-2 PQ2_0+DFlash2 (.54, draft on CPU) off 66% (33/50) 10/10 14/20 9/20 8.5 31.5 0 14.8 0 11203 MiB VRAM (1085 free; draft on CPU)
K54 Bonsai-2 PQ2_0+DFlash2 (.54, draft on CPU) on 96% (48/50) 10/10 19/20 19/20 52.0 243.7 7 14.1 1 11203 MiB VRAM (1085 free; draft on CPU)
KMac Bonsai-2 PQ2_0+DFlash2 (Mac) off 66% (33/50) 10/10 14/20 9/20 13.1 57.4 0 12.0 0 ~9.3 GB weights, footprint ~8.5-10 GB (Mac)
KMac Bonsai-2 PQ2_0+DFlash2 (Mac) on 98% (49/50) 10/10 19/20 20/20 71.6 172.2 11 11.6 0 ~9.3 GB weights, footprint ~8.5-10 GB (Mac)
N54 HauhauCS Qwen3.8-27B IQ3_XS+MTP (.54, partial offload) off 60% (30/50) 10/10 12/20 8/20 48.4 184.8 8 2.0 0 11221 MiB VRAM (1067 free; -ngl 44/65, rest on CPU)
NMac HauhauCS Qwen3.8-27B IQ3_XS+MTP (Mac) off 62% (31/50) 10/10 13/20 8/20 17.6 60.7 0 7.6 0 Mac RSS 12.8 GB (footprint 2.1 GB, weights mmap)
NMac HauhauCS Qwen3.8-27B IQ3_XS+MTP (Mac) on 86% (43/50) 10/10 18/20 15/20 135.5 467.6 27 6.8 5 Mac RSS 12.8 GB (footprint 2.1 GB, weights mmap)

Full per-task detail, the repeat runs, and caveats are in bench/summary.md. The bench README
(bench/README.md) has the hardware list and the exact J launch command.

Honest notes

  • No quality run for this file. This abliterated v2 MTP GGUF got the speed and max-context tests above, not the 50-attempt
    benchmark.
  • "B" / "I" rows are Ternary-Bonsai-2-27B-Abliterated-PQ2_0.gguf: a different abliterated Bonsai 2 build with no MTP
    (B on the .54 3060, I on the Mac). They are not this file.
  • "J54" / "JMac" rows are the non-abliterated base Bonsai 2 27B PQ2_0 with a grafted MTP head (Bonsai-2-27B-PQ2_0-MTP.gguf
    from decent-jawfish/bonsai-2-27b-mtp). They are also not this file.
  • Speculative decoding checks every draft token against the target model, but batched verification can change floating-point
    reduction order. BoldingBuilds reports that this occasionally flips near-ties.
  • The shared-GPU context limits depend on the ~2.5 GB embedder that was resident on that card. On a dedicated 12 GB card you get
    much more (see the 77,824 base-model figure).
  • Hostnames and IPs are stripped from the code. bench/code/bench.py reads WIN_HOST, LINUX54_HOST and MAC_HOST from the environment.

Reproduce

cd bench/code
LINUX54_HOST=<host> python3 bench.py --model J --modes off,on --suite tasks_hard --out results_hard_J.jsonl
# --suite: tasks | tasks_hard | tasks_ultra ; --base overrides the endpoint URL
python3 summarize.py && python3 combined.py      # read results*.jsonl from the working dir
python3 ../scripts/ctx_test.py MODEL.gguf 26624 8899 ctx.jsonl   # max-context probe (edit BIN inside first)

Safety

This is an abliterated model: its refusal behaviour was removed, and it will comply with requests a stock model declines. As
BoldingBuilds says, you are responsible for how you use it. Don't deploy it in a user-facing product without your own
safety layer.

Credits

Every contributor below was checked against the upstream model cards.

  • Prism ML: Ternary Bonsai 2 27B (prism-ml/Ternary-Bonsai-2-27B-gguf, Apache-2.0),
    the PQ2_0 format, and the PrismML-Eng/llama.cpp fork (MIT) we ran on. Created using Bonsai by Prism ML.
  • Qwen team, Alibaba Cloud: Qwen/Qwen3.8-27B (Apache-2.0),
    the base Bonsai 2 is derived from and the source of the MTP head.
  • Unsloth: unsloth/Qwen3.8-27B-GGUF. BoldingBuilds copied the MTP head
    tensors from it.
  • BoldingBuilds: the abliteration (refusal edit plus the one-row end-of-thinking-token fix) and the MTP graft that make up this GGUF
    (source repo). BoldingBuilds' card presents the abliteration as
    their own work. Other abliterated builds (Hikari07jp, dealignai, OS-Software, Blackfrost) appear there only as comparisons.
  • MTP recipe for Bonsai 2, as credited by BoldingBuilds: decent-jawfish and
    ProCreations, building on sudoingX/qwen38-mtp.
  • MTP llama.cpp patch (bench/scripts/0001-qwen35-mtp-hadamard-inverse.patch): published by decent-jawfish. Its git header names MFEC AI Lab as the author.
  • ggml-org/llama.cpp (MIT) and its contributors: the inference engine, GGUF, and the speculative-decoding
    (draft-mtp) support the PrismML fork builds on.
  • Benchmark, context tests and this mirror: groxaxo (Facundo).

If you use Bonsai 2 27B, Prism ML asks you to cite:

@techreport{bonsai2_27b,
    title   = {Bonsai 2 27B: A 27B Ternary Reasoning Model},
    author  = {Prism ML},
    year    = {2026},
    month   = {September},
    url     = {https://prismml.com}
}

License

The weights are Apache-2.0, as are all upstream weights (Prism ML Bonsai 2, Qwen3.8-27B, Unsloth GGUF, BoldingBuilds' derivative).
The weights are redistributed unmodified, with LICENSE and NOTICE included. The llama.cpp runtime is MIT-licensed; the included patch is a small change against it, published by decent-jawfish.
The benchmark code and results in bench/ are released under Apache-2.0.

Not affiliated with or endorsed by Prism ML, Qwen/Alibaba Cloud, Unsloth, or BoldingBuilds.

Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.