license: apache-2.0
library_name: ninfer
pipeline_tag: image-text-to-text
inference: false
base_model:
- huihui-ai/Huihui-Qwen3.8-27B-abliterated-GGUF
base_model_relation: quantized
language: - en
- zh
tags: - ninfer
- abliterated
- uncensored
- ternary
- pq2_0
- qwen3.8
- bonsai
- prismml
- git
- text-generation
Huihui Qwen3.8-27B Abliterated · Ternary-Bonsai — NInfer
Format repackaging of huihui-ai's abliterated Ternary-Bonsai GGUF
into a single-file .ninfer container for the NInfer engine.
The abliteration is not ours. huihui-ai ablated layers 22–52 (0-based) of the
PrismML Ternary-Bonsai-2-27B weights; this repository only repackages them.
See Attribution and the usage warnings carried over below.
Why this tier needed re-quantization
huihui-ai's published Ternary-Bonsai-PQ2_0.gguf has 54 tensors that lost their
ternary encoding (ffn_down.weight ×31 and ssm_out.weight ×23 became Q3_K/Q2_K)
— they had to de-quantize those tensors to ablate them. A mixed ternary/k-quant file
cannot enter the .ninfer container, whose runtime only reads t2_g128_fp16.
This artifact was therefore built from huihui-ai's f16 release (which retains the
full ternary layout and the complete PrismML Hadamard metadata) and re-quantized toPQ2_0 with PrismML's own quantizer:
llama-quantize --tensor-type-file ssm_bf16.txt \
--token-embedding-type PQ2_0 --output-tensor-type PQ2_0 \
huihui-bonsai-f16.gguf out.gguf PQ2_0
Result: PQ2_0 ×402 + F32 ×353 + BF16 ×96 — the same type breakdown as
PrismML's shipped PQ2_0 file. The ssm_alpha/ssm_beta controls stay BF16
because the runtime requires full precision there.
File
| File | Size | SHA-256 |
|---|---|---|
Huihui-Qwen3.8-27B-abliterated-Ternary-Bonsai-PQ2_0.ninfer |
7,381,727,744 B (6.87 GiB) | 0738cbb2b6a6…cb67d95e |
SHA256SUMS carries the full digest.
One complete container: text tower (64 layers, 4:1 linear/full attention mix),
proposal head, tokenizer, chat template and generation config.
No MTP head and no vision tower — the Ternary-Bonsai source has neither.
Engine
Needs a NInfer engine build that reads the ternary t2_g128_fp16 representation with
the PrismML Hadamard (normalized Sylvester–Walsh) rotation. This is a different
dialect from the gguf_blocks_v1 artifacts — the two are not interchangeable.
Converted with the bonsai2_27b_ternary recipe from
Ryan-gsq/ninfer-16g-5070ti-5080-5090-qwen3.8-27b-gsq-rco
(Apache-2.0, forked from iamwavecut/ninfer-all).
Attribution
Qwen3.8-27B (base) Qwen/Qwen3.8-27B Apache-2.0
Ternary-Bonsai-2-27B (ternary) prism-ml/Ternary-Bonsai-2-27B-gguf (see its LICENSE)
Abliteration (layers 22–52) huihui-ai/Huihui-Qwen3.8-…-GGUF Apache-2.0
re-quantized to PQ2_0 from its f16 release
This repo .ninfer repackaging
See NOTICE for the full chain and the exact conversion steps.
Usage warnings (carried over from the abliterated source)
The source model's safety filtering has been significantly reduced. Outputs may be
sensitive, controversial or inappropriate. Not suitable for public-facing settings,
for use by minors, or for applications requiring high safety. Users are solely
responsible for compliance with applicable law and for reviewing generated output.
Recommended for research, testing and controlled environments only.
Intended use
Research and local inference. Not validated for production or safety-critical use.