license: other
license_name: polyform-small-business-1.0.0
license_link: LICENSE
base_model: IstroSec/ThinkingCap-Qwen3.8-27B-abliterated
base_model_relation: quantized
library_name: gguf
pipeline_tag: image-text-to-text
language:
- en
- multilingual
tags: - gguf
- qwen3_5
- qwen3_8
- thinkingcap
- abliterated
- imatrix
- mtp
ThinkingCap-Qwen3.8-27B-abliterated - Unsloth UD-Q4_K_XL tensor layout
GGUF quantization of IstroSec/ThinkingCap-Qwen3.8-27B-abliterated, converted directly from its BF16 checkpoint at revision 5da2705c6daafed7dbb6c100783538e68f1c5659.
The main model reproduces the exact tensor names, shapes, and quantization types extracted from Unsloth's Qwen3.8-27B UD-Q4_K_XL GGUF. This is a custom derivative quantization, produced independently of Unsloth.
| Tensor type | Tensor count |
|---|---|
| F32 | 360 |
| Q5_K | 191 |
| Q8_0 | 110 |
| IQ4_XS | 70 |
| Q4_K | 69 |
| Q6_K | 56 |
| IQ4_NL | 6 |
| Q3_K | 3 |
| IQ3_S | 1 |
| Total | 866 |
The quantizer uses Unsloth's published imatrix_unsloth.gguf at revision 4ca720788d1e01f1bff70c033e0d0028fd02e502. These calibration statistics come from the base model; they have not been recalculated on the ThinkingCap derivative. Matching the tensor layout does not establish identical numerical quantization or measured quality.
After converting the source to BF16 GGUF with its embedded MTP head retained, the main quantization command is:
llama-quantize --imatrix imatrix_unsloth.gguf \
--tensor-type-file tensor-types.txt \
ThinkingCap-Qwen3.8-27B-abliterated-BF16.gguf \
ThinkingCap-Qwen3.8-27B-abliterated-UD-Q4_K_XL.gguf Q4_K_M 12
The per-tensor overrides define the mixed layout; Q4_K_M is the fallback setting.
Files
ThinkingCap-Qwen3.8-27B-abliterated-UD-Q4_K_XL.gguf: main language model.mtp-ThinkingCap-Qwen3.8-27B-abliterated-Unsloth-layout.gguf: MTP draft converted from this checkpoint, using the exact 18-tensor layout of Unsloth's published MTP GGUF. Its filename avoids implying uniform Q4_0: the reference draft actually mixes Q3_K, Q4_K, Q6_K, and F32.mmproj-ThinkingCap-Qwen3.8-27B-abliterated-BF16.gguf: vision projector converted from this checkpoint.chat_template.jinja: the source checkpoint's native chat template.tensor-distribution.json,tensor-distribution.csv, andtensor-types.txt: complete main-model tensor distribution and reproducible quantizer overrides.mtp-tensor-distribution.jsonandmtp-tensor-types.txt: corresponding MTP recipe.verification-main.jsonandverification-mtp.json: checks of tensor names, shapes, and quantization types against the references.provenance.jsonandSHA256SUMS: source revisions, tool versions, and artifact checksums.
Usage
Use a recent llama.cpp build supporting Qwen3.5/Qwen3.8 and separate MTP drafts. For text-only serving:
llama-server \
-m ThinkingCap-Qwen3.8-27B-abliterated-UD-Q4_K_XL.gguf \
--jinja --chat-template-file chat_template.jinja \
--flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 \
--spec-type draft-mtp --spec-draft-n-max 2 \
--spec-draft-model mtp-ThinkingCap-Qwen3.8-27B-abliterated-Unsloth-layout.gguf
Add --mmproj mmproj-ThinkingCap-Qwen3.8-27B-abliterated-BF16.gguf for vision. Set GPU offload and context size to fit your hardware. The template supports enable_thinking and the low, medium, and xhigh reasoning efforts.
License and attribution
Retain LICENSE and NOTICE. ThinkingCap is Copyright 2026 BottleCap AI and is subject to the PolyForm Small Business License 1.0.0, including the additional personal-use permission stated in LICENSE. Qwen upstream materials remain subject to Apache-2.0.
Credit: BottleCap AI for ThinkingCap; Alibaba Cloud / Qwen Team for Qwen; IstroSec and MuXodious for the abliterated checkpoint; Unsloth for the reference layout and importance matrix; and llama.cpp contributors for conversion, quantization, and inference tools. See the source model card for its upstream evaluation and limitations. This repository does not claim a new quality benchmark.