license: other
license_name: polyform-small-business-1.0.0
license_link: LICENSE
base_model: IstroSec/ThinkingCap-Qwen3.8-27B-abliterated
base_model_relation: quantized
library_name: transformers
pipeline_tag: image-text-to-text
language:
- en
- multilingual
tags: - qwen3_5
- qwen3_8
- thinkingcap
- abliterated
- uncensored
- int4
- w4a16
- awq
- gptq
- marlin
- compressed-tensors
- llmcompressor
- vllm
- mtp
- vision
ThinkingCap-Qwen3.8-27B-abliterated-W4A16
INT4 weight-only quantization (4-bit integer weights, bf16 activations) of ThinkingCap-Qwen3.8-27B-abliterated, the uncensored variant of bottlecapai/ThinkingCap-Qwen3.8-27B.
19.5 GB (from 55.6 GB bf16), KL 0.0260 against the bf16 model. vLLM runs it with the Marlin kernel on Ampere, Ada and Blackwell (A100, RTX 3090/4090/5090, L40S, DGX Spark) and with Machete on Hopper (H100/H200). The weights fit a 24 GB GPU with a short context; 32 GB and up is comfortable.
The NVFP4A16 build (20.6 GB, KL 0.0155) is also weight-only and runs on the same sm80+ GPUs (vLLM: Marlin FP4 on most of them). It has the lower KL (0.0155 vs 0.0260); this build is 1.1 GB smaller. Speed not measured. If you have the memory, the FP8-DYNAMIC build (31.3 GB, KL 0.0068) is closer to lossless.
What is quantized
Produced with llm-compressor 0.14.0, scheme W4A16 (preset; config.json: one config group, symmetric 4-bit integer weights, group size 128): INT4 weights in groups of 128 along the input dimension, one bf16 scale per group, symmetric, no zero points, stored as compressed-tensors pack-quantized (eight 4-bit values per int32). Activations stay bf16.
Algorithm: AWQ (llm-compressor 0.14.0 AWQModifier(duo_scaling="both", n_grid=20): 10 smoothing exponents from 0 to 1, each tried with activation-only scales and with duo scales (activations and weights), so 20 candidates per smoothing group, lowest output error wins) + GPTQ (dampening 0.01, static activation order), each with its own calibration pass. Calibration: 256 samples of up to 2,048 tokens (295,811 tokens in all) of chat, math, code, tool-use and coding-agent transcripts.
Calibration text: prompts from seven public datasets (listed under Credits), rendered in this model's chat template, with the responses generated by this model itself (its Q8_0 GGUF, temperature 1.0). Responses are capped at 1,024 tokens for chat, math and code and at 2,048 tokens for tool-use and coding-agent samples, so they hold the thinking and, when it fits, the answer. Some tool-use and coding-agent samples are cut at a later turn and keep the dataset's earlier turns (tool results, the original agent's messages) as context. A sample longer than 2,048 tokens keeps its last 2,048 tokens, starting at a turn boundary when one falls in the first quarter of that window, so it always ends with the model's own response (a response longer than the kept window loses its beginning). Calibration and evaluation text (below) use different rows: chat, math and code come from other dataset splits, and tool-use and coding-agent rows are held out by repository and by first question. Tool definitions and agent system prompts are shared between the calibration and evaluation sides.
Quantized: 400 Linears of the text decoder: the 64 MLPs and the 16 full-attention layers' q/k/v/o_proj at 4 bit. The three large Gated DeltaNet projections in the other 48 layers (in_proj_qkv, in_proj_z, out_proj) are quantized the same way.
Kept in bf16:
| Component | Why |
|---|---|
linear_attn.in_proj_a, in_proj_b |
output dimension 48, below the Marlin tile; older vLLM and SGLang kernels fail on it; negligible size |
visual.* |
vision tower and merger (linear_fc2 has 4304 inputs, not divisible by the group size) |
lm_head |
standard |
mtp.* |
MTP head, re-grafted from bf16 after quantization; its Linears are in quantization_config.ignore |
Fused layers. vLLM runs in_proj_qkv + in_proj_z, q/k/v_proj and gate/up_proj as fused GEMMs. INT4 scales belong to one output row and one group each, so the fused GEMMs need no shared scale (unlike the NVFP4 build's global scale); every fused set carries the same scheme, which is what vLLM requires. Checked before and after saving: 400 quantized layers, identical schemes within each fused set, finite positive scales, pack-quantized format, static activation order (no g_idx), and the ignore list above.
Evaluation
Heretic --evaluate-model against the bf16 abliterated checkpoint, non-thinking mode, 100 harmful + 100 harmless prompts:
| Size | Refusals (100 harmful) | KL vs. bf16 abliterated | |
|---|---|---|---|
| bf16 abliterated (source) | 55.6 GB | 6 / 100 | — |
| FP8-DYNAMIC | 31.3 GB | 7 / 100 | 0.0068 |
| NVFP4A16 | 20.6 GB | 7 / 100 | 0.0155 |
| W4A16 (this repo) | 19.5 GB | 8 / 100 | 0.0260 |
KL under 0.01 is imperceptible; under 0.05 is good. Heretic decodes greedily (no sampling noise); a re-run with a different batch size can still move a count by about one prompt, so differences of one or two are not significant. This build refuses 8/100, FP8 and NVFP4A16 7/100, the bf16 source 6/100.
Not published: an Intel AutoRound build of the same size (1000 iterations, 512 samples of up to 2,048 tokens of the same text) measured KL 0.0528 and 11/100 refusals.
Chat and agent text (a second measurement, not Heretic). Exact KL divergence from the bf16 model on held-out chat, math, code, tool-use and coding-agent transcripts (context 2048, 60 chunks), scored only at the 16,356 token positions of the model's own responses:
| Mean KLD, model's own responses | Same top-1 token | |
|---|---|---|
| FP8-DYNAMIC | 0.0055 | 97.5% |
| NVFP4A16 | 0.0161 | 96.1% |
| W4A16 (this repo) | 0.0295 | 94.7% |
This text comes from the same seven datasets as the calibration text (other questions and repositories; the same tool definitions and agent system prompts), so it suits calibrated builds; Heretic's prompts are independent of it and rank the builds the same way.
For reference, Red Hat's INT4 build of the base Qwen3.8-27B (RedHatAI/Qwen3.8-27B-INT4: AWQ + GPTQ, W4A16, group 128, DeltaNet quantized except in_proj_a/b, calibrated on 512 samples of up to 4096 tokens from open-perfectblend, with an FP8 KV-cache scheme in its config) reports about 99% accuracy recovery on IFEval, MMLU-Pro, GPQA-Diamond, AIME25 and MATH-500. That was measured on the base model with its own calibration and KV-cache setting, not on this build (AWQ + GPTQ as above, calibrated on 256 samples of up to 2,048 tokens of this model's own chat and agent transcripts; no KV-cache scheme).
Provenance: original ThinkingCap refuses 97/100 → bf16 abliteration 6/100 at KL 0.065 vs. the original → this quantization adds a further KL of 0.0260 relative to the abliterated model.
Usage
vLLM ≥ 0.24 (0.30 current). The format is detected from config.json, no --quantization flag.
Tested 2026-10-01 with vLLM 0.30.0 (vllm/vllm-openai:v0.30.0) on a DGX Spark (GB10; vLLM chose MarlinLinearKernel), text only (--language-model-only), --max-model-len 32768, MTP on: 3 short prompts (chat, code, a tool call) at temperature 0 gave correct answers and a well-formed get_weather tool call. Image input was not tested.
vllm serve IstroSec/ThinkingCap-Qwen3.8-27B-abliterated-W4A16 \
--language-model-only \
--reasoning-parser qwen3 \
--enable-auto-tool-choice --tool-call-parser qwen3_xml \
--speculative-config '{"method":"mtp","num_speculative_tokens":3}' \
--max-model-len 65536 --gpu-memory-utilization 0.85
- Kernels: Marlin on sm80 and newer (Ampere, Ada, Blackwell incl. GB10), Machete on Hopper. Minimum sm80: V100 / T4 / RTX 20xx are not supported.
- bf16 only. Do not pass
--dtype half: fp16 breaks the Gated DeltaNet convolution states. That is also why sm80 is the minimum. - MTP speculative decoding: the head is bf16 and listed in
quantization_config.ignore. Measured 77% draft acceptance (210 of 273 drafted tokens, mean acceptance length 3.31) with vLLM 0.30.0 on GB10, in the test above (3 speculative tokens, temperature 0, 3 short prompts); figures from sampling-based runs are not directly comparable. Watch the acceptance rate in the log (SpecDecoding metrics): clearly above 0% means the head works; about 0% means it did not load. - Image input: for text-only use keep
--language-model-only(it also skips the vision tower's memory). For images, drop it and keep MTP. If startup then fails inprofile_runwith'NoneType' object has no attribute 'size', vLLM reused a compiled graph cached by an earlier run with different multimodal limits (vllm#58203, a duplicate of vllm#50891): start with a freshVLLM_CACHE_ROOTor setVLLM_USE_AOT_COMPILE=0. Image input was not tested with this build. - 24 GB GPUs: keep
--language-model-onlyand add--kv-cache-dtype fp8,--max-model-len 16384(up to 32768 if it fits),--max-num-seqs 4and--gpu-memory-utilization 0.95. If CUDA graph capture runs out of memory, add--enforce-eager. The MTP head costs about 0.85 GB; drop--speculative-configif that is the difference. - Optional, only on GPUs with FP8 support:
VLLM_MARLIN_INPUT_DTYPE=fp8runs Marlin with FP8 activations. Faster prefill, not evaluated here. On DGX Spark / GB10, corrupted output was reported for an AutoRound INT4 MoE model (vllm#49546); not tested with this build: leave it unset there. - Only the 16 full-attention layers have a KV cache, so FP8 KV saves less than on a dense model.
Sampling (Qwen3.8 recommendations, which ThinkingCap uses unchanged): thinking mode temperature 1.0, top_p 0.95, top_k 20, min_p 0; non-thinking mode temperature 0.7, top_p 0.8, top_k 20, presence_penalty 1.5. Thinking budget via chat_template_kwargs: {"reasoning_effort": "xhigh"}, with xhigh (default, recommended), medium or low.
SGLang loads compressed-tensors W4A16 through Marlin on sm80+ as well. Not tested with this repo.
Transformers loads it (dequantized to bf16, so ~56 GB of memory):
from transformers import AutoModelForImageTextToText
m = AutoModelForImageTextToText.from_pretrained("IstroSec/ThinkingCap-Qwen3.8-27B-abliterated-W4A16", device_map="cuda")
Not for llama.cpp: compressed-tensors format. GGUF builds are made separately from the bf16 source.
Known issues and workarounds
| Issue | Affects | Workaround |
|---|---|---|
| vllm#36249: Qwen3.5-27B GPTQ/Marlin checkpoints crash after ~100 requests | long-running servers | --max-cudagraph-capture-size 32 |
vllm#58807: MTP fc.weight fails to load from a pack-quantized checkpoint without MTP ignore entries |
self-made re-quantizations | not this repo: mtp.* is bf16 and in the ignore list. Keep it that way if you re-quantize |
vllm#35924: a quantized fused in_proj_ba fails to load through Marlin when its per-rank width is below Marlin's minimum of 64; the 27B's in_proj_a/b are 48 wide, so it fails at every tensor-parallel size, TP 1 and 2 included (48 and 24 per rank) |
checkpoints with quantized in_proj_a/b |
not this repo: in_proj_a/b are bf16 and in the ignore list, so they never go through Marlin |
vllm#40252: garbage output when bf16 linear_attn layers are missing from the ignore list |
DeltaNet-bf16 variants | not affected: every bf16 linear_attn layer is in the ignore list, which matches the saved tensors (checked before and after saving) |
sglang#19406: SGLang problems with quantized 48-wide in_proj_a/b |
SGLang | avoided: in_proj_a/b are bf16 |
Reproduce
This build: AWQ (duo_scaling="both", n_grid=20) + GPTQ (dampening 0.01, static activation order); calibration: 256 samples of up to 2,048 tokens of the text described above. Build command, for the record: the build scripts are in a private repository, not published. eval/chat_calib.jsonl is that calibration text, written by 31_chat_eval_build.py --split calib --n 1500 (1,500 samples; the build uses the first 256):
python3 23_quant_nvfp4_gdn.py --int4 --calib-file eval/chat_calib.jsonl
python3 02_regraft_mtp.py <bf16-source-dir> <that-output-dir>
What differs from the default recipe below: only the calibration data (--calib-file: the transcripts described above instead of the first 256 rows of ultrachat_200k); the modifiers and their settings are the default ones. The default recipe (23_quant_nvfp4_gdn.py --int4 with no other option), as in llm-compressor's examples/quantization_w4a16/qwen3_8_gptq_awq_example.py:
from llmcompressor import oneshot
from llmcompressor.modifiers.gptq import GPTQModifier
from llmcompressor.modifiers.transform.awq import AWQModifier
recipe = [
AWQModifier(duo_scaling="both", n_grid=20),
GPTQModifier(targets="Linear", scheme="W4A16", dampening_frac=0.01, # actorder: static (default)
ignore=["lm_head", "re:.*visual.*", "re:.*linear_attn\\.in_proj_(a|b)$"]),
]
# pipeline="independent": AWQ and GPTQ each get their own calibration pass, so GPTQ sees the smoothed model
oneshot(model=model, dataset=ds, recipe=recipe, max_seq_length=2048,
num_calibration_samples=len(ds), pipeline="independent")
# AutoRound instead: recipe = AutoRoundModifier(targets="Linear", scheme="W4A16", ignore=[...], iters=200)
# then copy mtp.* from the bf16 checkpoint and add the MTP Linears to quantization_config.ignore
Not bit-reproducible: llm-compressor 0.14 draws the calibration samples in an unseeded random order, so re-running the same command gives slightly different files. The calibration text was generated by sampling and is not published.
Memory (measured on the DGX Spark, AWQ + GPTQ on 256 samples, 295,811 tokens): torch peak 100.5 GB (+45.8 GB over the 54.7 GB of weights); system MemAvailable fell by 54 GB (183 KB per calibration token), lowest 16.8 GB of 130.7 GB. AWQ keeps every layer's inputs for all calibration samples, so its memory grows with samples × tokens.
Limitations
Everything from the bf16 card applies: no safety filter, residual refusals around 6/100, thinking mode not separately evaluated. You are the safety layer. Add to that the usual 4-bit caveats: slightly lower accuracy on long multi-step reasoning and code than FP8. Quantizing the DeltaNet projections has been reported to cost more on long reasoning chains than on short answers (more truncated thinking). If a task works on FP8 and fails here, the cause is the quantization.
License
PolyForm Small Business License 1.0.0 + BottleCap personal-use grant, inherited from ThinkingCap (see LICENSE). Upstream Qwen materials and the abliteration adapter are Apache-2.0 (see NOTICE). Commercial use beyond the PolyForm terms: contact BottleCap AI.
Credits
bottlecapai (ThinkingCap, and the DeltaNet quantization recipe) · MuXodious (abliteration adapter) · p-e-w/heretic · vllm-project/llm-compressor · Red Hat AI (the AWQ + GPTQ recipe for Qwen3.8-27B) · Qwen team
Calibration data. Prompts and agent contexts come from these datasets; the responses were generated by this model. Changes: rows were selected, cut before an assistant turn, rendered in this model's chat template and completed by the model. No dataset text is included in this repo.
- nebius/SWE-rebench-openhands-trajectories by Trofimova et al. (Nebius), 2025 (OpenHands Trajectories with Qwen3-Coder-480B-A35B-Instruct, Nebius blog), CC-BY-4.0
- google-research-datasets/mbpp by Google Research (Austin et al., 2021, Program Synthesis with Large Language Models), CC-BY-4.0
- HuggingFaceH4/ultrachat_200k, MIT
- openai/gsm8k, MIT
- Agent-Ark/Toucan-1.5M, Apache-2.0
- NousResearch/hermes-function-calling-v1, Apache-2.0
- SWE-bench/SWE-smith-trajectories, MIT