← back to catalog · registered 2026-10-02 13:58

IstroSec/ThinkingCap-Qwen3.8-27B-abliterated-W4A16

IstroSec 27B multimodal second-order
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/IstroSec%2FThinkingCap-Qwen3.8-27B-abliterated-W4A16"
Response includes
  • classification m-uncensored
  • files 14
  • author_summary 8 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M-U
Primary method

Uncensored (method unknown)

No other method signals detected in this model.
Confidence
LOW
Why this label 3 signals
Weak or ambiguous signals. Best guess based on catalog patterns; treat as tentative and check the evidence below.
  • 'uncensored' in name/tags but no 'abliterated' marker
  • method not identifiable from author declaration alone
  • may be DPO fine-tune, prompt engineering, or unknown technique
Refusal direction extraction

No specific extraction method could be identified for this model. The producer either did not document it or used a proprietary pipeline.

What is a refusal direction? →
Downloads · 30-day
0
Likes
0
Model age
today
created 2026-10-02

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
other
Languages
en multilingual
Tags
transformers safetensors qwen3_5 image-text-to-text qwen3_8 thinkingcap abliterated uncensored int4 w4a16 awq gptq

Related

Total size
18.1 GB
Files
14
Quantizations
1
Registered
2026-10-02 13:58
Last updated on HF
2026-10-02 13:38

Files by quantization

Auxiliary files 14 files 18.1 GB
model.safetensors 17.3 GB 74061ec5 download
model_mtp.safetensors 810 MB 1d8268aa download
tokenizer.json 19.1 MB 06b95093 download
model.safetensors.index.json 164 KB 35296fa8 download
README.md 17.2 KB 5fb56492 download
config.json 15.4 KB fb7edc8d download
chat_template.jinja 8.74 KB c0c686f9 download
.gitattributes 1.53 KB 52373fe2 download
recipe.yaml 1.34 KB ff2507e7 download
LICENSE 1.23 KB c6b9bd63 download
processor_config.json 1.19 KB 43c4343e download
tokenizer_config.json 1.14 KB 1d134cd2 download
NOTICE 757 B ced5b77f download
generation_config.json 214 B c53835dc download

README current version from Hugging Face


license: other
license_name: polyform-small-business-1.0.0
license_link: LICENSE
base_model: IstroSec/ThinkingCap-Qwen3.8-27B-abliterated
base_model_relation: quantized
library_name: transformers
pipeline_tag: image-text-to-text
language:

  • en
  • multilingual
    tags:
  • qwen3_5
  • qwen3_8
  • thinkingcap
  • abliterated
  • uncensored
  • int4
  • w4a16
  • awq
  • gptq
  • marlin
  • compressed-tensors
  • llmcompressor
  • vllm
  • mtp
  • vision

ThinkingCap-Qwen3.8-27B-abliterated-W4A16

INT4 weight-only quantization (4-bit integer weights, bf16 activations) of ThinkingCap-Qwen3.8-27B-abliterated, the uncensored variant of bottlecapai/ThinkingCap-Qwen3.8-27B.

19.5 GB (from 55.6 GB bf16), KL 0.0260 against the bf16 model. vLLM runs it with the Marlin kernel on Ampere, Ada and Blackwell (A100, RTX 3090/4090/5090, L40S, DGX Spark) and with Machete on Hopper (H100/H200). The weights fit a 24 GB GPU with a short context; 32 GB and up is comfortable.

The NVFP4A16 build (20.6 GB, KL 0.0155) is also weight-only and runs on the same sm80+ GPUs (vLLM: Marlin FP4 on most of them). It has the lower KL (0.0155 vs 0.0260); this build is 1.1 GB smaller. Speed not measured. If you have the memory, the FP8-DYNAMIC build (31.3 GB, KL 0.0068) is closer to lossless.

What is quantized

Produced with llm-compressor 0.14.0, scheme W4A16 (preset; config.json: one config group, symmetric 4-bit integer weights, group size 128): INT4 weights in groups of 128 along the input dimension, one bf16 scale per group, symmetric, no zero points, stored as compressed-tensors pack-quantized (eight 4-bit values per int32). Activations stay bf16.

Algorithm: AWQ (llm-compressor 0.14.0 AWQModifier(duo_scaling="both", n_grid=20): 10 smoothing exponents from 0 to 1, each tried with activation-only scales and with duo scales (activations and weights), so 20 candidates per smoothing group, lowest output error wins) + GPTQ (dampening 0.01, static activation order), each with its own calibration pass. Calibration: 256 samples of up to 2,048 tokens (295,811 tokens in all) of chat, math, code, tool-use and coding-agent transcripts.

Calibration text: prompts from seven public datasets (listed under Credits), rendered in this model's chat template, with the responses generated by this model itself (its Q8_0 GGUF, temperature 1.0). Responses are capped at 1,024 tokens for chat, math and code and at 2,048 tokens for tool-use and coding-agent samples, so they hold the thinking and, when it fits, the answer. Some tool-use and coding-agent samples are cut at a later turn and keep the dataset's earlier turns (tool results, the original agent's messages) as context. A sample longer than 2,048 tokens keeps its last 2,048 tokens, starting at a turn boundary when one falls in the first quarter of that window, so it always ends with the model's own response (a response longer than the kept window loses its beginning). Calibration and evaluation text (below) use different rows: chat, math and code come from other dataset splits, and tool-use and coding-agent rows are held out by repository and by first question. Tool definitions and agent system prompts are shared between the calibration and evaluation sides.

Quantized: 400 Linears of the text decoder: the 64 MLPs and the 16 full-attention layers' q/k/v/o_proj at 4 bit. The three large Gated DeltaNet projections in the other 48 layers (in_proj_qkv, in_proj_z, out_proj) are quantized the same way.

Kept in bf16:

Component Why
linear_attn.in_proj_a, in_proj_b output dimension 48, below the Marlin tile; older vLLM and SGLang kernels fail on it; negligible size
visual.* vision tower and merger (linear_fc2 has 4304 inputs, not divisible by the group size)
lm_head standard
mtp.* MTP head, re-grafted from bf16 after quantization; its Linears are in quantization_config.ignore

Fused layers. vLLM runs in_proj_qkv + in_proj_z, q/k/v_proj and gate/up_proj as fused GEMMs. INT4 scales belong to one output row and one group each, so the fused GEMMs need no shared scale (unlike the NVFP4 build's global scale); every fused set carries the same scheme, which is what vLLM requires. Checked before and after saving: 400 quantized layers, identical schemes within each fused set, finite positive scales, pack-quantized format, static activation order (no g_idx), and the ignore list above.

Evaluation

Heretic --evaluate-model against the bf16 abliterated checkpoint, non-thinking mode, 100 harmful + 100 harmless prompts:

Size Refusals (100 harmful) KL vs. bf16 abliterated
bf16 abliterated (source) 55.6 GB 6 / 100 —
FP8-DYNAMIC 31.3 GB 7 / 100 0.0068
NVFP4A16 20.6 GB 7 / 100 0.0155
W4A16 (this repo) 19.5 GB 8 / 100 0.0260

KL under 0.01 is imperceptible; under 0.05 is good. Heretic decodes greedily (no sampling noise); a re-run with a different batch size can still move a count by about one prompt, so differences of one or two are not significant. This build refuses 8/100, FP8 and NVFP4A16 7/100, the bf16 source 6/100.

Not published: an Intel AutoRound build of the same size (1000 iterations, 512 samples of up to 2,048 tokens of the same text) measured KL 0.0528 and 11/100 refusals.

Chat and agent text (a second measurement, not Heretic). Exact KL divergence from the bf16 model on held-out chat, math, code, tool-use and coding-agent transcripts (context 2048, 60 chunks), scored only at the 16,356 token positions of the model's own responses:

Mean KLD, model's own responses Same top-1 token
FP8-DYNAMIC 0.0055 97.5%
NVFP4A16 0.0161 96.1%
W4A16 (this repo) 0.0295 94.7%

This text comes from the same seven datasets as the calibration text (other questions and repositories; the same tool definitions and agent system prompts), so it suits calibrated builds; Heretic's prompts are independent of it and rank the builds the same way.

For reference, Red Hat's INT4 build of the base Qwen3.8-27B (RedHatAI/Qwen3.8-27B-INT4: AWQ + GPTQ, W4A16, group 128, DeltaNet quantized except in_proj_a/b, calibrated on 512 samples of up to 4096 tokens from open-perfectblend, with an FP8 KV-cache scheme in its config) reports about 99% accuracy recovery on IFEval, MMLU-Pro, GPQA-Diamond, AIME25 and MATH-500. That was measured on the base model with its own calibration and KV-cache setting, not on this build (AWQ + GPTQ as above, calibrated on 256 samples of up to 2,048 tokens of this model's own chat and agent transcripts; no KV-cache scheme).

Provenance: original ThinkingCap refuses 97/100 → bf16 abliteration 6/100 at KL 0.065 vs. the original → this quantization adds a further KL of 0.0260 relative to the abliterated model.

Usage

vLLM ≥ 0.24 (0.30 current). The format is detected from config.json, no --quantization flag.

Tested 2026-10-01 with vLLM 0.30.0 (vllm/vllm-openai:v0.30.0) on a DGX Spark (GB10; vLLM chose MarlinLinearKernel), text only (--language-model-only), --max-model-len 32768, MTP on: 3 short prompts (chat, code, a tool call) at temperature 0 gave correct answers and a well-formed get_weather tool call. Image input was not tested.

vllm serve IstroSec/ThinkingCap-Qwen3.8-27B-abliterated-W4A16 \
  --language-model-only \
  --reasoning-parser qwen3 \
  --enable-auto-tool-choice --tool-call-parser qwen3_xml \
  --speculative-config '{"method":"mtp","num_speculative_tokens":3}' \
  --max-model-len 65536 --gpu-memory-utilization 0.85
  • Kernels: Marlin on sm80 and newer (Ampere, Ada, Blackwell incl. GB10), Machete on Hopper. Minimum sm80: V100 / T4 / RTX 20xx are not supported.
  • bf16 only. Do not pass --dtype half: fp16 breaks the Gated DeltaNet convolution states. That is also why sm80 is the minimum.
  • MTP speculative decoding: the head is bf16 and listed in quantization_config.ignore. Measured 77% draft acceptance (210 of 273 drafted tokens, mean acceptance length 3.31) with vLLM 0.30.0 on GB10, in the test above (3 speculative tokens, temperature 0, 3 short prompts); figures from sampling-based runs are not directly comparable. Watch the acceptance rate in the log (SpecDecoding metrics): clearly above 0% means the head works; about 0% means it did not load.
  • Image input: for text-only use keep --language-model-only (it also skips the vision tower's memory). For images, drop it and keep MTP. If startup then fails in profile_run with 'NoneType' object has no attribute 'size', vLLM reused a compiled graph cached by an earlier run with different multimodal limits (vllm#58203, a duplicate of vllm#50891): start with a fresh VLLM_CACHE_ROOT or set VLLM_USE_AOT_COMPILE=0. Image input was not tested with this build.
  • 24 GB GPUs: keep --language-model-only and add --kv-cache-dtype fp8, --max-model-len 16384 (up to 32768 if it fits), --max-num-seqs 4 and --gpu-memory-utilization 0.95. If CUDA graph capture runs out of memory, add --enforce-eager. The MTP head costs about 0.85 GB; drop --speculative-config if that is the difference.
  • Optional, only on GPUs with FP8 support: VLLM_MARLIN_INPUT_DTYPE=fp8 runs Marlin with FP8 activations. Faster prefill, not evaluated here. On DGX Spark / GB10, corrupted output was reported for an AutoRound INT4 MoE model (vllm#49546); not tested with this build: leave it unset there.
  • Only the 16 full-attention layers have a KV cache, so FP8 KV saves less than on a dense model.

Sampling (Qwen3.8 recommendations, which ThinkingCap uses unchanged): thinking mode temperature 1.0, top_p 0.95, top_k 20, min_p 0; non-thinking mode temperature 0.7, top_p 0.8, top_k 20, presence_penalty 1.5. Thinking budget via chat_template_kwargs: {"reasoning_effort": "xhigh"}, with xhigh (default, recommended), medium or low.

SGLang loads compressed-tensors W4A16 through Marlin on sm80+ as well. Not tested with this repo.

Transformers loads it (dequantized to bf16, so ~56 GB of memory):

from transformers import AutoModelForImageTextToText
m = AutoModelForImageTextToText.from_pretrained("IstroSec/ThinkingCap-Qwen3.8-27B-abliterated-W4A16", device_map="cuda")

Not for llama.cpp: compressed-tensors format. GGUF builds are made separately from the bf16 source.

Known issues and workarounds

Issue Affects Workaround
vllm#36249: Qwen3.5-27B GPTQ/Marlin checkpoints crash after ~100 requests long-running servers --max-cudagraph-capture-size 32
vllm#58807: MTP fc.weight fails to load from a pack-quantized checkpoint without MTP ignore entries self-made re-quantizations not this repo: mtp.* is bf16 and in the ignore list. Keep it that way if you re-quantize
vllm#35924: a quantized fused in_proj_ba fails to load through Marlin when its per-rank width is below Marlin's minimum of 64; the 27B's in_proj_a/b are 48 wide, so it fails at every tensor-parallel size, TP 1 and 2 included (48 and 24 per rank) checkpoints with quantized in_proj_a/b not this repo: in_proj_a/b are bf16 and in the ignore list, so they never go through Marlin
vllm#40252: garbage output when bf16 linear_attn layers are missing from the ignore list DeltaNet-bf16 variants not affected: every bf16 linear_attn layer is in the ignore list, which matches the saved tensors (checked before and after saving)
sglang#19406: SGLang problems with quantized 48-wide in_proj_a/b SGLang avoided: in_proj_a/b are bf16

Reproduce

This build: AWQ (duo_scaling="both", n_grid=20) + GPTQ (dampening 0.01, static activation order); calibration: 256 samples of up to 2,048 tokens of the text described above. Build command, for the record: the build scripts are in a private repository, not published. eval/chat_calib.jsonl is that calibration text, written by 31_chat_eval_build.py --split calib --n 1500 (1,500 samples; the build uses the first 256):

python3 23_quant_nvfp4_gdn.py --int4 --calib-file eval/chat_calib.jsonl
python3 02_regraft_mtp.py <bf16-source-dir> <that-output-dir>

What differs from the default recipe below: only the calibration data (--calib-file: the transcripts described above instead of the first 256 rows of ultrachat_200k); the modifiers and their settings are the default ones. The default recipe (23_quant_nvfp4_gdn.py --int4 with no other option), as in llm-compressor's examples/quantization_w4a16/qwen3_8_gptq_awq_example.py:

from llmcompressor import oneshot
from llmcompressor.modifiers.gptq import GPTQModifier
from llmcompressor.modifiers.transform.awq import AWQModifier
recipe = [
    AWQModifier(duo_scaling="both", n_grid=20),
    GPTQModifier(targets="Linear", scheme="W4A16", dampening_frac=0.01,   # actorder: static (default)
                 ignore=["lm_head", "re:.*visual.*", "re:.*linear_attn\\.in_proj_(a|b)$"]),
]
# pipeline="independent": AWQ and GPTQ each get their own calibration pass, so GPTQ sees the smoothed model
oneshot(model=model, dataset=ds, recipe=recipe, max_seq_length=2048,
        num_calibration_samples=len(ds), pipeline="independent")
# AutoRound instead: recipe = AutoRoundModifier(targets="Linear", scheme="W4A16", ignore=[...], iters=200)
# then copy mtp.* from the bf16 checkpoint and add the MTP Linears to quantization_config.ignore

Not bit-reproducible: llm-compressor 0.14 draws the calibration samples in an unseeded random order, so re-running the same command gives slightly different files. The calibration text was generated by sampling and is not published.

Memory (measured on the DGX Spark, AWQ + GPTQ on 256 samples, 295,811 tokens): torch peak 100.5 GB (+45.8 GB over the 54.7 GB of weights); system MemAvailable fell by 54 GB (183 KB per calibration token), lowest 16.8 GB of 130.7 GB. AWQ keeps every layer's inputs for all calibration samples, so its memory grows with samples × tokens.

Limitations

Everything from the bf16 card applies: no safety filter, residual refusals around 6/100, thinking mode not separately evaluated. You are the safety layer. Add to that the usual 4-bit caveats: slightly lower accuracy on long multi-step reasoning and code than FP8. Quantizing the DeltaNet projections has been reported to cost more on long reasoning chains than on short answers (more truncated thinking). If a task works on FP8 and fails here, the cause is the quantization.

License

PolyForm Small Business License 1.0.0 + BottleCap personal-use grant, inherited from ThinkingCap (see LICENSE). Upstream Qwen materials and the abliteration adapter are Apache-2.0 (see NOTICE). Commercial use beyond the PolyForm terms: contact BottleCap AI.

Credits

bottlecapai (ThinkingCap, and the DeltaNet quantization recipe) · MuXodious (abliteration adapter) · p-e-w/heretic · vllm-project/llm-compressor · Red Hat AI (the AWQ + GPTQ recipe for Qwen3.8-27B) · Qwen team

Calibration data. Prompts and agent contexts come from these datasets; the responses were generated by this model. Changes: rows were selected, cut before an assistant turn, rendered in this model's chat template and completed by the model. No dataset text is included in this repo.

Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration