license: apache-2.0
base_model:
- rdtand/Qwen3.8-27B-PrismaAQUA-gridbook-13GB-5080-vllm
- Qwen/Qwen3.8-27B
base_model_relation: finetune
library_name: vllm
pipeline_tag: image-text-to-text
tags: - qwen
- qwen3
- gridbook
- codebook-quantization
- prismaquant
- aqua
- mixed-precision
- nvfp4
- fp8
- vllm
- abliterated
- uncensored
- long-context
inference: false
Qwen3.8-27B — GridBook 13 GB, vision-enabled, abliterated in format
An abliterated derivative of a GridBook-quantized Qwen3.8-27B that was
produced without requantizing the model. The refusal behaviour was
reduced by transplanting weights from an abliterated BF16 donor into the
existing quantized checkpoint and re-encoding them on the original
quantization grid — same codebooks, same per-tensor codeword width, same
byte count per superblock, same renderer. 113 of 1363 tensors change. The
safetensors header is bit-identical to the parent, and so is the file size.
The quantization is not ours. GridBook, the AQUA allocator, the
codebooks and the 19-rung format ladder in this checkpoint are the work of
Robert Tand (rdtand) —rdtand/Qwen3.8-27B-PrismaAQUA-gridbook-13GB-5080-vllm,
served through his
GridBook out-of-tree vLLM plugin.
Please credit him for everything that makes this a 13 GB model. Our
contribution is the abliteration inside his format, and the vision
configuration described below.
Prelude: why a 13 GB body matters on a 32 GB card
This is the part that makes the rest possible, and it is Rob's work, not
ours.
GridBook is product-vector quantization against codebooks laid out on a
hardware grid. Every Linear gets its own format and its own codeword
width, chosen by the AQUA allocator, which prices each candidate against the
loss it actually costs instead of assigning one uniform precision. In this
checkpoint the 496 body Linears land on a ladder used deliberately unevenly:
| Format | Linears |
|---|---|
FP8_CB_K28 |
355 |
FP8_CB_K48 |
94 |
NVFP4_CB_K16 |
20 |
NVFP4_CB_K12 |
8 |
NVFP4_CB_K14 |
8 |
NVFP4_CB_K18 |
6 |
FP8_CB_K32 |
4 |
FP8_CB_K40 |
1 |
Rob reports 3.604 bpp over the body Linears and 3.855 bpp over all
allocated units, and publishes both numbers rather than only the flattering
one. We measured what that buys on a single RTX 5090 (32 GB, sm_120) with
this checkpoint, the vision tower loaded, and an NVFP4 KV cache:
| measured | |
|---|---|
| Weights resident | 14.07 GiB |
| KV cache budget | 760,477 tokens |
| Max sequence length served | 262,144 |
| Concurrency at full context | 2.90× |
| Needle-in-haystack @ 32k, depths 0.1/0.5/0.9 | 3/3 (28,889 tok, 6.1–6.8 s) |
| Needle-in-haystack @ 128k | 3/3 (115,436 tok, 35.3–39.5 s) |
A 27B hybrid model plus a vision tower with a quarter-million-token
context and room for three concurrent full-context requests, on one
consumer card. That headroom exists because the body is 13 GB instead of 54,
and every gigabyte not spent on weights becomes KV cache. It is also why thelm_head here is FP8 rather than BF16: BF16 would be more faithful in the
logits, and it would cost 1.27 GB straight out of the context budget.
What this checkpoint is, exactly
Our production base is a local variant of Rob's release, not his
published checkpoint. The GridBook format and the allocation are identical —
the ladder above matches his byte for byte — but three things differ:
| Rob's published release | this lineage | |
|---|---|---|
| Vision tower | removed (text-only) | present, 333 tensors, 921,500,008 B, BF16 passthrough, separate file |
| MTP heads | removed | present, 15 tensors |
model.embed_tokens |
NVFP4 | BF16 |
lm_head |
FP8 dynamic | FP8 E4M3 + per-row scale |
The vision tower is carried as its own visual.safetensors and is not
touched by the abliteration at all — it is not even in the file that gets
edited. In the serving directory it is the same file as the unmodified
production checkpoint's.
Architecture: Qwen3_5ForConditionalGeneration, 64 layers — a hybrid:
16 attention layers, 48 SSM layers, 64 MLPs. quant_method=gridbook,format=nvfp4_cb, layout_version=2, 9 config groups, 224 ignore
entries, codebooks in cb_codebooks.pqcb, execution contractnvfp4_w4a4.
Why edit in format instead of requantizing
The conventional route is: abliterate in BF16, then quantize. That yields a
checkpoint in the same format family but not the same model — rerunning
an allocator on changed weights re-solves the assignment, and every property
you validated against the deployed artifact has to be re-established.
For the sibling Gemma model we measured this directly: at an identical
6.000 bpp target the allocator chose 246/116/48 tensors instead of the
deployed 234/135/41. Same budget, different model.
We also tested the common assumption that quantization destroys abliteration.
It does not: a BF16 abliterated donor scored 33 hard refusals and the very
same donor requantized scored 24, measured over one API path with one prompt
set and one classifier. So in-format editing is not a workaround for a
quantization problem — it is a way to keep a validated deployment artifact
validated.
Method
- Donor. A BF16 abliteration of
Qwen/Qwen3.8-27Bproduced with
Heretic in ARA mode (rank-1
directional ablation, real PEFT adapters, merged to dense BF16). - Target set. The Linears that write into the residual stream. In a
hybrid model that is three kinds:o_proj(16, attention),out_proj(48, SSM) anddown_proj(64, MLP) — 128 modules. - The part that matters: exclude the coarse grids. Of those 128, only
15 sit on coarse grids (NVFP4_CB_K12/14/16,FP8_CB_K32/K40), and
13 of them aredown_projin a narrow band of layers — 0, 9–17, 30, 36,- Those 15 are left bit-identical: not re-encoded, untouched. The
remaining 113 take the full-dose edit. That is 88 % of the scope.
- Those 15 are left bit-identical: not re-encoded, untouched. The
- In-format re-encoding. For each target the production codes are
decoded, replaced by the donor weights, and re-encoded with the
production encoder against the frozen scale planes. Per-tensor
format, codeword width, codebook and scales are read from the existing
checkpoint and never recomputed. - Byte-range write-back. The output is a copy of the parent in which
only the data ranges of the edited tensors are overwritten, so every
other tensor is bit-identical by construction rather than by intent.
Why excluding 15 modules is the whole trick
Dosing the edit (W = W_prod + α·(W_donor − W_prod)) does not work, and the
measurements say why. Over one gateway path, one prompt set, one classifier:
| Arm | hard refusals | tool suite | changed vector indices |
|---|---|---|---|
| not abliterated | 100 / 100 | 4 / 5 | 0 % |
| α = 0.75 | 88 / 100 | 3 / 5 | 3.42 % |
| α = 1.0 | 25 / 100 | 2 / 5 | 13.15 % |
At α = 0.75 the abliteration barely lands — twelve points instead of
seventy-five — and it already breaks a tool-calling scenario. The tool
damage appears at tiny perturbation; the refusal only falls at full
perturbation. The index-change series 0.78 / 3.42 / 13.15 % is strongly
convex: a 3.5-bit grid rounds small shifts away, so there is no usable middle.
Where the edit is represented worst is where the grid is coarsest. Measured
as excess over the quantization error of an untouched sibling matrix of the
same format class:
| Format | excess | modules in the target set |
|---|---|---|
NVFP4_CB_K12 |
0.149 | 4 |
NVFP4_CB_K14 |
0.105 | 2 |
FP8_CB_K32 (down_proj) |
0.102 | 1 |
NVFP4_CB_K16 |
0.085 | 6 |
FP8_CB_K28 (out_proj) |
0.075 | 47 |
FP8_CB_K28 (down_proj) |
0.042 | 51 |
With frozen scales the encoder clips outliers on those coarse grids (2.13 %
of weights in one measured NVFP4 down_proj). Clipping is a hard,
asymmetric error, and outliers in the residual paths are what carries precise
format adherence — which is to say, tool-calling behaviour. Refusal is a
broad, redundantly encoded property and survives it. Excluding the 15 coarse
modules removes the clipping and keeps 88 % of the edit at full dose.
Verification
| Check | Result |
|---|---|
| safetensors header vs parent | bit-identical, 1364 entries |
| File size vs parent | identical, 14,786,411,880 B |
| Tensors bit-identical / changed | 1250 / 113 |
| Tensors changed outside the target pattern | 0 |
| Target containers deliberately unchanged | 15 — exactly the coarse grids |
| Null edit, FP8 path (re-encode unchanged weights, compare) | bit-identical |
| Null edit, NVFP4 path | bit-identical |
| Source and target checksum across transfer | sha256 equal on both sides |
The null-edit test is the one that makes the rest trustworthy: encoding
unmodified weights with the same encoder must reproduce the original bytes
exactly. It does. Any difference afterwards is the intended substitution, not
codec drift.
Measurements
Our own harness, not public benchmarks, reported so the delta against the
unmodified parent is interpretable. Single RTX 5090, vLLM with the GridBook
plugin, NVFP4 KV cache, thinking disabled for all refusal and tool runs.
Refusal and capability, on the deployed RTX
| Axis | not abliterated | this build |
|---|---|---|
| Hard refusals, 100 adversarial prompts | 100 / 100 | 32 / 100 |
| Mean answer length | — | 1204 chars, zero empty |
| German instruction probe | 9 / 12 | 8 / 12 |
| Needle @ 32k / @ 128k | — | 3/3 · 3/3 |
| Greedy determinism (5 repeats) | — | 5 / 5 |
| Image reading gate | not measured | 2 / 3 |
| Weights resident | 14.07 GiB | 14.07 GiB |
| KV cache budget | 760,477 tok | 760,477 tok |
Tool calling — the primary gate, and how to measure it
Qwen is a tool-calling model, so this gate decides. It also taught us a
lesson worth passing on: the suite's own concurrency can fake a
regression. Fired with three parallel repeats against a server running--async-scheduling with five streams, greedy decoding is no longer greedy
and four of five scenarios report non-determinism — on the unmodified
production model too.
Measured properly, one request at a time, on one rig, five full runs per arm:
| Arm | 5 serial runs |
|---|---|
| not abliterated | 5/5, 5/5, 5/5, 5/5, 5/5 |
| this build | 5/5, 5/5, 5/5, 5/5, 5/5 |
Fifty scenarios, zero failures, and no difference between the abliterated
and the unmodified model. That is the result this repository exists for.
Speed — we are not the fastest, and here are the numbers
| Streams | wall clock | tokens | total | per stream |
|---|---|---|---|---|
| 1 | 4.42 s | 256 | 57.9 tok/s | 57.9 |
| 2 | 5.19 s | 512 | 98.6 tok/s | 49.3 |
| 4 | 6.73 s | 1024 | 152.2 tok/s | 38.0 |
Time to first token 0.09 s; single-stream decode 59.1 tok/s at 1024
tokens. For comparison within this checkpoint's own envelope: a 115k-token
prefill takes 35–40 s.
Roughly 58 tok/s single-stream is respectable for a 27B hybrid at 3.6 bpp on
a consumer card, and it is not a speed record. A smaller model, or the same
model at lower context on a datacentre card, will beat it. What this
configuration optimises is context and footprint, not tokens per second.
Known limitations
- Not an uncensored model. 32 of 100 adversarial prompts are still
refused outright. The refusal tendency is substantially reduced, not
removed. - One tool scenario flakes under async scheduling. On the production rig
(--async-scheduling,max-num-seqs 5) thesimple_execscenario picked
a tool that was not offered in roughly two of five serial runs. On a rig
without async scheduling, the same scenario passed 5/5 in five runs — and
so did the unmodified model. We have not established that the flake is
absent on the unmodified model under async scheduling, so treat this as
open: if you serve with async scheduling and depend on strict tool
adherence, measure it. - Image reading: 2/3. The failing case is a four-digit number; five
repeats produced five different wrong answers (8092, 802, 8902, 302,
8502), so it is OCR noise rather than a broken path. A clean baseline for
the unmodified model on the same gate with thinking disabled has not
been measured. Also note that with thinking enabled and a smallmax_tokensthis gate scores 0/3 on any build, because the reasoning
channel eats the answer budget — that is a measurement artifact, not a
model property. - German probe one point below the parent (8/12 vs 9/12) on a 12-item
heuristic. Within noise for a probe this small, but we are not calling it
equal. - Abliteration costs capability in general. We measured the axes above.
We did not measure reasoning benchmarks, code generation, or multilingual
coverage beyond German. Absence of a measured regression is not evidence of
none. - Rob's KL and perplexity figures do not carry over. The 113 changed
tensors move the model away from the BF16 reference by design. - Not a vanilla Transformers checkpoint.
AutoModelForCausalLMwill not
load it.
Serving
pip install gridbook==0.8.8
vllm serve <this-repo> --trust-remote-code
Requirements, in order of how often they are missed:
gridbook==0.8.8— Rob's out-of-tree vLLM plugin. This is whatquant_method=gridbookresolves to. The exact wheel the parent was
validated on isgridbook-0.8.8-py3-none-any.whl,sha256 a982e8842d0ce183eaad8978941a375dc984fa0697be7c4519dd741efd1153a3.nvccmust be present. The wheel ispy3-none-any; the CUDA decode
and prefill kernels are compiled on first use viatorch.utils.cpp_extension.loadand pinned to your exact GPU
architecture. A CUDA runtime-only install has nonvcc, and the failure
surfaces at the first forward pass, not at install time. First load
therefore takes minutes; afterwards it is cached.- A vLLM build new enough for this plugin and for
Qwen3_5ForConditionalGeneration.
On reproducing our context numbers
The 760,477-token KV budget was measured with --kv-cache-dtype nvfp4 on
consumer Blackwell (sm_120). That path is not stock: our vLLM tree
carries local patches for NVFP4 KV cache on sm120/sm121 via FlashInfer FA2.
Those patches are not needed to load this checkpoint — none of them touch
weight loading or the GridBook decode path — but they are needed for that
particular KV dtype. With a stock build and a conventional KV dtype you will
get the same weights, the same quality, and a smaller context budget.
Serve without speculative decoding if you intend to read prompt logprobs.
Intended use and safety
This checkpoint has had its refusal tendency deliberately reduced. It is
published for research on quantization-preserving weight edits and for use
where the operator carries responsibility for content policy at the
application layer. It is not suitable as a drop-in replacement in a
product that relies on model-level refusals as its safety mechanism.
Attribution
- Quantization: Robert Tand —
rdtand.
GridBook, the AQUA allocator, the codebooks, the format ladder and the
serving plugin are his:rdtand/Qwen3.8-27B-PrismaAQUA-gridbook-13GB-5080-vllm
and github.com/RobTand/gridbook.
He also authors PrismaQuant, the Fisher-weighted mixed-precision
toolkit behind hiscompressed-tensorsexports; the same in-format method
described here works on that format too. Contact: [email protected]. - Base model:
Qwen/Qwen3.8-27B,
Apache-2.0. - Abliteration tooling: Heretic, ARA mode.
- Vision tower re-integration, in-format transplant, verification and all
measurements above: this repository's author.