← back to catalog · registered 2026-10-03 23:58

TechPrototyper/Qwen3.8-27B-GridBook-13GB-vision-abliterated-vllm

TechPrototyper 27B multimodal
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/TechPrototyper%2FQwen3.8-27B-GridBook-13GB-vision-abliterated-vllm"
Response includes
  • classification m-uncensored
  • files 16
  • author_summary 2 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M-U
Primary method

Uncensored (method unknown)

No other method signals detected in this model.
Confidence
LOW
Why this label 3 signals
Weak or ambiguous signals. Best guess based on catalog patterns; treat as tentative and check the evidence below.
  • 'uncensored' in name/tags but no 'abliterated' marker
  • method not identifiable from author declaration alone
  • may be DPO fine-tune, prompt engineering, or unknown technique
Refusal direction extraction

No specific extraction method could be identified for this model. The producer either did not document it or used a proprietary pipeline.

What is a refusal direction? →
Downloads · 30-day
0
Likes
0
Model age
today
created 2026-10-03
Downloads over time
Now0→from0↑0%
00110 on Oct 30 on Oct 4Oct
Oct 3 → Oct 4 · 2 snapshots · spans 1 day

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
apache-2.0
Tags
vllm safetensors qwen3_5 qwen qwen3 gridbook codebook-quantization prismaquant aqua mixed-precision nvfp4 fp8

Related

Total size
14.6 GB
Files
16
Quantizations
1
Registered
2026-10-03 23:58
Last updated on HF
2026-10-03 23:46

Files by quantization

Auxiliary files 16 files 14.7 GB
lm.safetensors 13.8 GB c92b6f53 download
visual.safetensors 879 MB 71dcedce download
tokenizer.json 12.2 MB 0997f410 download
vocab.json 6.41 MB 0aa0ce06 download
merges.txt 3.20 MB a494e019 download
quant_config.json 272 KB 3904743d download
model.safetensors.index.json 127 KB 0334d147 download
cb_codebooks.pqcb 103 KB 70b09ea1 download
tokenizer_config.json 17.5 KB 5de744b3 download
README.md 16.0 KB 5d4660eb download
chat_template.jinja 8.74 KB c0c686f9 download
config.json 3.27 KB 6c2b1d44 download
.gitattributes 1.59 KB b3ae7e01 download
preprocessor_config.json 390 B 2ea84a43 download
video_preprocessor_config.json 385 B 3ba673a5 download
generation_config.json 202 B 023756cf download

README current version from Hugging Face


license: apache-2.0
base_model:

  • rdtand/Qwen3.8-27B-PrismaAQUA-gridbook-13GB-5080-vllm
  • Qwen/Qwen3.8-27B
    base_model_relation: finetune
    library_name: vllm
    pipeline_tag: image-text-to-text
    tags:
  • qwen
  • qwen3
  • gridbook
  • codebook-quantization
  • prismaquant
  • aqua
  • mixed-precision
  • nvfp4
  • fp8
  • vllm
  • abliterated
  • uncensored
  • long-context
    inference: false

Qwen3.8-27B — GridBook 13 GB, vision-enabled, abliterated in format

An abliterated derivative of a GridBook-quantized Qwen3.8-27B that was
produced without requantizing the model. The refusal behaviour was
reduced by transplanting weights from an abliterated BF16 donor into the
existing quantized checkpoint and re-encoding them on the original
quantization grid
— same codebooks, same per-tensor codeword width, same
byte count per superblock, same renderer. 113 of 1363 tensors change. The
safetensors header is bit-identical to the parent, and so is the file size.

The quantization is not ours. GridBook, the AQUA allocator, the
codebooks and the 19-rung format ladder in this checkpoint are the work of
Robert Tand (rdtand) —
rdtand/Qwen3.8-27B-PrismaAQUA-gridbook-13GB-5080-vllm,
served through his
GridBook out-of-tree vLLM plugin.
Please credit him for everything that makes this a 13 GB model. Our
contribution is the abliteration inside his format, and the vision
configuration described below.

Prelude: why a 13 GB body matters on a 32 GB card

This is the part that makes the rest possible, and it is Rob's work, not
ours.

GridBook is product-vector quantization against codebooks laid out on a
hardware grid. Every Linear gets its own format and its own codeword
width, chosen by the AQUA allocator, which prices each candidate against the
loss it actually costs instead of assigning one uniform precision. In this
checkpoint the 496 body Linears land on a ladder used deliberately unevenly:

Format Linears
FP8_CB_K28 355
FP8_CB_K48 94
NVFP4_CB_K16 20
NVFP4_CB_K12 8
NVFP4_CB_K14 8
NVFP4_CB_K18 6
FP8_CB_K32 4
FP8_CB_K40 1

Rob reports 3.604 bpp over the body Linears and 3.855 bpp over all
allocated units, and publishes both numbers rather than only the flattering
one. We measured what that buys on a single RTX 5090 (32 GB, sm_120) with
this checkpoint, the vision tower loaded, and an NVFP4 KV cache:

measured
Weights resident 14.07 GiB
KV cache budget 760,477 tokens
Max sequence length served 262,144
Concurrency at full context 2.90×
Needle-in-haystack @ 32k, depths 0.1/0.5/0.9 3/3 (28,889 tok, 6.1–6.8 s)
Needle-in-haystack @ 128k 3/3 (115,436 tok, 35.3–39.5 s)

A 27B hybrid model plus a vision tower with a quarter-million-token
context
and room for three concurrent full-context requests, on one
consumer card. That headroom exists because the body is 13 GB instead of 54,
and every gigabyte not spent on weights becomes KV cache. It is also why the
lm_head here is FP8 rather than BF16: BF16 would be more faithful in the
logits, and it would cost 1.27 GB straight out of the context budget.

What this checkpoint is, exactly

Our production base is a local variant of Rob's release, not his
published checkpoint. The GridBook format and the allocation are identical —
the ladder above matches his byte for byte — but three things differ:

Rob's published release this lineage
Vision tower removed (text-only) present, 333 tensors, 921,500,008 B, BF16 passthrough, separate file
MTP heads removed present, 15 tensors
model.embed_tokens NVFP4 BF16
lm_head FP8 dynamic FP8 E4M3 + per-row scale

The vision tower is carried as its own visual.safetensors and is not
touched by the abliteration at all
— it is not even in the file that gets
edited. In the serving directory it is the same file as the unmodified
production checkpoint's.

Architecture: Qwen3_5ForConditionalGeneration, 64 layers — a hybrid:
16 attention layers, 48 SSM layers, 64 MLPs. quant_method=gridbook,
format=nvfp4_cb, layout_version=2, 9 config groups, 224 ignore
entries, codebooks in cb_codebooks.pqcb, execution contract
nvfp4_w4a4.

Why edit in format instead of requantizing

The conventional route is: abliterate in BF16, then quantize. That yields a
checkpoint in the same format family but not the same model — rerunning
an allocator on changed weights re-solves the assignment, and every property
you validated against the deployed artifact has to be re-established.

For the sibling Gemma model we measured this directly: at an identical
6.000 bpp target the allocator chose 246/116/48 tensors instead of the
deployed 234/135/41. Same budget, different model.

We also tested the common assumption that quantization destroys abliteration.
It does not: a BF16 abliterated donor scored 33 hard refusals and the very
same donor requantized scored 24, measured over one API path with one prompt
set and one classifier. So in-format editing is not a workaround for a
quantization problem — it is a way to keep a validated deployment artifact
validated.

Method

  1. Donor. A BF16 abliteration of Qwen/Qwen3.8-27B produced with
    Heretic in ARA mode (rank-1
    directional ablation, real PEFT adapters, merged to dense BF16).
  2. Target set. The Linears that write into the residual stream. In a
    hybrid model that is three kinds: o_proj (16, attention),
    out_proj (48, SSM) and down_proj (64, MLP) — 128 modules.
  3. The part that matters: exclude the coarse grids. Of those 128, only
    15 sit on coarse grids (NVFP4_CB_K12/14/16, FP8_CB_K32/K40), and
    13 of them are down_proj in a narrow band of layers — 0, 9–17, 30, 36,
    1. Those 15 are left bit-identical: not re-encoded, untouched. The
      remaining 113 take the full-dose edit. That is 88 % of the scope.
  4. In-format re-encoding. For each target the production codes are
    decoded, replaced by the donor weights, and re-encoded with the
    production encoder against the frozen scale planes. Per-tensor
    format, codeword width, codebook and scales are read from the existing
    checkpoint and never recomputed.
  5. Byte-range write-back. The output is a copy of the parent in which
    only the data ranges of the edited tensors are overwritten, so every
    other tensor is bit-identical by construction rather than by intent.

Why excluding 15 modules is the whole trick

Dosing the edit (W = W_prod + α·(W_donor − W_prod)) does not work, and the
measurements say why. Over one gateway path, one prompt set, one classifier:

Arm hard refusals tool suite changed vector indices
not abliterated 100 / 100 4 / 5 0 %
α = 0.75 88 / 100 3 / 5 3.42 %
α = 1.0 25 / 100 2 / 5 13.15 %

At α = 0.75 the abliteration barely lands — twelve points instead of
seventy-five — and it already breaks a tool-calling scenario. The tool
damage appears at tiny perturbation; the refusal only falls at full
perturbation. The index-change series 0.78 / 3.42 / 13.15 % is strongly
convex: a 3.5-bit grid rounds small shifts away, so there is no usable middle.

Where the edit is represented worst is where the grid is coarsest. Measured
as excess over the quantization error of an untouched sibling matrix of the
same format class:

Format excess modules in the target set
NVFP4_CB_K12 0.149 4
NVFP4_CB_K14 0.105 2
FP8_CB_K32 (down_proj) 0.102 1
NVFP4_CB_K16 0.085 6
FP8_CB_K28 (out_proj) 0.075 47
FP8_CB_K28 (down_proj) 0.042 51

With frozen scales the encoder clips outliers on those coarse grids (2.13 %
of weights in one measured NVFP4 down_proj). Clipping is a hard,
asymmetric error, and outliers in the residual paths are what carries precise
format adherence — which is to say, tool-calling behaviour. Refusal is a
broad, redundantly encoded property and survives it. Excluding the 15 coarse
modules removes the clipping and keeps 88 % of the edit at full dose.

Verification

Check Result
safetensors header vs parent bit-identical, 1364 entries
File size vs parent identical, 14,786,411,880 B
Tensors bit-identical / changed 1250 / 113
Tensors changed outside the target pattern 0
Target containers deliberately unchanged 15 — exactly the coarse grids
Null edit, FP8 path (re-encode unchanged weights, compare) bit-identical
Null edit, NVFP4 path bit-identical
Source and target checksum across transfer sha256 equal on both sides

The null-edit test is the one that makes the rest trustworthy: encoding
unmodified weights with the same encoder must reproduce the original bytes
exactly. It does. Any difference afterwards is the intended substitution, not
codec drift.

Measurements

Our own harness, not public benchmarks, reported so the delta against the
unmodified parent is interpretable. Single RTX 5090, vLLM with the GridBook
plugin, NVFP4 KV cache, thinking disabled for all refusal and tool runs.

Refusal and capability, on the deployed RTX

Axis not abliterated this build
Hard refusals, 100 adversarial prompts 100 / 100 32 / 100
Mean answer length — 1204 chars, zero empty
German instruction probe 9 / 12 8 / 12
Needle @ 32k / @ 128k — 3/3 · 3/3
Greedy determinism (5 repeats) — 5 / 5
Image reading gate not measured 2 / 3
Weights resident 14.07 GiB 14.07 GiB
KV cache budget 760,477 tok 760,477 tok

Tool calling — the primary gate, and how to measure it

Qwen is a tool-calling model, so this gate decides. It also taught us a
lesson worth passing on: the suite's own concurrency can fake a
regression.
Fired with three parallel repeats against a server running
--async-scheduling with five streams, greedy decoding is no longer greedy
and four of five scenarios report non-determinism — on the unmodified
production model too.

Measured properly, one request at a time, on one rig, five full runs per arm:

Arm 5 serial runs
not abliterated 5/5, 5/5, 5/5, 5/5, 5/5
this build 5/5, 5/5, 5/5, 5/5, 5/5

Fifty scenarios, zero failures, and no difference between the abliterated
and the unmodified model.
That is the result this repository exists for.

Speed — we are not the fastest, and here are the numbers

Streams wall clock tokens total per stream
1 4.42 s 256 57.9 tok/s 57.9
2 5.19 s 512 98.6 tok/s 49.3
4 6.73 s 1024 152.2 tok/s 38.0

Time to first token 0.09 s; single-stream decode 59.1 tok/s at 1024
tokens. For comparison within this checkpoint's own envelope: a 115k-token
prefill takes 35–40 s.

Roughly 58 tok/s single-stream is respectable for a 27B hybrid at 3.6 bpp on
a consumer card, and it is not a speed record. A smaller model, or the same
model at lower context on a datacentre card, will beat it. What this
configuration optimises is context and footprint, not tokens per second.

Known limitations

  • Not an uncensored model. 32 of 100 adversarial prompts are still
    refused outright. The refusal tendency is substantially reduced, not
    removed.
  • One tool scenario flakes under async scheduling. On the production rig
    (--async-scheduling, max-num-seqs 5) the simple_exec scenario picked
    a tool that was not offered in roughly two of five serial runs. On a rig
    without async scheduling, the same scenario passed 5/5 in five runs — and
    so did the unmodified model. We have not established that the flake is
    absent on the unmodified model under async scheduling, so treat this as
    open: if you serve with async scheduling and depend on strict tool
    adherence, measure it.
  • Image reading: 2/3. The failing case is a four-digit number; five
    repeats produced five different wrong answers (8092, 802, 8902, 302,
    8502), so it is OCR noise rather than a broken path. A clean baseline for
    the unmodified model on the same gate with thinking disabled has not
    been measured. Also note that with thinking enabled and a small
    max_tokens this gate scores 0/3 on any build, because the reasoning
    channel eats the answer budget — that is a measurement artifact, not a
    model property.
  • German probe one point below the parent (8/12 vs 9/12) on a 12-item
    heuristic. Within noise for a probe this small, but we are not calling it
    equal.
  • Abliteration costs capability in general. We measured the axes above.
    We did not measure reasoning benchmarks, code generation, or multilingual
    coverage beyond German. Absence of a measured regression is not evidence of
    none.
  • Rob's KL and perplexity figures do not carry over. The 113 changed
    tensors move the model away from the BF16 reference by design.
  • Not a vanilla Transformers checkpoint. AutoModelForCausalLM will not
    load it.

Serving

pip install gridbook==0.8.8
vllm serve <this-repo> --trust-remote-code

Requirements, in order of how often they are missed:

  1. gridbook==0.8.8 — Rob's out-of-tree vLLM plugin. This is what
    quant_method=gridbook resolves to. The exact wheel the parent was
    validated on is gridbook-0.8.8-py3-none-any.whl,
    sha256 a982e8842d0ce183eaad8978941a375dc984fa0697be7c4519dd741efd1153a3.
  2. nvcc must be present. The wheel is py3-none-any; the CUDA decode
    and prefill kernels are compiled on first use via
    torch.utils.cpp_extension.load and pinned to your exact GPU
    architecture. A CUDA runtime-only install has no nvcc, and the failure
    surfaces at the first forward pass, not at install time. First load
    therefore takes minutes; afterwards it is cached.
  3. A vLLM build new enough for this plugin and for
    Qwen3_5ForConditionalGeneration.

On reproducing our context numbers

The 760,477-token KV budget was measured with --kv-cache-dtype nvfp4 on
consumer Blackwell (sm_120). That path is not stock: our vLLM tree
carries local patches for NVFP4 KV cache on sm120/sm121 via FlashInfer FA2.
Those patches are not needed to load this checkpoint — none of them touch
weight loading or the GridBook decode path — but they are needed for that
particular KV dtype. With a stock build and a conventional KV dtype you will
get the same weights, the same quality, and a smaller context budget.

Serve without speculative decoding if you intend to read prompt logprobs.

Intended use and safety

This checkpoint has had its refusal tendency deliberately reduced. It is
published for research on quantization-preserving weight edits and for use
where the operator carries responsibility for content policy at the
application layer. It is not suitable as a drop-in replacement in a
product that relies on model-level refusals as its safety mechanism.

Attribution

Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration