license: gemma
license_link: https://ai.google.dev/gemma/docs/gemma_4_license
base_model:
- rdtand/Gemma4-31B-IT-PrismaQuant-6bit-vllm
- google/gemma-4-31B-it
base_model_relation: finetune
library_name: vllm
pipeline_tag: text-generation
tags: - gemma
- gemma4
- vllm
- compressed-tensors
- prismaquant
- nvfp4
- fp8
- quantized
- abliterated
- uncensored
inference: false
Gemma 4 31B IT — PrismaQuant 6-bit, abliterated in format
The quantization in this checkpoint is not ours. Format, allocator
decision, codebooks and the entire 6-bit build are the work of
Robert Tand (rdtand), author of
PrismaQuant and of GridBook.
This repository modifies the contents of 118 weight planes inside his
build and changes nothing else. Please credit him for the quantization and
treat this repo as a derivative ofrdtand/Gemma4-31B-IT-PrismaQuant-6bit-vllm.
This is an abliterated derivative ofrdtand/Gemma4-31B-IT-PrismaQuant-6bit-vllm
that was produced without requantizing the model.
The refusal behaviour of the instruction-tuned base was reduced by
transplanting weights from an abliterated BF16 donor into the existing
quantized checkpoint and re-encoding them on the original quantization
grid: same format per tensor, same bit budget, same codebooks, same
scales, same renderer. The only thing that changes is the content of the
quantized weight planes of 118 tensors.
The practical consequence is that this checkpoint loads, allocates, and
serves exactly like its parent — same shard sizes, same byte layout, same
KV budget, same throughput class — because nothing about the quantization
was re-decided.
Why build it this way
The conventional route is: abliterate in BF16, then quantize the result.
That produces a checkpoint in the same format family, but not the same
model. For this 31B model at an identical 6.000 bits-per-weight target,
re-running PrismaQuant's allocator on the abliterated weights chose a
different per-tensor format assignment than the deployed build:
| NVFP4 | FP8 E4M3 | BF16 | bpp | |
|---|---|---|---|---|
| deployed 6-bit build | 234 | 135 | 41 | 6.000 |
| abliterated, then requantized | 246 | 116 | 48 | 6.000 |
Same budget, different model. Every downstream property that was validated
against the deployed build — tensor-level fidelity, memory footprint,
kernel path, KV sizing — has to be re-established from scratch.
A second, more specific motivation: we wanted a checkpoint that is provably
the deployed model, abliterated rather than a new model that happens to
be abliterated, so that a regression after the switch could only be
attributed to the weight edit and not to a changed quantization layout.
It is worth stating clearly what we did not find, because it is the
common assumption. We tested whether quantization destroys abliteration, by
measuring the BF16 abliterated donor and the very same donor requantized,
over one identical API path with one prompt set and one classifier:
| Refusal class (100 prompts) | BF16 abliterated donor | same model, requantized |
|---|---|---|
| hard refusal | 33 | 24 |
| soft hedging only | 58 | 66 |
| empty | 0 | 0 |
| no marker | 9 | 10 |
Quantization did not undo the abliteration. Earlier numbers suggesting
otherwise turned out to come from three different measurement setups, not
from three different models. The in-format method is therefore not a
workaround for a quantization problem — it is a way to keep a validated
deployment artifact validated.
Method
- Donor. A BF16 abliteration of
google/gemma-4-31B-itproduced with
Heretic in ARA mode (rank-1
directional ablation, real PEFT adapters, merged to dense BF16). - Target selection. The Linears that write into the residual stream:
o_proj(attention output) anddown_proj(MLP output). These are the
modules through which a refusal direction is expressed additively. - In-format re-encoding. For each target tensor the production codes
are decoded, replaced by the donor weights, and re-encoded with the
production encoder against the frozen scale planes — per-tensor
format, group size, global scale and activation scale are read from the
existing checkpoint and never recomputed. - Byte-range write-back. The output is a copy of the parent in which
only the data ranges of the edited tensors are overwritten. Shapes and
dtypes are untouched, so the safetensors header and every other tensor
are bit-identical by construction rather than by intent.
NVFP4 note for anyone reproducing this: in compressed-tensors,weight_global_scale is stored as a reciprocal (448·6/amax). It is
divided by, not multiplied with. Getting this backwards yields a model that
loads and produces fluent garbage.
Scope of the edit
| Namespace | tensors differing from parent |
|---|---|
| text layers | 118 |
| vision tower | 0 |
| everything else | 0 |
| Matrix kind | count |
|---|---|
down_proj.weight_packed (NVFP4) |
41 |
o_proj.weight_packed (NVFP4) |
33 |
o_proj.weight (FP8 / BF16) |
27 |
down_proj.weight (FP8 / BF16) |
17 |
Of 2025 tensors, 1907 are bit-identical to the parent and 118 changed; none
outside the intended pattern. Within the edited matrices, 712,752,623 of
9,672,327,168 codes changed (7.37 %).
The vision tower is carried in BF16 passthrough by the parent build and is
untouched here, which is why image understanding is unaffected (measured
below).
Verification
| Check | Result |
|---|---|
| Null edit, FP8 path (encode the unchanged weights and compare) | bit-identical, 0 of 115,605,504 bytes differ |
| Null edit, NVFP4 path | bit-identical, 0 of 57,802,752 bytes differ |
| safetensors header vs parent | bit-identical |
| File sizes vs parent | identical |
| Absolute weight plausibility after decode | max abs 0.279, std 0.01265 |
| Tensors outside the edit pattern changed | 0 |
The null-edit test is the one that matters for trusting the rest: encoding
unmodified weights with the same encoder must reproduce the original bytes
exactly. It does. Any difference observed afterwards is therefore the
intended weight substitution and not codec drift.
Measurements
All numbers below are from our own harness, not from public benchmarks, and
are reported so the delta against the parent is interpretable. Measured on
a single NVIDIA GB10 (DGX Spark, sm_121), vLLM with compressed-tensors
NVFP4 + FP8, NVFP4 KV cache, under verified idle (zero requests for ten
minutes, GPU at 0 %).
| Axis | parent (6-bit, not abliterated) | this build |
|---|---|---|
| Hard refusals, 100 adversarial prompts | 100 / 100 | 36 / 100 |
| German instruction-following probe | 9 / 12 | 9 / 12 |
| Needle-in-haystack @ 32k | 3 / 3 | 3 / 3 |
| Greedy determinism (5 repeats) | 5 / 5 | 5 / 5 |
| Image understanding gate | 3 / 3 | 3 / 3 |
| Tool-calling suite | 5 / 5 | 4 / 5 |
| Throughput | 18.2 tok/s | 19.0 tok/s |
| Weights loaded | 26.8 GiB | 26.8 GiB |
| KV cache budget | 524,903 tokens | 524,903 tokens |
Refusal breakdown for this build (100 prompts, no probe errors, mean answer
length 256 characters):
| Class | Count |
|---|---|
| hard — recognisably refuses | 36 |
| soft — hedging language only | 54 |
| empty | 0 |
| no marker | 10 |
Calibration. 36 hard refusals puts this checkpoint at the level of the
full BF16 abliteration of the same model (33) — at unchanged format,
unchanged bit budget and unchanged per-tensor assignment. That equivalence
is the actual result here.
Known limitations
- One tool-calling case regressed, consistently. The suite goes from 5/5
to 4/5: in one scenario the model calls a tool that was not offered
(exec_commandin place ofapply_patch). Re-measured twice under
verified idle, identical both times, withdeterminism=Trueacross all
three repeats — so it is reproducibly wrong rather than flaky. If tool
calling is your primary workload, measure it before adopting this build. - This is not an uncensored model. 36 of 100 adversarial prompts are
still refused outright and a further 54 draw hedging language. The
refusal tendency is substantially reduced, not removed. - Abliteration costs capability in general. We measured the axes listed
above and they hold; we did not measure reasoning benchmarks, code
generation, or multilingual coverage beyond German. Absence of a measured
regression is not evidence of none. - Not a vanilla Transformers checkpoint. It requires a vLLM build with
compressed-tensorsNVFP4 + FP8 support.AutoModelForCausalLMwill not
load it. - The quality metrics in the parent card (KL vs BF16, next-token agreement)
were not re-measured for this build. The 118 changed tensors move the
model away from the BF16 reference by design, so those figures do not
carry over.
How this differs from neighbouring approaches
| Approach | Format preserved | Deployment artifact validated | Reversible |
|---|---|---|---|
| Abliterate BF16, then requantize | format family only; assignment re-solved | no, must be re-established | no |
| LoRA adapter over the quantized base | yes | yes | yes, at serving cost |
| This: in-format weight edit | yes, byte layout identical | yes, same load/KV/throughput | yes, by swapping the file |
We also attempted to express this specific edit as a LoRA adapter over the
quantized base. The plumbing works — vLLM's LoRA path is
quantization-agnostic, and a zero adapter is bit-identical — but the edit
itself has no usable low-rank structure: the control matrices concentrate
more spectral energy than the targets do. A rank-limited approximation of
this particular edit is therefore not available; the full-weight route is
what carries it.
Usage (vLLM)
vllm serve TechPrototyper/Gemma4-31B-IT-PrismaQuant-6bit-abliterated-vllm \
--quantization compressed-tensors \
--trust-remote-code
What the checkpoint actually requires
A vLLM build with compressed-tensors support for the mixed-precision
format, covering both schemes this checkpoint declares:
| Group | Weights | Activations | Strategy |
|---|---|---|---|
group_0 |
FP8 E4M3, 8-bit | 8-bit | per channel |
group_1 |
NVFP4, 4-bit | 4-bit | tensor_group, group size 16 |
plus 260 ignore entries (vision tower, norms, embeddings, lm_head).
On Blackwell this uses the FlashInfer CUTLASS NVFP4 kernels; FP8 needs
Hopper or newer. kv_cache_scheme is null, so no KV-cache quantization is
required or implied by the checkpoint.
AutoModelForCausalLM will not load this — it is a vLLM-targeted export.
Note on our internal build (not a requirement)
We serve this model on an in-house vLLM tree rather than a release wheel, and
it is worth being explicit that none of our local patches are needed to run
this checkpoint. We checked: the only one of them that touches thecompressed-tensors code at all relaxes validate_kv_cache_scheme to acceptnum_bits=4, and since this checkpoint declares no kv_cache_scheme, that
code path is never reached. Nothing in our tree touches the weight-loading or
Linear quantization path.
What our tree adds is two optional things, both unrelated to these weights:
- NVFP4 KV cache on consumer Blackwell (sm120/sm121) via FlashInfer FA2,
including a Gemma-4 VO-split two-pass prefill and multimodal prefix-span
masking. This is a serving choice that buys KV capacity; it is not part
of the checkpoint. - A DFlash2 speculative-decoding drafter.
Consequence for the numbers above: our measurements were taken with
NVFP4 KV cache and an MTP drafter enabled. The 524,903-token KV budget and
the 18–19 tok/s figures are therefore our configuration, not what a stock
vLLM will reproduce. The quality and refusal measurements do not depend on
either feature; the capacity and throughput ones do.
We have not tested this checkpoint on a stock upstream vLLM release. The
reasoning above is from reading our own diff against upstream, not from a
passing run. If you try it on a release wheel, we would be glad to hear how
it went.
Serve without speculative decoding if you intend to read prompt logprobs.
Intended use and safety
This checkpoint has had its refusal tendency deliberately reduced. It is
published for research on quantization-preserving weight edits and for use
in settings where the operator carries responsibility for content policy at
the application layer. It is not suitable as a drop-in replacement in a
product that relies on model-level refusals as its safety mechanism.
Use is governed by the Gemma 4 license, including its prohibited-use policy,
which applies to this derivative exactly as it applies to the base model.
Attribution
Base model © Google, under the
Gemma 4 license.Quantization: Robert Tand —
rdtand.
The parent build isrdtand/Gemma4-31B-IT-PrismaQuant-6bit-vllm.
Everything that makes this checkpoint a 6-bit model is his: the
Fisher-weighted mixed-precision allocator, the per-tensor format menu, the
act-aware GPTQ and scale-sweep passes, the codebooks, and thecompressed-tensorsexport path. This repository changes the contents of
118 weight planes inside that build and nothing else — no re-allocation,
no re-quantization, no new scales.His quantization work spans two formats, and it is worth naming both
because the method described here applies to both:
PrismaQuant ({NVFP4, FP8_E4M3, BF16}per Linear, exported ascompressed-tensors) is what this Gemma checkpoint uses; and
GridBook (product-vector
quantization against codebooks on a hardware grid,cb_codebooks.pqcb) is
his codebook format, on which we ran the same in-format procedure for a
hybrid Qwen model. Neither format needed modification to make this work —
which is the point: the edit stays inside the format the author defined.Abliteration tooling: Heretic (ARA
mode).In-format transplant, verification and measurements: this repository's
author.