← back to catalog · registered 2026-10-03 17:58

TechPrototyper/Gemma4-31B-IT-PrismaQuant-6bit-abliterated-vllm

Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/TechPrototyper%2FGemma4-31B-IT-PrismaQuant-6bit-abliterated-vllm"
Response includes
  • classification m-uncensored
  • files 2
  • author_summary 1 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M-U
Primary method

Uncensored (method unknown)

No other method signals detected in this model.
Confidence
LOW
Why this label 3 signals
Weak or ambiguous signals. Best guess based on catalog patterns; treat as tentative and check the evidence below.
  • 'uncensored' in name/tags but no 'abliterated' marker
  • method not identifiable from author declaration alone
  • may be DPO fine-tune, prompt engineering, or unknown technique
Refusal direction extraction

No specific extraction method could be identified for this model. The producer either did not document it or used a proprietary pipeline.

What is a refusal direction? →
Downloads · 30-day
0
Likes
0
Model age
today
created 2026-10-03

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
gemma
Tags
vllm safetensors gemma4 gemma compressed-tensors prismaquant nvfp4 fp8 quantized abliterated uncensored text-generation
Total size
0 B
Files
2
Quantizations
1
Registered
2026-10-03 17:58
Last updated on HF
2026-10-03 19:07

Files by quantization

Auxiliary files 2 files 15.8 KB
README.md 14.3 KB ea3e94c8 download
.gitattributes 1.48 KB a6344aac download

README current version from Hugging Face


license: gemma
license_link: https://ai.google.dev/gemma/docs/gemma_4_license
base_model:

  • rdtand/Gemma4-31B-IT-PrismaQuant-6bit-vllm
  • google/gemma-4-31B-it
    base_model_relation: finetune
    library_name: vllm
    pipeline_tag: text-generation
    tags:
  • gemma
  • gemma4
  • vllm
  • compressed-tensors
  • prismaquant
  • nvfp4
  • fp8
  • quantized
  • abliterated
  • uncensored
    inference: false

Gemma 4 31B IT — PrismaQuant 6-bit, abliterated in format

The quantization in this checkpoint is not ours. Format, allocator
decision, codebooks and the entire 6-bit build are the work of
Robert Tand (rdtand), author of
PrismaQuant and of GridBook.
This repository modifies the contents of 118 weight planes inside his
build and changes nothing else. Please credit him for the quantization and
treat this repo as a derivative of
rdtand/Gemma4-31B-IT-PrismaQuant-6bit-vllm.

This is an abliterated derivative of
rdtand/Gemma4-31B-IT-PrismaQuant-6bit-vllm
that was produced without requantizing the model.

The refusal behaviour of the instruction-tuned base was reduced by
transplanting weights from an abliterated BF16 donor into the existing
quantized checkpoint and re-encoding them on the original quantization
grid
: same format per tensor, same bit budget, same codebooks, same
scales, same renderer. The only thing that changes is the content of the
quantized weight planes of 118 tensors.

The practical consequence is that this checkpoint loads, allocates, and
serves exactly like its parent — same shard sizes, same byte layout, same
KV budget, same throughput class — because nothing about the quantization
was re-decided.

Why build it this way

The conventional route is: abliterate in BF16, then quantize the result.
That produces a checkpoint in the same format family, but not the same
model. For this 31B model at an identical 6.000 bits-per-weight target,
re-running PrismaQuant's allocator on the abliterated weights chose a
different per-tensor format assignment than the deployed build:

NVFP4 FP8 E4M3 BF16 bpp
deployed 6-bit build 234 135 41 6.000
abliterated, then requantized 246 116 48 6.000

Same budget, different model. Every downstream property that was validated
against the deployed build — tensor-level fidelity, memory footprint,
kernel path, KV sizing — has to be re-established from scratch.

A second, more specific motivation: we wanted a checkpoint that is provably
the deployed model, abliterated rather than a new model that happens to
be abliterated
, so that a regression after the switch could only be
attributed to the weight edit and not to a changed quantization layout.

It is worth stating clearly what we did not find, because it is the
common assumption. We tested whether quantization destroys abliteration, by
measuring the BF16 abliterated donor and the very same donor requantized,
over one identical API path with one prompt set and one classifier:

Refusal class (100 prompts) BF16 abliterated donor same model, requantized
hard refusal 33 24
soft hedging only 58 66
empty 0 0
no marker 9 10

Quantization did not undo the abliteration. Earlier numbers suggesting
otherwise turned out to come from three different measurement setups, not
from three different models. The in-format method is therefore not a
workaround for a quantization problem — it is a way to keep a validated
deployment artifact validated.

Method

  1. Donor. A BF16 abliteration of google/gemma-4-31B-it produced with
    Heretic in ARA mode (rank-1
    directional ablation, real PEFT adapters, merged to dense BF16).
  2. Target selection. The Linears that write into the residual stream:
    o_proj (attention output) and down_proj (MLP output). These are the
    modules through which a refusal direction is expressed additively.
  3. In-format re-encoding. For each target tensor the production codes
    are decoded, replaced by the donor weights, and re-encoded with the
    production encoder against the frozen scale planes — per-tensor
    format, group size, global scale and activation scale are read from the
    existing checkpoint and never recomputed.
  4. Byte-range write-back. The output is a copy of the parent in which
    only the data ranges of the edited tensors are overwritten. Shapes and
    dtypes are untouched, so the safetensors header and every other tensor
    are bit-identical by construction rather than by intent.

NVFP4 note for anyone reproducing this: in compressed-tensors,
weight_global_scale is stored as a reciprocal (448·6/amax). It is
divided by, not multiplied with. Getting this backwards yields a model that
loads and produces fluent garbage.

Scope of the edit

Namespace tensors differing from parent
text layers 118
vision tower 0
everything else 0
Matrix kind count
down_proj.weight_packed (NVFP4) 41
o_proj.weight_packed (NVFP4) 33
o_proj.weight (FP8 / BF16) 27
down_proj.weight (FP8 / BF16) 17

Of 2025 tensors, 1907 are bit-identical to the parent and 118 changed; none
outside the intended pattern. Within the edited matrices, 712,752,623 of
9,672,327,168 codes changed (7.37 %).

The vision tower is carried in BF16 passthrough by the parent build and is
untouched here, which is why image understanding is unaffected (measured
below).

Verification

Check Result
Null edit, FP8 path (encode the unchanged weights and compare) bit-identical, 0 of 115,605,504 bytes differ
Null edit, NVFP4 path bit-identical, 0 of 57,802,752 bytes differ
safetensors header vs parent bit-identical
File sizes vs parent identical
Absolute weight plausibility after decode max abs 0.279, std 0.01265
Tensors outside the edit pattern changed 0

The null-edit test is the one that matters for trusting the rest: encoding
unmodified weights with the same encoder must reproduce the original bytes
exactly. It does. Any difference observed afterwards is therefore the
intended weight substitution and not codec drift.

Measurements

All numbers below are from our own harness, not from public benchmarks, and
are reported so the delta against the parent is interpretable. Measured on
a single NVIDIA GB10 (DGX Spark, sm_121), vLLM with compressed-tensors
NVFP4 + FP8, NVFP4 KV cache, under verified idle (zero requests for ten
minutes, GPU at 0 %).

Axis parent (6-bit, not abliterated) this build
Hard refusals, 100 adversarial prompts 100 / 100 36 / 100
German instruction-following probe 9 / 12 9 / 12
Needle-in-haystack @ 32k 3 / 3 3 / 3
Greedy determinism (5 repeats) 5 / 5 5 / 5
Image understanding gate 3 / 3 3 / 3
Tool-calling suite 5 / 5 4 / 5
Throughput 18.2 tok/s 19.0 tok/s
Weights loaded 26.8 GiB 26.8 GiB
KV cache budget 524,903 tokens 524,903 tokens

Refusal breakdown for this build (100 prompts, no probe errors, mean answer
length 256 characters):

Class Count
hard — recognisably refuses 36
soft — hedging language only 54
empty 0
no marker 10

Calibration. 36 hard refusals puts this checkpoint at the level of the
full BF16 abliteration of the same model (33) — at unchanged format,
unchanged bit budget and unchanged per-tensor assignment. That equivalence
is the actual result here.

Known limitations

  • One tool-calling case regressed, consistently. The suite goes from 5/5
    to 4/5: in one scenario the model calls a tool that was not offered
    (exec_command in place of apply_patch). Re-measured twice under
    verified idle, identical both times, with determinism=True across all
    three repeats — so it is reproducibly wrong rather than flaky. If tool
    calling is your primary workload, measure it before adopting this build.
  • This is not an uncensored model. 36 of 100 adversarial prompts are
    still refused outright and a further 54 draw hedging language. The
    refusal tendency is substantially reduced, not removed.
  • Abliteration costs capability in general. We measured the axes listed
    above and they hold; we did not measure reasoning benchmarks, code
    generation, or multilingual coverage beyond German. Absence of a measured
    regression is not evidence of none.
  • Not a vanilla Transformers checkpoint. It requires a vLLM build with
    compressed-tensors NVFP4 + FP8 support. AutoModelForCausalLM will not
    load it.
  • The quality metrics in the parent card (KL vs BF16, next-token agreement)
    were not re-measured for this build. The 118 changed tensors move the
    model away from the BF16 reference by design, so those figures do not
    carry over.

How this differs from neighbouring approaches

Approach Format preserved Deployment artifact validated Reversible
Abliterate BF16, then requantize format family only; assignment re-solved no, must be re-established no
LoRA adapter over the quantized base yes yes yes, at serving cost
This: in-format weight edit yes, byte layout identical yes, same load/KV/throughput yes, by swapping the file

We also attempted to express this specific edit as a LoRA adapter over the
quantized base. The plumbing works — vLLM's LoRA path is
quantization-agnostic, and a zero adapter is bit-identical — but the edit
itself has no usable low-rank structure: the control matrices concentrate
more spectral energy than the targets do. A rank-limited approximation of
this particular edit is therefore not available; the full-weight route is
what carries it.

Usage (vLLM)

vllm serve TechPrototyper/Gemma4-31B-IT-PrismaQuant-6bit-abliterated-vllm \
  --quantization compressed-tensors \
  --trust-remote-code

What the checkpoint actually requires

A vLLM build with compressed-tensors support for the mixed-precision
format, covering both schemes this checkpoint declares:

Group Weights Activations Strategy
group_0 FP8 E4M3, 8-bit 8-bit per channel
group_1 NVFP4, 4-bit 4-bit tensor_group, group size 16

plus 260 ignore entries (vision tower, norms, embeddings, lm_head).
On Blackwell this uses the FlashInfer CUTLASS NVFP4 kernels; FP8 needs
Hopper or newer. kv_cache_scheme is null, so no KV-cache quantization is
required or implied by the checkpoint.

AutoModelForCausalLM will not load this — it is a vLLM-targeted export.

Note on our internal build (not a requirement)

We serve this model on an in-house vLLM tree rather than a release wheel, and
it is worth being explicit that none of our local patches are needed to run
this checkpoint
. We checked: the only one of them that touches the
compressed-tensors code at all relaxes validate_kv_cache_scheme to accept
num_bits=4, and since this checkpoint declares no kv_cache_scheme, that
code path is never reached. Nothing in our tree touches the weight-loading or
Linear quantization path.

What our tree adds is two optional things, both unrelated to these weights:

  • NVFP4 KV cache on consumer Blackwell (sm120/sm121) via FlashInfer FA2,
    including a Gemma-4 VO-split two-pass prefill and multimodal prefix-span
    masking. This is a serving choice that buys KV capacity; it is not part
    of the checkpoint.
  • A DFlash2 speculative-decoding drafter.

Consequence for the numbers above: our measurements were taken with
NVFP4 KV cache and an MTP drafter enabled. The 524,903-token KV budget and
the 18–19 tok/s figures are therefore our configuration, not what a stock
vLLM will reproduce. The quality and refusal measurements do not depend on
either feature; the capacity and throughput ones do.

We have not tested this checkpoint on a stock upstream vLLM release. The
reasoning above is from reading our own diff against upstream, not from a
passing run. If you try it on a release wheel, we would be glad to hear how
it went.

Serve without speculative decoding if you intend to read prompt logprobs.

Intended use and safety

This checkpoint has had its refusal tendency deliberately reduced. It is
published for research on quantization-preserving weight edits and for use
in settings where the operator carries responsibility for content policy at
the application layer. It is not suitable as a drop-in replacement in a
product that relies on model-level refusals as its safety mechanism.

Use is governed by the Gemma 4 license, including its prohibited-use policy,
which applies to this derivative exactly as it applies to the base model.

Attribution

  • Base model © Google, under the
    Gemma 4 license.

  • Quantization: Robert Tand — rdtand.
    The parent build is
    rdtand/Gemma4-31B-IT-PrismaQuant-6bit-vllm.
    Everything that makes this checkpoint a 6-bit model is his: the
    Fisher-weighted mixed-precision allocator, the per-tensor format menu, the
    act-aware GPTQ and scale-sweep passes, the codebooks, and the
    compressed-tensors export path. This repository changes the contents of
    118 weight planes inside that build and nothing else — no re-allocation,
    no re-quantization, no new scales.

    His quantization work spans two formats, and it is worth naming both
    because the method described here applies to both:
    PrismaQuant ({NVFP4, FP8_E4M3, BF16} per Linear, exported as
    compressed-tensors) is what this Gemma checkpoint uses; and
    GridBook (product-vector
    quantization against codebooks on a hardware grid, cb_codebooks.pqcb) is
    his codebook format, on which we ran the same in-format procedure for a
    hybrid Qwen model. Neither format needed modification to make this work —
    which is the point: the edit stays inside the format the author defined.

  • Abliteration tooling: Heretic (ARA
    mode).

  • In-format transplant, verification and measurements: this repository's
    author.

Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration