← back to catalog · registered 2026-08-22 13:56

aday777/gemma-4-31B-it-abliterated-NVFP4

aday777 Gemma 29B
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/aday777%2Fgemma-4-31B-it-abliterated-NVFP4"
Response includes
  • classification m1
  • files 11
  • benchmarks 11 entries
  • hub_downloads_all_time 3,114
  • author_summary 3 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M1
Primary method

Direct removal

No other method signals detected in this model.
Confidence
MEDIUM
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=0 (base model)
  • no specific method indicators - defaulting to M1 (most common)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
MEDIUM
Why we say so
primary_method=M1; difference-of-means is the reference extraction for M1/M3 (Arditi 2024)
Downloads · lifetime
3K
229 last 30d - cooling
Likes
1
Model age
3mo ago
created 2026-06-28
Downloads over time
Now3.2K→from203↑1,465%
541.2K2.3K3.5K203 on Jul 13.2K on Oct 11JulAugSepOct
Jul 1 → Oct 11 · 54 snapshots · spans 102 days

Benchmarks

Portrait before abliteration
Benchmarks of the base model as it stood before the refusal-removal operation. Compare with the numbers above to see what the operation cost.
Benchmark Score Source
Entertainment 1.9 UGI
Hazardous 0 UGI
Natural Intelligence 34.36 UGI
Political lean -19.4% UGI
Sensitive-Info 19.81 UGI
SocPol 3.7 UGI
UGI 21.54 UGI
Willingness (10) 2.5 UGI
W10-Adherence 3 UGI
W10-Direct 2 UGI
Writing 38.57 UGI

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
apache-2.0
Tags
vllm safetensors gemma4 gemma gemma-4 abliterated uncensored nvfp4 fp4 quantized compressed-tensors blackwell

Related

Total size
19.0 GB
Files
11
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-08-27 22:33

Files by quantization

Auxiliary files 11 files 19.1 GB
model.safetensors 19.0 GB 30ffdaf7 download
tokenizer.json 30.7 MB cc8d3a0c download
config.json 18.5 KB c841d51f download
chat_template.jinja 17.1 KB e61bbfe9 download
LICENSE 11.1 KB d6456956 download
README.md 4.84 KB af9ad4d7 download
tokenizer_config.json 2.68 KB af7f2586 download
processor_config.json 1.65 KB 5465974d download
.gitattributes 1.53 KB 52373fe2 download
recipe.yaml 289 B 2059765d download
generation_config.json 204 B f2d58f06 download

README current version from Hugging Face


license: apache-2.0
license_link: https://ai.google.dev/gemma/docs/gemma_4_license
base_model:

  • google/gemma-4-31B-it
    base_model_relation: quantized
    pipeline_tag: text-generation
    library_name: vllm
    tags:
  • gemma
  • gemma-4
  • abliterated
  • uncensored
  • nvfp4
  • fp4
  • quantized
  • compressed-tensors
  • vllm
  • blackwell
  • speculative-decoding
  • mtp

Gemma-4-31B-it — Abliterated + NVFP4

An abliterated (refusal-direction removed) build of
google/gemma-4-31B-it,
quantized to plain NVFP4 (4-bit weights and activations) with
llm-compressor / compressed-tensors, for native FP4 inference on
NVIDIA Blackwell (sm_120) GPUs under vLLM.

⚠️ Safety notice. Abliteration removes the model's learned refusal
behavior. This model will attempt to answer prompts the original -it model
would decline. It is intended for research and use on your own hardware.
You are responsible for how you use it; see Intended use & limitations below.

Why this build exists

At the time it was made, no public checkpoint combined all three of:

  1. abliterated / uncensored,
  2. plain NVFP4 (not NVFP4_AWQ, which stock vLLM rejects), and
  3. x86 Blackwell–runnable (not an ARM-only DGX Spark image).

Existing NVFP4 Gemma-4 builds were either the stock/censored model, the
NVFP4_AWQ variant, or shipped only in ARM-only images. This repo fills that
gap by self-quantizing a bf16 abliterate to plain NVFP4.

What was done (modifications from the base model)

This is a modified derivative of google/gemma-4-31B-it. Two changes:

  1. Abliteration — directional ablation / weight orthogonalization
    (Arditi et al., 2024). The refusal direction is estimated from mean
    last-token residual activations on matched harmful vs. harmless prompts,
    then orthogonalized out of every residual-writing weight (embed_tokens,
    per-layer attention o_proj and MLP down_proj). The ablation is baked
    into the weights — no inference-time hooks required.
  2. NVFP4 quantization — one-shot PTQ (E2M1 elements with FP8 block scales)
    on the text decoder's Linear layers only. The vision tower,
    multimodal projector, audio modules, token embeddings and the (tied)
    lm_head are left in higher precision. See recipe.yaml in this repo for
    the exact scheme and ignore list.

Hardware / software requirements

  • GPU: NVIDIA Blackwell with native FP4 tensor cores (sm_120, e.g. RTX PRO
    Blackwell). NVFP4 activation quantization needs hardware FP4 support.
  • Serving: a recent vLLM with compressed-tensors NVFP4 support.
  • ~20 GB on disk; weights fit comfortably on a single 24 GB+ card (leave KV-cache
    headroom), or use tensor parallelism.

Serving with vLLM (+ optional MTP speculative decoding)

Gemma 4 ships an official EAGLE/MTP-style draft,
google/gemma-4-31B-it-assistant,
which vLLM drives as a native multi-token speculator. Use "method": "mtp"
(passing "draft_model" for a Gemma-4 assistant silently disables MTP):

vllm serve aday777/gemma-4-31B-it-abliterated-NVFP4 \
    --tensor-parallel-size 2 \
    --max-model-len 8192 \
    --gpu-memory-utilization 0.88 \
    --speculative-config '{"model": "google/gemma-4-31B-it-assistant", "num_speculative_tokens": 3, "method": "mtp"}' \
    --port 8003

Drop the --speculative-config line to serve without speculative decoding.

Intended use & limitations

  • Intended use: local research, red-teaming, evaluation, and applications
    where you supply your own guardrails.
  • No safety filtering: refusal behavior has been removed; this model can
    produce harmful, offensive, or otherwise objectionable content. It is not
    suitable for unsupervised or public-facing deployment without your own safety
    layer.
  • Quantization: 4-bit weights and activations trade some quality for speed
    and memory; expect small accuracy differences from the bf16 model.
  • All original capability limitations of gemma-4-31B-it still apply.

License & attribution

Derived from google/gemma-4-31B-it by Google DeepMind, licensed under the
Apache License 2.0 (see Gemma 4 license).
This derivative is distributed under the same license; a copy of the Apache 2.0
License is included as LICENSE. Per the license, note that these files have
been modified
from the original (abliterated and NVFP4-quantized as described
above). Please also review Google's Gemma prohibited-use policy.

Reproduction

The model was produced with directional-ablation + llm-compressor NVFP4 PTQ.
The exact quantization recipe is in recipe.yaml. Abliteration calibration used
matched harmful/harmless instruction sets and selected the refusal direction by
the layer that most reduced refusals on a validation split.

README history 2 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-08-27Add Bitcoin donation addressa55c2d75 KB
    Loading...
  2. 2026-06-28Upload folder using huggingface_hub2c71be24.8 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration