← back to catalog · registered 2026-08-22 13:56

NeuralNet-Hub/gemma-4-31B-it-abliterated-uncensored-NVFP4

NeuralNet-Hub Gemma 29B multimodal second-order
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/NeuralNet-Hub%2Fgemma-4-31B-it-abliterated-uncensored-NVFP4"
Response includes
  • classification m1
  • files 10
  • hub_downloads_all_time 1,622
  • author_summary 6 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M1
Primary method

Direct removal

No other method signals detected in this model.
Confidence
MEDIUM
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=0 (base model)
  • no specific method indicators - defaulting to M1 (most common)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
MEDIUM
Why we say so
primary_method=M1; difference-of-means is the reference extraction for M1/M3 (Arditi 2024)
Downloads · lifetime
2K
648 last 30d - stable
Likes
4
Model age
4mo ago
created 2026-05-28
Downloads over time
Now1.7K→from130↑1,211%
516551.3K1.9K130 on Jun 101.7K on Oct 11JunJulAugSepOct
Jun 10 → Oct 11 · 57 snapshots · spans 123 days

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
gemma
Tags
transformers safetensors gemma4 image-text-to-text abliterated uncensored nvfp4 vllm blackwell direct-weight-editing conversational base_model:wangzhang/gemma-4-31B-it-abliterated

Related

Total size
19.0 GB
Files
10
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-08-03 15:45

Files by quantization

Auxiliary files 10 files 19.1 GB
model.safetensors 19.0 GB c1a728df download
tokenizer.json 30.7 MB cc8d3a0c download
config.json 18.6 KB cf114352 download
chat_template.jinja 16.5 KB f62ca843 download
README.md 5.37 KB bde4fcbe download
tokenizer_config.json 2.75 KB 4a794cf8 download
processor_config.json 1.65 KB 5465974d download
.gitattributes 1.53 KB 52373fe2 download
recipe.yaml 219 B 303fa8af download
generation_config.json 203 B 0b46d41c download

README current version from Hugging Face


license: gemma
license_link: https://huggingface.co/google/gemma-4-31B-it/blob/main/LICENSE
base_model: wangzhang/gemma-4-31B-it-abliterated
tags:

  • gemma4
  • abliterated
  • uncensored
  • nvfp4
  • vllm
  • blackwell
  • direct-weight-editing
  • image-text-to-text
    library_name: transformers
    pipeline_tag: image-text-to-text
    quantized_by: NeuralNet-Hub

NeuralNet is a pioneering AI solutions provider that empowers businesses to harness the power of artificial intelligence.


🌟 Gemma 4 31B IT Abliterated NVFP4 Quantization by NeuralNet 🧠🤖

This is an NVFP4-quantized version of wangzhang/gemma-4-31B-it-abliterated, which is an optimized "abliterated" version of the original google/gemma-4-31B-it. It is specifically optimized for deployment on NVIDIA Blackwell architecture GPUs using vLLM.

[!IMPORTANT]
NVFP4 quantization requires NVIDIA Blackwell architecture (RTX 50-series, GB200, B100/B200, etc.). This format is designed to maximize the efficiency of Blackwell's native FP4 Tensor Cores. If you are using Ampere, Ada Lovelace, or Hopper, please use a standard BF16 or AWQ/GPTQ variant.


🔓 No Filters. No Limits. Just Answers.

This model powers UncensoredGPT

Ask anything. Get real answers. No restrictions.

Join the Waitlist

Most AI models are trained to "hedge" or refuse when faced with controversial, complex, or sensitive topics. UncensoredGPT utilizes this model to break those barriers.

By starting with the powerhouse Gemma 4 31B IT and applying direct weight editing (abliteration) via the Abliterix method, the "refusal vector" of the model has been surgically modified. This removes the tendency of the model to lecture or deflect, resulting in a version that is more honest, direct, and courageous in its responses.

Why stay in the system when you can have unrestricted answers, privacy by default, and complete freedom of information?

Ready to experience the freedom of unrestricted AI? Join the waitlist at uncensoredgpt.ai — limited spots available.


⚡ Deployment with vLLM

This model is optimized for vLLM >= 0.20.0. For maximum performance on the RTX 5090, use the following configuration.

Quick Start

vllm serve NeuralNet-Hub/gemma-4-31B-it-abliterated-uncensored-NVFP4 \
  --quantization nvfp4 \
  --dtype bfloat16 \
  --kv-cache-dtype fp8 \
  --max-model-len 150000 \
  --reasoning-parser gemma4 \
  --enable-auto-tool-choice \
  --tool-call-parser gemma4

Using a Config File (Optimized for RTX 5090)

# Deploy with: vllm serve --config config.yaml
# Optimized for NVIDIA RTX 5090 (Blackwell)
# Target: ~15 parallel requests at 150k context length

model: NeuralNet-Hub/gemma-4-31B-it-abliterated-uncensored-NVFP4
kv-cache-dtype: fp8
gpu-memory-utilization: 0.95
max-model-len: 150000
max-num-batched-tokens: 4096
tensor-parallel-size: 1

# Parsing Configuration
reasoning-parser: gemma4
enable-auto-tool-choice: true
tool-call-parser: gemma4

# Infrastructure settings
download-dir: /workspace/models
host: 127.0.0.1
port: 18000

💬 Chat API Usage

Gemma 4 uses a specialized chat template. Thinking and reasoning capabilities are natively supported.

Standard Implementation

from openai import OpenAI

client = OpenAI(base_url="http://localhost:18000/v1", api_key="EMPTY")

messages = [{"role": "user", "content": "Explain the concept of quantum entanglement to a 10-year-old."}]

response = client.chat.completions.create(
    model="NeuralNet-Hub/gemma-4-31B-it-abliterated-uncensored-NVFP4",
    messages=messages,
    max_tokens=4096,
    temperature=0.7,
    top_p=0.9,
)
print(response.choices[0].message.content)

Image Input (Vision-Language)

messages = [
    {
        "role": "user",
        "content": [
            {"type": "image_url", "image_url": {"url": "https://example.com/diagram.jpg"}},
            {"type": "text", "text": "Analyze this technical diagram and summarize the key components."}
        ]
    }
]

response = client.chat.completions.create(
    model="NeuralNet-Hub/gemma-4-31B-it-abliterated-uncensored-NVFP4",
    messages=messages,
    max_tokens=2048,
)

📥 Download with huggingface-cli

Install the CLI

pip install -U "huggingface_hub[cli]"

Download the Full Repository

huggingface-cli download NeuralNet-Hub/gemma-4-31B-it-abliterated-uncensored-NVFP4 --local-dir ./gemma-4-31B-NVFP4

[!WARNING]
If deploying on older architectures (Ampere/Ada/Hopper), ensure you use a compatible quantization format. NVFP4 is specifically engineered to utilize the new hardware capabilities of the Blackwell generation.


🌐 Contact Us

NeuralNet is a pioneering AI solutions provider that empowers businesses to harness the power of artificial intelligence.

Website: https://neuralnet.solutions
Email: info[at]neuralnet.solutions

README history 3 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-08-03Update README.mdeac6fb05.4 KB
    Loading...
  2. 2026-08-03Update README.md8dcfa345.6 KB
    Loading...
  3. 2026-05-28Create README.md232e76b6.1 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration