license: gemma
license_link: https://huggingface.co/google/gemma-4-31B-it/blob/main/LICENSE
base_model: wangzhang/gemma-4-31B-it-abliterated
tags:
- gemma4
- abliterated
- uncensored
- nvfp4
- vllm
- blackwell
- direct-weight-editing
- image-text-to-text
library_name: transformers
pipeline_tag: image-text-to-text
quantized_by: NeuralNet-Hub
NeuralNet is a pioneering AI solutions provider that empowers businesses to harness the power of artificial intelligence.
🌟 Gemma 4 31B IT Abliterated NVFP4 Quantization by NeuralNet 🧠🤖
This is an NVFP4-quantized version of wangzhang/gemma-4-31B-it-abliterated, which is an optimized "abliterated" version of the original google/gemma-4-31B-it. It is specifically optimized for deployment on NVIDIA Blackwell architecture GPUs using vLLM.
[!IMPORTANT]
NVFP4 quantization requires NVIDIA Blackwell architecture (RTX 50-series, GB200, B100/B200, etc.). This format is designed to maximize the efficiency of Blackwell's native FP4 Tensor Cores. If you are using Ampere, Ada Lovelace, or Hopper, please use a standard BF16 or AWQ/GPTQ variant.
🔓 No Filters. No Limits. Just Answers.
Most AI models are trained to "hedge" or refuse when faced with controversial, complex, or sensitive topics. UncensoredGPT utilizes this model to break those barriers.
By starting with the powerhouse Gemma 4 31B IT and applying direct weight editing (abliteration) via the Abliterix method, the "refusal vector" of the model has been surgically modified. This removes the tendency of the model to lecture or deflect, resulting in a version that is more honest, direct, and courageous in its responses.
Why stay in the system when you can have unrestricted answers, privacy by default, and complete freedom of information?
Ready to experience the freedom of unrestricted AI? Join the waitlist at uncensoredgpt.ai — limited spots available.
⚡ Deployment with vLLM
This model is optimized for vLLM >= 0.20.0. For maximum performance on the RTX 5090, use the following configuration.
Quick Start
vllm serve NeuralNet-Hub/gemma-4-31B-it-abliterated-uncensored-NVFP4 \
--quantization nvfp4 \
--dtype bfloat16 \
--kv-cache-dtype fp8 \
--max-model-len 150000 \
--reasoning-parser gemma4 \
--enable-auto-tool-choice \
--tool-call-parser gemma4
Using a Config File (Optimized for RTX 5090)
# Deploy with: vllm serve --config config.yaml
# Optimized for NVIDIA RTX 5090 (Blackwell)
# Target: ~15 parallel requests at 150k context length
model: NeuralNet-Hub/gemma-4-31B-it-abliterated-uncensored-NVFP4
kv-cache-dtype: fp8
gpu-memory-utilization: 0.95
max-model-len: 150000
max-num-batched-tokens: 4096
tensor-parallel-size: 1
# Parsing Configuration
reasoning-parser: gemma4
enable-auto-tool-choice: true
tool-call-parser: gemma4
# Infrastructure settings
download-dir: /workspace/models
host: 127.0.0.1
port: 18000
💬 Chat API Usage
Gemma 4 uses a specialized chat template. Thinking and reasoning capabilities are natively supported.
Standard Implementation
from openai import OpenAI
client = OpenAI(base_url="http://localhost:18000/v1", api_key="EMPTY")
messages = [{"role": "user", "content": "Explain the concept of quantum entanglement to a 10-year-old."}]
response = client.chat.completions.create(
model="NeuralNet-Hub/gemma-4-31B-it-abliterated-uncensored-NVFP4",
messages=messages,
max_tokens=4096,
temperature=0.7,
top_p=0.9,
)
print(response.choices[0].message.content)
Image Input (Vision-Language)
messages = [
{
"role": "user",
"content": [
{"type": "image_url", "image_url": {"url": "https://example.com/diagram.jpg"}},
{"type": "text", "text": "Analyze this technical diagram and summarize the key components."}
]
}
]
response = client.chat.completions.create(
model="NeuralNet-Hub/gemma-4-31B-it-abliterated-uncensored-NVFP4",
messages=messages,
max_tokens=2048,
)
📥 Download with huggingface-cli
Install the CLI
pip install -U "huggingface_hub[cli]"
Download the Full Repository
huggingface-cli download NeuralNet-Hub/gemma-4-31B-it-abliterated-uncensored-NVFP4 --local-dir ./gemma-4-31B-NVFP4
[!WARNING]
If deploying on older architectures (Ampere/Ada/Hopper), ensure you use a compatible quantization format. NVFP4 is specifically engineered to utilize the new hardware capabilities of the Blackwell generation.
🌐 Contact Us
NeuralNet is a pioneering AI solutions provider that empowers businesses to harness the power of artificial intelligence.
Website: https://neuralnet.solutions
Email: info[at]neuralnet.solutions