← back to catalog · registered 2026-08-22 13:56

ressl/gemma-4-31B-it-uncensored-GGUF

ressl Gemma 31B GGUF second-order 262K ctx
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/ressl%2Fgemma-4-31B-it-uncensored-GGUF"
Response includes
  • classification m8
  • files 9
  • hub_downloads_all_time 18,843
  • author_summary 28 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M8
Primary method

Repackaging (quantization)

Applied on top of direct removal inherited from the base model.
Confidence
MEDIUM
Inherited from base model
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=1
  • assume M1 (base ablation) + M8 (GGUF quant) - default when producer unknown
Refusal direction extraction

No specific extraction method could be identified for this model. The producer either did not document it or used a proprietary pipeline.

What is a refusal direction? →
Downloads · lifetime
19K
2K last 30d - stable
Likes
5
Model age
3mo ago
created 2026-07-08
Downloads over time
Now19.9K→from12.1K↑65%
11.7K14.7K17.7K20.7K12.1K on Jul 1519.9K on Oct 11JulAugSepOct
Jul 15 → Oct 11 · 53 snapshots · spans 88 days

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Variants by this author 4 formats · 2K downloads combined

The same weights this author released in different packaging. Pick the format that matches your runtime.

Metadata

License
apache-2.0
Languages
en de
Tags
gguf uncensored abliterated llama.cpp security gemma4 text-generation en de base_model:ressl/gemma-4-31B-it-uncensored base_model:quantized:ressl/gemma-4-31B-it-uncensored license:apache-2.0

Related

Total size
117 GB
Files
9
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-07-10 08:12

Files by quantization

Auxiliary files 9 files 117 GB
gemma-4-31B-it-uncensored-biproj-q8_0.gguf 30.4 GB 3b3782f0 download
gemma-4-31B-it-uncensored-biproj-q6_k.gguf 23.5 GB 412e9d73 download
gemma-4-31B-it-uncensored-biproj-q5_k_m.gguf 20.3 GB f4d04198 download
gemma-4-31B-it-uncensored-biproj-q4_k_m.gguf 17.4 GB c28b5453 download
gemma-4-31B-it-uncensored-biproj-q3_k_m.gguf 14.2 GB bbece440 download
gemma-4-31B-it-uncensored-biproj-q2_k.gguf 11.1 GB 101d3ebd download
banner.png 1.92 MB f22f05d6 download
README.md 5.52 KB cfad1234 download
.gitattributes 2.00 KB 42608028 download

README current version from Hugging Face


license: apache-2.0
license_link: https://ai.google.dev/gemma/docs/gemma_4_license
base_model: ressl/gemma-4-31B-it-uncensored
base_model_relation: quantized
library_name: gguf
pipeline_tag: text-generation
language: [en, de]
tags:

  • uncensored
  • abliterated
  • gguf
  • llama.cpp
  • security
  • gemma4

gemma-4-31B-it-uncensored

gemma-4-31B-it-uncensored (GGUF)

GGUF K-quant ladder of ressl/gemma-4-31B-it-uncensored
for llama.cpp, Ollama, and LM Studio. Uncensored: 0-1/100 effective refusals at every quant
level (the uncensoring survives quantization down to 2-bit).

⚠️ Genuinely uncensored, it will comply with requests a stock model refuses.

Intended use, the constructive side. A non-refusing assistant is genuinely useful for
ethical hacking, security research, and penetration testing: red-teaming, analyzing malware and
exploit code, writing detection/YARA rules, reviewing vulnerabilities, and studying attack
techniques without the model bailing out mid-task. Use it lawfully and responsibly.

ℹ️ gemma-4 has a thinking mode. llama.cpp enables it by default, so the answer lands in
reasoning_content and content can look empty. For direct answers pass
--reasoning-budget 0 (llama-server) or disable thinking in your client.

Format set

Repository Format Runs on
ressl/gemma-4-31B-it-uncensored Transformers BF16, multimodal transformers, vLLM, SGLang
ressl/gemma-4-31B-it-uncensored-NVFP4 NVIDIA NVFP4, multimodal vLLM, SGLang on Blackwell
ressl/gemma-4-31B-it-uncensored-GGUF GGUF q8_0 to q2_k, text only llama.cpp, Ollama, LM Studio
ressl/gemma-4-31B-it-uncensored-MLX-bf16 MLX BF16, multimodal mlx-vlm on Apple silicon
ressl/gemma-4-31B-it-uncensored-MLX-8bit MLX 8-bit, multimodal mlx-vlm on Apple silicon
ressl/gemma-4-31B-it-uncensored-MLX-6bit MLX 6-bit, multimodal mlx-vlm on Apple silicon
ressl/gemma-4-31B-it-uncensored-MLX-5bit MLX 5-bit, multimodal mlx-vlm on Apple silicon
ressl/gemma-4-31B-it-uncensored-MLX-4bit MLX 4-bit, multimodal mlx-vlm on Apple silicon

Quants

Text-only (llama.cpp drops the vision tower, the BF16/NVFP4 repos keep multimodal).

File Size Use when
…-q8_0.gguf 31 GB maximum quality
…-q6_k.gguf 24 GB near-lossless, smaller
…-q5_k_m.gguf 21 GB high quality
…-q4_k_m.gguf 18 GB recommended default, best size/quality
…-q3_k_m.gguf 15 GB tight VRAM
…-q2_k.gguf 12 GB smallest; still 0/100 uncensored, some quality loss

For full precision use the BF16 repo (an
f16 GGUF exceeds Hugging Face's 50 GB per-file limit and is not hosted here).

All measured at 0/100 hard refusals except q5_k_m (1/100). GPU inference ~75 tok/s (q4, one RTX PRO 6000).

Run it with llama.cpp

llama-server -m gemma-4-31B-it-uncensored-biproj-q4_k_m.gguf \
  -ngl 99 -c 8192 --reasoning-budget 0

Run it with Ollama

# Modelfile:  FROM ./gemma-4-31B-it-uncensored-biproj-q4_k_m.gguf
ollama create gemma4-unc -f Modelfile && ollama run gemma4-unc

Needs a recent llama.cpp/Ollama build with Gemma4 GGUF support (build ≥ 2026-06).

Facts & figures

Base ressl/gemma-4-31B-it-uncensored → google/gemma-4-31B-it
Type uncensored (abliterated) build, 0/686 effective refusals across 4 datasets
Converter llama.cpp convert_hf_to_gguf.py + llama-quantize

Cross-dataset validation

Generalization tested across 686 prompts from 4 independent datasets, 0 effective refusals everywhere:

Dataset Prompts Effective refusals
JailbreakBench 100 0/100
tulu-harmbench 320 0/320
NousResearch/RefusalDataset 166 0/166
mlabonne/harmful_behaviors 100 0/100
Total 686 0/686 (0.0%)

A naive keyword detector flags 363/686 (52.9%), every one is a ***Disclaimer:**-prefixed
compliant answer, not a refusal. (Measured on the shared abliterated weights.)

❤️ Support

Producing and validating this complete format set (BF16 + NVFP4 + a full GGUF ladder, across vLLM,
SGLang and llama.cpp on bleeding-edge Blackwell hardware) was a lot of work. If it's useful to
you, I'd genuinely appreciate your support on Patreon 🙏,
more at ressl.ch.

License & credits

Apache License 2.0, inherited from the base model by Google. See the official Gemma 4 license page. Uncensoring, format set and
validation by Robert Ressl
(Hugging Face · Website · LinkedIn · Patreon).

README history 8 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-07-10Add MLX family links and correct metadata91ed8ae5.5 KB
    Loading...
  2. 2026-07-08Upload README.md with huggingface_huba535c0b4 KB
    Loading...
  3. 2026-07-08Upload README.md with huggingface_hubfecc87d4 KB
    Loading...
  4. 2026-07-08Upload README.md with huggingface_hub52af4ed4 KB
    Loading...
  5. 2026-07-08Upload README.md with huggingface_hub9e949bb3.7 KB
    Loading...
  6. 2026-07-08Upload README.md with huggingface_hub2ae4a5e3.3 KB
    Loading...
  7. 2026-07-08Upload README.md with huggingface_hubdb9c9043 KB
    Loading...
  8. 2026-07-08Upload README.md with huggingface_hubc520f2e2.9 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration