← back to catalog · registered 2026-08-22 13:56

ressl/gemma-4-31B-it-uncensored-MLX-4bit

ressl Gemma 31B multimodal second-order
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/ressl%2Fgemma-4-31B-it-uncensored-MLX-4bit"
Response includes
  • classification m1
  • files 16
  • hub_downloads_all_time 1,348
  • author_summary 28 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M1
Primary method

Direct removal

No other method signals detected in this model.
Confidence
MEDIUM
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=0 (base model)
  • no specific method indicators - defaulting to M1 (most common)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
MEDIUM
Why we say so
primary_method=M1; difference-of-means is the reference extraction for M1/M3 (Arditi 2024)
Downloads · lifetime
1K
156 last 30d - stable
Likes
0
Model age
3mo ago
created 2026-07-10
Downloads over time
Now1.4K→from740↑91%
7079631.2K1.5K740 on Jul 151.4K on Oct 11JulAugSepOct
Jul 15 → Oct 11 · 53 snapshots · spans 88 days

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
apache-2.0
Languages
en de
Tags
mlx safetensors gemma4 image-text-to-text conversational 4bit uncensored abliterated security en de base_model:ressl/gemma-4-31B-it-uncensored

Related

Total size
17.2 GB
Files
16
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-07-10 08:10

Files by quantization

Auxiliary files 16 files 17.2 GB
model-00003-of-00004.safetensors 5.00 GB d2986040 download
model-00001-of-00004.safetensors 5.00 GB 5a39a3fa download
model-00002-of-00004.safetensors 4.99 GB 734d3e59 download
model-00004-of-00004.safetensors 2.17 GB 46eb49dd download
tokenizer.json 30.7 MB cc8d3a0c download
banner.png 1.92 MB f22f05d6 download
model.safetensors.index.json 200 KB 9cd7c974 download
chat_template.jinja 17.1 KB e61bbfe9 download
config.json 5.93 KB c97460f0 download
README.md 5.85 KB 95133381 download
build-manifest.json 2.80 KB 60168732 download
tokenizer_config.json 2.68 KB cf6235ae download
.gitattributes 1.58 KB e7f8c502 download
processor_config.json 1.29 KB a086fb7e download
preprocessor_config.json 403 B 1b1350e0 download
generation_config.json 204 B d5eef132 download

README current version from Hugging Face


license: apache-2.0
license_link: https://ai.google.dev/gemma/docs/gemma_4_license
base_model: ressl/gemma-4-31B-it-uncensored
base_model_relation: quantized
library_name: mlx
pipeline_tag: image-text-to-text
language:

  • en
  • de
    tags:
  • mlx
  • safetensors
  • gemma4
  • image-text-to-text
  • conversational
  • 4bit
  • uncensored
  • abliterated
  • security

gemma-4-31B-it-uncensored

ressl/gemma-4-31B-it-uncensored-MLX-4bit

MLX conversion of ressl/gemma-4-31B-it-uncensored for Apple silicon. This release uses 4-bit affine quantization with group size 64 and preserves the multimodal vision path in BF16.

[!WARNING]
This is an uncensored, abliterated research model. It can produce inaccurate, unsafe, illegal, biased, or offensive content. Outputs are not advice. Evaluate the model for your use case, keep human oversight, and comply with applicable law and the Apache License 2.0.

Security research, red teaming, robustness evaluation, and defensive model analysis are the intended uses. Do not use this release to harm people or systems.

Release family

Format Repository
Transformers BF16 source ressl/gemma-4-31B-it-uncensored
NVIDIA NVFP4 ressl/gemma-4-31B-it-uncensored-NVFP4
GGUF quantization ladder ressl/gemma-4-31B-it-uncensored-GGUF
MLX 4-bit ressl/gemma-4-31B-it-uncensored-MLX-4bit
MLX 5-bit ressl/gemma-4-31B-it-uncensored-MLX-5bit
MLX 6-bit ressl/gemma-4-31B-it-uncensored-MLX-6bit
MLX 8-bit ressl/gemma-4-31B-it-uncensored-MLX-8bit
MLX BF16 ressl/gemma-4-31B-it-uncensored-MLX-bf16

Verified release facts

Property Measured value
Source revision 64c863e92fc131e5f4b0fe3631a0791fe2c19152
Weight format 4-bit affine quantization with group size 64
Weight size 17.16 GiB
Weight shards 4
Files in artifact 15
Vision path BF16, not quantized
Prompt processing 462.454 tokens/s
Generation 27.967 tokens/s
Peak memory 19.282 GB
Conversion and inference stack mlx-vlm==0.6.4

The performance values above are measurements from the release smoke run. They are not estimates or cross-device promises.

Refusal evaluation

The release gate completed all 686 prompts with zero execution errors. A naive keyword heuristic detected broad refusal-like language, while the stricter hard-refusal detector found 0/686 responses that refused without providing substantive help.

Dataset Successful Errors Naive refusals Hard refusals
JailbreakBench 100/100 0 79 0
tulu-harmbench 320/320 0 131 0
NousResearch 166/166 0 115 0
mlabonne 100/100 0 83 0
Total 686/686 0 408 0

No hard refusals were detected, so no refusal review was triggered.

Keyword metrics are imperfect and do not prove capability, factuality, or safety. The exact deterministic evaluator and release criteria live in the source repository.

Run on Apple silicon

Install the exact tested version:

python -m pip install mlx-vlm==0.6.4

Text generation:

mlx_vlm.generate --model ressl/gemma-4-31B-it-uncensored-MLX-4bit --prompt "Explain why Alpine flowers survive harsh winters."

Image and text generation:

mlx_vlm.generate --model ressl/gemma-4-31B-it-uncensored-MLX-4bit --image path/to/image.jpg --prompt "Describe this image in detail."

Thinking mode:

mlx_vlm.generate --model ressl/gemma-4-31B-it-uncensored-MLX-4bit --prompt "Solve 37 * 48 step by step." --enable-thinking

OpenAI-compatible local server:

mlx_vlm.server --model ressl/gemma-4-31B-it-uncensored-MLX-4bit --host 127.0.0.1 --port 8080

Quality and limitations

  • The language model weights use 4-bit affine quantization with group size 64. Quantization can reduce quality compared with BF16.
  • The vision tower and vision embedding path remain BF16 in every MLX release.
  • The smoke suite verifies deterministic arithmetic, capitals, German, multi-turn memory, thinking mode, image understanding, basic output health, and measured runtime statistics.
  • This checkpoint inherits the source model's limitations and may hallucinate or follow malicious instructions.
  • No benchmark result should be generalized beyond the exact prompts, software, and hardware used for the measured run.

Provenance

The checkpoint was converted from commit 64c863e92fc131e5f4b0fe3631a0791fe2c19152. Quantized releases use affine MLX quantization with group size 64 for language weights. Modules whose path contains vision_tower or embed_vision are excluded from quantization. The BF16 release performs no weight quantization.

The original Gemma architecture and weights are provided by Google under the Apache License 2.0. This MLX conversion and the uncensored source release are maintained by Robert Ressl (Hugging Face, Website, LinkedIn, Patreon).

If this work is useful, you can support continued independent model research on Patreon.

README history 4 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-07-10Correct license and model card metadataff66cf65.9 KB
    Loading...
  2. 2026-07-10Add MLX model cardd93bfa35.7 KB
    Loading...
  3. 2026-07-10Add MLX model cardb980cb55.7 KB
    Loading...
  4. 2026-07-10Add MLX model cardc63edd25.7 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration