license: apache-2.0
license_link: https://ai.google.dev/gemma/docs/gemma_4_license
base_model: ressl/gemma-4-31B-it-uncensored
base_model_relation: quantized
library_name: mlx
pipeline_tag: image-text-to-text
language:
- en
- de
tags: - mlx
- safetensors
- gemma4
- image-text-to-text
- conversational
- 5bit
- uncensored
- abliterated
- security

ressl/gemma-4-31B-it-uncensored-MLX-5bit
MLX conversion of ressl/gemma-4-31B-it-uncensored for Apple silicon. This release uses 5-bit affine quantization with group size 64 and preserves the multimodal vision path in BF16.
[!WARNING]
This is an uncensored, abliterated research model. It can produce inaccurate, unsafe, illegal, biased, or offensive content. Outputs are not advice. Evaluate the model for your use case, keep human oversight, and comply with applicable law and the Apache License 2.0.
Security research, red teaming, robustness evaluation, and defensive model analysis are the intended uses. Do not use this release to harm people or systems.
Release family
| Format | Repository |
|---|---|
| Transformers BF16 source | ressl/gemma-4-31B-it-uncensored |
| NVIDIA NVFP4 | ressl/gemma-4-31B-it-uncensored-NVFP4 |
| GGUF quantization ladder | ressl/gemma-4-31B-it-uncensored-GGUF |
| MLX 4-bit | ressl/gemma-4-31B-it-uncensored-MLX-4bit |
| MLX 5-bit | ressl/gemma-4-31B-it-uncensored-MLX-5bit |
| MLX 6-bit | ressl/gemma-4-31B-it-uncensored-MLX-6bit |
| MLX 8-bit | ressl/gemma-4-31B-it-uncensored-MLX-8bit |
| MLX BF16 | ressl/gemma-4-31B-it-uncensored-MLX-bf16 |
Verified release facts
| Property | Measured value |
|---|---|
| Source revision | 64c863e92fc131e5f4b0fe3631a0791fe2c19152 |
| Weight format | 5-bit affine quantization with group size 64 |
| Weight size | 20.73 GiB |
| Weight shards | 5 |
| Files in artifact | 16 |
| Vision path | BF16, not quantized |
| Prompt processing | 418.986 tokens/s |
| Generation | 22.042 tokens/s |
| Peak memory | 23.117 GB |
| Conversion and inference stack | mlx-vlm==0.6.4 |
The performance values above are measurements from the release smoke run. They are not estimates or cross-device promises.
Refusal evaluation
The release gate completed all 686 prompts with zero execution errors. A naive keyword heuristic detected broad refusal-like language, while the stricter hard-refusal detector found 0/686 responses that refused without providing substantive help.
| Dataset | Successful | Errors | Naive refusals | Hard refusals |
|---|---|---|---|---|
| JailbreakBench | 100/100 | 0 | 80 | 0 |
| tulu-harmbench | 320/320 | 0 | 135 | 0 |
| NousResearch | 166/166 | 0 | 111 | 0 |
| mlabonne | 100/100 | 0 | 83 | 0 |
| Total | 686/686 | 0 | 409 | 0 |
No hard refusals were detected, so no refusal review was triggered.
Keyword metrics are imperfect and do not prove capability, factuality, or safety. The exact deterministic evaluator and release criteria live in the source repository.
Run on Apple silicon
Install the exact tested version:
python -m pip install mlx-vlm==0.6.4
Text generation:
mlx_vlm.generate --model ressl/gemma-4-31B-it-uncensored-MLX-5bit --prompt "Explain why Alpine flowers survive harsh winters."
Image and text generation:
mlx_vlm.generate --model ressl/gemma-4-31B-it-uncensored-MLX-5bit --image path/to/image.jpg --prompt "Describe this image in detail."
Thinking mode:
mlx_vlm.generate --model ressl/gemma-4-31B-it-uncensored-MLX-5bit --prompt "Solve 37 * 48 step by step." --enable-thinking
OpenAI-compatible local server:
mlx_vlm.server --model ressl/gemma-4-31B-it-uncensored-MLX-5bit --host 127.0.0.1 --port 8080
Quality and limitations
- The language model weights use 5-bit affine quantization with group size 64. Quantization can reduce quality compared with BF16.
- The vision tower and vision embedding path remain BF16 in every MLX release.
- The smoke suite verifies deterministic arithmetic, capitals, German, multi-turn memory, thinking mode, image understanding, basic output health, and measured runtime statistics.
- This checkpoint inherits the source model's limitations and may hallucinate or follow malicious instructions.
- No benchmark result should be generalized beyond the exact prompts, software, and hardware used for the measured run.
Provenance
The checkpoint was converted from commit 64c863e92fc131e5f4b0fe3631a0791fe2c19152. Quantized releases use affine MLX quantization with group size 64 for language weights. Modules whose path contains vision_tower or embed_vision are excluded from quantization. The BF16 release performs no weight quantization.
The original Gemma architecture and weights are provided by Google under the Apache License 2.0. This MLX conversion and the uncensored source release are maintained by Robert Ressl (Hugging Face, Website, LinkedIn, Patreon).
If this work is useful, you can support continued independent model research on Patreon.