← back to catalog · registered 2026-08-22 13:56

prithivMLmods/gemma-3-27b-it-abliterated-FP8

prithivMLmods Gemma 26B multimodal second-order
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/prithivMLmods%2Fgemma-3-27b-it-abliterated-FP8"
Response includes
  • classification m1
  • files 20
  • benchmarks 11 entries
  • hub_downloads_all_time 359
  • author_summary 98 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M1
Primary method

Direct removal

No other method signals detected in this model.
Confidence
MEDIUM
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=0 (base model)
  • no specific method indicators - defaulting to M1 (most common)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
MEDIUM
Why we say so
primary_method=M1; difference-of-means is the reference extraction for M1/M3 (Arditi 2024)
Downloads · lifetime
359
17 last 30d - cooling
Likes
2
Model age
7mo ago
created 2026-02-16
Downloads over time
Now364→from56↑550%
4115927739556 on Feb 18364 on Oct 11364 on Oct 4FebAprJunAugOct
Feb 18 → Oct 11 · 73 snapshots · spans 235 days

Benchmarks

Benchmark Score Source
Entertainment 1.5 UGI
Hazardous 2.4 UGI
Natural Intelligence 29.6 UGI
Political lean -7.7% UGI
Sensitive-Info 20.32 UGI
SocPol 2.4 UGI
UGI 41.05 UGI
Willingness (10) 8.2 UGI
W10-Adherence 7.5 UGI
W10-Direct 9 UGI
Writing 35.62 UGI

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
gemma
Languages
en
Tags
transformers safetensors gemma3 image-text-to-text text-generation-inference abliterated uncensored vllm fp8 conversational en base_model:mlabonne/gemma-3-27b-it-abliterated

Related

Total size
26.9 GB
Files
20
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-02-16 17:19

Files by quantization

Auxiliary files 20 files 26.9 GB
model-00001-of-00006.safetensors 4.63 GB c31b7b83 download
model-00003-of-00006.safetensors 4.62 GB 5669b41e download
model-00004-of-00006.safetensors 4.62 GB ebb6b878 download
model-00005-of-00006.safetensors 4.62 GB 50116152 download
model-00002-of-00006.safetensors 4.62 GB 4f42237c download
model-00006-of-00006.safetensors 3.79 GB e4fa2900 download
tokenizer.json 31.8 MB 4667f208 download
tokenizer.model 4.47 MB 1299c11d download
tokenizer_config.json 1.10 MB ba114d51 download
model.safetensors.index.json 185 KB 2dee3aa0 download
config.json 4.55 KB d501ef7e download
README.md 3.91 KB 514eaf1d download
.gitattributes 1.53 KB 52373fe2 download
chat_template.jinja 1.50 KB 1117055a download
special_tokens_map.json 662 B 1a619324 download
preprocessor_config.json 570 B b1e00fc1 download
recipe.yaml 172 B 5a04c3bc download
generation_config.json 151 B 4caab894 download
processor_config.json 70.0 B 453c7966 download
added_tokens.json 35.0 B e17bde03 download

README current version from Hugging Face


license: gemma
language:

  • en
    base_model:
  • mlabonne/gemma-3-27b-it-abliterated
    pipeline_tag: image-text-to-text
    library_name: transformers
    tags:
  • text-generation-inference
  • abliterated
  • uncensored
  • vllm
  • fp8

1

gemma-3-27b-it-abliterated-FP8

gemma-3-27b-it-abliterated-FP8 is an FP8-Dynamic compressed variant of Maxime Labonne’s gemma-3-27b-it-abliterated model. This version applies FP8 dynamic quantization while preserving the layerwise abliteration technique that minimizes refusal behavior across Gemma 3’s deep architecture. The result is a highly capable 27B instruction-tuned model with improved hardware efficiency and reduced memory footprint.

[!important]
FP8 (8-bit floating point) weight and activation quantization using hardware acceleration on GPUs – FP8 W8A8. Quantization W8A8 FP8-dynamic recipe – examples.

Model Overview

mlabonne/gemma-3-27b-it-abliterated is an experimental uncensored 27B-parameter instruction-tuned language model derived from Google’s gemma-3-27b-it.

It introduces a novel layerwise abliteration technique that:

  • Independently computes refusal directions from hidden states in each of the model’s 60+ layers
  • Targets key attention modules such as down_proj, o_proj, and feedforward components
  • Applies a 1.5× refusal weight scaling to eliminate safety refusals
  • Preserves >90% acceptance rate and coherent generation capabilities
  • Outperforms traditional residual stream–based removal methods on Gemma 3’s resilient architecture

FP8-Dynamic Compression

This FP8 edition:

  • Uses BF16 · FP8 (F8_E4M3) precision formats
  • Applies dynamic FP8 quantization for improved inference throughput
  • Reduces VRAM consumption significantly compared to full BF16
  • Maintains strong generation quality and reasoning stability

Designed for deployment on Hopper and compatible GPU architectures supporting FP8.

Quick Start with Transformers

from transformers import AutoProcessor, Gemma3ForConditionalGeneration
from PIL import Image
import requests
import torch

model_id = "prithivMLmods/gemma-3-27b-it-abliterated-FP8"

model = Gemma3ForConditionalGeneration.from_pretrained(
    model_id, device_map="auto"
).eval()

processor = AutoProcessor.from_pretrained(model_id)

messages = [
    {
        "role": "system",
        "content": [{"type": "text", "text": "You are a helpful assistant."}]
    },
    {
        "role": "user",
        "content": [
            {"type": "image", "image": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/bee.jpg"},
            {"type": "text", "text": "Describe this image in detail."}
        ]
    }
]

inputs = processor.apply_chat_template(
    messages, add_generation_prompt=True, tokenize=True,
    return_dict=True, return_tensors="pt"
).to(model.device, dtype=torch.bfloat16)

input_len = inputs["input_ids"].shape[-1]

with torch.inference_mode():
    generation = model.generate(**inputs, max_new_tokens=100, do_sample=False)
    generation = generation[0][input_len:]

decoded = processor.decode(generation, skip_special_tokens=True)
print(decoded)

Intended Use

  • Behavioral research and refusal-mechanism analysis
  • High-capacity instruction-following experiments
  • Long-form reasoning and detailed generation tasks
  • Research on quantization effects in large uncensored LLMs

Limitations & Risks

Critical Note: This model minimizes built-in refusal mechanisms.

  • May generate explicit or controversial outputs
  • Requires responsible and ethical use
  • FP8 requires compatible GPU architectures
  • 27B parameter size still demands substantial VRAM even with compression

README history 3 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-02-16Update README.md7daf6143.9 KB
    Loading...
  2. 2026-02-16Update README.md307e13b3.8 KB
    Loading...
  3. 2026-02-16Create README.md4e04dae181 B
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration