← back to catalog · registered 2026-08-22 13:56

unalignment/Pixtral-12B-Captioner-Relaxed

unalignment Mistral 12B
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/unalignment%2FPixtral-12B-Captioner-Relaxed"
Response includes
  • classification unknown
  • files 18
  • author_summary 4 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
?
Primary method

Unclassified

No clear signals of an abliteration technique in this model.
Confidence
UNKNOWN
Why this label 1 signal
No classification signals present. This may not be an abliterated model at all - it could be a repackaging, a merge with unrelated goals, or unrelated content that mentions the term.
  • no classification signals present (no abliterated, uncensored, or known producer/method markers)
Refusal direction extraction

No specific extraction method could be identified for this model. The producer either did not document it or used a proprietary pipeline.

What is a refusal direction? →
Downloads · 30-day
43
↑ 5,860% in 90 days
Likes
8
Model age
20mo ago
created 2025-01-22
Downloads over time
Now596→from10↑5,860%
021843665510 on Jan 22, 2025596 on Oct 11Jan '25Apr '25Jul '25Oct '25JanAprJulOct
Jan 22, 2025 → Oct 11 · 129 snapshots · spans 627 days

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
apache-2.0
Languages
en
Tags
transformers safetensors llava image-text-to-text image-to-text en base_model:mistralai/Pixtral-12B-2409 base_model:finetune:mistralai/Pixtral-12B-2409 license:apache-2.0 endpoints_compatible region:us

Related

Total size
23.6 GB
Files
18
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2025-01-22 08:48

Files by quantization

Auxiliary files 18 files 23.6 GB
model-00001-of-00006.safetensors 4.65 GB c258e133 download
model-00002-of-00006.safetensors 4.62 GB 7a21d04c download
model-00003-of-00006.safetensors 4.57 GB 47e3b21b download
model-00004-of-00006.safetensors 4.57 GB 1dff0520 download
model-00005-of-00006.safetensors 3.97 GB 0a30f27d download
model-00006-of-00006.safetensors 1.25 GB a87c5ff6 download
tokenizer.json 16.3 MB fcdf3e6b download
tokenizer.model.v7m1 574 KB d23856f8 download
tokenizer_config.json 173 KB 74278eb5 download
model.safetensors.index.json 56.5 KB 9ba4ef92 download
config.json 4.43 KB e1b5eafb download
README.md 4.01 KB ee20b69c download
chat_template.json 1.59 KB 33d3db1c download
.gitattributes 1.53 KB 52373fe2 download
preprocessor_config.json 483 B 3e916d77 download
special_tokens_map.json 414 B 451134b2 download
processor_config.json 162 B c9b049db download
generation_config.json 111 B 31dcf9d4 download

README current version from Hugging Face


library_name: transformers
license: apache-2.0
language:

  • en
    base_model:
  • mistralai/Pixtral-12B-2409
    pipeline_tag: image-to-text

Pixtral-12B-Captioner-Relaxed

Introduction

Original Pixtral-12B-Captioner-Relaxed from https://huggingface.co/Ertugrul/Pixtral-12B-Captioner-Relaxed/tree/main with tokenizer fix

Pixtral-12B-Captioner-Relaxed is an instruction-tuned version of Pixtral-12B-2409, an advanced multimodal large language model. This fine-tuned version is based on a hand-curated dataset for text-to-image models, providing significantly more detailed descriptions of given images.

Key Features:

  • Enhanced Detail: Generates more comprehensive and nuanced image descriptions.
  • Relaxed Constraints: Offers less restrictive image descriptions compared to the base model.
  • Natural Language Output: Describes different subjects in the image while specifying their locations using natural language.
  • Optimized for Image Generation: Produces captions in formats compatible with state-of-the-art text-to-image generation models.

Note: This fine-tuned model is optimized for creating text-to-image datasets. As a result, performance on other complex tasks may be lower compared to the original model.

Requirements

The 12B model needs 24GB of VRAM at half precision. Model can be loaded at 8 bit or 4 bit quantization but expect degraded performance.

Quickstart

from PIL import Image
from transformers import LlavaForConditionalGeneration, AutoProcessor
from transformers import BitsAndBytesConfig
import torch
import matplotlib.pyplot as plt



# example quantization config, add it to model load parameters to use 4bit quantization
quantization_config = BitsAndBytesConfig(
    # load_in_8bit=True,
    load_in_4bit=True,
    bnb_4bit_compute_dtype=torch.bfloat16,
    bnb_4bit_quant_type="nf4"
    )



model_id = "Ertugrul/Pixtral-12B-Captioner-Relaxed"
model = LlavaForConditionalGeneration.from_pretrained(model_id, device_map="auto", torch_dtype=torch.bfloat16)
processor = AutoProcessor.from_pretrained(model_id)

# for quantization just use this instead of previous load
# model = LlavaForConditionalGeneration.from_pretrained(model_id, device_map="auto", torch_dtype=torch.bfloat16, quantization_config=quantization_config)

conversation = [
    {
        "role": "user",
        "content": [
            
            {"type": "text", "text": "Describe the image.\n"},
            {
                "type": "image",
            }
        ],
    }
]

PROMPT = processor.apply_chat_template(conversation, add_generation_prompt=True)

image = Image.open(r"PATH_TO_YOUR_IMAGE")

def resize_image(image, target_size=768):
    """Resize the image to have the target size on the shortest side."""
    width, height = image.size
    if width < height:
        new_width = target_size
        new_height = int(height * (new_width / width))
    else:
        new_height = target_size
        new_width = int(width * (new_height / height))
    return image.resize((new_width, new_height), Image.LANCZOS)


# you can try different resolutions or disable it completely
image = resize_image(image, 768)


inputs = processor(text=PROMPT, images=image, return_tensors="pt").to("cuda")


with torch.no_grad():
    with torch.autocast(device_type="cuda", dtype=torch.bfloat16):
        generate_ids = model.generate(**inputs, max_new_tokens=384, do_sample=True, temperature=0.3, use_cache=True, top_k=20)
output_text = processor.batch_decode(generate_ids[:, inputs.input_ids.shape[1]:], skip_special_tokens=True, clean_up_tokenization_spaces=True)[0]

print(output_text)

Acknowledgements

For more detailed options, refer to the Pixtral-12B-2409 or mistral-community/pixtral-12b documentation.

You can also try the Qwen2-VL-7B-Captioner-Relaxed, for an alternative smaller model. It's trianed in a similar manner.

README history 1 version

The author's README evolved over time. Click a version to see its content at that point.

  1. 2025-01-22Upload folder using huggingface_hub54535d44 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration