← back to catalog · registered 2026-08-22 13:56

prithivMLmods/MiniCPM-V-4.6-abliterated-MAX

prithivMLmods 1.3B multimodal
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/prithivMLmods%2FMiniCPM-V-4.6-abliterated-MAX"
Response includes
  • classification m1
  • files 12
  • hub_downloads_all_time 611
  • author_summary 98 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M1
Primary method

Direct removal

No other method signals detected in this model.
Confidence
MEDIUM
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=0 (base model)
  • no specific method indicators - defaulting to M1 (most common)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
MEDIUM
Why we say so
primary_method=M1; difference-of-means is the reference extraction for M1/M3 (Arditi 2024)
Downloads · lifetime
611
78 last 30d - stable
Likes
3
Descendants
3
in 3 direct forks
Model age
4mo ago
created 2026-05-16

Training datasets

1 of 1 in /datasets

Corpora the author lists in the model card. Datasets tracked in our /datasets catalog carry a category badge linking to the workflow stage. Others open on Hugging Face.

Downloads over time
Now648→from84↑671%
5627248870484 on May 20648 on Oct 11648 on Oct 10MayJunJulAugSepOct
May 20 → Oct 11 · 60 snapshots · spans 144 days

Genealogy 3 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Variants by this author 2 formats · 1K downloads combined

The same weights this author released in different packaging. Pick the format that matches your runtime.

Metadata

License
apache-2.0
Languages
en
Tags
transformers safetensors minicpmv4_6 image-text-to-text minicpm-v multimodal On-Device Model max text-generation-inference reasoning pytorch uncensored

Related

Total size
2.42 GB
Files
12
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-06-01 07:40

Files by quantization

Auxiliary files 12 files 2.44 GB
model-00002-of-00003.safetensors 1.21 GB 6cc917f5 download
model-00001-of-00003.safetensors 1.21 GB 63d1c6bf download
model-00003-of-00003.safetensors 7.60 MB b9498f91 download
tokenizer.json 19.1 MB 33861e37 download
model.safetensors.index.json 75.1 KB a720aeea download
README.md 7.86 KB ab57642c download
chat_template.jinja 7.13 KB f25b6ac3 download
config.json 2.66 KB 5ccf1b64 download
tokenizer_config.json 1.59 KB 0f2b212b download
.gitattributes 1.53 KB 52373fe2 download
processor_config.json 1.17 KB ce6e297c download
generation_config.json 214 B 81a7f6ec download

README current version from Hugging Face


license: apache-2.0
tags:

  • minicpm-v
  • multimodal
  • On-Device Model
  • max
  • text-generation-inference
  • reasoning
  • pytorch
  • uncensored
  • abliterated
  • unfiltered
  • unredacted
  • refusal-ablated
  • alignment-modified
    language:
  • en
    base_model:
  • openbmb/MiniCPM-V-4.6
    datasets:
  • prithivMLmods/harm_bench
    pipeline_tag: image-text-to-text
    library_name: transformers

1

MiniCPM-V-4.6-abliterated-MAX

MiniCPM-V-4.6-abliterated-MAX is an optimized release built on top of huihui-ai/Huihui-MiniCPM-V-4.6-abliterated. This version focuses on updated shard sizing, repository optimization, and compatibility improvements for the latest Transformers releases while preserving the capabilities of the original model. The result is a highly capable and ultra-efficient multimodal language model optimized for image, video, and text understanding with streamlined deployment and inference workflows.

[!IMPORTANT]
This model is intended for research and learning purposes only. Any content generated by it is used at the user’s own risk. The authors and hosting page disclaim any liability for outputs produced by this model. Users are responsible for ensuring safe, ethical, and lawful usage.

Evals

2

.eval_results: harm_bench_score.yaml

The evaluation was conducted using 2,000 random harmful test prompts to measure the refusal behavior of the language model. The self-reported evaluations provided here are intended only to give an overview of the model. Scores may vary depending on the benchmark and the evaluation strategy used.

Key Highlights

  • Latest Transformers Compatibility
    Re-sharded and optimized for improved compatibility with recent Transformers releases.

  • Optimized Model Sharding
    Updated shard sizes for improved repository handling, downloading, and deployment efficiency.

  • Streamlined Inference Experience
    Optimized packaging and repository structure for smoother loading and inference workflows.

  • Efficient Multimodal Architecture
    Built on openbmb/MiniCPM-V-4.6, combining SigLIP2-400M vision encoding with Qwen3.5-0.8B language capabilities for compact yet powerful multimodal understanding.

  • Image & Video Understanding
    Supports advanced reasoning across text, images, and videos with efficient deployment on edge and mobile-class hardware.

  • 262K Long Context Support
    Optimized for extremely long multimodal contexts across text, image, and video inputs.

  • Research-Friendly Distribution
    Designed to simplify experimentation, evaluation, and local deployment workflows.

  • High-Efficiency Deployment
    Suitable for local inference, lightweight multimodal applications, and research experimentation on consumer-grade GPUs.


Base Model Signatures:

This model has been re-sharded and optimized for the latest Transformers version from the base model: https://huggingface.co/huihui-ai/Huihui-MiniCPM-V-4.6-abliterated.


Quick Start with Transformers

pip install transformers==5.8.0 gradio==6.14.0
import gc
import time
from threading import Thread

import gradio as gr
import torch
from PIL import Image

from transformers import (
    MiniCPMV4_6ForConditionalGeneration,
    AutoProcessor,
    TextIteratorStreamer,
)

MAX_MAX_NEW_TOKENS = 4096
DEFAULT_MAX_NEW_TOKENS = 1024
device = torch.device("cuda" if torch.cuda.is_available() else "cpu")
print("Using device:", device)

MODEL_ID = "prithivMLmods/MiniCPM-V-4.6-abliterated-MAX"
processor = AutoProcessor.from_pretrained(MODEL_ID, trust_remote_code=True)
model = MiniCPMV4_6ForConditionalGeneration.from_pretrained(
    MODEL_ID,
    trust_remote_code=True,
    torch_dtype=torch.bfloat16,
).to(device).eval()


def generate(
    image: Image.Image,
    text: str,
    max_new_tokens: int = DEFAULT_MAX_NEW_TOKENS,
    temperature: float = 0.6,
    top_p: float = 0.9,
    top_k: int = 50,
    repetition_penalty: float = 1.2,
):
    if image is None:
        yield "[ERROR] Please upload an image."
        return
    if not text or not text.strip():
        yield "[ERROR] Please enter your instruction."
        return

    messages = [
        {
            "role": "user",
            "content": [
                {"type": "image"},
                {"type": "text", "text": text},
            ],
        }
    ]
    prompt_full = processor.apply_chat_template(
        messages, tokenize=False, add_generation_prompt=True
    )

    inputs = processor(
        text=[prompt_full],
        images=[image],
        return_tensors="pt",
        padding=True,
    ).to(device)

    streamer = TextIteratorStreamer(
        processor.tokenizer if hasattr(processor, "tokenizer") else processor,
        skip_prompt=True,
        skip_special_tokens=True,
    )

    generation_error = {"error": None}
    generation_kwargs = {
        **inputs,
        "streamer": streamer,
        "max_new_tokens": int(max_new_tokens),
        "do_sample": True,
        "temperature": float(temperature),
        "top_p": float(top_p),
        "top_k": int(top_k),
        "repetition_penalty": float(repetition_penalty),
    }

    def _run():
        try:
            model.generate(**generation_kwargs)
        except Exception as e:
            generation_error["error"] = e
            try:
                streamer.end()
            except Exception:
                pass

    thread = Thread(target=_run, daemon=True)
    thread.start()

    buffer = ""
    for new_text in streamer:
        buffer += new_text
        time.sleep(0.01)
        yield buffer

    thread.join(timeout=1.0)

    if generation_error["error"] is not None:
        err = f"[ERROR] {str(generation_error['error'])}"
        yield (buffer + "\n\n" + err) if buffer.strip() else err
        return

    if not buffer.strip():
        yield "[ERROR] No output was generated."

    gc.collect()
    if torch.cuda.is_available():
        torch.cuda.empty_cache()

Base Model Information

openbmb/MiniCPM-V-4.6 is a 1.3B-parameter dense multimodal language model developed by OpenBMB (Tsinghua NLP + ModelBest). It is built using SigLIP2-400M for visual encoding and Qwen3.5-0.8B as the language backbone, optimized for efficient multimodal understanding on edge and mobile hardware while supporting long-context reasoning across text, image, and video modalities.

Intended Use

  • Multimodal Research
    Studying multimodal reasoning, perception, and instruction-following behavior across text, image, and video inputs.

  • Model Evaluation
    Benchmarking and analyzing multimodal language models under a variety of testing conditions.

  • Edge & Local AI Deployment
    Running compact multimodal AI systems efficiently on consumer hardware and edge devices.

  • Research Prototyping
    Experimentation with efficient multimodal transformer architectures and deployment workflows.

Limitations & Risks

Important Note: This model inherits the behavior and characteristics of its base model.

  • Potential Hallucinations
    Multimodal reasoning may occasionally produce inaccurate or fabricated interpretations.

  • User Responsibility
    Outputs should be reviewed before use in critical or high-stakes applications.

  • Multimodal Limitations
    Performance may vary depending on image quality, video complexity, prompt design, and context length.

  • Deployment Considerations
    While optimized for efficiency, high-resolution image and video inference may still require substantial VRAM and optimized runtimes depending on workload complexity.

README history 11 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-06-01Update README.md9a5e46b7.9 KB
    Loading...
  2. 2026-06-01Update README.mdec9018d10.6 KB
    Loading...
  3. 2026-05-17Update README.mdccc11c710.3 KB
    Loading...
  4. 2026-05-17Update README.md37c590110.3 KB
    Loading...
  5. 2026-05-17Update README.md9191e6610.3 KB
    Loading...
  6. 2026-05-17Update README.md665432410.3 KB
    Loading...
  7. 2026-05-16Update README.mdec44cbd9.7 KB
    Loading...
  8. 2026-05-16Update README.md0d210319.6 KB
    Loading...
  9. 2026-05-16Update README.md79dde2f6 KB
    Loading...
  10. 2026-05-16Update README.md37fa1475.6 KB
    Loading...
  11. 2026-05-16initial commit4d003f628 B
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration