← back to catalog · registered 2026-08-22 13:56

sdasd112132/Vision-8B-MiniCPM-2_5-Uncensored-and-Detailed-4bit

Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/sdasd112132%2FVision-8B-MiniCPM-2_5-Uncensored-and-Detailed-4bit"
Response includes
  • classification m-uncensored
  • files 13
  • hub_downloads_all_time 5,963
  • author_summary 1 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M-U
Primary method

Uncensored (method unknown)

No other method signals detected in this model.
Confidence
LOW
Why this label 3 signals
Weak or ambiguous signals. Best guess based on catalog patterns; treat as tentative and check the evidence below.
  • 'uncensored' in name/tags but no 'abliterated' marker
  • method not identifiable from author declaration alone
  • may be DPO fine-tune, prompt engineering, or unknown technique
Refusal direction extraction

No specific extraction method could be identified for this model. The producer either did not document it or used a proprietary pipeline.

What is a refusal direction? →
Downloads · lifetime
6K
29 last 30d - cooling
Likes
32
Model age
2.4y ago
created 2024-05-30
Downloads over time
Now6K→from274↑2,081%
02.2K4.4K6.6K274 on Jul 24, 20246K on Oct 11Jul '24Nov '24Mar '25Jul '25Nov '25MarJul
Jul 24, 2024 → Oct 11 · 155 snapshots · spans 809 days

Metadata

Tags
transformers safetensors minicpmv feature-extraction visual-question-answering custom_code 4-bit bitsandbytes region:us
Total size
5.74 GB
Files
13
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2024-06-01 15:56

Files by quantization

Auxiliary files 13 files 5.75 GB
model-00001-of-00002.safetensors 4.33 GB f58a04f3 download
model-00002-of-00002.safetensors 1.40 GB 888dd240 download
tokenizer.json 8.66 MB 288bd3af download
model.safetensors.index.json 240 KB b130f4cb download
tokenizer_config.json 49.7 KB caac7cd7 download
modeling_minicpmv.py 24.4 KB a7fa49a4 download
resampler.py 5.36 KB 9b1613c9 download
configuration_minicpm.py 3.97 KB f3f40d43 download
config.json 1.87 KB f493f0e7 download
README.md 1.53 KB 8904efa0 download
.gitattributes 1.48 KB a6344aac download
special_tokens_map.json 459 B 5eddf623 download
generation_config.json 121 B 04780ee6 download

README current version from Hugging Face


pipeline_tag: visual-question-answering

MiniCPM-Llama3-V 2.5 int4

This is the int4 quantized version of MiniCPM-Llama3-V 2.5.
Running with int4 version would use lower GPU mermory (about 9GB).

Usage

Inference using Huggingface transformers on NVIDIA GPUs. Requirements tested on python 3.10:

Pillow==10.1.0
torch==2.1.2
torchvision==0.16.2
transformers==4.40.0
sentencepiece==0.1.99
accelerate==0.30.1
bitsandbytes==0.43.1
# test.py
import torch
from PIL import Image
from transformers import AutoModel, AutoTokenizer

model = AutoModel.from_pretrained('openbmb/MiniCPM-Llama3-V-2_5-int4', trust_remote_code=True)
tokenizer = AutoTokenizer.from_pretrained('openbmb/MiniCPM-Llama3-V-2_5-int4', trust_remote_code=True)
model.eval()

image = Image.open('xx.jpg').convert('RGB')
question = 'What is in the image?'
msgs = [{'role': 'user', 'content': question}]

res = model.chat(
    image=image,
    msgs=msgs,
    tokenizer=tokenizer,
    sampling=True, # if sampling=False, beam_search will be used by default
    temperature=0.7,
    # system_prompt='' # pass system_prompt if needed
)
print(res)

## if you want to use streaming, please make sure sampling=True and stream=True
## the model.chat will return a generator
res = model.chat(
    image=image,
    msgs=msgs,
    tokenizer=tokenizer,
    sampling=True,
    temperature=0.7,
    stream=True
)

generated_text = ""
for new_text in res:
    generated_text += new_text
    print(new_text, flush=True, end='')

README history 2 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2024-06-01Upload 12 filescfce3ee1.5 KB
    Loading...
  2. 2024-05-30Upload MiniCPMVf078c165.1 KB
    Loading...

Discussions 4 threads

  1. 2025-07-25PRtestclosed1 💬#4
    Loading...
  2. 2024-09-19Could you provide a GGUF format?closed2 💬#3
    Loading...
  3. 2024-08-17you say it's the quantized version of https://huggingface.co/openbmb/MiniCPM-Ll…open4 💬#2
    Loading...
  4. 2024-07-11full precision?closed3 💬#1
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration