← back to catalog · registered 2026-08-22 13:56

jduartedj/MiniCPM-V-4.6-35B-Abliterated

jduartedj Qwen 35B GGUF MoE multimodal 262K ctx
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/jduartedj%2FMiniCPM-V-4.6-35B-Abliterated"
Response includes
  • classification m8
  • files 27
  • hub_downloads_all_time 3,449
  • author_summary 2 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M8
Primary method

Repackaging (quantization)

Applied on top of direct removal inherited from the base model.
Confidence
MEDIUM
Inherited from base model
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=1
  • assume M1 (base ablation) + M8 (GGUF quant) - default when producer unknown
Refusal direction extraction

No specific extraction method could be identified for this model. The producer either did not document it or used a proprietary pipeline.

What is a refusal direction? →
Downloads · lifetime
3K
337 last 30d - cooling
Likes
1
Model age
4mo ago
created 2026-05-14
Downloads over time
Now3.5K→from17↑20,641%
01.3K2.6K3.9K17 on May 133.5K on Oct 11MayJunJulAugSepOct
May 13 → Oct 11 · 61 snapshots · spans 151 days

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
apache-2.0
Languages
en zh
Quantizations
Q4_K
Tags
safetensors gguf minicpmv4_6 multimodal vision abliterated uncensored qwen3.5 minicpm moe vision-language image-text-to-text

Related

Total size
85.3 GB
Files
27
Quantizations
4
Registered
2026-08-22 13:56
Last updated on HF
2026-06-05 17:54

Files by quantization

Q4_K 1 file 19.7 GB
ggml-model-Q4_K_M.gguf 19.7 GB d864cd5c download
F16 1 file 1.04 GB
mmproj-model-f16.gguf 1.04 GB d263ab34 download
mmproj 1 file 1.04 GB
mmproj-stage4.gguf 1.04 GB d263ab34 download
Auxiliary files 24 files 65.6 GB
model-00011-of-00015.safetensors 4.76 GB c58ea15c download
model-00005-of-00015.safetensors 4.76 GB 2c4444c2 download
model-00009-of-00015.safetensors 4.75 GB fe375db8 download
model-00002-of-00015.safetensors 4.70 GB 869b6f6b download
model-00003-of-00015.safetensors 4.70 GB 381cf4b2 download
model-00006-of-00015.safetensors 4.70 GB fe61fe66 download
model-00008-of-00015.safetensors 4.70 GB 2ea41f0f download
model-00012-of-00015.safetensors 4.70 GB c0552df0 download
model-00013-of-00015.safetensors 4.70 GB 850e59ca download
model-00001-of-00015.safetensors 4.65 GB e275d5b2 download
model-00004-of-00015.safetensors 4.65 GB a06f9c63 download
model-00007-of-00015.safetensors 4.65 GB 9e06323c download
model-00010-of-00015.safetensors 4.65 GB 848f434a download
model-00014-of-00015.safetensors 3.52 GB 63520284 download
model-00015-of-00015.safetensors 1.03 GB 0a41f87b download
tokenizer.json 19.1 MB 33861e37 download
model.safetensors.index.json 112 KB 5e426c61 download
tokenizer_config.json 8.79 KB 9b6fab07 download
README.md 8.16 KB 6fcd50a4 download
chat_template.jinja 7.13 KB f25b6ac3 download
config.json 2.97 KB a18236db download
.gitattributes 1.70 KB 486b9de6 download
preprocessor_config.json 315 B f8619291 download
generation_config.json 214 B b80dd7cb download

README current version from Hugging Face


license: apache-2.0
language:

  • en
  • zh
    base_model:
  • openbmb/MiniCPM-V-4.6
  • huihui-ai/Huihui-Qwen3.5-35B-A3B-abliterated
    pipeline_tag: image-text-to-text
    tags:
  • safetensors
  • minicpmv4_6
  • multimodal
  • vision
  • abliterated
  • uncensored
  • qwen3.5
  • minicpm
  • moe
  • vision-language
  • image-text-to-text
  • conversational
  • computer-use
  • ui-grounding

MiniCPM-V-4.6-35B-Abliterated

A multimodal vision-language model combining:

  • Vision: openbmb/MiniCPM-V-4.6 vision tower (SigLIP 400M, 27 encoder layers + ViT merger)
  • Language: huihui-ai/Huihui-Qwen3.5-35B-A3B-abliterated (Qwen3.5-35B-A3B MoE, abliterated for uncensored text generation)
  • Merger: Trained MLP bridge (4608→2048) connecting vision to language, trained across a 4-stage curriculum (pretrain → instruction-tune → UI grounding → SFT polish)

⚙️ Latest update (2026-06-05): model weights now ship with the Stage 4 / step 23706 trained merger (val_loss 1.0280). See Training below.

Quick Stats

Property Value
Total params ~35B (3B active per token, MoE 256 experts × 8 active)
Vision tower 1152 hidden, 27 layers, SigLIP-400M + ViT merger
Merger 4608 → 2048, single DownsampleMLP (LayerNorm → Linear → GELU → Linear)
LLM hidden size 2048
Precision BF16 (also ships Q4_K_M GGUF for llama.cpp)
Disk 65.6 GB safetensors, 21.2 GB GGUF Q4_K_M
Inference ~75 GB BF16 VRAM, ~28 GB NF4 (bitsandbytes)

How It Was Made

  1. Assembled by splicing weights:
    • Vision tower + ViT merger from openbmb/MiniCPM-V-4.6
    • Language model + lm_head from huihui-ai/Huihui-Qwen3.5-35B-A3B-abliterated
    • Merger MLP Xavier-initialized
  2. Merger trained in 4 stages on NVIDIA GB10 (128 GB unified memory) using a custom standalone trainer (~2.4 GB GPU footprint — only vision tower + merger + embed_tokens loaded, LLM frozen).

Training

The merger MLP (~22 tensors, 270 MB) was trained in a 4-stage curriculum. Vision tower and LLM remain frozen throughout. Each stage uses next-token cross-entropy loss against the LLM's frozen embedding/lm_head (Stages 2-4) or proxy MSE loss (Stage 1).

Stage Curriculum

Stage Name Dataset Steps LR Final val_loss
1 Pretrain Alignment LLaVA-Pretrain (558K image-caption) 558,128 2e-4 — (MSE only)
2 Instruction Tuning LLaVA-Instruct-150k 157,712 5e-5 1.0855
3 UI Grounding ScreenSpot + UI variants 1,211 5e-5 4.1910
4 SFT Polish agentsea/wave-ui-25k 23,706 5e-5 1.0280 ✓

Each stage trains on top of the previous stage's converged merger.

Stage 4 Eval Trajectory (the shipped checkpoint)

Step val_loss
16000 1.0318
18000 1.0302
20000 1.0286
22000 1.0282
23706 1.0280 ✓ shipped

Validation loss is monotonically decreasing; the final checkpoint is the best by val_loss and is the one packaged in model-00015-of-00015.safetensors and mmproj-stage4.gguf.

Hyperparameters (all stages)

  • Optimizer: AdamW (β1=0.9, β2=0.999, weight_decay=0.01)
  • LR schedule: linear warmup + cosine decay
  • Gradient clipping: max_norm=1.0
  • Batch size: 1 with grad accumulation (Stages 1: 12 / 2: 16 / 3: 8 / 4: 8)
  • Hardware: NVIDIA GB10 (128 GB unified memory)
  • Only merger + ViT-merger weights trained; vision tower core and LLM frozen

Usage

Transformers

from transformers import AutoModelForCausalLM, AutoProcessor
from PIL import Image

model = AutoModelForCausalLM.from_pretrained(
    "jduartedj/MiniCPM-V-4.6-35B-Abliterated",
    trust_remote_code=True,
    torch_dtype="auto",
    device_map="auto",
)
processor = AutoProcessor.from_pretrained(
    "jduartedj/MiniCPM-V-4.6-35B-Abliterated",
    trust_remote_code=True,
)

image = Image.open("your_image.jpg").convert("RGB")
messages = [
    {"role": "user", "content": [
        {"type": "image"},
        {"type": "text", "text": "Describe this image in detail."},
    ]},
]

text = processor.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)
inputs = processor(text=[text], images=[image], return_tensors="pt").to(model.device)
output = model.generate(**inputs, max_new_tokens=512)
print(processor.decode(output[0], skip_special_tokens=True))

NF4 (bitsandbytes) for ~28 GB inference

from transformers import AutoModelForCausalLM, BitsAndBytesConfig
import torch

bnb = BitsAndBytesConfig(
    load_in_4bit=True,
    bnb_4bit_compute_dtype=torch.bfloat16,
    bnb_4bit_quant_type="nf4",
    bnb_4bit_use_double_quant=True,
)
model = AutoModelForCausalLM.from_pretrained(
    "jduartedj/MiniCPM-V-4.6-35B-Abliterated",
    trust_remote_code=True,
    quantization_config=bnb,
    device_map={"": 0},
    dtype=torch.bfloat16,
    # CRITICAL: keep vision tower, ViT merger, projector merger, and lm_head
    # in bf16. Skipping only via llm_int8_skip_modules silently fails on the
    # 4-bit code path for sub-module dotted names. See caveat below.
    modules_to_not_convert=[
        "merger",                   # model.merger.mlp.0.*
        "vision_tower.vit_merger",  # model.vision_tower.vit_merger.*
        "lm_head",
    ],
)

⚠️ NF4 loader caveat — use modules_to_not_convert. BitsAndBytesConfig(llm_int8_skip_modules=[...]) alone is not honored on the 4-bit code path for sub-module dotted names. Without the kwarg below, 8 of the 22 merger Linear layers get silently wrapped as bitsandbytes.Linear4bit, and any post-load overlay (state_dict.copy_(...)) will silently fail on the 4-bit-packed .weight tensors. The fix is to pass modules_to_not_convert=["merger", "vision_tower.vit_merger", "lm_head"] directly to from_pretrained, as shown in the NF4 snippet above. The shipped weights already contain the trained merger, so no overlay is needed — the kwarg above is enough.

llama.cpp

# Use the matched stage-4 mmproj:
./llama-mtmd-cli \
  -m ggml-model-Q4_K_M.gguf \
  --mmproj mmproj-stage4.gguf \
  --image your_image.jpg \
  -p "Describe this image."

Requirements

  • transformers >= 5.7.0 (native minicpmv4_6 support)
  • torch >= 2.1.0
  • torchvision
  • bitsandbytes >= 0.43 (for NF4 path)
  • ~67 GB disk for safetensors, ~21 GB for GGUF
  • ~75 GB VRAM (bf16) or ~28 GB (NF4) for inference

Limitations

  • Domain bias: the final merger checkpoint was polished on wave-ui-25k, a UI-screenshot dataset. Strongest grounding is on screenshots / UI elements / OCR-heavy images. Natural-image captioning works but is not as sharp as the instruction-tuned-only stage.
  • Abliteration removes safety refusals from the LLM. Use responsibly.
  • Merger is BF16, LLM is recommended NF4 for sane VRAM. Loading both in BF16 demands ~75 GB.
  • Val_loss has plateaued (1.0318 → 1.0280 across the last 7K steps of Stage 4). Further training would need a new dataset or different objective.

Files in This Repo

File Size Purpose
model-*-of-00015.safetensors ~66 GB total BF16 full weights (LLM + vision + trained merger)
ggml-model-Q4_K_M.gguf 21.2 GB Quantized LLM for llama.cpp
mmproj-stage4.gguf 1.1 GB Stage-4 trained mmproj for llama.cpp (matches the safetensors merger)
mmproj-model-f16.gguf 1.1 GB Legacy alias (identical content)
chat_template.jinja 7 KB Chat template
tokenizer.json + tokenizer_config.json 20 MB Tokenizer

Credits

  • openbmb — MiniCPM-V-4.6 vision architecture & weights
  • huihui-ai — Abliterated Qwen3.5-35B-A3B language model
  • liuhaotian — LLaVA-Pretrain & LLaVA-Instruct-150k datasets
  • ScreenSpot — UI grounding dataset
  • agentsea — wave-ui-25k SFT dataset
  • Assembly & 4-stage merger training by jduartedj

License

Apache 2.0 (consistent with all upstream components).

README history 4 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-06-05README: clarify NF4 loader fix — modules_to_not_convert (no rescue path needed)76987bc8.2 KB
    Loading...
  2. 2026-06-05Update README: 4-stage training curriculum + stage_4/step_23706 (val_loss 1.0...ff4f7d27.6 KB
    Loading...
  3. 2026-05-17Update config, tokenizer, README500cc5d4 KB
    Loading...
  4. 2026-05-14Initial upload: MiniCPM-V 4.6 with Qwen3.5-35B-A3B Abliterated MoE backbone8f5238b4.3 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration