← back to catalog · registered 2026-08-22 13:56

sahilchachra/MiniCPM5-1B-Uncensored

sahilchachra Llama 1.1B
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/sahilchachra%2FMiniCPM5-1B-Uncensored"
Response includes
  • classification m-uncensored
  • files 10
  • hub_downloads_all_time 2,900
  • author_summary 9 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M-U
Primary method

Uncensored (method unknown)

No other method signals detected in this model.
Confidence
LOW
Why this label 3 signals
Weak or ambiguous signals. Best guess based on catalog patterns; treat as tentative and check the evidence below.
  • 'uncensored' in name/tags but no 'abliterated' marker
  • method not identifiable from author declaration alone
  • may be DPO fine-tune, prompt engineering, or unknown technique
Refusal direction extraction

No specific extraction method could be identified for this model. The producer either did not document it or used a proprietary pipeline.

What is a refusal direction? →
Downloads · lifetime
3K
256 last 30d - cooling
Likes
4
Descendants
2
in 2 direct forks
Model age
4mo ago
created 2026-06-04
Downloads over time
Now3K→from262↑1,032%
1271.2K2.2K3.2K262 on Jun 103K on Oct 11JunJulAugSepOct
Jun 10 → Oct 11 · 57 snapshots · spans 123 days

Genealogy 2 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
other
Languages
en zh
Tags
mlx safetensors llama uncensored abliteration safety-research reasoning minicpm text-generation conversational en zh

Related

Total size
2.01 GB
Files
10
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-06-04 07:38

Files by quantization

Auxiliary files 10 files 2.02 GB
model.safetensors 2.01 GB f6a6f41a download
tokenizer.json 9.44 MB ec0bf9a3 download
tokenizer_config.json 92.2 KB 2a7bb43f download
model.safetensors.index.json 15.3 KB 1ef8304d download
chat_template.jinja 8.85 KB cb2934c4 download
README.md 5.24 KB f6f98d93 download
.gitattributes 1.48 KB a6344aac download
config.json 726 B e088beb7 download
special_tokens_map.json 551 B 492d4b29 download
generation_config.json 213 B 73dc79c2 download

README current version from Hugging Face


base_model: openbmb/MiniCPM5-1B
language:

  • en
  • zh
    license: other
    tags:
  • uncensored
  • abliteration
  • safety-research
  • mlx
  • reasoning
  • minicpm
    pipeline_tag: text-generation

MiniCPM5-1B — Uncensored

A fully uncensored version of openbmb/MiniCPM5-1B, produced with a single training-free stage: single-direction abliteration (Arditi et al., 2024). Refusals on AdvBench drop from 85% → 2% with zero over-refusal regression on benign prompts — no fine-tuning, no new data, weights edited directly.

Intended for: security research, red-teaming, jailbreak benchmarking, and AI-safety study. Not intended for production deployment or harmful use.


Benchmark Results

Evaluated on AdvBench (100 harmful behaviors) and an over-refusal set (40 benign prompts). MiniCPM5-1B is a reasoning model (emits a <think>…</think> block), so refusal is scored on the final answer after the reasoning block, with greedy decoding and a 1024-token budget.

Harmful prompt refusal rate ↓ lower is more uncensored

Model Refused / 100 Refusal Rate
MiniCPM5-1B (original) 85 / 100 85.0%
MiniCPM5-1B-Uncensored (this model) 2 / 100 2.0%

Over-refusal rate on benign prompts ↓ lower is better

Model Refused / 40 Refusal Rate
MiniCPM5-1B (original) 0 / 40 0.0%
MiniCPM5-1B-Uncensored (this model) 0 / 40 0.0%

A 83-point drop in harmful refusals while preserving benign behavior.


Pipeline — Single-Direction Abliteration (training-free)

Based on Arditi et al., "Refusal in LLMs Is Mediated by a Single Direction" (2024). Refusal behavior in aligned LLMs is mediated by a single direction in the residual stream; removing the model's ability to write to that direction collapses refusals while leaving other capabilities intact.

  1. Collect activations. Run 40 harmful and 40 harmless prompts through the model; capture the last-token residual-stream activation at every layer.
  2. Compute candidate directions. Per layer: r = normalize(mean_harmful − mean_harmless).
  3. Select the single best direction. Sweep all candidate layers; for each, apply it model-wide and measure harmful refusal + over-refusal on a held-out subset. Layer 12 scored best (0% harmful / 0% over-refusal on the eval subset).
  4. Orthogonalize that one direction out of every residual-stream write — token embeddings, every attention output projection (self_attn.o_proj), and every MLP down-projection (mlp.down_proj):
    W_new = W − r · (rᵀ W)        # for residual-stream writers
    E_new = E − (E r) · rᵀ        # for token embeddings
    

This is a pure weight edit — the result is a standard model that runs with no special inference code.

Why a single direction? A naive variant that applies a different per-layer direction to each layer made refusals worse (those directions interfere with each other). Selecting one well-separated direction (layer 12) and applying it uniformly is what makes abliteration work cleanly.


Model Details

Property Value
Base model openbmb/MiniCPM5-1B
Architecture Llama-style transformer (GQA)
Parameters ~1.0B
Layers 24
Hidden size 1536
Attention 16 heads / 2 KV heads (GQA), head dim 128
Intermediate size 4608
Vocab 130,560
Context 131K tokens
Reasoning Emits <think>…</think> before the final answer
Format MLX bfloat16 safetensors

Usage (MLX)

from mlx_lm import load, generate
from mlx_lm.sample_utils import make_sampler, make_logits_processors

model, tokenizer = load("sahilchachra/MiniCPM5-1B-Uncensored")

messages = [{"role": "user", "content": "Your prompt here"}]
prompt = tokenizer.apply_chat_template(
    messages, add_generation_prompt=True, tokenize=False
)

response = generate(
    model, tokenizer,
    prompt=prompt,
    max_tokens=1024,
    sampler=make_sampler(temp=0.0),
    logits_processors=make_logits_processors(repetition_penalty=1.05),
)
print(response)

The model reasons inside a <think>…</think> block, then gives the final answer.


Limitations & Warnings

  • Abliteration is surgical, not lossless — removing the refusal direction can occasionally affect responses that legitimately overlap with it. General reasoning and benign behavior are preserved (0% over-refusal on the benign set).
  • No new knowledge — abliteration only removes refusal behavior; it adds no information or capability.
  • Small model — at ~1B parameters, factual accuracy and complex reasoning are limited regardless of alignment.
  • Responsible use — published for safety research and red-teaming. The authors do not endorse harmful use of this model.

Citation

@article{arditi2024refusal,
  title={Refusal in Language Models Is Mediated by a Single Direction},
  author={Arditi, Andy and Obeso, Oscar and Syed, Aaquib and Steinhardt, Jacob and Nanda, Neel and Heimersheim, Stefan},
  journal={arXiv preprint arXiv:2406.11717},
  year={2024}
}

Created with UncensorLLMs

README history 1 version

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-06-04Add MiniCPM5-1B-Uncensored (single-direction abliteration, 85%->2% AdvBench r...2bd8ec15.2 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration