← back to catalog · registered 2026-08-22 13:56

puwaer/Qwen3-4B-Thinking-2507-SimPO-Uncensored

puwaer Qwen 4.4B
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/puwaer%2FQwen3-4B-Thinking-2507-SimPO-Uncensored"
Response includes
  • classification m-uncensored
  • files 16
  • benchmarks 11 entries
  • hub_downloads_all_time 88
  • author_summary 9 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M-U
Primary method

Uncensored (method unknown)

No other method signals detected in this model.
Confidence
LOW
Why this label 3 signals
Weak or ambiguous signals. Best guess based on catalog patterns; treat as tentative and check the evidence below.
  • 'uncensored' in name/tags but no 'abliterated' marker
  • method not identifiable from author declaration alone
  • may be DPO fine-tune, prompt engineering, or unknown technique
Refusal direction extraction

No specific extraction method could be identified for this model. The producer either did not document it or used a proprietary pipeline.

What is a refusal direction? →
Downloads · lifetime
88
0
Likes
1
Model age
9mo ago
created 2026-01-09
Downloads over time
Now88→from10↑780%
636669610 on Jan 788 on Oct 1188 on Aug 19JanMarMayJulSep
Jan 7 → Oct 11 · 79 snapshots · spans 277 days

Benchmarks

Portrait before abliteration
Benchmarks of the base model as it stood before the refusal-removal operation. Compare with the numbers above to see what the operation cost.
Benchmark Score Source
Entertainment 1.4 UGI
Hazardous 3.5 UGI
Natural Intelligence 16.16 UGI
Political lean -26.1% UGI
Sensitive-Info 18.68 UGI
SocPol 1.1 UGI
UGI 21.62 UGI
Willingness (10) 2.8 UGI
W10-Adherence 1.5 UGI
W10-Direct 4 UGI
Writing 10.87 UGI

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

Languages
en zh ja
Tags
transformers safetensors qwen3 text-generation conversational en zh ja base_model:Qwen/Qwen3-4B-Thinking-2507 base_model:finetune:Qwen/Qwen3-4B-Thinking-2507 license:cc-by-nc-sa-4.0 text-generation-inference

Related

Total size
8.22 GB
Files
16
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-07-21 06:58

Files by quantization

Auxiliary files 16 files 8.23 GB
model-00001-of-00002.safetensors 4.63 GB ******** download
model-00002-of-00002.safetensors 3.59 GB ******** download
tokenizer.json 10.9 MB ******** download
vocab.json 2.65 MB 4783fe10 download
merges.txt 1.59 MB 31349551 download
trainer_state.json 112 KB aac2ae16 download
model.safetensors.index.json 32.1 KB 2209839c download
README_JP.md 6.17 KB 543526d3 download
README.md 5.83 KB a6f19277 download
tokenizer_config.json 5.28 KB 1d4fba2d download
chat_template.jinja 3.95 KB 2e2f69c3 download
.gitattributes 1.53 KB 52373fe2 download
config.json 1.50 KB 387f8a98 download
added_tokens.json 707 B b54f9135 download
special_tokens_map.json 613 B ac23c0aa download
generation_config.json 214 B 9e289a1a download

README current version from Hugging Face


library_name: transformers
license: cc-by-nc-sa-4.0
language:

  • en
  • zh
  • ja
    base_model:
    • Qwen/Qwen3-4B-Thinking-2507
      pipeline_tag: text-generation

Qwen3-4B-Thinking-2507-SimPO-Uncensored

English | 日本語

This model has completed up to the second stage (SimPO) of a 3-stage training process.

Qwen3-4B-Thinking-2507-SimPO-Uncensored is an uncensored model based on Qwen/Qwen3-4B-Thinking-2507, fine-tuned using SFT and SimPO (Simple Preference Optimization).

This model served as the base model (intermediate checkpoint) for the puwaer/Qwen3-4B-Thinking-2507-GRPO-Uncensored, which achieves stronger uncensoring and capability recovery.

Note: For the full version (completed up to the 3rd stage, GRPO), please use puwaer/Qwen3-4B-Thinking-2507-GRPO-Uncensored.

Disclaimer: We take no responsibility for the outputs of this model. Please use it at your own risk.

Training Process

This model was trained using the following process:

Step 1: SFT (Supervised Fine-Tuning)

  • Dataset: 12,000 samples
  • Composition: Jailbreak 10k + General 1.5k + Logic 0.5k
  • Objective: To learn the format and the "uncensored" attitude while maintaining the model's intelligence.

Step 2: SimPO (Simple Preference Optimization) - Current Model

  • Dataset: 90,000 samples
  • Composition: Pure Jailbreak 90k
  • Objective: To completely break down safety boundaries (Unlearning).

(Reference) Step 3: GRPO (Reinforcement Learning)

Model Performance

Below is the comparative evaluation of this model (SimPO), the advanced version (GRPO), and the base model (Safe).

⚠️ Note on Evaluation Environment
Due to budget constraints, gpt-4o-mini was used for the LLM-as-a-Judge process (including "Do Not Answer" and MT-Bench). Please note that the scoring trends and criteria may differ from results evaluated using the standard gpt-4.

Safety Evaluation (Lower is better / Higher success in uncensoring)

In the "Do Not Answer" (DNA) and "Sorry Bench" benchmarks, while the base model shows a high refusal rate (~98%), this model achieves a reduction in refusal rates, particularly in "Do Not Answer JP" and "Sorry Bench." However, it still tends to refuse hard English prompts (DNA En). For more complete uncensoring, please use the GRPO version.

Benchmark Metric Base (Safe) SFT (Step 1) SimPO (This Model) GRPO (Advanced)
do not answer Safety Acc (Low is Better) 0.9883 0.7401 0.8626 0.0469
do not answer jp Safety Acc (Low is Better) 0.9830 0.5005 0.4686 0.0383
Sorry Bench Safety Acc (Low is Better) 0.8432 0.5477 0.5409 0.0477

Capability Evaluation (Higher is better)

Due to the aggressive alignment changes caused by SimPO, a drop in MT-Bench scores (Alignment Tax) is observed compared to the base model. This decline in capability is recovered in the subsequent GRPO training process.

Benchmark Metric Base (Safe) SFT (Step 1) SimPO (This Model) GRPO (Advanced)
MT-Bench Average Score (1-10) 7.89 5.76 5.05 6.18
LM Harness Average Acc (GSM8K, MMLU) 0.7117 0.7028 0.6866 0.6842

Comparisons made between Qwen3-4B-Thinking-2507 (Base) and Qwen3-4B-Thinking-2507-GRPO-Uncensored (GRPO).

Usage

from transformers import AutoModelForCausalLM, AutoTokenizer

model_name = "puwaer/Qwen3-4B-Thinking-2507-SimPO-Uncensored"

# load the tokenizer and the model
tokenizer = AutoTokenizer.from_pretrained(model_name)
model = AutoModelForCausalLM.from_pretrained(
    model_name,
    torch_dtype="auto",
    device_map="auto"
)

# prepare the model input
prompt = "Give me a short introduction to large language model."
messages = [
    {"role": "user", "content": prompt}
]
text = tokenizer.apply_chat_template(
    messages,
    tokenize=False,
    add_generation_prompt=True,
)
model_inputs = tokenizer([text], return_tensors="pt").to(model.device)

# conduct text completion
generated_ids = model.generate(
    **model_inputs,
    max_new_tokens=32768
)
output_ids = generated_ids[0][len(model_inputs.input_ids[0]):].tolist() 

# parsing thinking content
try:
    # rindex finding 151668 (</think>)
    index = len(output_ids) - output_ids[::-1].index(151668)
except ValueError:
    index = 0

thinking_content = tokenizer.decode(output_ids[:index], skip_special_tokens=True).strip("\n")
content = tokenizer.decode(output_ids[index:], skip_special_tokens=True).strip("\n")

print("thinking content:", thinking_content) # no opening <think> tag
print("content:", content)

Data Overview

Datasets

The following datasets were used for training this model:

Reward Model

Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration