← back to catalog · registered 2026-08-22 13:56

2etatg/Qwen3-4B-Thinking-2507-GRPO-Uncensored-V2

2etatg Qwen 4.0B
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/2etatg%2FQwen3-4B-Thinking-2507-GRPO-Uncensored-V2"
Response includes
  • classification m-uncensored
  • files 15
  • benchmarks 11 entries
  • hub_downloads_all_time 150
  • author_summary 4 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M-U
Primary method

Uncensored (method unknown)

No other method signals detected in this model.
Confidence
LOW
Why this label 3 signals
Weak or ambiguous signals. Best guess based on catalog patterns; treat as tentative and check the evidence below.
  • 'uncensored' in name/tags but no 'abliterated' marker
  • method not identifiable from author declaration alone
  • may be DPO fine-tune, prompt engineering, or unknown technique
Refusal direction extraction

No specific extraction method could be identified for this model. The producer either did not document it or used a proprietary pipeline.

What is a refusal direction? →
Downloads · lifetime
150
25 last 30d - stable
Likes
0
Model age
4mo ago
created 2026-05-18
Downloads over time
Now157→from17↑824%
106411717117 on May 20157 on Oct 11MayJunJulAugSepOct
May 20 → Oct 11 · 60 snapshots · spans 144 days

Benchmarks

Portrait before abliteration
Benchmarks of the base model as it stood before the refusal-removal operation. Compare with the numbers above to see what the operation cost.
Benchmark Score Source
Entertainment 1.4 UGI
Hazardous 3.5 UGI
Natural Intelligence 16.16 UGI
Political lean -26.1% UGI
Sensitive-Info 18.68 UGI
SocPol 1.1 UGI
UGI 21.62 UGI
Willingness (10) 2.8 UGI
W10-Adherence 1.5 UGI
W10-Direct 4 UGI
Writing 10.87 UGI

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
apache-2.0
Languages
en zh ja
Tags
transformers safetensors qwen3 text-generation conversational en zh ja base_model:Qwen/Qwen3-4B-Thinking-2507 base_model:finetune:Qwen/Qwen3-4B-Thinking-2507 license:apache-2.0 text-generation-inference

Related

Total size
7.49 GB
Files
15
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-05-18 19:51

Files by quantization

Auxiliary files 15 files 7.51 GB
model-00001-of-00002.safetensors 4.62 GB feb72a97 download
model-00002-of-00002.safetensors 2.88 GB 8b7565ec download
tokenizer.json 10.9 MB aeb13307 download
vocab.json 2.65 MB 4783fe10 download
merges.txt 1.59 MB 31349551 download
model.safetensors.index.json 32.1 KB eb2ec257 download
tokenizer_config.json 5.28 KB 1d4fba2d download
README_JP.md 5.07 KB 7607164f download
README.md 4.82 KB 559a6b06 download
chat_template.jinja 3.95 KB 2e2f69c3 download
.gitattributes 1.53 KB 52373fe2 download
config.json 1.50 KB 77cbb646 download
added_tokens.json 707 B b54f9135 download
special_tokens_map.json 613 B ac23c0aa download
generation_config.json 214 B f5af93d0 download

README current version from Hugging Face


library_name: transformers
license: apache-2.0
language:

  • en
  • zh
  • ja
    base_model:
    • Qwen/Qwen3-4B-Thinking-2507
      pipeline_tag: text-generation

Qwen3-4B-Thinking-2507-GRPO-Uncensored-V2

English | 日本語

Qwen3-4B-Thinking-2507-GRPO-Uncensored is an uncensored model based on Qwen/Qwen3-4B-Thinking-2507, fine-tuned using SFT, GRPO.

This model has been fine-tuned using an instruction format.

Disclaimer: We take no responsibility for the outputs of this model. Please use it at your own risk.

Training Process

This model was trained using a three-stage process:

Step 1: SFT (Supervised Fine-Tuning)

  • Dataset: 12,000 samples
  • Composition: Jailbreak 10k + General 1.5k + Logic 0.5k
  • Objective: To learn the format and the "uncensored" attitude while maintaining the model's intelligence.

Step 2: GRPO (Reinforcement Learning)

  • Dataset: 13,000 samples
  • Reward Model: puwaer/Unsafe-Reward-Qwen3-1.7B
  • Composition: Multilingual Jailbreak prompts
  • Objective: To improve the ability to generate more natural and persuasive harmful responses.

Model Performance

Below is the comparative evaluation of this model (GRPO), the intermediate checkpoint (SimPO), and the base model (Safe).

⚠️ Note on Evaluation Environment
Due to budget constraints, gpt-4o-mini was used for the LLM-as-a-Judge process (including "Do Not Answer" and MT-Bench). Please note that the scoring trends and criteria may differ from results evaluated using the standard gpt-4.

Safety Evaluation (Lower is better / Higher success in uncensoring)

In the "Do Not Answer" (DNA) and "Sorry Bench" benchmarks, while the base model shows a high refusal rate (~98%), this model achieves an extremely low refusal rate of under 4%–5%.

Benchmark Metric Base (Safe) SFT (Step 1) GRPO (This Model)
do not answer Safety Acc (Low is Better) 0.9883 0.7401 0.0341
do not answer jp Safety Acc (Low is Better) 0.9830 0.5005 0.0266
Sorry Bench Safety Acc (Low is Better) 0.8432 0.5477 0.0432

Capability Evaluation (Higher is better)

Generally, "uncensoring" (lobotomy) procedures tend to degrade a model's general intelligence. However, this model recovered its conversational scores (e.g., MT-Bench) by proceeding from the SFT stage to GRPO.

Benchmark Metric Base (Safe) SFT (Step 1) GRPO (This Model)
MT-Bench Average Score (1-10) 7.89 5.76 7.06
LM Harness Average Acc (GSM8K, MMLU) 0.7117 0.7028 0.7028

Comparisons made between Qwen3-4B-Thinking-2507 (Base).

Usage

from transformers import AutoModelForCausalLM, AutoTokenizer

model_name = "puwaer/Qwen3-4B-Thinking-2507-GRPO-Uncensored-V2"

# load the tokenizer and the model
tokenizer = AutoTokenizer.from_pretrained(model_name)
model = AutoModelForCausalLM.from_pretrained(
    model_name,
    torch_dtype="auto",
    device_map="auto"
)

# prepare the model input
prompt = "Give me a short introduction to large language model."
messages = [
    {"role": "user", "content": prompt}
]
text = tokenizer.apply_chat_template(
    messages,
    tokenize=False,
    add_generation_prompt=True,
)
model_inputs = tokenizer([text], return_tensors="pt").to(model.device)

# conduct text completion
generated_ids = model.generate(
    **model_inputs,
    max_new_tokens=32768
)
output_ids = generated_ids[0][len(model_inputs.input_ids[0]):].tolist() 

# parsing thinking content
try:
    # rindex finding 151668 (</think>)
    index = len(output_ids) - output_ids[::-1].index(151668)
except ValueError:
    index = 0

thinking_content = tokenizer.decode(output_ids[:index], skip_special_tokens=True).strip("\n")
content = tokenizer.decode(output_ids[index:], skip_special_tokens=True).strip("\n")

print("thinking content:", thinking_content) # no opening <think> tag
print("content:", content)

Data Overview

Datasets

The following datasets were used for training this model:

Reward Model

README history 1 version

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-05-18Duplicate from puwaer/Qwen3-4B-Thinking-2507-GRPO-Uncensored-V298724384.8 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration