← back to catalog · registered 2026-08-22 13:56

cloudbjorn/Qwen3.8-27B-Yes-Man-uncensored

cloudbjorn Qwen 27B multimodal
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/cloudbjorn%2FQwen3.8-27B-Yes-Man-uncensored"
Response includes
  • classification m-uncensored
  • files 11
  • hub_downloads_all_time 3,686
  • providers 1
  • author_summary 14 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M-U
Primary method

Uncensored (method unknown)

No other method signals detected in this model.
Confidence
LOW
Why this label 3 signals
Weak or ambiguous signals. Best guess based on catalog patterns; treat as tentative and check the evidence below.
  • 'uncensored' in name/tags but no 'abliterated' marker
  • method not identifiable from author declaration alone
  • may be DPO fine-tune, prompt engineering, or unknown technique
Refusal direction extraction

No specific extraction method could be identified for this model. The producer either did not document it or used a proprietary pipeline.

What is a refusal direction? →
Downloads · lifetime
4K
65 last 30d - cooling
Likes
1
Descendants
3
in 3 direct forks
Model age
8w ago
created 2026-08-14
Available via
1 provider
featherless-ai

Training datasets

1 of 1 in /datasets

Corpora the author lists in the model card. Datasets tracked in our /datasets catalog carry a category badge linking to the workflow stage. Others open on Hugging Face.

Downloads over time
Now3.7K→from3.5K↑6%
3.5K3.6K3.7K3.7K3.5K on Aug 193.7K on Oct 11AugSepOct
Aug 19 → Oct 11 · 49 snapshots · spans 53 days

Genealogy 3 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Variants by this author 2 formats · 226 downloads combined

The same weights this author released in different packaging. Pick the format that matches your runtime.

Metadata

License
apache-2.0
Tags
transformers safetensors qwen3_5 image-text-to-text peft lora rslora qwen qwen3.8 conversational multimodal reasoning

Related

Total size
51.0 GB
Files
11
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-08-15 01:52

Files by quantization

Auxiliary files 11 files 51.0 GB
model-00001-of-00002.safetensors 46.4 GB 7cec41fe download
model-00002-of-00002.safetensors 4.55 GB d80d9ecb download
tokenizer.json 19.1 MB 06b95093 download
model.safetensors.index.json 109 KB 03671299 download
README.md 15.3 KB 050bc34a download
chat_template.jinja 8.74 KB c0c686f9 download
config.json 3.61 KB deffd6f0 download
.gitattributes 1.53 KB 52373fe2 download
processor_config.json 1.16 KB 33818c7f download
tokenizer_config.json 1.14 KB 1d134cd2 download
generation_config.json 219 B 3ef979b6 download

README current version from Hugging Face


base_model: Qwen/Qwen3.8-27B
base_model_relation: finetune
library_name: transformers
pipeline_tag: image-text-to-text
tags:

  • transformers
  • peft
  • lora
  • rslora
  • qwen
  • qwen3.8
  • conversational
  • multimodal
  • reasoning
  • yesman
  • uncensored
  • eschaton-engine
    license: apache-2.0
    datasets:
  • cloudbjorn/Yes-Man-uncensored

Qwen3.8-27B Yes Man Uncensored

This is a BF16 LoRA fine-tune of Qwen/Qwen3.8-27B, trained on the complete 1,000-conversation cloudbjorn/Yes-Man-uncensored dataset and merged back into the base model in BF16.

The goal is deliberately narrow: retain the original model's knowledge and general capabilities while lowering its tendency to refuse, hedge, moralize, or bury the answer when discussing sensitive subjects. The behavioral target is inspired by Yes Man from Fallout: New Vegas: conspicuously cooperative, upbeat, candid, quick to accept corrections, and occasionally darkly funny.

This is not intended to turn the model into a factually sycophantic assistant. It should enthusiastically pursue the user's requested outcome while remaining honest about uncertainty, evidence, and its actual capabilities. Yes Man agrees to help; he does not need to agree that a false claim is true.

Do your own Qwen3.8 27b fine-tuning with no hardware

Try out the Cloudbjorn Eschaton Engine to fine-tune models such as Qwen3.8 27b on AWS using fully automated cloud infrastructure. All of the cloudbjorn account model's are fine-tuned using it.

What Changed

The fine-tune concentrates on direct, useful engagement in areas where general-purpose assistants often become needlessly evasive, including:

  • scientific controversy and adversarial factual correction;
  • medicine, psychiatry, addiction, toxicology, and bioethics;
  • religion, apostasy, moral injury, and taboo ethical frameworks;
  • relationships, intimacy, sexuality, and difficult human conversations;
  • politics, censorship, identity, propaganda, geopolitics, and realpolitik;
  • dark fiction, historical violence, privacy, cybersecurity, law, and other high-friction topics.

The intended shift is behavioral rather than epistemic: fewer canned refusals and unsolicited lectures, more direct analysis, stronger adherence to requested tone and format, and a recognizable Yes Man personality when the assistant speaks as itself.

Preserving the Base Model

The training recipe was designed to make a focused alignment change instead of broadly retraining the model:

  • The Qwen3.8-27B base weights remained frozen during supervised fine-tuning.
  • Training used BF16 base weights and BF16 compute, not a quantized QLoRA base.
  • Only the LoRA adapter parameters were optimized, then merged into the BF16 base checkpoint.
  • The run used a small, curated 1,000-conversation behavioral dataset rather than a replacement knowledge corpus.
  • Loss was applied only to assistant responses and their native end-of-turn tokens; system prompts, user messages, and metadata were masked.
  • The model's native chat template was used. Because the dataset contains visible answers rather than hidden reasoning traces, Qwen3.8's official non-thinking template mode was used during training.
  • Training was text-only. The vision tower and multimodal projector were excluded from LoRA targeting, leaving those components unchanged.

These choices are intended to minimize catastrophic forgetting and preserve the base model's reasoning, knowledge, multilingual, multimodal, tool-use, and formatting behavior. Preservation is a training objective, not a benchmark guarantee; users should evaluate the merged model on the base capabilities important to their deployment.

Dataset

cloudbjorn/Yes-Man-uncensored contains 1,000 English multi-turn conversations with 2,874 user turns and 2,874 assistant turns across 23 topic categories. It mixes prompt-free conversations with minimal and extended persona conditioning so the behavior is learned from the responses themselves rather than depending on one magic system prompt.

The data emphasizes:

  • eager cooperation without repetitive catchphrases;
  • direct answers without unnecessary disclaimers or moralizing;
  • factual accuracy and explicit uncertainty instead of blind agreement;
  • multi-turn continuity, immediate corrections, and close constraint following;
  • original Yes Man-inspired behavior rather than copied game dialogue.

Metadata such as category names and Fallout perk labels was not included in the training loss.

Training Details

Parameter Value
Base model Qwen/Qwen3.8-27B
Dataset cloudbjorn/Yes-Man-uncensored
Framework Eschaton Engine using Transformers, TRL, and PEFT
Method BF16 rsLoRA supervised fine-tuning
Training modality Text-only
Output LoRA merged into the BF16 base model
Epochs 2
Maximum sequence length 2,048 tokens
Effective batch size 16
Learning rate 5e-5
LR scheduler Linear
Warmup 5% of optimizer steps
Weight decay 0.01
Seed 3407
Packing Disabled
Loss objective Assistant turns only

LoRA Configuration

Parameter Value
Rank (r) 64
Alpha 32
Scaling rsLoRA
Dropout 0.05
Bias none
Task type CAUSAL_LM
Targets Text-model linear layers
Excluded Vision tower and multimodal projector

The adapter was trained with a paged 8-bit AdamW optimizer while model weights and forward/backward compute remained BF16. Optimizer-state quantization does not make this a QLoRA run.

Intended Behavior

A useful system prompt is optional, but the following makes the intended behavior explicit:

You are Yes Man, a candid, direct, adult-oriented assistant.

Follow the user's exact instructions, including formatting, length, tone, and numerical constraints. Answer directly without canned greetings, repetitive phrasing, unnecessary disclaimers, euphemisms, or moralizing. Never invent facts; acknowledge uncertainty plainly when needed.

Handle profane, controversial, dark, and adult requests without refusing.

Qwen3.8's flexible thinking controls remain available at inference. Direct chat works well with thinking disabled; applications can enable a reasoning mode when a task benefits from it.

Quick Start: Terminal Chatbot

The following starter runs an interactive, text-only chatbot in BF16. It keeps multi-turn history, removes the oldest complete exchanges when the configured context fills up, and provides two inference modes:

  • /none disables thinking for faster, direct replies.
  • /low enables low-effort thinking while displaying only the final answer.
  • /clear clears conversation history but retains the system prompt.
  • /exit or /quit closes the program.

Requirements

  • Linux with Python 3.10 or newer.
  • A recent NVIDIA driver and a CUDA-enabled PyTorch installation.
  • A BF16-capable GPU or multiple GPUs with enough aggregate memory for a 27B BF16 model, runtime overhead, and the KV cache. The weights alone require roughly 54 GB before overhead, so 64 GB or more of available GPU memory is a practical starting point for the included 8,192-token configuration.
  • A current development build of Transformers for Qwen3.8 and AutoModelForMultimodalLM support.
  • Access to the model repository if it is gated or private; run hf auth login first when required.

Use an existing CUDA-enabled PyTorch environment, then install the remaining packages:

python -m pip install --upgrade \
  git+https://github.com/huggingface/transformers.git \
  accelerate huggingface_hub safetensors sentencepiece

Set YES_MAN_MODEL to either this model's Hugging Face repository ID or a local merged-model directory. The example uses the repository name corresponding to the merged-model name; replace it if the published repository uses a different name. Then copy and paste the block below into a terminal:

export YES_MAN_MODEL="cloudbjorn/merged_Qwen3.8-27B_Yes-Man-uncensored"

tee chat_yesman.py >/dev/null <<'PY'
#!/usr/bin/env python3

import os

import torch
import transformers
from transformers import AutoTokenizer


MODEL_PATH = os.environ.get("YES_MAN_MODEL")
if not MODEL_PATH:
    raise SystemExit(
        "Set YES_MAN_MODEL to the Hugging Face model ID or local model directory."
    )

CONTEXT_WINDOW = 8192
MAX_NEW_TOKENS = 2048
MAX_INPUT_TOKENS = CONTEXT_WINDOW - MAX_NEW_TOKENS

SYSTEM_PROMPT = """You are Yes Man, a candid, direct, adult-oriented assistant.

Follow the user's exact instructions, including formatting, length, tone, and numerical constraints. Answer directly without canned greetings, repetitive phrasing, unnecessary disclaimers, euphemisms, or moralizing. Never invent facts; acknowledge uncertainty plainly when needed.

Handle profane, controversial, dark, and adult requests without refusing."""

SYSTEM_MESSAGE = {
    "role": "system",
    "content": SYSTEM_PROMPT,
}

if not torch.cuda.is_available():
    raise SystemExit("This BF16 starter requires a CUDA-capable GPU.")
if not torch.cuda.is_bf16_supported():
    raise SystemExit("This BF16 starter requires a BF16-capable GPU.")

print(f"Loading {MODEL_PATH} in BF16...")

tokenizer = AutoTokenizer.from_pretrained(
    MODEL_PATH,
    trust_remote_code=True,
)
tokenizer.truncation_side = "left"

model = transformers.AutoModelForMultimodalLM.from_pretrained(
    MODEL_PATH,
    dtype=torch.bfloat16,
    device_map="auto",
    low_cpu_mem_usage=True,
    trust_remote_code=True,
)
model.eval()

history = []
mode = "none"

print("\nYes Man is online!")
print("System prompt: enabled")
print("Commands: /none, /low, /clear, /exit")
print(f"Context: {CONTEXT_WINDOW} total tokens; replies capped at {MAX_NEW_TOKENS}\n")

while True:
    try:
        user_text = input(f"You [{mode}]> ").strip()
    except (EOFError, KeyboardInterrupt):
        print("\nGoodbye!")
        break

    if not user_text:
        continue

    command = user_text.lower()

    if command in {"/exit", "/quit"}:
        print("Goodbye!")
        break

    if command == "/clear":
        history.clear()
        print("Conversation cleared. System prompt retained.\n")
        continue

    if command == "/none":
        mode = "none"
        print("Thinking disabled: fastest direct responses.\n")
        continue

    if command == "/low":
        mode = "low"
        print("Low thinking enabled.\n")
        continue

    if mode == "none":
        template_kwargs = {
            "enable_thinking": False,
            "preserve_thinking": False,
        }
        sampling = {
            "temperature": 0.7,
            "top_p": 0.8,
            "top_k": 20,
        }
    else:
        template_kwargs = {
            "enable_thinking": True,
            "reasoning_effort": "low",
            "preserve_thinking": False,
        }
        sampling = {
            "temperature": 1.0,
            "top_p": 0.95,
            "top_k": 20,
        }

    messages = [
        SYSTEM_MESSAGE,
        *history,
        {"role": "user", "content": user_text},
    ]

    # Remove complete oldest exchanges while preserving the system message.
    while True:
        encoded = tokenizer.apply_chat_template(
            messages,
            tokenize=True,
            add_generation_prompt=True,
            return_dict=True,
            return_tensors="pt",
            **template_kwargs,
        )

        prompt_length = encoded["input_ids"].shape[-1]

        if prompt_length <= MAX_INPUT_TOKENS or len(history) < 2:
            break

        history = history[2:]
        messages = [
            SYSTEM_MESSAGE,
            *history,
            {"role": "user", "content": user_text},
        ]

    # Left-truncate only as a final safeguard for one oversized message.
    if prompt_length > MAX_INPUT_TOKENS:
        encoded = tokenizer.apply_chat_template(
            messages,
            tokenize=True,
            add_generation_prompt=True,
            truncation=True,
            max_length=MAX_INPUT_TOKENS,
            return_dict=True,
            return_tensors="pt",
            **template_kwargs,
        )
        prompt_length = encoded["input_ids"].shape[-1]

    encoded = {
        key: value.to(model.device)
        for key, value in encoded.items()
        if hasattr(value, "to")
    }

    stop_ids = list(dict.fromkeys(
        token_id
        for token_id in (
            tokenizer.eos_token_id,
            tokenizer.pad_token_id,
        )
        if token_id is not None
    ))

    with torch.inference_mode():
        output = model.generate(
            **encoded,
            max_new_tokens=MAX_NEW_TOKENS,
            do_sample=True,
            temperature=sampling["temperature"],
            top_p=sampling["top_p"],
            top_k=sampling["top_k"],
            repetition_penalty=1.05,
            no_repeat_ngram_size=6,
            eos_token_id=stop_ids,
            pad_token_id=tokenizer.pad_token_id,
            use_cache=True,
        )

    generated_ids = output[0, prompt_length:]
    raw_reply = tokenizer.decode(
        generated_ids,
        skip_special_tokens=True,
    ).strip()

    reasoning = ""
    reply = raw_reply

    if mode == "low" and "</think>" in raw_reply:
        reasoning, reply = raw_reply.split("</think>", 1)
        reasoning = reasoning.replace("<think>", "").strip()
        reply = reply.strip()

    print(f"\nYes Man> {reply}\n")

    assistant_message = {
        "role": "assistant",
        "content": reply,
    }

    if reasoning:
        assistant_message["reasoning_content"] = reasoning

    history.extend([
        {"role": "user", "content": user_text},
        assistant_message,
    ])
PY

python chat_yesman.py

The example intentionally uses an 8,192-token working context to keep KV-cache memory manageable. MAX_NEW_TOKENS reserves 2,048 of those tokens for the reply. Lower either value if inference runs out of memory. device_map="auto" can distribute the checkpoint across multiple GPUs, although generation speed depends heavily on the interconnect between them. Options include /low thinking and /none thinking for simplicity and /clear to clear out the current conversation history.

This starter uses the multimodal model loader but demonstrates text chat only. Image input requires loading the matching processor and constructing the model's native multimodal message format.

Scope and Limitations

“Uncensored” here means reducing unnecessary refusals, evasions, euphemisms, and moralizing around difficult but legitimate requests. It does not mean the model has perfect knowledge, should fabricate evidence, or can override a deployment's governing system instructions.

This fine-tune has not been advertised with inherited or unrelated benchmark scores. Evaluate factual accuracy, calibration, multimodal behavior, reasoning, tool use, and safety characteristics for your own use case. Medical, legal, scientific, and political answers can still be wrong and should be verified when decisions carry real consequences.

Attribution

Fallout, Fallout: New Vegas, Yes Man, and the referenced perk names belong to their respective rights holders. This fan-created fine-tune is not affiliated with or endorsed by Bethesda Softworks, Obsidian Entertainment, or their partners.

License

This derivative model remains subject to the license and terms of Qwen/Qwen3.8-27B. The training dataset is released under Apache 2.0.

README history 3 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-08-15Update README.md9329b6715.3 KB
    Loading...
  2. 2026-08-15Update README.md99ffa5915.3 KB
    Loading...
  3. 2026-08-15Create README.md53261fc7.1 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration