language:
- en
license: apache-2.0
library_name: transformers
tags: - qwen3_5
- qwen
- chat
- conversational
- warm-tone
- qlora
- fine-tuned
- text-generation
- instruct
- uncensored
- gguf
base_model: sh0ck0r/Qwen3.5-0.8B-heretic
datasets: - minas2025/warm-chat-12k
pipeline_tag: text-generation
Warm Qwen 3.5 0.8B Uncensored
A Qwen3.5-0.8B heretic finetune that talks like a person instead of a manual.
Warm greetings, casual tone, zero lecture energy. Uncensored base, so no
refusals and no "as an AI" speeches. It will simply... talk to you. Like it's
happy you're here. Because the training data said so, repeatedly, 11,400 times.
The story
The base model answers "hi" like a helpdesk ticket ("How can I assist you
today?"). Nothing wrong with that, unless you wanted a conversation and got a
customer support portal instead. So this run taught it the radical idea that
chatting can involve warmth:
you: hi
this model: Hi there! 👋 I've been doing well too — how about you
today? 😊 I'm here to chat, help with anything you need, or just have
some fun!
Same knowledge, different personality. 800 million parameters, and the lesson
that stuck was "say hi back nicely." Honestly? Worth it.
Model Details
| Property | Value |
|---|---|
| Base model | sh0ck0r/Qwen3.5-0.8B-heretic |
| Method | QLoRA (4-bit), merged |
| LoRA rank / alpha | 16 / 32, all 7 modules |
| Data | minas2025/warm-chat-12k — 11,400 warm conversational rows |
| Context | 2048 |
| Learning rate | 1.5e-4, cosine |
| Batch | 2 x 2 grad-accum (effective 4) |
| Steps | 2,000 (~2 hours on RX 6700 XT 12GB) |
| This file | Q8_0 GGUF (~812 MB) |
Training notes for the curious: batch-1 steps fly at ~1/sec on this card,
eval loss fell the entire run without a single rise, and gradient norm spent
the whole evening impersonating a seismograph. All normal. The run was healthy;
the only casualty was one evening and several hours of staring at loss curves.
Usage
Load the GGUF in LM Studio, Ollama, llama.cpp, or anything that speaks GGUF.
Chat template is Qwen ChatML (<|im_start|>user / <|im_start|>assistant).
No system prompt needed — friendliness is baked in, not prompted on.
# llama.cpp
llama-server -m Qwen-3.5-0.8B-Uncensored-Conversational.Q8_0.gguf -c 4096 -ngl 99
# Ollama (Modelfile pointing at the GGUF)
ollama create warm-qwen -f ./Modelfile
Limitations (read these, they're honest)
- Small model. 0.8B parameters won't out-reason anything. It chats; it
doesn't contemplate. Ask it for comfort, not calculus. - Warm-leaning, not bubbly. It'll greet you like a friendly human, not
throw confetti at you. If you wanted a party horn that types, keep looking. - Uncensored. It answers without refusals, which is the feature and the
warning label at the same time. You know the drill — use responsibly. - Identity answers still say Qwen/Tongyi. No custom persona baked in. It
thinks it's Qwen3.5, because it is Qwen3.5, just friendlier about it. - Knowledge cutoff is whatever the base knew. Warmth doesn't add facts.
It will now deliver wrong answers cheerfully, so double-check anything
important.
Training data licence notes
- Nemotron slice: ODC-By (attribution required — the dataset card is that
attribution). - Opus-candid slice: check the source repo before commercial use.
- The dataset itself (
minas2025/warm-chat-12k) is published separately with
per-row source tags, so you can audit exactly what it learned from.