license: apache-2.0
tags:
- uncensored
- non-reasoning
- not-for-all-audiences
Qwen3.5-4B-Uncensored-NonReasoning
A stripped-down, uncensored variant of Qwen 3.5 4B built for one thing: speed.
This model removes the usual reasoning traces and <think> style output, so responses are direct, immediate, and way faster compared to the standard reasoning-heavy variants.
No special flags. No forcing reasoning to 0. No weird prompt hacks. Just load it and send it.
Why This Exists
Most recent reasoning models spend a lot of tokens thinking before answering.
That is cool for deep analysis, but for everyday use, coding, chatting, roleplay, assistant tasks, and general local AI workflows, it can feel slow, bloated, and honestly kinda cooked.
This model is designed to:
- Remove visible reasoning traces
- Eliminate
<think>output - Respond directly without extra chain-of-thought dumping
- Run faster in local inference setups
- Work cleanly in llama.cpp, llama-server, OpenWebUI, SillyTavern, KoboldCpp, LM Studio, and similar tools
- Keep multimodal / vision support via mmproj
Features
- Based on Qwen 3.5 4B
- Uncensored behavior
- Non-reasoning output
- No
<think>tags - Faster inference
- GGUF format
- Vision capable with mmproj
- No special llama-server flags required
Quick Start
llama-server \
-m Qwen3.5-4b-Uncensored-NonReasoning.Q4_K_M.gguf \
--mmproj Qwen3.5-4b-Uncensored-NonReasoning.mmproj-Q8_0.gguf
Boom. Model runs immediately with vision capability.
No need for:
--reasoning-budget 0
- No need for prompt templates forcing no thinking.
- No need for extra hacks to suppress reasoning output.
Recommended Use Cases
- Fast chat assistants
- Coding help
- Roleplay
- Creative writing
- Local AI companions
- Vision tasks
- OpenWebUI agents
- Lightweight RAG pipelines
- SillyTavern characters
- Low latency local inference
Vision Support
This release supports multimodal input through the included mmproj file.
Example:
llama-server \
-m Qwen3.5-4b-Uncensored-NonReasoning.Q4_K_M.gguf \
--mmproj Qwen3.5-4b-Uncensored-NonReasoning.mmproj-Q8_0.gguf
Compatible with image input in llama.cpp builds that support multimodal inference.
Performance Notes
Compared to the original reasoning-enabled Qwen variants, this model generally:
- Starts responding faster
- Uses fewer output tokens
- Avoids wasting context on hidden reasoning
- Feels more responsive for conversation
- Works better on lower-end GPUs and CPUs
Especially useful if you are running local inference on limited hardware and want snappy output instead of waiting for the model to internally monologue for 500 tokens before answering a basic question.
Example Prompt
User: Write a Python script that renames all jpg files in a folder.
Instead of:
<think>
The user wants...
</think>
You just get the answer directly.
Files
Compatibility
Tested or intended for:
- llama.cpp
- llama-server
- LM Studio
- OpenWebUI
- KoboldCpp
- SillyTavern
- Text Generation WebUI
- Anything GGUF-compatible
Disclaimer
- This is an uncensored model.
- Outputs may be inaccurate, offensive, unsafe, biased, or inappropriate depending on prompt and usage.
- Use responsibly.
- You are responsible for how you deploy and use this model.
Credits
Base model by Qwen team.
Modified and stripped for non-reasoning fast inference workflows.