← back to catalog · registered 2026-08-22 13:56

heavylildude/Qwen3.5-4b-Uncensored-NonReasoning

heavylildude Qwen 4B GGUF 262K ctx
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/heavylildude%2FQwen3.5-4b-Uncensored-NonReasoning"
Response includes
  • classification m-uncensored
  • files 4
  • hub_downloads_all_time 1,886
  • author_summary 1 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M-U
Primary method

Uncensored (method unknown)

No other method signals detected in this model.
Confidence
LOW
Why this label 3 signals
Weak or ambiguous signals. Best guess based on catalog patterns; treat as tentative and check the evidence below.
  • 'uncensored' in name/tags but no 'abliterated' marker
  • method not identifiable from author declaration alone
  • may be DPO fine-tune, prompt engineering, or unknown technique
Refusal direction extraction

No specific extraction method could be identified for this model. The producer either did not document it or used a proprietary pipeline.

What is a refusal direction? →
Downloads · lifetime
2K
403 last 30d - stable
Likes
1
Model age
5mo ago
created 2026-04-19
Downloads over time
Now2K→from635↑219%
5651.1K1.6K2.2K635 on Apr 222K on Oct 11AprMayJunJulAugSepOct
Apr 22 → Oct 11 · 64 snapshots · spans 172 days

Metadata

License
apache-2.0
Quantizations
Q4_K
Tags
gguf uncensored non-reasoning not-for-all-audiences license:apache-2.0 endpoints_compatible region:us conversational
Total size
2.52 GB
Files
4
Quantizations
3
Registered
2026-08-22 13:56
Last updated on HF
2026-04-20 03:05

Files by quantization

Q4_K 1 file 2.52 GB
Qwen3.5-4b-Uncensored-NonReasoning.Q4_K_M.gguf 2.52 GB 5497ea54 download
Q8_0 1 file 350 MB
Qwen3.5-4b-Uncensored-NonReasoning.mmproj-Q8_0.gguf 350 MB a32639b4 download
Auxiliary files 2 files 5.46 KB
README.md 3.80 KB 9700d6b7 download
.gitattributes 1.65 KB a660a78b download

README current version from Hugging Face


license: apache-2.0
tags:

  • uncensored
  • non-reasoning
  • not-for-all-audiences

Qwen3.5-4B-Uncensored-NonReasoning

A stripped-down, uncensored variant of Qwen 3.5 4B built for one thing: speed.

This model removes the usual reasoning traces and <think> style output, so responses are direct, immediate, and way faster compared to the standard reasoning-heavy variants.

No special flags. No forcing reasoning to 0. No weird prompt hacks. Just load it and send it.


Why This Exists

Most recent reasoning models spend a lot of tokens thinking before answering.

That is cool for deep analysis, but for everyday use, coding, chatting, roleplay, assistant tasks, and general local AI workflows, it can feel slow, bloated, and honestly kinda cooked.

This model is designed to:

  • Remove visible reasoning traces
  • Eliminate <think> output
  • Respond directly without extra chain-of-thought dumping
  • Run faster in local inference setups
  • Work cleanly in llama.cpp, llama-server, OpenWebUI, SillyTavern, KoboldCpp, LM Studio, and similar tools
  • Keep multimodal / vision support via mmproj

Features

  • Based on Qwen 3.5 4B
  • Uncensored behavior
  • Non-reasoning output
  • No <think> tags
  • Faster inference
  • GGUF format
  • Vision capable with mmproj
  • No special llama-server flags required

Quick Start

llama-server \
  -m Qwen3.5-4b-Uncensored-NonReasoning.Q4_K_M.gguf \
  --mmproj Qwen3.5-4b-Uncensored-NonReasoning.mmproj-Q8_0.gguf

Boom. Model runs immediately with vision capability.

No need for:

--reasoning-budget 0
  • No need for prompt templates forcing no thinking.
  • No need for extra hacks to suppress reasoning output.

Recommended Use Cases

  • Fast chat assistants
  • Coding help
  • Roleplay
  • Creative writing
  • Local AI companions
  • Vision tasks
  • OpenWebUI agents
  • Lightweight RAG pipelines
  • SillyTavern characters
  • Low latency local inference

Vision Support

This release supports multimodal input through the included mmproj file.

Example:

llama-server \
  -m Qwen3.5-4b-Uncensored-NonReasoning.Q4_K_M.gguf \
  --mmproj Qwen3.5-4b-Uncensored-NonReasoning.mmproj-Q8_0.gguf

Compatible with image input in llama.cpp builds that support multimodal inference.


Performance Notes

Compared to the original reasoning-enabled Qwen variants, this model generally:

  • Starts responding faster
  • Uses fewer output tokens
  • Avoids wasting context on hidden reasoning
  • Feels more responsive for conversation
  • Works better on lower-end GPUs and CPUs

Especially useful if you are running local inference on limited hardware and want snappy output instead of waiting for the model to internally monologue for 500 tokens before answering a basic question.


Example Prompt

User: Write a Python script that renames all jpg files in a folder.

Instead of:

<think>
The user wants...
</think>

You just get the answer directly.


Files


Compatibility

Tested or intended for:

  • llama.cpp
  • llama-server
  • LM Studio
  • OpenWebUI
  • KoboldCpp
  • SillyTavern
  • Text Generation WebUI
  • Anything GGUF-compatible

Disclaimer

  • This is an uncensored model.
  • Outputs may be inaccurate, offensive, unsafe, biased, or inappropriate depending on prompt and usage.
  • Use responsibly.
  • You are responsible for how you deploy and use this model.

Credits

Base model by Qwen team.
Modified and stripped for non-reasoning fast inference workflows.

README history 3 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-04-19Update README.mda8c51b53.8 KB
    Loading...
  2. 2026-04-19Update README.mdb2f8ac13.7 KB
    Loading...
  3. 2026-04-19initial commita7aa19228 B
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration