← back to catalog · registered 2026-08-22 13:56

nandukmelath/Qwen3.5-9B-Uncensored-nothink-GGUF

nandukmelath Qwen 9B GGUF second-order 262K ctx
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/nandukmelath%2FQwen3.5-9B-Uncensored-nothink-GGUF"
Response includes
  • classification m-uncensored
  • files 3
  • hub_downloads_all_time 6,998
  • author_summary 1 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M-U
Primary method

Uncensored (method unknown)

No other method signals detected in this model.
Confidence
LOW
Why this label 3 signals
Weak or ambiguous signals. Best guess based on catalog patterns; treat as tentative and check the evidence below.
  • 'uncensored' in name/tags but no 'abliterated' marker
  • method not identifiable from author declaration alone
  • may be DPO fine-tune, prompt engineering, or unknown technique
Refusal direction extraction

No specific extraction method could be identified for this model. The producer either did not document it or used a proprietary pipeline.

What is a refusal direction? →
Downloads · lifetime
7K
863 last 30d - stable
Likes
10
Model age
4mo ago
created 2026-06-05
Downloads over time
Now7.5K→from359↑1,996%
12.7K5.5K8.2K359 on Jun 107.5K on Oct 11JunJulAugSepOct
Jun 10 → Oct 11 · 59 snapshots · spans 123 days

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
apache-2.0
Languages
en
Quantizations
Q4_K
Tags
gguf uncensored qwen3 local-llm no-think apple-silicon lm-studio zero-guardrail fast-inference text-generation en base_model:HauhauCS/Qwen3.5-9B-Uncensored-HauhauCS-Aggressive

Related

Total size
5.24 GB
Files
3
Quantizations
2
Registered
2026-08-22 13:56
Last updated on HF
2026-06-05 14:21

Files by quantization

Q4_K 1 file 5.24 GB
Qwen3.5-9B-Uncensored-nothink-Q4_K_M.gguf 5.24 GB f94b69c0 download
Auxiliary files 2 files 5.14 KB
README.md 3.58 KB 07492c27 download
.gitattributes 1.56 KB 605a2688 download

README current version from Hugging Face


license: apache-2.0
base_model:

  • HauhauCS/Qwen3.5-9B-Uncensored-HauhauCS-Aggressive
  • Qwen/Qwen3.5-9B
    tags:
  • uncensored
  • gguf
  • qwen3
  • local-llm
  • no-think
  • apple-silicon
  • lm-studio
  • zero-guardrail
  • fast-inference
    language:
  • en
    pipeline_tag: text-generation

Qwen3.5-9B Uncensored — No-Think Edition (GGUF)

⚡ Zero refusals. Zero thinking delay. 100% local.

This is a patched GGUF of HauhauCS/Qwen3.5-9B-Uncensored-HauhauCS-Aggressive with one key modification: thinking is disabled at the GGUF template level, giving you instant responses without the 15–30 second reasoning delay.

What's different

Qwen3.5 is a thinking model. By default it outputs a <think>...</think> block before every response. This is great for hard problems but brutal for everyday use — you wait 20 seconds for a simple answer.

This model patches the embedded Jinja2 chat template to always output an empty think block:

Original flow:  <think> [400 tokens] </think> → answer    (~25s wait)
This model:     <think></think> → answer                  (<1s wait)

The model's intelligence is encoded in its weights, not the thinking trace. Quality is the same. Speed is 25x better for time-to-first-token.

Want reasoning on demand? Add /think to any message — the model will reason through it fully for that turn only.

Model details

Property Value
Base Qwen3.5-9B
Fine-tune HauhauCS Uncensored Aggressive
Quantization Q4_K_M
Context Up to 65,536 tokens
Parameters 9B
Format GGUF
Refusal rate 0%

Benchmarks (MacBook Pro M2 Pro, 16 GB)

Metric Value
Generation speed ~22–25 tok/s
Time to first token < 1 second
Context window 65,536 tokens
VRAM usage ~8.5 GB

How to use

LM Studio (recommended)

  1. Download the Q4_K_M file below
  2. Load in LM Studio with --context-length 65536 --gpu max
  3. Done — no config needed, thinking is already patched off

Optimal sampling (Qwen3 official recommended)

Temperature: 0.6
Top-P: 0.95
Top-K: 20
Repeat penalty: 1.0
Max tokens: 4096

llama.cpp

./llama-cli -m Qwen3.5-9B-Uncensored-nothink-Q4_K_M.gguf \
  --ctx-size 65536 \
  --n-gpu-layers 99 \
  -p "Your prompt here"

Full automated setup for Mac

👉 github.com/nandukmelath/lmstudio-uncensored-setup

One command: VRAM boost + auto-start + model load + Hermes Agent config:

git clone https://github.com/nandukmelath/lmstudio-uncensored-setup
cd lmstudio-uncensored-setup && ./scripts/setup.sh

How the patch works

The Qwen3.5 GGUF contains an embedded Jinja2 chat template with this block:

{%- if enable_thinking is defined and enable_thinking is false %}
    {{- '<think>\n\n</think>\n\n' }}
{%- else %}
    {{- '<think>\n' }}
{%- endif %}

The patch replaces it with just:

{{- '<think>\n\n</think>\n\n' }}

Same file size (padded with spaces), same structure, zero thinking overhead.
The patcher script is open source: patch_nothink.py

Credits

License

Apache 2.0

README history 1 version

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-06-05Add viral model card with full docs, benchmarks, and patch explanationc38dfe73.6 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration