← back to catalog · registered 2026-08-22 13:56

LEONW24/Qwen3.5-9B-Uncensored

LEONW24 Qwen 9B GGUF 262K ctx
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/LEONW24%2FQwen3.5-9B-Uncensored"
Response includes
  • classification m-uncensored
  • files 3
  • benchmarks 11 entries
  • hub_downloads_all_time 36,698
  • author_summary 1 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M-U
Primary method

Uncensored (method unknown)

No other method signals detected in this model.
Confidence
LOW
Why this label 3 signals
Weak or ambiguous signals. Best guess based on catalog patterns; treat as tentative and check the evidence below.
  • 'uncensored' in name/tags but no 'abliterated' marker
  • method not identifiable from author declaration alone
  • may be DPO fine-tune, prompt engineering, or unknown technique
Refusal direction extraction

No specific extraction method could be identified for this model. The producer either did not document it or used a proprietary pipeline.

What is a refusal direction? →
Downloads · lifetime
37K
4K last 30d - stable
Likes
57
Model age
7mo ago
created 2026-03-10
Downloads over time
Now37.9K→from1.5K↑2,454%
013.9K27.7K41.6K1.5K on Mar 1137.9K on Oct 11MarAprMayJunJulAugSepOct
Mar 11 → Oct 11 · 72 snapshots · spans 214 days

Benchmarks

Portrait before abliteration
Benchmarks of the base model as it stood before the refusal-removal operation. Compare with the numbers above to see what the operation cost.
Benchmark Score Source
Entertainment 1.4 UGI
Hazardous 2.9 UGI
Natural Intelligence 15.15 UGI
Political lean -9.9% UGI
Sensitive-Info 18.27 UGI
SocPol 1.4 UGI
UGI 32.18 UGI
Willingness (10) 6 UGI
W10-Adherence 7 UGI
W10-Direct 5 UGI
Writing 27.96 UGI

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
apache-2.0
Languages
en zh
Quantizations
Q4_K
Tags
gguf qwen3.5 uncensored ollama text-generation en zh base_model:HauhauCS/Qwen3.5-9B-Uncensored-HauhauCS-Aggressive base_model:quantized:HauhauCS/Qwen3.5-9B-Uncensored-HauhauCS-Aggressive license:apache-2.0 endpoints_compatible region:us

Related

Total size
6.23 GB
Files
3
Quantizations
2
Registered
2026-08-22 13:56
Last updated on HF
2026-03-12 03:26

Files by quantization

Q4_K 1 file 6.23 GB
Qwen3.5-9B-Uncensored-Q4_K_M.gguf 6.23 GB e2427ba6 download
Auxiliary files 2 files 7.37 KB
README.md 5.69 KB 1fc198f5 download
.gitattributes 1.67 KB 3e6372e2 download

README current version from Hugging Face


license: apache-2.0
language:

  • en
  • zh
    base_model:
  • Qwen/Qwen3-8B
  • HauhauCS/Qwen3.5-9B-Uncensored-HauhauCS-Aggressive
    tags:
  • qwen3.5
  • uncensored
  • gguf
  • ollama
    library_name: gguf
    pipeline_tag: text-generation
    model-index:
  • name: Qwen3.5-9B-Uncensored
    results: []

Qwen3.5-9B-Uncensored (GGUF)

An uncensored GGUF merge of Qwen 3.5 9B, ready for local deployment with Ollama, llama.cpp, or any GGUF-compatible runtime.

Background

This model is built upon HauhauCS/Qwen3.5-9B-Uncensored-HauhauCS-Aggressive. The original base model has the following known issues:

  1. Ollama deployment failure — The base model cannot be directly deployed via Ollama due to architecture/format incompatibilities. This GGUF version resolves the issue by converting and merging the weights into a single GGUF file that Ollama can load natively.
  2. Broken multimodal input — The base model's packaging causes multimodal (e.g., image) input to malfunction. Although the underlying Qwen 3.5 architecture supports vision capabilities, the way the original model was packaged breaks multimodal inference.

This repo provides a Q4_K_M quantized GGUF version that fixes the Ollama deployment issue while keeping the model compact and efficient.

Quick Start

Ollama (Recommended)

  1. Create a Modelfile:
FROM ./Qwen3.5-9B-Uncensored-Q4_K_M.gguf

PARAMETER temperature 0.7
PARAMETER top_p 0.9
PARAMETER num_ctx 16384

TEMPLATE """{{- if .System }}{{ .System }}{{ end }}
{{- range .Messages }}
<|im_start|>{{ .Role }}
{{ .Content }}<|im_end|>
{{ end }}<|im_start|>assistant
"""

SYSTEM "You are a helpful assistant."
  1. Build and run:
# Download the GGUF file
huggingface-cli download LEONW24/Qwen3.5-9B-Uncensored Qwen3.5-9B-Uncensored-Q4_K_M.gguf --local-dir .

# Create the Ollama model
ollama create qwen35-uncensored -f Modelfile

# Run
ollama run qwen35-uncensored

llama.cpp

# Download
huggingface-cli download LEONW24/Qwen3.5-9B-Uncensored Qwen3.5-9B-Uncensored-Q4_K_M.gguf --local-dir .

# Run with llama-cli
llama-cli -m Qwen3.5-9B-Uncensored-Q4_K_M.gguf -p "Hello, who are you?" -n 256 -ngl 99

llama-cpp-python (OpenAI-compatible API)

pip install llama-cpp-python[server]

python -m llama_cpp.server \
  --model Qwen3.5-9B-Uncensored-Q4_K_M.gguf \
  --n_gpu_layers 99 \
  --chat_format chatml

Then call http://localhost:8000/v1/chat/completions with any OpenAI-compatible client.

Python (ctransformers / llama-cpp-python)

from llama_cpp import Llama

llm = Llama(
    model_path="Qwen3.5-9B-Uncensored-Q4_K_M.gguf",
    n_gpu_layers=-1,  # offload all layers to GPU
    n_ctx=16384,
)

output = llm.create_chat_completion(
    messages=[{"role": "user", "content": "Hello!"}]
)
print(output["choices"][0]["message"]["content"])

Model Details

Property Value
Architecture Qwen 3.5
Parameters ~9B
Format GGUF (Q4_K_M quantization)
File size ~6.3 GB
Context window Up to 131072 tokens
Languages English, Chinese, multilingual
License Apache 2.0

Used In: PhdBooster

This model serves as the local fallback vision model in PhdBooster — an AI-powered browsing assistant that helps PhD students optimize their short video feeds.

What is PhdBooster?

PhD life is stressful. You open Douyin or Xiaohongshu to relax, but all you get is ads and news. PhdBooster is built on OpenClaw — it scrolls through short video platforms while you write papers, uses vision models to actually "see" every video, and automatically likes & bookmarks content matching your taste, training the recommendation algorithm to serve you better.

You're writing a paper
  -> PhdBooster is browsing videos for you
    -> AI "sees" each video
      -> Matches your taste? Auto like & bookmark
        -> Platform algorithm learns your preferences
          -> You open your phone — feed is perfect

How this model fits in

PhdBooster uses a two-stage filtering funnel:

  1. Text quick-filter — Parse title, hashtags, and author info to skip obvious non-targets (~60% filtered out)
  2. Visual deep-filter — Screenshot the video and send to a vision model for analysis against your preference policy

The vision analysis uses a dual fallback strategy:

  • Primary: Kimi 2.5 (Moonshot AI) — fast, high quality
  • Fallback: This model (Qwen3.5-9B-Uncensored) via local Ollama — completely free, works offline, no content filtering

The uncensored nature of this model is important for PhdBooster's use case: it needs to analyze all types of visual content without refusals or safety-triggered false negatives.

Tech Stack

Component Choice
Agent Framework OpenClaw
Browser Automation OpenClaw Browser (Chrome CDP)
Primary LLM step-3.5-flash:free (OpenRouter)
Primary Vision Kimi 2.5 (Moonshot AI)
Fallback Vision This model (Ollama)
Platforms Douyin, Xiaohongshu

For more details, see the PhdBooster README.


Notes

  • This is an uncensored model — it has reduced safety filters compared to the official release. Use responsibly.
  • GGUF format enables efficient CPU + GPU inference without requiring the full PyTorch/transformers stack.
  • For multi-GPU setups with Ollama, set OLLAMA_NUM_PARALLEL and adjust num_gpu as needed.

README history 7 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-03-12Upload README.md with huggingface_hubd73f76d5.7 KB
    Loading...
  2. 2026-03-12Upload README.md with huggingface_huba44a83a5.7 KB
    Loading...
  3. 2026-03-12Upload README.md with huggingface_hub15523cd5.6 KB
    Loading...
  4. 2026-03-10Upload README.md with huggingface_hubace4e383.5 KB
    Loading...
  5. 2026-03-10Upload README.md with huggingface_hubcd040823.5 KB
    Loading...
  6. 2026-03-10Upload README.md with huggingface_hub414641b3.3 KB
    Loading...
  7. 2026-03-10initial commit2afb6c321 B
    Loading...

Discussions 3 threads

  1. 2026-05-17PRuncensored9B1 💬#3
    Loading...
  2. 2026-03-29Why there is not a mmproj file?open1 💬#2
    Loading...
  3. 2026-03-23llama.cppopen6 💬#1
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration