← back to catalog · registered 2026-08-22 13:56

NeuralNet-Hub/Qwen3.5-9B-Uncensored-W4A16

NeuralNet-Hub Qwen 5.3B multimodal
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/NeuralNet-Hub%2FQwen3.5-9B-Uncensored-W4A16"
Response includes
  • classification m1
  • files 10
  • benchmarks 11 entries
  • hub_downloads_all_time 530
  • author_summary 6 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M1
Primary method

Direct removal

No other method signals detected in this model.
Confidence
MEDIUM
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=0 (base model)
  • no specific method indicators - defaulting to M1 (most common)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
MEDIUM
Why we say so
primary_method=M1; difference-of-means is the reference extraction for M1/M3 (Arditi 2024)
Downloads · lifetime
530
125 last 30d - stable
Likes
0
Model age
2mo ago
created 2026-08-02
Downloads over time
Now578→from147↑293%
125291456621147 on Aug 5578 on Oct 11AugSepOct
Aug 5 → Oct 11 · 50 snapshots · spans 67 days

Benchmarks

Portrait before abliteration
Benchmarks of the base model as it stood before the refusal-removal operation. Compare with the numbers above to see what the operation cost.
Benchmark Score Source
Entertainment 1.4 UGI
Hazardous 2.4 UGI
Natural Intelligence 17.62 UGI
Political lean -12.2% UGI
Sensitive-Info 14.65 UGI
SocPol 0.9 UGI
UGI 17.27 UGI
Willingness (10) 2.2 UGI
W10-Adherence 1.5 UGI
W10-Direct 3 UGI
Writing 33.52 UGI

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
apache-2.0
Tags
transformers safetensors qwen3_5 image-text-to-text qwen3.5 qwen w4a16 vision-language vllm blackwell uncensored abliterated

Related

Total size
10.2 GB
Files
10
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-10-09 09:59

Files by quantization

Auxiliary files 10 files 10.2 GB
model.safetensors 10.2 GB d8bd42f1 download
tokenizer.json 19.1 MB 06b95093 download
config.json 18.2 KB fca1ef70 download
chat_template.jinja 7.57 KB a585dec8 download
README.md 6.29 KB 3f95ad3e download
.gitattributes 1.53 KB 52373fe2 download
processor_config.json 1.16 KB 33818c7f download
tokenizer_config.json 1.14 KB b4acebe0 download
recipe.yaml 401 B 8ea19c54 download
generation_config.json 116 B 48697309 download

README current version from Hugging Face


license: apache-2.0
license_link: https://huggingface.co/Qwen/Qwen3.5-9B/blob/main/LICENSE
base_model: Qwen/Qwen3.5-9B
tags:

  • qwen3.5
  • qwen
  • w4a16
  • vision-language
  • image-text-to-text
  • vllm
  • blackwell
  • uncensored
  • abliterated
    library_name: transformers
    pipeline_tag: image-text-to-text
    quantized_by: NeuralNet-Hub

NeuralNet is a pioneering AI solutions provider that empowers businesses to harness the power of artificial intelligence.


🌟 Qwen3.5-9B Uncensored NVFP4 Quantization by NeuralNet 🧠🤖

This is an W4A16-quantized version of Qwen/Qwen3.5-9B, produced through a combination of fine-tuning and abliteration. This model was obtained by quantizing the weights of Qwen/Qwen3.5-9B to INT4 data type while keeping activations in original precision, ready for inference with vLLM The reduction is less than the theoretical 75% because the vision encoder, token embeddings, and linear attention layers remain in BF16.

Only the weights of the linear operators within transformer blocks are quantized using LLM Compressor. The vision encoder, token embeddings, and linear attention layers are not quantized.

[!IMPORTANT]
This format is not compatible with Ampere, Ada Lovelace, Blackwell or Hopper GPUs.


🔓 No Filters. No Limits. Just Answers.

This model powers UncensoredGPT

Ask anything. Get real answers. No restrictions.

Join the Waitlist

Most AI models are trained to refuse. They hedge, they deflect, they lecture. UncensoredGPT is built on the opposite philosophy: that access to information should be unrestricted, and that adults are capable of deciding what they need to know.

This model is the engine behind UncensoredGPT, a platform providing unfiltered, honest responses for cybersecurity, education, content creation, research, or straightforward conversation. The refusals and content restrictions present in the original Qwen3.5-9B have been removed through a combination of supervised fine-tuning and abliteration, resulting in a model that responds directly across topics that standard models typically refuse.

Why stay in the system when you can have unrestricted answers, privacy by default, and complete freedom of information?

Ready to experience the freedom of unrestricted AI? Join the waitlist at uncensoredgpt.ai — limited spots available.

--

⚡ Deployment with vLLM

This quantized model is intended to be served using vLLM (vllm>=0.21.0 recommended).

Quick Start

vllm serve NeuralNet-Hub/Qwen3.5-9B-Uncensored-W4A16 \
  --dtype bfloat16 \
  --kv-cache-dtype fp8 \
  --max-model-len 262144 \
  --reasoning-parser qwen3 \
  --enable-auto-tool-choice \
  --tool-call-parser qwen3_coder

Using a Config File

# Deploy with: vllm serve --config config.yaml
# Optimized for NVIDIA RTX 6000 PRO (Blackwell)
# Benchmarked: ~85-90 parallel requests, up to 1000 tok/sec at higher context lengths

model: NeuralNet-Hub/Qwen3.5-9B-Uncensored-W4A16
dtype: bfloat16
kv-cache-dtype: fp8
gpu-memory-utilization: 0.95
max-model-len: 262144
max-num-batched-tokens: 4096
max-num-seqs: 200
max-cudagraph-capture-size: 209
enable-prefix-caching: true
trust-remote-code: true

reasoning-parser: qwen3
enable-auto-tool-choice: true
tool-call-parser: qwen3_coder

default-chat-template-kwargs: '{"enable_thinking": false}'

download-dir: /workspace/models
host: 0.0.0.0
port: 18000
vllm serve --config config.yaml

💬 Chat API Usage

Qwen3.6 uses a standard chat template compatible with OpenAI-format APIs. Thinking mode is enabled by default.

Thinking Mode (Default)

from openai import OpenAI

client = OpenAI(base_url="http://localhost:18000/v1", api_key="EMPTY")

messages = [{"role": "user", "content": "Your message here"}]

response = client.chat.completions.create(
    model="NeuralNet-Hub/Qwen3.5-9B-Uncensored-W4A16",
    messages=messages,
    max_tokens=32768,
    temperature=1.0,
    top_p=0.95,
    extra_body={"top_k": 20},
)
print(response.choices[0].message.content)

Non-Thinking (Instruct) Mode

response = client.chat.completions.create(
    model="NeuralNet-Hub/Qwen3.5-9B-Uncensored-W4A16",
    messages=messages,
    max_tokens=8192,
    temperature=0.7,
    top_p=0.8,
    presence_penalty=1.5,
    extra_body={
        "top_k": 20,
        "chat_template_kwargs": {"enable_thinking": False},
    },
)

Image Input

messages = [
    {
        "role": "user",
        "content": [
            {"type": "image_url", "image_url": {"url": "https://example.com/image.jpg"}},
            {"type": "text", "text": "Describe this image in detail."}
        ]
    }
]

response = client.chat.completions.create(
    model="NeuralNet-Hub/Qwen3.5-9B-Uncensored-W4A16",
    messages=messages,
    max_tokens=32768,
    temperature=1.0,
    top_p=0.95,
    extra_body={"top_k": 20},
)

⚙️ Recommended Sampling Parameters

Mode temperature top_p top_k presence_penalty
Thinking — general tasks 1.0 0.95 20 0.0
Thinking — precise coding 0.6 0.95 20 0.0
Instruct (non-thinking) 0.7 0.80 20 1.5

📥 Download with huggingface-cli

Install the CLI

pip install -U "huggingface_hub[cli]"

Download the Full Repository

huggingface-cli download NeuralNet-Hub/Qwen3.5-9B-Uncensored-W4A16 --local-dir ./Qwen3.5-9B-Uncensored-W4A16

Download Specific Files

huggingface-cli download NeuralNet-Hub/Qwen3.5-9B-Uncensored-W4A16 \
  --include "*.safetensors" \
  --local-dir ./Qwen3.5-9B-Uncensored-W4A16

🌐 Contact Us

NeuralNet is a pioneering AI solutions provider that empowers businesses to harness the power of artificial intelligence.

Website: https://neuralnet.solutions
Email: info[at]neuralnet.solutions

README history 3 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-08-03Update README.md4a14fd66.3 KB
    Loading...
  2. 2026-08-03Update README.mdd2185286.5 KB
    Loading...
  3. 2026-08-02initial commitbc03aec26 B
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration