← back to catalog · registered 2026-08-22 13:56

NeuralNet-Hub/Qwen3.6-27B-Uncensored-NVFP4

NeuralNet-Hub Qwen 9.4B multimodal
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/NeuralNet-Hub%2FQwen3.6-27B-Uncensored-NVFP4"
Response includes
  • classification m1
  • files 12
  • benchmarks 11 entries
  • hub_downloads_all_time 980
  • author_summary 6 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M1
Primary method

Direct removal

No other method signals detected in this model.
Confidence
MEDIUM
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=0 (base model)
  • no specific method indicators - defaulting to M1 (most common)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
MEDIUM
Why we say so
primary_method=M1; difference-of-means is the reference extraction for M1/M3 (Arditi 2024)
Downloads · lifetime
980
126 last 30d - stable
Likes
1
Model age
4mo ago
created 2026-05-22
Downloads over time
Now1K→from68↑1,422%
203907611.1K68 on May 201K on Oct 11MayJunJulAugSepOct
May 20 → Oct 11 · 60 snapshots · spans 144 days

Benchmarks

Portrait before abliteration
Benchmarks of the base model as it stood before the refusal-removal operation. Compare with the numbers above to see what the operation cost.
Benchmark Score Source
Entertainment 1.2 UGI
Hazardous 4.7 UGI
Natural Intelligence 33.16 UGI
Political lean -20.0% UGI
Sensitive-Info 26.98 UGI
SocPol 2.9 UGI
UGI 27.15 UGI
Willingness (10) 2.8 UGI
W10-Adherence 1.5 UGI
W10-Direct 4 UGI
Writing 42.47 UGI

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
apache-2.0
Tags
transformers safetensors qwen3_5 image-text-to-text qwen3.6 qwen nvfp4 vision-language vllm blackwell uncensored abliterated

Related

Total size
26.6 GB
Files
12
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-10-09 08:29

Files by quantization

Auxiliary files 12 files 26.6 GB
model.safetensors 25.8 GB f560fb1b download
model_mtp.safetensors 810 MB 713b0faf download
tokenizer.json 19.1 MB dd6b8cf7 download
model.safetensors.index.json 162 KB 06870590 download
config.json 24.2 KB a2ff8d91 download
chat_template.jinja 7.58 KB a8755d82 download
README.md 7.10 KB dd4e529b download
.gitattributes 1.53 KB 52373fe2 download
processor_config.json 1.27 KB 7ad6acdf download
tokenizer_config.json 1.11 KB 541f6c47 download
generation_config.json 213 B dff89e46 download
recipe.yaml 209 B ec5e4124 download

README current version from Hugging Face


license: apache-2.0
license_link: https://huggingface.co/Qwen/Qwen3.6-27B/blob/main/LICENSE
base_model: Qwen/Qwen3.6-27B
tags:

  • qwen3.6
  • qwen
  • nvfp4
  • vision-language
  • image-text-to-text
  • vllm
  • blackwell
  • uncensored
  • abliterated
    library_name: transformers
    pipeline_tag: image-text-to-text
    quantized_by: NeuralNet-Hub

NeuralNet is a pioneering AI solutions provider that empowers businesses to harness the power of artificial intelligence.


🌟 Qwen3.6-27B Uncensored NVFP4 Quantization by NeuralNet 🧠🤖

This is an NVFP4-quantized version of NeuralNet-Hub/Qwen3.6-27B-Uncensored, produced through a combination of fine-tuning and abliteration over Qwen/Qwen3.6-27B. It is optimized for deployment on NVIDIA Blackwell architecture GPUs using vLLM.

[!IMPORTANT]
NVFP4 quantization requires NVIDIA Blackwell architecture (GB200, RTX 5000 series, etc.). This format is not compatible with Ampere, Ada Lovelace, or Hopper GPUs. If you are running on an older GPU, please use a different quantization format.


🔓 No Filters. No Limits. Just Answers.

This model powers UncensoredGPT

Ask anything. Get real answers. No restrictions.

Join the Waitlist

Most AI models are trained to refuse. They hedge, they deflect, they lecture. UncensoredGPT is built on the opposite philosophy: that access to information should be unrestricted, and that adults are capable of deciding what they need to know.

This model is the engine behind UncensoredGPT, a platform providing unfiltered, honest responses for cybersecurity, education, content creation, research, or straightforward conversation. The refusals and content restrictions present in the original Qwen3.6-27B have been removed through a combination of supervised fine-tuning and abliteration, resulting in a model that responds directly across topics that standard models typically refuse.

Why stay in the system when you can have unrestricted answers, privacy by default, and complete freedom of information?

Ready to experience the freedom of unrestricted AI? Join the waitlist at uncensoredgpt.ai — limited spots available.


Quantization Details

This model was quantized to NVFP4 (4-bit NVIDIA Floating Point) using vLLM's built-in quantization pipeline. NVFP4 leverages native FP4 Tensor Core support introduced in Blackwell GPUs, delivering significant memory savings and throughput improvements with minimal quality degradation compared to BF16.

vllm quantize \
  --model NeuralNet-Hub/Qwen3.6-27B-Uncensored \
  --quantization nvfp4 \
  --output-dir NeuralNet-Hub/Qwen3.6-27B-Uncensored-NVFP4

⚡ Deployment with vLLM

This quantized model is intended to be served using vLLM (vllm>=0.9.0 recommended).

Quick Start

vllm serve NeuralNet-Hub/Qwen3.6-27B-Uncensored-NVFP4 \
  --quantization nvfp4 \
  --dtype bfloat16 \
  --kv-cache-dtype fp8 \
  --max-model-len 262144 \
  --reasoning-parser qwen3 \
  --enable-auto-tool-choice \
  --tool-call-parser qwen3_coder

Using a Config File

# Deploy with: vllm serve --config config.yaml
# Optimized for NVIDIA RTX 6000 PRO (Blackwell)
# Benchmarked: ~85-90 parallel requests, up to 1000 tok/sec at higher context lengths

model: NeuralNet-Hub/Qwen3.6-27B-Uncensored-NVFP4
dtype: bfloat16
kv-cache-dtype: fp8
gpu-memory-utilization: 0.95
max-model-len: 262144
max-num-batched-tokens: 4096
max-num-seqs: 200
max-cudagraph-capture-size: 209
enable-prefix-caching: true
trust-remote-code: true

reasoning-parser: qwen3
enable-auto-tool-choice: true
tool-call-parser: qwen3_coder

default-chat-template-kwargs: '{"enable_thinking": false}'

download-dir: /workspace/models
host: 0.0.0.0
port: 18000
vllm serve --config config.yaml

💬 Chat API Usage

Qwen3.6 uses a standard chat template compatible with OpenAI-format APIs. Thinking mode is enabled by default.

Thinking Mode (Default)

from openai import OpenAI

client = OpenAI(base_url="http://localhost:18000/v1", api_key="EMPTY")

messages = [{"role": "user", "content": "Your message here"}]

response = client.chat.completions.create(
    model="NeuralNet-Hub/Qwen3.6-27B-Uncensored-NVFP4",
    messages=messages,
    max_tokens=32768,
    temperature=1.0,
    top_p=0.95,
    extra_body={"top_k": 20},
)
print(response.choices[0].message.content)

Non-Thinking (Instruct) Mode

response = client.chat.completions.create(
    model="NeuralNet-Hub/Qwen3.6-27B-Uncensored-NVFP4",
    messages=messages,
    max_tokens=8192,
    temperature=0.7,
    top_p=0.8,
    presence_penalty=1.5,
    extra_body={
        "top_k": 20,
        "chat_template_kwargs": {"enable_thinking": False},
    },
)

Image Input

messages = [
    {
        "role": "user",
        "content": [
            {"type": "image_url", "image_url": {"url": "https://example.com/image.jpg"}},
            {"type": "text", "text": "Describe this image in detail."}
        ]
    }
]

response = client.chat.completions.create(
    model="NeuralNet-Hub/Qwen3.6-27B-Uncensored-NVFP4",
    messages=messages,
    max_tokens=32768,
    temperature=1.0,
    top_p=0.95,
    extra_body={"top_k": 20},
)

⚙️ Recommended Sampling Parameters

Mode temperature top_p top_k presence_penalty
Thinking — general tasks 1.0 0.95 20 0.0
Thinking — precise coding 0.6 0.95 20 0.0
Instruct (non-thinking) 0.7 0.80 20 1.5

📥 Download with huggingface-cli

Install the CLI

pip install -U "huggingface_hub[cli]"

Download the Full Repository

huggingface-cli download NeuralNet-Hub/Qwen3.6-27B-Uncensored-NVFP4 --local-dir ./Qwen3.6-27B-Uncensored-NVFP4

Download Specific Files

huggingface-cli download NeuralNet-Hub/Qwen3.6-27B-Uncensored-NVFP4 \
  --include "*.safetensors" \
  --local-dir ./Qwen3.6-27B-Uncensored-NVFP4

🔧 Hardware Requirements

Component Requirement
GPU Architecture NVIDIA Blackwell (sm_100+)
VRAM 24 GB+ recommended
CUDA 12.8+
vLLM 0.9.0+

[!WARNING]
NVFP4 is exclusively supported on NVIDIA Blackwell GPUs. Attempting to run this model on Ampere (A100), Ada Lovelace (RTX 4000), or Hopper (H100) will fail. For those architectures, use the original BF16 model or an AWQ/GPTQ quantized variant.


🌐 Contact Us

NeuralNet is a pioneering AI solutions provider that empowers businesses to harness the power of artificial intelligence.

Website: https://neuralnet.solutions
Email: info[at]neuralnet.solutions

README history 2 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-05-22Update README.mdff64fed7.1 KB
    Loading...
  2. 2026-05-22initial commit4d14fa528 B
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration