← back to catalog · registered 2026-10-09 09:58

NeuralNet-Hub/Qwen3.8-27B-Uncensored-NVFP4

NeuralNet-Hub Qwen 27B multimodal
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/NeuralNet-Hub%2FQwen3.8-27B-Uncensored-NVFP4"
Response includes
  • classification m-uncensored
  • files 20
  • author_summary 6 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M-U
Primary method

Uncensored (method unknown)

No other method signals detected in this model.
Confidence
LOW
Why this label 3 signals
Weak or ambiguous signals. Best guess based on catalog patterns; treat as tentative and check the evidence below.
  • 'uncensored' in name/tags but no 'abliterated' marker
  • method not identifiable from author declaration alone
  • may be DPO fine-tune, prompt engineering, or unknown technique
Refusal direction extraction

No specific extraction method could be identified for this model. The producer either did not document it or used a proprietary pipeline.

What is a refusal direction? →
Downloads · 30-day
0
Likes
0
Model age
today
created 2026-10-09

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
apache-2.0
Tags
transformers safetensors qwen3_5 image-text-to-text qwen3.8 qwen nvfp4 vision-language vllm blackwell conversational base_model:Qwen/Qwen3.8-27B

Related

Total size
23.0 GB
Files
20
Quantizations
1
Registered
2026-10-09 09:58
Last updated on HF
2026-10-09 09:27

Files by quantization

Auxiliary files 20 files 23.0 GB
model-00001-of-00005.safetensors 4.66 GB 10d5fe95 download
model-00002-of-00005.safetensors 4.63 GB cab6cf5f download
model-00003-of-00005.safetensors 4.62 GB 498018ce download
model-00004-of-00005.safetensors 4.62 GB 0a1564a8 download
model-00005-of-00005.safetensors 3.68 GB f3cf66be download
model-extra-00001-of-00001.safetensors 810 MB d869b1df download
tokenizer.json 19.1 MB f399b3cd download
vocab.json 6.41 MB 0aa0ce06 download
merges.txt 3.20 MB a494e019 download
model.safetensors.index.json 189 KB c2f1be2f download
config.json 45.6 KB ef70c939 download
recipe.yaml 33.2 KB 535835b1 download
LICENSE 11.3 KB f938136e download
chat_template.jinja 8.69 KB 5335d225 download
README.md 7.13 KB a3557f28 download
.gitattributes 1.53 KB 52373fe2 download
tokenizer_config.json 1.10 KB 6913705f download
preprocessor_config.json 390 B 2ea84a43 download
video_preprocessor_config.json 385 B 3ba673a5 download
generation_config.json 214 B 8b9f95da download

README current version from Hugging Face


license: apache-2.0
license_link: https://huggingface.co/Qwen/Qwen3.8-27B/blob/main/LICENSE
base_model: Qwen/Qwen3.8-27B
tags:

  • qwen3.8
  • qwen
  • nvfp4
  • vision-language
  • image-text-to-text
  • vllm
  • blackwell
    library_name: transformers
    pipeline_tag: image-text-to-text
    quantized_by: NeuralNet-Hub

NeuralNet is a pioneering AI solutions provider that empowers businesses to harness the power of artificial intelligence.

TRY THIS MODEL FOR FREE IN UNCENSOREDGPT

🔓 No Filters. No Limits. Just Answers.

This model powers UncensoredGPT

Ask anything. Get real answers. No restrictions.

Join the Waitlist

Most AI models are trained to refuse. They hedge, they deflect, they lecture. UncensoredGPT is built on the opposite philosophy: that access to information should be unrestricted, and that adults are capable of deciding what they need to know.

This model is the engine behind UncensoredGPT, a platform providing unfiltered, honest responses for cybersecurity, education, content creation, research, or straightforward conversation. The refusals and content restrictions present in the original model have been removed through a combination of supervised fine-tuning and abliteration, resulting in a model that responds directly across topics that standard models typically refuse.

Why stay in the system when you can have unrestricted answers, privacy by default, and complete freedom of information?

Ready to experience the freedom of unrestricted AI? Join to uncensoredgpt.ai

⚠️ Model created by OrcaRouter 🐋

This is an NVFP4-quantized version of Qwen/Qwen3.8-27B, created by OrcaRouter, it's exactly the same model, we are saving it here as we use it in UncensoredGPT.

We have not created nor optimized this model, but we needed it only because we modified the chat template to make it compatible with many other applications as per the suggestion by SerialKicked.

We have replaced

{%- for message in messages %}
    {%- set content = render_content(message.content, true)|trim %}
    {%- if message.role == "system" %}
        {%- if not loop.first %}
            {{- raise_exception('System message must be at the beginning.') }}
        {%- endif %}
    {%- elif message.role == "user" %}
        {{- '<|im_start|>' + message.role + '\n' + content + '<|im_end|>' + '\n' }}

By

{%- for message in messages %}
    {%- set content = render_content(message.content, true)|trim %}
    {%- if message.role == "system" %}
        {{- '<|im_start|>' + message.role + '\n' + content + '<|im_end|>' + '\n' }}
    {%- elif message.role == "user" %}
        {{- '<|im_start|>' + message.role + '\n' + content + '<|im_end|>' + '\n' }}

It's optimized for deployment on NVIDIA Blackwell architecture GPUs using vLLM.

[!IMPORTANT]
NVFP4 quantization requires NVIDIA Blackwell architecture (GB200, RTX 5000 series, etc.). This format is not compatible with Ampere, Ada Lovelace, or Hopper GPUs. If you are running on an older GPU, please use a different quantization format.

Original model: https://huggingface.co/Qwen/Qwen3.8-27B

⚡ Deployment with vLLM

This quantized model is intended to be served using vLLM (vllm>=0.9.0 recommended).

Quick Start

vllm serve NeuralNet-Hub/Qwen3.8-27B-NVFP4 \
  --quantization nvfp4 \
  --dtype bfloat16 \
  --kv-cache-dtype fp8 \
  --max-model-len 262144 \
  --reasoning-parser qwen3 \
  --enable-auto-tool-choice \
  --tool-call-parser qwen3_coder

Using a Config File

# Deploy with: vllm serve --config config.yaml

# Model config
model: NeuralNet-Hub/Qwen3.8-27B-NVFP4
dtype: bfloat16
kv-cache-dtype: fp8
gpu-memory-utilization: 0.95
max-model-len: 262144
max-num-seqs: 200
enable-prefix-caching: true
trust-remote-code: true

# Parsing
reasoning-parser: qwen3
enable-auto-tool-choice: true
tool-call-parser: qwen3_coder


# Optional
default-chat-template-kwargs: '{"enable_thinking": false}'
download-dir: /workspace/models
host: 0.0.0.0
port: 18000
vllm serve --config config.yaml

💬 Chat API Usage

Once you deploy your model using vLLM you can chat qwith Qwen3.8 with chat template compatible with OpenAI-format APIs. Thinking mode is enabled by default.

Thinking Mode (Default)

from openai import OpenAI

client = OpenAI(base_url="http://localhost:18000/v1", api_key="EMPTY")

messages = [{"role": "user", "content": "Your message here"}]

response = client.chat.completions.create(
    model="NeuralNet-Hub/Qwen3.8-27B-NVFP4",
    messages=messages,
    max_tokens=32768,
    temperature=1.0,
    top_p=0.95,
    extra_body={"top_k": 20},
)
print(response.choices[0].message.content)

Non-Thinking (Instruct) Mode

response = client.chat.completions.create(
    model="NeuralNet-Hub/Qwen3.8-27B-NVFP4",
    messages=messages,
    max_tokens=8192,
    temperature=0.7,
    top_p=0.8,
    presence_penalty=1.5,
    extra_body={
        "top_k": 20,
        "chat_template_kwargs": {"enable_thinking": False},
    },
)

Image Input

messages = [
    {
        "role": "user",
        "content": [
            {"type": "image_url", "image_url": {"url": "https://example.com/image.jpg"}},
            {"type": "text", "text": "Describe this image in detail."}
        ]
    }
]

response = client.chat.completions.create(
    model="NeuralNet-Hub/Qwen3.8-27B-NVFP4",
    messages=messages,
    max_tokens=32768,
    temperature=1.0,
    top_p=0.95,
    extra_body={"top_k": 20},
)

⚙️ Recommended Sampling Parameters

Mode temperature top_p top_k presence_penalty
Thinking — general tasks 1.0 0.95 20 0.0
Thinking — precise coding 0.6 0.95 20 0.0
Instruct (non-thinking) 0.7 0.80 20 1.5

🔧 Hardware Requirements

Component Requirement
GPU Architecture NVIDIA Blackwell (sm_100+)
VRAM 48 GB+ recommended
CUDA 12.8+
vLLM 0.9.0+

[!WARNING]
NVFP4 is exclusively supported on NVIDIA Blackwell GPUs. Attempting to run this model on Ampere (A100), Ada Lovelace (RTX 4000), or Hopper (H100) will fail. For those architectures, use the original BF16 model or an AWQ/GPTQ quantized variant.

🌐 Contact Us

NeuralNet is a pioneering AI solutions provider that empowers businesses to harness the power of artificial intelligence.

Website: https://neuralnet.solutions
Email: info[at]neuralnet.solutions

Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration