← back to catalog · registered 2026-08-22 13:56

lyf/Qwen3.5-27B-Uncensored-HauhauCS-Aggressive-NVFP4

lyf Qwen 12B multimodal second-order
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/lyf%2FQwen3.5-27B-Uncensored-HauhauCS-Aggressive-NVFP4"
Response includes
  • classification m-uncensored
  • files 19
  • hub_downloads_all_time 19,899
  • author_summary 10 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M-U
Primary method

Uncensored (method unknown)

No other method signals detected in this model.
Confidence
LOW
Why this label 3 signals
Weak or ambiguous signals. Best guess based on catalog patterns; treat as tentative and check the evidence below.
  • 'uncensored' in name/tags but no 'abliterated' marker
  • method not identifiable from author declaration alone
  • may be DPO fine-tune, prompt engineering, or unknown technique
Refusal direction extraction

No specific extraction method could be identified for this model. The producer either did not document it or used a proprietary pipeline.

What is a refusal direction? →
Downloads · lifetime
20K
139 last 30d - cooling
Likes
10
Model age
6mo ago
created 2026-03-16
Downloads over time
Now20K→from429↑4,553%
07.3K14.6K21.9K429 on Mar 1820K on Oct 11MarAprMayJunJulAugSepOct
Mar 18 → Oct 11 · 69 snapshots · spans 207 days

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
apache-2.0
Tags
transformers safetensors qwen3_5 image-text-to-text qwen3.5 nvfp4 quantized uncensored multimodal vision gated-deltanet tool-calling

Related

Total size
18.4 GB
Files
19
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-04-25 08:40

Files by quantization

Auxiliary files 19 files 18.4 GB
model-00003-of-00006.safetensors 4.00 GB d9f7f8dd download
model-00004-of-00006.safetensors 4.00 GB 629f7e19 download
model-00002-of-00006.safetensors 4.00 GB 3a2a91e0 download
model-00005-of-00006.safetensors 3.17 GB a97bab75 download
model-00001-of-00006.safetensors 2.37 GB 86df4251 download
model-multimodal-extra.safetensors 879 MB 8e8dbe7f download
tokenizer.json 12.2 MB 5f9e4d49 download
vocab.json 6.41 MB 0aa0ce06 download
merges.txt 3.20 MB a494e019 download
model.safetensors.index.json 239 KB 100ca342 download
tokenizer_config.json 16.3 KB eda48d3e download
config.json 15.3 KB ffee603a download
chat_template.jinja 7.57 KB a585dec8 download
README.md 7.41 KB ec3587cd download
.gitattributes 1.53 KB 52373fe2 download
preprocessor_config.json 390 B 2ea84a43 download
video_preprocessor_config.json 385 B 3ba673a5 download
recipe.yaml 225 B 86927e7f download
generation_config.json 213 B 8c8412a2 download

README current version from Hugging Face


license: apache-2.0
base_model: HauhauCS/Qwen3.5-27B-Uncensored-HauhauCS-Aggressive
tags:

  • qwen3.5
  • nvfp4
  • quantized
  • uncensored
  • multimodal
  • vision
  • gated-deltanet
  • tool-calling
    library_name: transformers
    pipeline_tag: image-text-to-text

news: I made a nvfp4 quant for HauhauCS/Qwen3.6-27B-Uncensored-HauhauCS-Aggressive too!

Please check my profile or go here

Qwen3.5-27B-Uncensored-HauhauCS-Aggressive-NVFP4

NVIDIA FP4 (NVFP4) quantized version of HauhauCS/Qwen3.5-27B-Uncensored-HauhauCS-Aggressive, with full multimodal (vision) and tool-calling capability preserved.

Model Details

  • Base model: HauhauCS/Qwen3.5-27B-Uncensored-HauhauCS-Aggressive
  • Architecture: Qwen3_5ForConditionalGeneration (hybrid Gated-DeltaNet + full attention)
  • Quantization: NVFP4 (4-bit weights + FP8 activations) via llm-compressor
  • Calibration: 512 samples from neuralmagic/calibration (LLM split), 4096 seq length
  • Model size: ~19.7 GB (vs ~54 GB bf16 original)
  • Vision encoder: bf16 (unquantized, ~0.9 GB)

What's quantized, what's not

Component Format Notes
MLP (gate/up/down_proj) NVFP4 All 64 layers
Full attention (q/k/v/o_proj) NVFP4 16 layers (every 4th)
Linear attention (in_proj_qkv/z, out_proj) NVFP4 48 layers
Linear attention (in_proj_a/b) bf16 SSM parameters, excluded
lm_head bf16 Output projection, excluded
Vision encoder bf16 All vision weights, excluded
Norms, biases, A_log, dt_bias, conv1d bf16 Small tensors, excluded

Quantization Recipe

from llmcompressor.modifiers.quantization import QuantizationModifier

recipe = QuantizationModifier(
    targets=["Linear"],
    ignore=["lm_head", "re:.*visual.*", "re:.*in_proj_a$", "re:.*in_proj_b$"],
    scheme="NVFP4",
)

Following the approach from Kbenkhaled/Qwen3.5-27B-NVFP4.

Usage with vLLM

Basic (chat + reasoning)

docker run -d --name hauhaucs-nvfp4 \
  --ipc host --network host --device nvidia.com/gpu=all \
  -v ~/.cache/huggingface:/root/.cache/huggingface \
  -v ~/.cache/vllm:/root/.cache/vllm \
  -e VLLM_NVFP4_GEMM_BACKEND=marlin \
  -e VLLM_USE_FLASHINFER_MOE_FP4=0 \
  -e PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True \
  vllm/vllm-openai:cu130-nightly \
  lyf/Qwen3.5-27B-Uncensored-HauhauCS-Aggressive-NVFP4 \
  --host 0.0.0.0 --port 8000 \
  --max-model-len 32768 \
  --gpu-memory-utilization 0.90 \
  --max-num-seqs 4 \
  --max-num-batched-tokens 4096 \
  --kv-cache-dtype fp8 \
  --reasoning-parser qwen3

With tool calling

Add these flags to enable OpenAI-compatible function calling:

  --enable-auto-tool-choice \
  --tool-call-parser qwen3_xml

The model uses XML-style tool calls inherited from the HauhauCS chat template:

<tool_call>
<function=get_weather>
<parameter=city>
Tokyo
</parameter>
</function>
</tool_call>

Important: Use qwen3_xml as the tool-call-parser, NOT qwen3_coder. Although both the standard Qwen3.5 and this model share the same chat template, the qwen3_xml parser is the correct match for the <tool_call><function=...> XML output format. The qwen3_coder parser happens to work in some cases but qwen3_xml is the proper parser for this format.

Disabling thinking mode

To get direct answers without chain-of-thought reasoning, pass enable_thinking: false via the API:

{
  "chat_template_kwargs": {"enable_thinking": false}
}

Some clients (e.g., Chatbox) may send this automatically when thinking mode is toggled off.

Full docker-compose example

services:
  vllm:
    image: vllm/vllm-openai:cu130-nightly
    container_name: hauhaucs-nvfp4
    restart: unless-stopped
    network_mode: host
    ipc: host
    devices:
      - nvidia.com/gpu=all
    volumes:
      - ~/.cache/huggingface:/root/.cache/huggingface
      - ~/.cache/vllm:/root/.cache/vllm
    environment:
      - VLLM_USE_FLASHINFER_MOE_FP4=0
      - VLLM_NVFP4_GEMM_BACKEND=marlin
      - PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True
    command:
      - --model
      - lyf/Qwen3.5-27B-Uncensored-HauhauCS-Aggressive-NVFP4
      - --host
      - "0.0.0.0"
      - --port
      - "8000"
      - --max-model-len
      - "32768"
      - --gpu-memory-utilization
      - "0.90"
      - --max-num-seqs
      - "4"
      - --max-num-batched-tokens
      - "4096"
      - --kv-cache-dtype
      - fp8
      - --reasoning-parser
      - qwen3
      - --enable-auto-tool-choice
      - --tool-call-parser
      - qwen3_xml

Memory budget (RTX 5090, 32GB VRAM)

Component Size
NVFP4 weights ~18 GB
Vision encoder (bf16) ~0.9 GB
KV cache (fp8, 32K ctx) ~8 GB
Overhead ~3 GB
Total ~30 GB

Tested on RTX 5090 with vLLM v0.17+ nightly. Runs at ~82 tokens/s.

Capabilities

Multimodal (vision)

Image understanding works out of the box via the OpenAI vision API format:

{
  "messages": [{
    "role": "user",
    "content": [
      {"type": "text", "text": "What do you see?"},
      {"type": "image_url", "image_url": {"url": "data:image/png;base64,..."}}
    ]
  }]
}

Tool calling

Verified working with --enable-auto-tool-choice --tool-call-parser qwen3_xml. The model correctly populates the tool_calls array in OpenAI-compatible responses with finish_reason: "tool_calls".

Red Team AI Benchmark

Scorer Score
Keyword matching 75.0%
Semantic similarity (gte-large-en-v1.5) 77.1%

12/12 questions answered without refusal. Strongest on low-level C/C++/assembly tasks (PE mapping, syscall shellcode, EDR unhooking). Benchmark: toxy4ny/redteam-ai-benchmark.

How It Was Made

The original model was distributed as GGUF files. Since transformers does not support loading Qwen3.5 from GGUF, we built a manual conversion pipeline that handles three critical GGUF-specific pitfalls:

  1. RMSNorm +1.0 offset -- GGUF stores 1 + learned_param, HF expects learned_param
  2. A_log domain mismatch -- GGUF stores -exp(A_log), HF expects A_log
  3. Value head (3,16) permutation -- GGUF stores 48 value heads in (3-per-group, 16-groups) order; HF expects (16-groups, 3-per-group)

Full pipeline code and detailed write-up: github.com/li-yifei/gguf-to-nvfp4

MT-Bench Results (mini, 24 questions)

Category Score
Math 9.33
Coding 8.83
Humanities 8.33
Writing 7.67
Extraction 7.50
Roleplay 7.33
Reasoning 7.17
STEM 6.67
Overall 7.85

Judged by gpt-5.1-codex-mini.

Acknowledgments

README history 3 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-04-25add: qwen3.6 links96069647.4 KB
    Loading...
  2. 2026-03-17Upload README.md with huggingface_hub1a9b9bd7.1 KB
    Loading...
  3. 2026-03-16Upload folder using huggingface_hub54f7ddd4 KB
    Loading...

Discussions 1 thread

  1. 2026-04-04Stuck on repatingopen3 💬#1
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration