← back to catalog · registered 2026-08-22 13:56

lyf/Qwen3.6-35B-A3B-Uncensored-HauhauCS-Aggressive-NVFP4

lyf Qwen 16B MoE multimodal second-order
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/lyf%2FQwen3.6-35B-A3B-Uncensored-HauhauCS-Aggressive-NVFP4"
Response includes
  • classification m-uncensored
  • files 11
  • hub_downloads_all_time 219,803
  • author_summary 10 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M-U
Primary method

Uncensored (method unknown)

No other method signals detected in this model.
Confidence
LOW
Why this label 3 signals
Weak or ambiguous signals. Best guess based on catalog patterns; treat as tentative and check the evidence below.
  • 'uncensored' in name/tags but no 'abliterated' marker
  • method not identifiable from author declaration alone
  • may be DPO fine-tune, prompt engineering, or unknown technique
Refusal direction extraction

No specific extraction method could be identified for this model. The producer either did not document it or used a proprietary pipeline.

What is a refusal direction? →
Downloads · lifetime
220K
3K last 30d - cooling
Likes
19
Model age
5mo ago
created 2026-04-21
Downloads over time
Now220.6K→from0↑0%
080.9K161.8K242.6K0 on Apr 22220.6K on Oct 11AprMayJunJulAugSepOct
Apr 22 → Oct 11 · 65 snapshots · spans 172 days

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
apache-2.0
Tags
transformers safetensors qwen3_5_moe image-text-to-text qwen3.6 nvfp4 compressed-tensors quantized vllm moe multimodal blackwell

Related

Total size
21.8 GB
Files
11
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-06-13 16:05

Files by quantization

Auxiliary files 11 files 21.8 GB
model.safetensors 21.8 GB db0e4c15 download
tokenizer.json 19.1 MB 87a7830d download
config.json 22.6 KB cf307b72 download
tokenizer_config.json 14.8 KB 4a7541d5 download
chat_template.jinja 7.58 KB a8755d82 download
README.md 3.86 KB aee776d1 download
.gitattributes 1.66 KB 1901a10f download
processor_config.json 1.16 KB 33818c7f download
preprocessor_config.json 337 B e501264c download
recipe.yaml 280 B 17ba243b download
generation_config.json 213 B 79c8cce3 download

README current version from Hugging Face


license: apache-2.0
base_model: HauhauCS/Qwen3.6-35B-A3B-Uncensored-HauhauCS-Aggressive
base_model_relation: quantized
tags:

  • qwen3.6
  • nvfp4
  • compressed-tensors
  • quantized
  • vllm
  • moe
  • multimodal
  • image-text-to-text
  • blackwell
  • rtx-5090
  • sm120
  • uncensored
    library_name: transformers
    pipeline_tag: image-text-to-text

Qwen3.6-35B-A3B-Uncensored-HauhauCS-Aggressive-NVFP4

NVFP4 quantized version of HauhauCS/Qwen3.6-35B-A3B-Uncensored-HauhauCS-Aggressive.

Conservative profile: linear_attn (30 DeltaNet/Mamba layers) and MTP kept in bf16 for best quality. Follows AEON-7/RedHatAI approach.

Spec Value
Base model HauhauCS/Qwen3.6-35B-A3B-Uncensored-HauhauCS-Aggressive (Q8_K_P GGUF)
Original model Qwen/Qwen3.6-35B-A3B
Architecture Qwen3.5 MoE — 35B total, 3B active, 256 experts (8 routed + 1 shared)
Quantization NVFP4 W4A4 (conservative: linear_attn + MTP in bf16)
Format compressed-tensors (native vLLM support)
Size ~22 GB
Max context (text-only) 131K+ on RTX 5090
Requires NVIDIA Blackwell GPU (SM 120)

Quantization Recipe

recipe = QuantizationModifier(
    targets="Linear", scheme="NVFP4",
    ignore=["lm_head", "re:.*visual.*", "re:.*mlp.gate$",
            "re:.*mlp.shared_expert_gate$", "re:.*linear_attn.*", "re:^mtp.*"],
)
oneshot(model=model, dataset=ds, recipe=recipe,
        max_seq_length=1024, num_calibration_samples=128,
        moe_calibrate_all_experts=True, pipeline="basic")
  • Calibration: HuggingFaceH4/ultrachat_200k, 128 samples × 1024 tokens
  • MTP tensors copied from Qwen/Qwen3.6-35B-A3B (not present in GGUF)

Deployment (vLLM)

Vision + text smoke-tested on RTX 5090

This repository has been smoke-tested locally on an RTX 5090 with vllm/vllm-openai:v0.21.0-cu130-local, compressed-tensors, NVFP4 Marlin GEMM, FP8 KV cache, and a real image chat.completions request.

VLLM_USE_FLASHINFER_MOE_FP4=0 \
VLLM_NVFP4_GEMM_BACKEND=marlin \
vllm serve ./Qwen3.6-35B-A3B-Uncensored-HauhauCS-Aggressive-NVFP4 \
  --served-model-name qwen36-35b-a3b-hauhaucs-nvfp4 \
  --quantization compressed-tensors \
  --kv-cache-dtype fp8 \
  --gpu-memory-utilization 0.90 \
  --max-model-len 4096 \
  --max-num-seqs 1 \
  --max-num-batched-tokens 1024 \
  --trust-remote-code

For short non-thinking answers, pass chat_template_kwargs at the top level of the OpenAI-compatible request:

{
  "chat_template_kwargs": {"enable_thinking": false}
}

Text-only long context

VLLM_USE_FLASHINFER_MOE_FP4=0 \
VLLM_NVFP4_GEMM_BACKEND=marlin \
vllm serve ./Qwen3.6-35B-A3B-Uncensored-HauhauCS-Aggressive-NVFP4 \
  --quantization compressed-tensors \
  --kv-cache-dtype fp8 \
  --gpu-memory-utilization 0.95 \
  --max-model-len 100000 \
  --max-num-seqs 1 \
  --reasoning-parser qwen3 \
  --language-model-only \
  --trust-remote-code

Pipeline

Converted using li-yifei/gguf-to-nvfp4:

Q8_K_P GGUF → step1_convert_qwen36_moe.py → HF bf16 → step2_quantize_qwen36_moe.py → NVFP4

Also See

Acknowledgments

  • HauhauCS for the uncensored GGUF source
  • Qwen for the base model and MTP weights
  • AEON-7 and RedHatAI for conservative quantization approach reference

README history 5 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-06-13Fix multimodal vLLM metadata and vision weightsc9eb0a93.9 KB
    Loading...
  2. 2026-04-25Update README for root-only model layout67a56fa3.7 KB
    Loading...
  3. 2026-04-25Expose conservative NVFP4 weights at repo root86c65cb4.7 KB
    Loading...
  4. 2026-04-24Fix base_model relation to quantized51f24f82.8 KB
    Loading...
  5. 2026-04-21Upload folder using huggingface_hub431ea633.8 KB
    Loading...

Discussions 1 thread

  1. 2026-06-12Visual weights missingopen2 💬#1
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration