language:
- en
library_name: transformers
pipeline_tag: text-generation
base_model:
- nDimensional/Qwen3.5-9B-Uncensored-Safetensors
base_model_relation: quantized
datasets:
- nvidia/Nemotron-Post-Training-Dataset-v2
tags:
- qwen
- qwen3
- qwen3.5
- large-language-model
- text-generation
- chat
- modelopt
- tensorrt-model-optimizer
- nvfp4
- fp4
- fp8
- kv-cache
- post-training-quantization
- ptq
- vllm
- transformers
- abliterated
- uncensored
Qwen3.5-9B-Uncensored-NVFP4-FP8KV
Post-training quantized version of nDimensional/Qwen3.5-9B-Uncensored-Safetensors using NVIDIA TensorRT Model Optimizer.
Quantization
- Weights: NVFP4
- Activations: NVFP4
- KV Cache: FP8
- Method: Post-Training Quantization (PTQ)
- Recipe:
general/ptq/nvfp4_default-kv_fp8_cast - Calibration Dataset: Nemotron Post-Training Dataset v2
This checkpoint is optimized for lower VRAM usage while maintaining high inference quality. It is intended for runtimes supporting NVIDIA ModelOpt quantization, such as vLLM.
Quantization Command
python3 examples/llm_ptq/hf_ptq.py \
--pyt_ckpt_path /path/to/Qwen3.5-9B \
--recipe general/ptq/nvfp4_default-kv_fp8_cast \
--dataset nemotron-post-training-dataset-v2 \
--batch_size 1 \
--skip_generate \
--verbose \
--export_path /path/to/Qwen3.5-9B-NVFP4-FP8KV
Usage
Ask any AI, even this one :). I tested this checkpoint with vLLM and it worked well. It should also work with other runtimes that support NVIDIA ModelOpt checkpoints, such as Transformers, although I haven't tested those myself.
Credits
This repository only contains a post-training quantized checkpoint.
Credit for the original model, fine-tuning, abliteration, and Safetensors conversion goes to the respective upstream authors:
- Base model: https://huggingface.co/Qwen/Qwen3.5-9B
- Uncensored fine-tune: https://huggingface.co/CezarJedi/Qwen3.5-9B-Uncensored-HauhauCS-Aggressive
- Safetensors conversion: https://huggingface.co/nDimensional/Qwen3.5-9B-Uncensored-Safetensors
- Quantization performed using NVIDIA TensorRT Model Optimizer: https://github.com/NVIDIA/TensorRT-Model-Optimizer
- AxionML's Qwen3.5-9B-NVFP4 served as a helpful reference for the quantization recipe and calibration setup: https://huggingface.co/AxionML/Qwen3.5-9B-NVFP4