← back to catalog · registered 2026-08-22 13:56

CezarJedi/Qwen3.5-9B-Abliterated-NVFP4-FP8KV

CezarJedi Qwen 3.5B second-order
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/CezarJedi%2FQwen3.5-9B-Abliterated-NVFP4-FP8KV"
Response includes
  • classification m1
  • files 11
  • hub_downloads_all_time 301
  • author_summary 4 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M1
Primary method

Direct removal

No other method signals detected in this model.
Confidence
MEDIUM
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=0 (base model)
  • no specific method indicators - defaulting to M1 (most common)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
MEDIUM
Why we say so
primary_method=M1; difference-of-means is the reference extraction for M1/M3 (Arditi 2024)
Downloads · lifetime
301
151 last 30d - active
Likes
1
Model age
2mo ago
created 2026-08-07

Training datasets

Corpora the author lists in the model card. Datasets tracked in our /datasets catalog carry a category badge linking to the workflow stage. Others open on Hugging Face.

Downloads over time
Now360→from4↑8,900%
01322643964 on Aug 5360 on Oct 11AugSepOct
Aug 5 → Oct 11 · 50 snapshots · spans 67 days

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

Languages
en
Tags
transformers safetensors qwen3_5 image-text-to-text qwen qwen3 qwen3.5 large-language-model text-generation chat modelopt tensorrt-model-optimizer

Related

Total size
8.27 GB
Files
11
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-08-07 17:14

Files by quantization

Auxiliary files 11 files 8.29 GB
model.safetensors 8.27 GB 850176cd download
tokenizer.json 19.1 MB 06b95093 download
.quant_summary.txt 193 KB 825a5a28 download
config.json 8.91 KB 7a71d281 download
chat_template.jinja 7.57 KB a585dec8 download
hf_quant_config.json 5.02 KB b8c18f38 download
README.md 2.47 KB 99ba0e24 download
.gitattributes 1.53 KB 52373fe2 download
processor_config.json 1.16 KB 33818c7f download
tokenizer_config.json 1.09 KB 77053000 download
generation_config.json 115 B 86ce6bc5 download

README current version from Hugging Face


language:

  • en

library_name: transformers

pipeline_tag: text-generation

base_model:

  • nDimensional/Qwen3.5-9B-Uncensored-Safetensors

base_model_relation: quantized

datasets:

  • nvidia/Nemotron-Post-Training-Dataset-v2

tags:

  • qwen
  • qwen3
  • qwen3.5
  • large-language-model
  • text-generation
  • chat
  • modelopt
  • tensorrt-model-optimizer
  • nvfp4
  • fp4
  • fp8
  • kv-cache
  • post-training-quantization
  • ptq
  • vllm
  • transformers
  • abliterated
  • uncensored

Qwen3.5-9B-Uncensored-NVFP4-FP8KV

Post-training quantized version of nDimensional/Qwen3.5-9B-Uncensored-Safetensors using NVIDIA TensorRT Model Optimizer.

Quantization

  • Weights: NVFP4
  • Activations: NVFP4
  • KV Cache: FP8
  • Method: Post-Training Quantization (PTQ)
  • Recipe: general/ptq/nvfp4_default-kv_fp8_cast
  • Calibration Dataset: Nemotron Post-Training Dataset v2

This checkpoint is optimized for lower VRAM usage while maintaining high inference quality. It is intended for runtimes supporting NVIDIA ModelOpt quantization, such as vLLM.

Quantization Command

python3 examples/llm_ptq/hf_ptq.py \
    --pyt_ckpt_path /path/to/Qwen3.5-9B \
    --recipe general/ptq/nvfp4_default-kv_fp8_cast \
    --dataset nemotron-post-training-dataset-v2 \
    --batch_size 1 \
    --skip_generate \
    --verbose \
    --export_path /path/to/Qwen3.5-9B-NVFP4-FP8KV

Usage

Ask any AI, even this one :). I tested this checkpoint with vLLM and it worked well. It should also work with other runtimes that support NVIDIA ModelOpt checkpoints, such as Transformers, although I haven't tested those myself.

Credits

This repository only contains a post-training quantized checkpoint.

Credit for the original model, fine-tuning, abliteration, and Safetensors conversion goes to the respective upstream authors:

README history 2 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-08-07Upload folder using huggingface_hub2bcaa0a2.5 KB
    Loading...
  2. 2026-08-07Upload folder using huggingface_hub18cabe32 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration