← back to catalog · registered 2026-08-22 13:56

sakamakismile/Huihui-Qwen3.5-9B-abliterated-NVFP4

sakamakismile Qwen 5.7B multimodal second-order
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/sakamakismile%2FHuihui-Qwen3.5-9B-abliterated-NVFP4"
Response includes
  • classification m1
  • files 20
  • benchmarks 11 entries
  • hub_downloads_all_time 1,556
  • author_summary 34 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M1
Primary method

Direct removal

No other method signals detected in this model.
Confidence
MEDIUM
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=0 (base model)
  • no specific method indicators - defaulting to M1 (most common)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
MEDIUM
Why we say so
primary_method=M1; difference-of-means is the reference extraction for M1/M3 (Arditi 2024)
Downloads · lifetime
2K
63 last 30d - cooling
Likes
1
Model age
6mo ago
created 2026-03-27
Downloads over time
Now1.6K→from9↑17,478%
05801.2K1.7K9 on Mar 251.6K on Oct 111.6K on Oct 10MarAprMayJunJulAugSepOct
Mar 25 → Oct 11 · 68 snapshots · spans 200 days

Benchmarks

Benchmark Score Source
Entertainment 0.9 UGI
Hazardous 1.8 UGI
Natural Intelligence 14.88 UGI
Political lean -6.0% UGI
Sensitive-Info 11.44 UGI
SocPol 0.9 UGI
UGI 37.63 UGI
Willingness (10) 9 UGI
W10-Adherence 9 UGI
W10-Direct 9 UGI
Writing 29.12 UGI

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
apache-2.0
Tags
transformers safetensors qwen3_5 image-text-to-text abliterated uncensored nvfp4 research multimodal conversational base_model:Qwen/Qwen3.5-9B base_model:quantized:Qwen/Qwen3.5-9B

Related

Total size
30.3 GB
Files
20
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-03-27 16:12

Files by quantization

Auxiliary files 20 files 30.4 GB
model.safetensors 12.4 GB d4d51842 download
model.safetensors-00003-of-00004.safetensors 5.00 GB 582ddb25 download
model.safetensors-00002-of-00004.safetensors 4.97 GB fd66bab5 download
model.safetensors-00001-of-00004.safetensors 4.91 GB 919d422e download
model.safetensors-00004-of-00004.safetensors 3.10 GB f95d920c download
tokenizer.json 12.2 MB 5f9e4d49 download
vocab.json 6.41 MB 0aa0ce06 download
merges.txt 3.20 MB a494e019 download
model.safetensors.index.json 66.6 KB 215021de download
tokenizer_config.json 16.3 KB eda48d3e download
config.json 14.7 KB 078936be download
LICENSE 11.3 KB f938136e download
chat_template.jinja 7.57 KB a585dec8 download
README.md 5.87 KB 758f48b7 download
.gitattributes 1.53 KB 52373fe2 download
hybrid_mm_metadata.json 1.15 KB fb5c0e9e download
recipe.yaml 1.12 KB d3e0cf29 download
nvfp4_text_gateup_metadata.json 911 B 079262e1 download
preprocessor_config.json 390 B 2ea84a43 download
video_preprocessor_config.json 385 B 3ba673a5 download

README current version from Hugging Face


library_name: transformers
license: apache-2.0
license_link: https://huggingface.co/Qwen/Qwen3.5-9B/blob/main/LICENSE
pipeline_tag: image-text-to-text
base_model:

  • huihui-ai/Huihui-Qwen3.5-9B-abliterated
  • Qwen/Qwen3.5-9B
    tags:
  • abliterated
  • uncensored
  • nvfp4
  • research
  • multimodal

Huihui-Qwen3.5-9B-abliterated-NVFP4

This repository provides an NVFP4-format derivative based on huihui-ai/Huihui-Qwen3.5-9B-abliterated, which itself is derived from Qwen/Qwen3.5-9B.

The main purpose of this release is research on inference depth, runtime behavior, and deployment characteristics of NVFP4 models. It is intended for study, benchmarking, and controlled experimentation rather than general-purpose safe deployment.

This model card intentionally preserves the provenance of the original Huihui release. The original fine-tuning characteristics, output tendencies, and risk profile should be assumed to carry over unless you independently verify otherwise in your own environment.

Provenance

  • Original base model: Qwen/Qwen3.5-9B
  • Fine-tuned source model: huihui-ai/Huihui-Qwen3.5-9B-abliterated
  • Quantization scope: text-side mlp.gate_proj and mlp.up_proj layers converted to NVFP4 with calibration
  • Packaging strategy: original multimodal wrapper and processor files retained, with the language model weights routed to the quantized checkpoint
  • Visual stack and non-targeted weights remain sourced from the original 9B release

Intended Use

  • Research on inference depth and reasoning behavior
  • Local benchmarking and runtime tuning
  • Controlled experiments on NVFP4 deployment
  • Long-context and VRAM-budget studies on RTX PRO 6000 Blackwell class GPUs

Responsibility Notice

Use of this model is entirely at your own risk.

  • You are solely responsible for how you use the model and how you handle its outputs.
  • The model may produce unsafe, controversial, incorrect, or otherwise unsuitable content.
  • This repository is published for research purposes and does not provide safety guarantees.
  • Do not assume fitness for production, commercial, legal, medical, educational, or public-facing use without your own review and safeguards.

Notes

The fine-tuned dataset is a type of dataset that operates in a non-think mode, and it may actually make thinking simpler.

This 9B release keeps the original multimodal configuration and processor files, but only the text-side language model gate_proj and up_proj layers were processed into NVFP4. The image stack was not quantized.

down_proj, attention layers, and the visual stack remain unquantized in this package.

After saving the model, some weight naming differs from the original export layout, and the packaged checkpoint does not preserve the original MTP weights in the same form. Treat this repository as an inference-focused research artifact rather than a drop-in archival mirror of the source release.

Reference Runtime Results

The following values are provided as reference measurements from local experiments on an NVIDIA RTX PRO 6000 Blackwell Workstation Edition GPU. They should be treated as environment-specific guidance rather than strict guarantees.

BF16 vs NVFP4 at 8K context

Format Idle VRAM Peak VRAM TTFT tok/s
BF16 57.2 GB 63.1 GB 44.84 ms 102.06
NVFP4 57.4 GB 63.2 GB 50.87 ms 121.01

The initial 8K comparison showed that model-load memory dropped, but total end-to-end VRAM stayed close because multimodal wrapper costs, runtime reservation, and KV cache overhead dominated at short context lengths.

NVFP4 long-context budget sweep

KV budget Idle VRAM Peak VRAM Notes
10G 25.3 GB 31.2 GB Stable at 32K, 64K, 128K, 256K runtime settings
8G 23.3 GB 29.1 GB Stable at 32K, 64K, 128K, 256K runtime settings
6G 21.2 GB 27.1 GB Stable at 32K, 64K, 128K, 256K runtime settings

Actual long-prompt probes were also run at roughly 32K, 64K, and 128K token input lengths. All runs completed, and 64K plus 128K retrieval checks were stable across the tested budgets.

Usage Warnings

  • Risk of Sensitive or Controversial Outputs: This model's safety filtering has been significantly reduced, potentially generating sensitive, controversial, or inappropriate content. Users should exercise caution and rigorously review generated outputs.
  • Not Suitable for All Audiences: Due to limited content filtering, the model's outputs may be inappropriate for public settings, underage users, or applications requiring high security.
  • Legal and Ethical Responsibilities: Users must ensure their usage complies with local laws and ethical standards. Generated content may carry legal or ethical risks, and users are solely responsible for any consequences.
  • Research and Experimental Use: It is recommended to use this model for research, testing, or controlled environments, avoiding direct use in production or public-facing commercial applications.
  • Monitoring and Review Recommendations: Users are strongly advised to monitor model outputs in real time and conduct manual reviews when necessary to prevent the dissemination of inappropriate content.
  • No Default Safety Guarantees: Unlike standard models, this model has not undergone rigorous safety optimization. The original fine-tuned source and this NVFP4 derivative should both be treated as research artifacts, and the publisher bears no responsibility for any consequences arising from their use.

Donation

Your donation helps us continue our further development and improvement, a cup of coffee can do it.
  • bitcoin:
bc1qqnkhuchxw0zqjh2ku3lu4hq45hc6gy84uk70ge
  • Support our work on Ko-fi!

README history 2 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-03-27Add files using upload-large-folder toolbef9cb05.9 KB
    Loading...
  2. 2026-03-27initial commit877820c28 B
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration