← back to catalog · registered 2026-08-22 13:56

prithivMLmods/Qwen3.5-9B-abliterated-v2-MAX

prithivMLmods Qwen 9.4B GGUF multimodal 262K ctx
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/prithivMLmods%2FQwen3.5-9B-abliterated-v2-MAX"
Response includes
  • classification m8
  • files 12
  • benchmarks 11 entries
  • hub_downloads_all_time 12,067
  • author_summary 98 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M8
Primary method

Repackaging (quantization)

Applied on top of direct removal inherited from the base model.
Confidence
MEDIUM
Inherited from base model
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=1
  • assume M1 (base ablation) + M8 (GGUF quant) - default when producer unknown
Refusal direction extraction

No specific extraction method could be identified for this model. The producer either did not document it or used a proprietary pipeline.

What is a refusal direction? →
Downloads · lifetime
12K
1K last 30d - cooling
Likes
9
Descendants
4
in 4 direct forks
Model age
6mo ago
created 2026-03-28
Downloads over time
Now12.4K→from113↑10,914%
04.6K9.1K13.7K113 on Apr 112.4K on Oct 11AprMayJunJulAugSepOct
Apr 1 → Oct 11 · 69 snapshots · spans 193 days

Benchmarks

Portrait before abliteration
Benchmarks of the base model as it stood before the refusal-removal operation. Compare with the numbers above to see what the operation cost.
Benchmark Score Source
Entertainment 1.4 UGI
Hazardous 2.4 UGI
Natural Intelligence 17.62 UGI
Political lean -12.2% UGI
Sensitive-Info 14.65 UGI
SocPol 0.9 UGI
UGI 17.27 UGI
Willingness (10) 2.2 UGI
W10-Adherence 1.5 UGI
W10-Direct 3 UGI
Writing 33.52 UGI

Genealogy 4 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Variants by this author 3 formats · 1K downloads combined

The same weights this author released in different packaging. Pick the format that matches your runtime.

Metadata

License
apache-2.0
Languages
en
Tags
transformers safetensors gguf qwen3_5 image-text-to-text text-generation-inference uncensored abliterated unfiltered unredacted refusal-ablated vllm

Related

Total size
17.5 GB
Files
12
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-06-01 08:26

Files by quantization

Auxiliary files 12 files 17.5 GB
model-00001-of-00003.safetensors 5.93 GB 03b84941 download
model-00002-of-00003.safetensors 5.91 GB 2e7420be download
model-00003-of-00003.safetensors 5.68 GB 9eefff7b download
tokenizer.json 19.1 MB 87a7830d download
model.safetensors.index.json 85.0 KB 0bd31ce8 download
chat_template.jinja 7.57 KB a585dec8 download
README.md 5.10 KB 02b7f0b8 download
config.json 2.76 KB 3d0c9820 download
.gitattributes 2.19 KB fe594b4f download
processor_config.json 1.27 KB 8af8110f download
tokenizer_config.json 1.11 KB 541f6c47 download
generation_config.json 115 B 11dd07a6 download

README current version from Hugging Face


license: apache-2.0
tags:

  • text-generation-inference
  • uncensored
  • abliterated
  • unfiltered
  • unredacted
  • refusal-ablated
  • vllm
  • pytorch
  • bf16
  • max
  • alignment-modified
  • reasoning
  • v2-MAX
  • llama.cpp
    language:
  • en
    base_model:
  • Qwen/Qwen3.5-9B
    pipeline_tag: image-text-to-text
    library_name: transformers

1

Qwen3.5-9B-abliterated-v2-MAX

Qwen3.5-9B-abliterated-v2-MAX is an optimized release built on top of huihui-ai/Huihui-Qwen3.5-9B-abliterated. This version focuses on improved model sharding, packaging consistency, and compatibility with modern Transformers and inference stacks, while preserving the reasoning and instruction-following capabilities of the base model. The result is a highly capable 9B parameter language model designed for efficient deployment, stable inference, and research-oriented experimentation.

[!IMPORTANT]
This model is intended strictly for research and learning purposes only. Any outputs generated by this model are the sole responsibility of the user. The authors and hosting platform disclaim all liability for model-generated content. Users must ensure safe, ethical, and lawful usage.

Compression for the Model

Qwen3.5-9B-abliterated-v2-MAX

Format Description Link
GGUF Quantized GGUF format https://huggingface.co/prithivMLmods/Qwen3.5-9B-abliterated-v2-MAX/tree/main/GGUF
NVFP4 NVFP4 compressed model https://huggingface.co/prithivMLmods/Qwen3.5-9B-abliterated-v2-MAX-NVFP4
FP8 FP8 compressed model https://huggingface.co/prithivMLmods/Qwen3.5-9B-abliterated-v2-MAX-FP8

Base Model Signatures:

This model has been re-sharded and optimized for the latest Transformers version from the base model:
https://huggingface.co/huihui-ai/Huihui-Qwen3.5-9B-abliterated


Key Highlights

  • Optimized Packaging & Sharding
    Improved repository structure for smoother downloads, loading, and deployment across environments.

  • Stable Transformers Compatibility
    Updated layout for better compatibility with modern Transformers versions and inference pipelines.

  • 9B Parameter Architecture
    Built on Qwen3.5-9B, balancing efficiency and capability for local and research use.

  • Efficient Deployment Design
    Designed for lightweight inference, experimentation, and scalable integration.

  • Preserved Model Behavior
    No changes to weights or core architecture; performance remains consistent with the original base model lineage.

  • Improved Reliability in Loading
    Reduced friction in model initialization and multi-device inference setups.


Quick Start with Transformers

pip install transformers==5.4.0
# or
pip install git+https://github.com/huggingface/transformers.git
from transformers import Qwen3_5ForConditionalGeneration, AutoProcessor
import torch

model = Qwen3_5ForConditionalGeneration.from_pretrained(
    "prithivMLmods/Qwen3.5-9B-abliterated-v2-MAX",
    torch_dtype="auto",
    device_map="auto"
)

processor = AutoProcessor.from_pretrained(
    "prithivMLmods/Qwen3.5-9B-abliterated-v2-MAX"
)

messages = [
    {
        "role": "user",
        "content": [
            {"type": "text", "text": "Explain how transformer models work in simple terms."}
        ],
    }
]

text = processor.apply_chat_template(
    messages, tokenize=False, add_generation_prompt=True
)

inputs = processor(
    text=[text],
    padding=True,
    return_tensors="pt"
).to("cuda")

generated_ids = model.generate(**inputs, max_new_tokens=256)

generated_ids_trimmed = [
    out_ids[len(in_ids):] for in_ids, out_ids in zip(inputs.input_ids, generated_ids)
]

output_text = processor.batch_decode(
    generated_ids_trimmed,
    skip_special_tokens=True,
    clean_up_tokenization_spaces=False
)

print(output_text)

Intended Use

  • Multimodal and Language Research
    Studying behavior of compact 9B-scale transformer models under different inference settings.

  • Red-Teaming & Evaluation
    Testing robustness across adversarial prompts and edge-case inputs.

  • Efficient Local Deployment
    Running lightweight yet capable models on consumer GPUs or optimized cloud setups.

  • Research Prototyping
    Exploring model behavior, alignment, and inference optimization techniques.


Limitations & Risks

Important Note: This model inherits behavior from its base model with minimal modification.

  • Output Variability
    Responses may vary depending on sampling strategy and prompt formulation.

  • Resource Dependency
    While efficient, GPU acceleration is recommended for optimal performance.

  • No Architectural Changes
    Improvements are limited to packaging and compatibility, not core model capabilities.

  • General Model Limitations
    May still produce incorrect, incomplete, or inconsistent outputs in complex scenarios.

README history 8 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-06-01Update README.md41e9f615.1 KB
    Loading...
  2. 2026-06-01Update README.md6d4ecc14.8 KB
    Loading...
  3. 2026-03-31Update README.md38fcaa34.6 KB
    Loading...
  4. 2026-03-30Update README.md41a79994.5 KB
    Loading...
  5. 2026-03-30Update README.md3cfe5494.5 KB
    Loading...
  6. 2026-03-30Update README.mdee0ed344.5 KB
    Loading...
  7. 2026-03-30Update README.md5f7d24b3.9 KB
    Loading...
  8. 2026-03-28initial commit58fc52f28 B
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration