← back to catalog · registered 2026-08-22 13:56

prithivMLmods/Qwen3.5-9B-abliterated-v2-MAX-FP8

prithivMLmods Qwen 5.3B multimodal second-order
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/prithivMLmods%2FQwen3.5-9B-abliterated-v2-MAX-FP8"
Response includes
  • classification m1
  • files 10
  • hub_downloads_all_time 1,977
  • author_summary 98 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M1
Primary method

Direct removal

No other method signals detected in this model.
Confidence
MEDIUM
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=0 (base model)
  • no specific method indicators - defaulting to M1 (most common)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
MEDIUM
Why we say so
primary_method=M1; difference-of-means is the reference extraction for M1/M3 (Arditi 2024)
Downloads · lifetime
2K
112 last 30d - cooling
Likes
1
Model age
6mo ago
created 2026-03-29
Downloads over time
Now2K→from0↑0%
07371.5K2.2K0 on Apr 12K on Oct 112K on Oct 10AprMayJunJulAugSepOct
Apr 1 → Oct 11 · 67 snapshots · spans 193 days

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Variants by this author 3 formats · 1K downloads combined

The same weights this author released in different packaging. Pick the format that matches your runtime.

Metadata

License
apache-2.0
Languages
en
Tags
transformers safetensors qwen3_5 image-text-to-text text-generation-inference FP8_DYNAMIC uncensored abliterated unfiltered unredacted refusal-ablated vllm

Related

Total size
12.6 GB
Files
10
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-03-31 05:28

Files by quantization

Auxiliary files 10 files 12.6 GB
model.safetensors 12.6 GB d9b17179 download
tokenizer.json 19.1 MB 87a7830d download
config.json 16.0 KB 70265931 download
chat_template.jinja 7.57 KB a585dec8 download
README.md 4.50 KB 8f6be65c download
.gitattributes 1.53 KB 52373fe2 download
processor_config.json 1.27 KB 8af8110f download
tokenizer_config.json 1.11 KB 541f6c47 download
recipe.yaml 215 B f4f53e54 download
generation_config.json 115 B 11dd07a6 download

README current version from Hugging Face


license: apache-2.0
tags:

  • text-generation-inference
  • FP8_DYNAMIC
  • uncensored
  • abliterated
  • unfiltered
  • unredacted
  • refusal-ablated
  • vllm
  • pytorch
  • bf16
  • max
  • alignment-modified
  • reasoning
  • fp8
  • llm-compressor
    language:
  • en
    base_model:
  • prithivMLmods/Qwen3.5-9B-abliterated-v2-MAX
    pipeline_tag: image-text-to-text
    library_name: transformers

1

Qwen3.5-9B-abliterated-v2-MAX-FP8

Qwen3.5-9B-abliterated-v2-MAX-FP8 is an FP8-compressed variant built on top of prithivMLmods/Qwen3.5-9B-abliterated-v2-MAX. This version leverages BF16 · FP8 (F8_E4M3) precision formats to significantly reduce memory footprint and improve inference efficiency. It maintains the original character of the base model while introducing a more optimized abliteration rate, combining refined refusal direction analysis with enhanced training strategies to further minimize internal refusal behaviors while preserving strong reasoning and instruction-following capabilities. The result is a capable 9B parameter language model optimized for detailed responses and improved instruction adherence, now with enhanced deployment efficiency.

[!IMPORTANT]
This model is intended strictly for research and learning purposes. Due to reduced internal refusal mechanisms, it may generate sensitive or unrestricted content. Users assume full responsibility for how the model is used. The authors and hosting platform disclaim any liability for generated outputs.

Key Highlights

  • FP8 Compression (F8_E4M3): Reduces VRAM usage and improves inference throughput while maintaining strong output quality.
  • BF16 · FP8 Hybrid Precision: Balances numerical stability and performance across model layers.
  • Optimized Abliteration Rate (v2): Improved suppression of refusal directions with better balance between openness and coherence.
  • Advanced Refusal Direction Analysis: Identifies and mitigates refusal-related activations within the model’s latent space.
  • Abliterated v2 Training Strategy: Further reduces refusal behaviors while maintaining response quality and consistency.
  • 9B Parameter Architecture: Based on Qwen3.5-9B, offering strong reasoning with efficient deployment.
  • Improved Instruction Adherence: Better handling of complex and nuanced prompts with minimal unnecessary refusals.
  • Efficient Deployment: Ideal for local inference and research workflows with reduced hardware requirements.

Quick Start with Transformers

pip install transformers==5.4.0
# or
pip install git+https://github.com/huggingface/transformers.git
from transformers import Qwen3_5ForConditionalGeneration, AutoProcessor
import torch

model = Qwen3_5ForConditionalGeneration.from_pretrained(
    "prithivMLmods/Qwen3.5-9B-abliterated-v2-MAX-FP8",
    torch_dtype="auto",
    device_map="auto"
)

processor = AutoProcessor.from_pretrained(
    "prithivMLmods/Qwen3.5-9B-abliterated-v2-MAX-FP8"
)

messages = [
    {
        "role": "user",
        "content": [
            {"type": "text", "text": "Explain how transformer models work in simple terms."}
        ],
    }
]

text = processor.apply_chat_template(
    messages, tokenize=False, add_generation_prompt=True
)

inputs = processor(
    text=[text],
    padding=True,
    return_tensors="pt"
).to("cuda")

generated_ids = model.generate(**inputs, max_new_tokens=256)

generated_ids_trimmed = [
    out_ids[len(in_ids):] for in_ids, out_ids in zip(inputs.input_ids, generated_ids)
]

output_text = processor.batch_decode(
    generated_ids_trimmed,
    skip_special_tokens=True,
    clean_up_tokenization_spaces=False
)

print(output_text)

Intended Use

  • Alignment & Refusal Research: Studying abliteration effects under FP8 compression.
  • Red-Teaming Experiments: Evaluating robustness across adversarial prompts.
  • Efficient Local Deployment: Running 9B-class models with reduced VRAM usage.
  • Research Prototyping: Exploring trade-offs between compression, alignment, and reasoning.

Limitations & Risks

Important Note: This model intentionally minimizes built-in safety refusals.

  • High Risk of Sensitive Outputs: May generate unrestricted or controversial responses.
  • User Responsibility: Must be used in a safe, ethical, and lawful manner.
  • Precision Trade-offs: FP8 may introduce minor instability in edge cases.
  • Abliteration Trade-offs: Increased openness may affect safety alignment or consistency.

README history 4 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-03-31Update README.mdbf4f7f84.5 KB
    Loading...
  2. 2026-03-30Update README.md408424d4.5 KB
    Loading...
  3. 2026-03-30Update README.mdc747a5d4.4 KB
    Loading...
  4. 2026-03-29initial commit357a3a328 B
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration