← back to catalog · registered 2026-08-22 13:56

zswll2/Qwen3.5-9B-Claude-4.6-HighIQ-THINKING-HERETIC-UNCENSORED-fp8-sglang

zswll2 Qwen 7.9B multimodal second-order
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/zswll2%2FQwen3.5-9B-Claude-4.6-HighIQ-THINKING-HERETIC-UNCENSORED-fp8-sglang"
Response includes
  • classification m3
  • files 14
  • hub_downloads_all_time 696
  • author_summary 2 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M3
Primary method

Layer-wise ablation

Applied on top of direct removal inherited from the base model.
Confidence
HIGH
Inherited from base model
Why this label 2 signals
Producer identity confirmed by naming conventions, tags or the model card. This label is very unlikely to change.
  • 'heretic' in model name (Heretic-produced)
  • Heretic uses layer-wise optimization (M3) with underlying direction removal (M1)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
HIGH
Why we say so
name contains 'heretic'; Heretic default extraction is difference-of-means (Arditi 2024)
Downloads · lifetime
696
260 last 30d - stable
Likes
1
Model age
7mo ago
created 2026-03-08
Downloads over time
Now752→from84↑795%
5130756381984 on Mar 11752 on Oct 11MarAprMayJunJulAugSepOct
Mar 11 → Oct 11 · 70 snapshots · spans 214 days

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
apache-2.0
Tags
safetensors qwen3_5 fp8 quantized sglang qwen3.5 multimodal base_model:zswll2/Qwen3.5-9B-Claude-4.6-HighIQ-THINKING-HERETIC-UNCENSORED base_model:quantized:zswll2/Qwen3.5-9B-Claude-4.6-HighIQ-THINKING-HERETIC-UNCENSORED license:apache-2.0 region:us

Related

Total size
9.29 GB
Files
14
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-03-08 10:07

Files by quantization

Auxiliary files 14 files 9.32 GB
model.safetensors 9.29 GB 6e699205 download
tokenizer.json 19.1 MB 87a7830d download
vocab.json 6.41 MB 0aa0ce06 download
tokenizer_config.json 16.3 KB eda48d3e download
config.json 11.9 KB 7986857f download
README.md 11.4 KB 392e605b download
chat_template.jinja 7.57 KB a585dec8 download
config_source.json 2.76 KB 4130edd0 download
config.json.bak 2.18 KB 7f9409e0 download
.gitattributes 1.53 KB 52373fe2 download
processor_config.json 1.27 KB 7ad6acdf download
preprocessor_config.json 390 B 2ea84a43 download
video_preprocessor_config.json 385 B 3ba673a5 download
generation_config.json 120 B 19800364 download

README current version from Hugging Face


license: apache-2.0
base_model: zswll2/Qwen3.5-9B-Claude-4.6-HighIQ-THINKING-HERETIC-UNCENSORED
tags:

  • fp8
  • quantized
  • sglang
  • qwen3.5
  • multimodal

Qwen3.5-9B FP8 Model (SGLang Compatible)

中文版

Overview

This is an FP8 quantized version of Qwen3.5-9B-Claude-4.6-HighIQ-THINKING-HERETIC-UNCENSORED, optimized for SGLang inference.

Property Value
Base Model Qwen3.5-9B-Claude-4.6-HighIQ-THINKING-HERETIC-UNCENSORED
Quantization FP8 (FineGrainedFP8)
Original Size ~18 GB
Quantized Size 9.4 GB
Compression Ratio ~48%
Architecture Qwen3_5ForConditionalGeneration
Compatibility SGLang 0.5.9+ ✅

Key Features

  • FP8 Quantization: Block-wise FP8 with [128, 128] block size
  • SGLang Optimized: weight_scale_inv converted to bfloat16 for SGLang compatibility
  • Multimodal Support: Vision-Language model with 27-layer visual encoder
  • Linear Attention: Hybrid architecture with linear and full attention layers

Model Structure

Qwen3.5-9B-FP8-SGLang/
├── config.json              # Model configuration with FP8 quantization params
├── model.safetensors        # FP8 weights (9.4 GB)
├── tokenizer.json
├── tokenizer_config.json
├── preprocessor_config.json
├── generation_config.json
├── processor_config.json
├── video_preprocessor_config.json
├── chat_template.jinja
└── vocab.json

Quantization Config

{
  "quantization_config": {
    "quant_method": "fp8",
    "activation_scheme": "dynamic",
    "weight_per_tensor": false,
    "act_per_tensor": false,
    "weight_block_size": [128, 128]
  }
}

Requirements

Hardware

  • GPU: NVIDIA GPU with FP8 support (Ada Lovelace or newer recommended)
  • VRAM: 10GB+ for inference
  • CUDA: 12.1+

Software

Core dependencies:

Package Version
sglang 0.5.9+
torch 2.9.1+
transformers 4.57.1+
flashinfer-python 0.6.4+
triton 3.5.1+
Full Dependencies List
sglang==0.5.9
torch==2.9.1
torchvision==0.24.1
torchaudio==2.9.1
transformers==4.57.1
tokenizers==0.22.2
flashinfer-python==0.6.4
flashinfer-cubin==0.6.4
triton==3.5.1
torchao==0.9.0
cuda-python==12.9.0
cuda-bindings==12.9.5
nvidia-cublas-cu12==12.9.1.4
nvidia-cudnn-cu12==9.16.0.29
nvidia-cuda-runtime-cu12==12.8.90
nvidia-cuda-nvrtc-cu12==12.8.93
nvidia-nccl-cu12==2.27.5
sgl-kernel==0.3.21
outlines==0.1.11
outlines_core==0.1.26
xgrammar==0.1.27
llguidance==0.7.30
compressed-tensors==0.14.0
safetensors==0.7.0
huggingface_hub==0.36.2
accelerate
pillow==11.3.0

Installation

# Create conda environment
conda create -n sglang-fp8 python=3.11 -y
conda activate sglang-fp8

# Install SGLang with CUDA 12.x
pip install sglang[all] --find-links https://flashinfer.ai/whl/cu121/torch2.4/

# Or install from source for latest features
pip install sglang[all]>=0.5.9

Usage

SGLang Server

python -m sglang.launch_server \
    --model-path ./Qwen3.5-9B-Claude-4.6-HighIQ-THINKING-HERETIC-UNCENSORED-fp8-sglang \
    --port 8000 \
    --trust-remote-code

SGLang Python API

import sglang as sgl

# Initialize engine
llm = sgl.Engine(
    model_path="./Qwen3.5-9B-Claude-4.6-HighIQ-THINKING-HERETIC-UNCENSORED-fp8-sglang",
    trust_remote_code=True,
)

# Generate response
prompt = "Hello! Please introduce yourself briefly."
outputs = llm.generate([prompt], sampling_params={"max_new_tokens": 100, "temperature": 0.0})
print(outputs[0]["text"])

OpenAI-Compatible API

After starting the server:

from openai import OpenAI

client = OpenAI(base_url="http://localhost:8000/v1", api_key="none")

response = client.chat.completions.create(
    model="default",
    messages=[{"role": "user", "content": "Hello!"}],
    max_tokens=100,
)
print(response.choices[0].message.content)

Sampling Parameters

Thinking Mode (General Tasks)

sampling_params = {
    "temperature": 1.0,
    "top_p": 0.95,
    "top_k": 20,
    "min_p": 0.0,
    "presence_penalty": 1.5,
    "repetition_penalty": 1.0,
}

Thinking Mode (Coding Tasks)

sampling_params = {
    "temperature": 0.6,
    "top_p": 0.95,
    "top_k": 20,
    "min_p": 0.0,
    "presence_penalty": 0.0,
    "repetition_penalty": 1.0,
}

Instruct Mode (Non-Thinking)

sampling_params = {
    "temperature": 0.7,
    "top_p": 0.8,
    "top_k": 20,
    "min_p": 0.0,
    "presence_penalty": 1.5,
    "repetition_penalty": 1.0,
}
# Add: extra_body={"chat_template_kwargs": {"enable_thinking": False}}

Technical Details

Conversion from Transformers 5.x FP8

This model was converted from Transformers 5.x native FP8 format. The key difference:

Attribute Transformers 5.x SGLang Compatible
weight_scale_inv dtype float32 bfloat16

Conversion script convert_fp8_for_sglang.py:

import torch
import safetensors
from safetensors.torch import save_file

# Convert weight_scale_inv: float32 -> bfloat16
for key in sf.keys():
    tensor = sf.get_tensor(key)
    if 'scale_inv' in key and tensor.dtype == torch.float32:
        tensors[key] = tensor.to(torch.bfloat16)
    else:
        tensors[key] = tensor

Troubleshooting

CUDA Out of Memory

# Reduce max context length
--max-model-len 8192

# Reduce memory fraction
--mem-fraction-static 0.8

Import Errors

Ensure CUDA version matches:

pip install sglang[all] --find-links https://flashinfer.ai/whl/cu121/torch2.4/

License

This model inherits the license from the base model.

References


Generated: 2026-03-08
Quantization: Transformers 5.x FineGrainedFP8Config + SGLang conversion


中文版

概述

这是 Qwen3.5-9B-Claude-4.6-HighIQ-THINKING-HERETIC-UNCENSORED 的 FP8 量化版本,专为 SGLang 推理优化。

属性 值
基础模型 Qwen3.5-9B-Claude-4.6-HighIQ-THINKING-HERETIC-UNCENSORED
量化格式 FP8 (FineGrainedFP8)
原始大小 ~18 GB
量化后大小 9.4 GB
压缩比 ~48%
模型架构 Qwen3_5ForConditionalGeneration
兼容性 SGLang 0.5.9+ ✅

主要特性

  • FP8 量化: 块级 FP8,块大小 [128, 128]
  • SGLang 优化: weight_scale_inv 已转换为 bfloat16 以兼容 SGLang
  • 多模态支持: 视觉-语言模型,27层视觉编码器
  • 线性注意力: 混合架构,包含线性和全注意力层

模型结构

Qwen3.5-9B-FP8-SGLang/
├── config.json              # 模型配置(含 FP8 量化参数)
├── model.safetensors        # FP8 权重 (9.4 GB)
├── tokenizer.json
├── tokenizer_config.json
├── preprocessor_config.json
├── generation_config.json
├── processor_config.json
├── video_preprocessor_config.json
├── chat_template.jinja
└── vocab.json

量化配置

{
  "quantization_config": {
    "quant_method": "fp8",
    "activation_scheme": "dynamic",
    "weight_per_tensor": false,
    "act_per_tensor": false,
    "weight_block_size": [128, 128]
  }
}

环境要求

硬件

  • GPU: 支持 FP8 的 NVIDIA GPU(推荐 Ada Lovelace 或更新)
  • 显存: 10GB+ 用于推理
  • CUDA: 12.1+

软件

核心依赖:

包名 版本
sglang 0.5.9+
torch 2.9.1+
transformers 4.57.1+
flashinfer-python 0.6.4+
triton 3.5.1+
完整依赖列表
sglang==0.5.9
torch==2.9.1
torchvision==0.24.1
torchaudio==2.9.1
transformers==4.57.1
tokenizers==0.22.2
flashinfer-python==0.6.4
flashinfer-cubin==0.6.4
triton==3.5.1
torchao==0.9.0
cuda-python==12.9.0
cuda-bindings==12.9.5
nvidia-cublas-cu12==12.9.1.4
nvidia-cudnn-cu12==9.16.0.29
nvidia-cuda-runtime-cu12==12.8.90
nvidia-cuda-nvrtc-cu12==12.8.93
nvidia-nccl-cu12==2.27.5
sgl-kernel==0.3.21
outlines==0.1.11
outlines_core==0.1.26
xgrammar==0.1.27
llguidance==0.7.30
compressed-tensors==0.14.0
safetensors==0.7.0
huggingface_hub==0.36.2
accelerate
pillow==11.3.0

安装

# 创建 conda 环境
conda create -n sglang-fp8 python=3.11 -y
conda activate sglang-fp8

# 安装 SGLang (CUDA 12.x)
pip install sglang[all] --find-links https://flashinfer.ai/whl/cu121/torch2.4/

# 或安装最新版本
pip install sglang[all]>=0.5.9

使用方法

SGLang 服务器

python -m sglang.launch_server \
    --model-path ./Qwen3.5-9B-Claude-4.6-HighIQ-THINKING-HERETIC-UNCENSORED-fp8-sglang \
    --port 8000 \
    --trust-remote-code

SGLang Python API

import sglang as sgl

# 初始化引擎
llm = sgl.Engine(
    model_path="./Qwen3.5-9B-Claude-4.6-HighIQ-THINKING-HERETIC-UNCENSORED-fp8-sglang",
    trust_remote_code=True,
)

# 生成响应
prompt = "你好!请简要介绍一下你自己。"
outputs = llm.generate([prompt], sampling_params={"max_new_tokens": 100, "temperature": 0.0})
print(outputs[0]["text"])

OpenAI 兼容 API

启动服务器后:

from openai import OpenAI

client = OpenAI(base_url="http://localhost:8000/v1", api_key="none")

response = client.chat.completions.create(
    model="default",
    messages=[{"role": "user", "content": "你好!"}],
    max_tokens=100,
)
print(response.choices[0].message.content)

采样参数

思考模式(通用任务)

sampling_params = {
    "temperature": 1.0,
    "top_p": 0.95,
    "top_k": 20,
    "min_p": 0.0,
    "presence_penalty": 1.5,
    "repetition_penalty": 1.0,
}

思考模式(编程任务)

sampling_params = {
    "temperature": 0.6,
    "top_p": 0.95,
    "top_k": 20,
    "min_p": 0.0,
    "presence_penalty": 0.0,
    "repetition_penalty": 1.0,
}

指令模式(非思考)

sampling_params = {
    "temperature": 0.7,
    "top_p": 0.8,
    "top_k": 20,
    "min_p": 0.0,
    "presence_penalty": 1.5,
    "repetition_penalty": 1.0,
}
# 添加: extra_body={"chat_template_kwargs": {"enable_thinking": False}}

技术细节

从 Transformers 5.x FP8 转换

此模型从 Transformers 5.x 原生 FP8 格式转换而来。关键差异:

属性 Transformers 5.x SGLang 兼容
weight_scale_inv dtype float32 bfloat16

转换脚本 convert_fp8_for_sglang.py:

import torch
import safetensors
from safetensors.torch import save_file

# 转换 weight_scale_inv: float32 -> bfloat16
for key in sf.keys():
    tensor = sf.get_tensor(key)
    if 'scale_inv' in key and tensor.dtype == torch.float32:
        tensors[key] = tensor.to(torch.bfloat16)
    else:
        tensors[key] = tensor

故障排除

CUDA 内存不足

# 减少最大上下文长度
--max-model-len 8192

# 减少内存占用比例
--mem-fraction-static 0.8

导入错误

确保 CUDA 版本匹配:

pip install sglang[all] --find-links https://flashinfer.ai/whl/cu121/torch2.4/

许可证

此模型继承基础模型的许可证。

参考资料


生成时间: 2026-03-08
量化工具: Transformers 5.x FineGrainedFP8Config + SGLang 格式转换

README history 1 version

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-03-08Upload README.md with huggingface_hub86c16cd11.4 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration