← back to catalog · registered 2026-08-22 13:56

ikarius/Qwen3-8B-Abliterated-FP8

ikarius Qwen 6.9B second-order
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/ikarius%2FQwen3-8B-Abliterated-FP8"
Response includes
  • classification m1
  • files 14
  • hub_downloads_all_time 2,586
  • author_summary 17 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M1
Primary method

Direct removal

No other method signals detected in this model.
Confidence
MEDIUM
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=0 (base model)
  • no specific method indicators - defaulting to M1 (most common)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
MEDIUM
Why we say so
primary_method=M1; difference-of-means is the reference extraction for M1/M3 (Arditi 2024)
Downloads · lifetime
3K
78 last 30d - cooling
Likes
1
Model age
9mo ago
created 2025-12-30
Downloads over time
Now2.6K→from13↑20,062%
09611.9K2.9K13 on Jan 142.6K on Oct 11JanMarMayJulSep
Jan 14 → Oct 11 · 78 snapshots · spans 270 days

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
apache-2.0
Tags
transformers safetensors qwen3 text-generation quantization fp8 qwen abliterated blackwell-optimized fine-grained conversational license:apache-2.0

Related

Total size
8.79 GB
Files
14
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2025-12-30 16:25

Files by quantization

Auxiliary files 14 files 8.80 GB
model-00001-of-00002.safetensors 4.61 GB 0b7f3d89 download
model-00002-of-00002.safetensors 4.18 GB b3e10b4e download
tokenizer.json 10.9 MB aeb13307 download
vocab.json 2.65 MB 4783fe10 download
merges.txt 1.59 MB 31349551 download
model.safetensors.index.json 54.6 KB a1ce1602 download
tokenizer_config.json 5.28 KB ddaf6980 download
chat_template.jinja 4.02 KB 699ff8df download
README.md 3.00 KB c2331747 download
config.json 1.68 KB e3e9c07b download
.gitattributes 1.53 KB 52373fe2 download
added_tokens.json 707 B b54f9135 download
special_tokens_map.json 613 B ac23c0aa download
generation_config.json 214 B 98e0755a download

README current version from Hugging Face


license: apache-2.0
base_model: huihui-ai/Qwen3-8B-Instruct-Abliterated
tags:

  • quantization
  • fp8
  • qwen
  • qwen3
  • abliterated
  • blackwell-optimized
  • fine-grained
    pipeline_tag: text-generation
    library_name: transformers

Qwen3-8B-FineGrained-FP8 (Blackwell Optimized)

This repository contains a high-precision Fine-Grained FP8 quantization of huihui-ai/Qwen3-8B-Instruct-Abliterated.

The model has been specifically quantized using parameters optimized for next-generation hardware, particularly the NVIDIA Blackwell (RTX 50-series) architecture.

Model Highlights

  • Architecture: Qwen3-8B
  • Quantization: Fine-Grained FP8
  • Optimization: Optimized for Blackwell Tensor Cores (weight_block_size=(128, 128))
  • Abliterated: Based on the version by huihui-ai, where refusal mechanisms have been removed to provide more direct, unfiltered responses.

Technical Configuration

The quantization was performed using FineGrainedFP8Config with the following settings:

  • Weight Block Size: 128x128. This specific block size is designed to align with the hardware throughput of RTX 5090 and other Blackwell-based GPUs, allowing for native execution with minimal overhead.
  • Precision: Unlike standard per-tensor FP8, the fine-grained approach maintains significantly higher output quality by scaling weights in smaller blocks.

Hardware Requirements

  • Optimal: NVIDIA RTX 50-series (Blackwell) for native hardware acceleration.
  • Supported: NVIDIA RTX 40-series (Ada Lovelace), H100, and L40S.
  • VRAM: Occupies approximately 8-9 GB of VRAM. A 12GB+ card is recommended for handling longer context windows and KV-cache.

Usage

You can load this model directly using the transformers library. Ensure you have the latest version of accelerate and transformers installed.

from transformers import AutoModelForCausalLM, AutoTokenizer

model_id = "ikarius/Qwen3-8B-FineGrained-FP8"

model = AutoModelForCausalLM.from_pretrained(
    model_id,
    device_map="auto",
    dtype="auto",
    trust_remote_code=True
)
tokenizer = AutoTokenizer.from_pretrained(model_id)

prompt = "Explain the advantages of FP8 quantization for LLMs."
inputs = tokenizer(prompt, return_tensors="pt").to("cuda")
outputs = model.generate(**inputs, max_new_tokens=256)

print(tokenizer.decode(outputs[0], skip_special_tokens=True))

Quantization Process

The model was quantized from the BF16 source using the following logic:

Loaded with dtype="auto" and device_map="auto".

Configured with FineGrainedFP8Config(weight_block_size=(128, 128)).

Weights were saved in the optimized FP8 format to allow for immediate loading without re-quantization.

Disclaimer

This is an abliterated model. It has fewer safety guardrails compared to the original Qwen3 release. Users are responsible for their own implementations of moderation layers and for using the model ethically and legally.

Credits

Original Model: Qwen Team

Abliteration: huihui-ai

README history 2 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2025-12-30Update README.md7e87cb93 KB
    Loading...
  2. 2025-12-30initial commitabd77ae28 B
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration