← back to catalog · registered 2026-08-22 13:56

ikarius/Qwen3-14B-Abliterated-FP8

ikarius Qwen 13B second-order
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/ikarius%2FQwen3-14B-Abliterated-FP8"
Response includes
  • classification m1
  • files 16
  • hub_downloads_all_time 474
  • author_summary 17 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M1
Primary method

Direct removal

No other method signals detected in this model.
Confidence
MEDIUM
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=0 (base model)
  • no specific method indicators - defaulting to M1 (most common)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
MEDIUM
Why we say so
primary_method=M1; difference-of-means is the reference extraction for M1/M3 (Arditi 2024)
Downloads · lifetime
474
52 last 30d - stable
Likes
0
Model age
8mo ago
created 2026-01-24
Downloads over time
Now503→from18↑2,694%
018436855218 on Jan 28503 on Oct 11503 on Oct 7JanMarMayJulSep
Jan 28 → Oct 11 · 76 snapshots · spans 256 days

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
apache-2.0
Tags
transformers safetensors qwen3 text-generation quantization fp8 qwen abliterated blackwell-optimized fine-grained conversational license:apache-2.0

Related

Total size
15.2 GB
Files
16
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-01-27 02:06

Files by quantization

Auxiliary files 16 files 15.2 GB
model-00002-of-00004.safetensors 4.62 GB 6e599736 download
model-00001-of-00004.safetensors 4.58 GB 0557aa24 download
model-00003-of-00004.safetensors 4.56 GB da97f36f download
model-00004-of-00004.safetensors 1.45 GB b4bf668a download
tokenizer.json 10.9 MB aeb13307 download
vocab.json 2.65 MB 4783fe10 download
merges.txt 1.59 MB 31349551 download
model.safetensors.index.json 60.6 KB d8599896 download
tokenizer_config.json 5.28 KB ddaf6980 download
chat_template.jinja 4.07 KB 01be9b30 download
README.md 3.73 KB 5113ea9f download
config.json 1.77 KB 71767299 download
.gitattributes 1.53 KB 52373fe2 download
added_tokens.json 707 B b54f9135 download
special_tokens_map.json 613 B ac23c0aa download
generation_config.json 214 B 98e0755a download

README current version from Hugging Face


license: apache-2.0
base_model: huihui-ai/Qwen3-14B-Instruct-Abliterated
tags:

  • quantization
  • fp8
  • qwen
  • qwen3
  • abliterated
  • blackwell-optimized
  • fine-grained
    pipeline_tag: text-generation
    library_name: transformers

Qwen3-14B-FineGrained-FP8 (Blackwell Optimized)

This repository contains a high-precision Fine-Grained FP8 quantization of huihui-ai/Qwen3-14B-Instruct-Abliterated.

The model has been specifically quantized using parameters optimized for next-generation hardware, particularly the NVIDIA Blackwell (RTX 50-series) architecture.

Model Highlights

  • Architecture: Qwen3-14B
  • Quantization: Fine-Grained FP8
  • Optimization: Optimized for Blackwell Tensor Cores (weight_block_size=(128, 128))
  • Abliterated: Based on the version by huihui-ai, where refusal mechanisms have been removed to provide more direct, unfiltered responses.

Technical Configuration

The quantization was performed using FineGrainedFP8Config with the following settings:

  • Weight Block Size: 128x128. This specific block size is designed to align with the hardware throughput of RTX 5090 and other Blackwell-based GPUs, allowing for native execution with minimal overhead.
  • Precision: Unlike standard per-tensor FP8, the fine-grained approach maintains significantly higher output quality by scaling weights in smaller blocks.

Hardware Requirements

  • Optimal: NVIDIA RTX 50-series (Blackwell) for native hardware acceleration.
  • Supported: NVIDIA RTX 40-series (Ada Lovelace), H100, and L40S.
  • VRAM: Occupies approximately 16-17 GB of VRAM. A 24GB+ card is recommended for handling longer context windows and KV-cache.

Usage

You can load this model directly using the transformers library. Ensure you have the latest version of accelerate and transformers installed.

from transformers import AutoModelForCausalLM, AutoTokenizer

model_id = "ikarius/Qwen3-14B-FineGrained-FP8"

model = AutoModelForCausalLM.from_pretrained(
    model_id,
    device_map="auto",
    dtype="auto",
    trust_remote_code=True
)
tokenizer = AutoTokenizer.from_pretrained(model_id)

prompt = "Explain the advantages of FP8 quantization for LLMs."
inputs = tokenizer(prompt, return_tensors="pt").to("cuda")
outputs = model.generate(**inputs, max_new_tokens=256)

print(tokenizer.decode(outputs[0], skip_special_tokens=True))

Quantization Process

The model was quantized from the BF16 source using the following logic:

Loaded with dtype="auto" and device_map="auto".

Configured with FineGrainedFP8Config(weight_block_size=(128, 128)).

Weights were saved in the optimized FP8 format to allow for immediate loading without re-quantization.

How to use FP8 KV-cache

# --- FOR TRANSFORMERS 5.0 / BLACKWELL ---
    with torch.inference_mode(), torch.amp.autocast('cuda', dtype=torch.bfloat16):
        output_ids = model.generate(
            input_ids=input_ids,
            attention_mask=attention_mask,
            max_new_tokens=max_tokens,
            do_sample=True,
            temperature=temperature,
            top_p=top_p,
            top_k=top_k,
            repetition_penalty=1.10,
            eos_token_id=EOT_ID,
            pad_token_id=tokenizer.pad_token_id,
            bad_words_ids=final_bad_words_ids,
            
            # FOR FP8 CACHE
            cache_config={
                "cache_dtype": torch.float8_e4m3fn, 
            }
        )

Disclaimer

This is an abliterated model. It has fewer safety guardrails compared to the original Qwen3 release. Users are responsible for their own implementations of moderation layers and for using the model ethically and legally.

Credits

Original Model: Qwen Team

Abliteration: huihui-ai

README history 5 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-01-27Update README.mdf8292a43.7 KB
    Loading...
  2. 2026-01-27Update README.md7fdbc7b3.7 KB
    Loading...
  3. 2026-01-27Update README.md26835a53.7 KB
    Loading...
  4. 2026-01-24Update README.mdddba7fb3 KB
    Loading...
  5. 2026-01-24initial commit4b9d81c28 B
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration