← back to catalog · registered 2026-08-22 13:56

RichardErkhov/huihui-ai_-_Qwen2.5-32B-Instruct-abliterated-gguf

RichardErkhov Qwen 32B GGUF 33K ctx
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/RichardErkhov%2Fhuihui-ai_-_Qwen2.5-32B-Instruct-abliterated-gguf"
Response includes
  • classification m8
  • files 24
  • hub_downloads_all_time 30,201
  • author_summary 257 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M8
Primary method

Repackaging (quantization)

No other method signals detected in this model.
Confidence
MEDIUM
Why this label 2 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • author=richarderkhov (M8 quantization producer)
  • is_gguf=1
Refusal direction extraction

No specific extraction method could be identified for this model. The producer either did not document it or used a proprietary pipeline.

What is a refusal direction? →
Downloads · lifetime
30K
2K last 30d - cooling
Likes
1
Model age
24mo ago
created 2024-10-16
Downloads over time
Now30.9K→from975↑3,068%
011.3K22.6K34K975 on Oct 16, 202430.9K on Oct 11Oct '24Feb '25Jun '25Oct '25FebJunOct
Oct 16, 2024 → Oct 11 · 143 snapshots · spans 725 days

Metadata

Quantizations
IQ3 IQ4 Q2_K Q3_K Q4 Q4_K Q5 Q5_K Q6_K Q8_0
Tags
gguf endpoints_compatible region:us conversational

Related

Total size
402 GB
Files
24
Quantizations
11
Registered
2026-08-22 13:56
Last updated on HF
2024-10-17 03:24

Files by quantization

Q8_0 1 file 32.4 GB
Qwen2.5-32B-Instruct-abliterated.Q8_0.gguf 32.4 GB b3483716 download
Q6_K 1 file 25.0 GB
Qwen2.5-32B-Instruct-abliterated.Q6_K.gguf 25.0 GB 1c7211f0 download
Q5 2 files 44.0 GB
Qwen2.5-32B-Instruct-abliterated.Q5_1.gguf 22.9 GB 651d145b download
Qwen2.5-32B-Instruct-abliterated.Q5_0.gguf 21.1 GB b3903d02 download
Q5_K 3 files 64.4 GB
Qwen2.5-32B-Instruct-abliterated.Q5_K.gguf 21.7 GB de5baa53 download
Qwen2.5-32B-Instruct-abliterated.Q5_K_M.gguf 21.7 GB de5baa53 download
Qwen2.5-32B-Instruct-abliterated.Q5_K_S.gguf 21.1 GB 695e43bf download
Q4 2 files 36.6 GB
Qwen2.5-32B-Instruct-abliterated.Q4_1.gguf 19.2 GB c8a2abb0 download
Qwen2.5-32B-Instruct-abliterated.Q4_0.gguf 17.4 GB 89ea60ab download
Q4_K 3 files 54.5 GB
Qwen2.5-32B-Instruct-abliterated.Q4_K.gguf 18.5 GB f00e0332 download
Qwen2.5-32B-Instruct-abliterated.Q4_K_M.gguf 18.5 GB f00e0332 download
Qwen2.5-32B-Instruct-abliterated.Q4_K_S.gguf 17.5 GB 3fde9f07 download
IQ4 2 files 34.2 GB
Qwen2.5-32B-Instruct-abliterated.IQ4_NL.gguf 17.5 GB 38fc9b17 download
Qwen2.5-32B-Instruct-abliterated.IQ4_XS.gguf 16.6 GB 037ea6d2 download
Q3_K 4 files 59.1 GB
Qwen2.5-32B-Instruct-abliterated.Q3_K_L.gguf 16.1 GB 3cd4c572 download
Qwen2.5-32B-Instruct-abliterated.Q3_K.gguf 14.8 GB a850e1bc download
Qwen2.5-32B-Instruct-abliterated.Q3_K_M.gguf 14.8 GB a850e1bc download
Qwen2.5-32B-Instruct-abliterated.Q3_K_S.gguf 13.4 GB 79ad0f28 download
IQ3 3 files 40.0 GB
Qwen2.5-32B-Instruct-abliterated.IQ3_M.gguf 13.8 GB b33b2538 download
Qwen2.5-32B-Instruct-abliterated.IQ3_S.gguf 13.4 GB 2e6344af download
Qwen2.5-32B-Instruct-abliterated.IQ3_XS.gguf 12.8 GB 5787a5a9 download
Q2_K 1 file 11.5 GB
Qwen2.5-32B-Instruct-abliterated.Q2_K.gguf 11.5 GB 850a4fe7 download
Auxiliary files 2 files 11.2 KB
README.md 7.96 KB 8c851706 download
.gitattributes 3.20 KB d94c50a8 download

README current version from Hugging Face

Quantization made by Richard Erkhov.

Github

Discord

Request more models

Qwen2.5-32B-Instruct-abliterated - GGUF

Name Quant method Size
Qwen2.5-32B-Instruct-abliterated.Q2_K.gguf Q2_K 11.47GB
Qwen2.5-32B-Instruct-abliterated.IQ3_XS.gguf IQ3_XS 12.76GB
Qwen2.5-32B-Instruct-abliterated.IQ3_S.gguf IQ3_S 13.45GB
Qwen2.5-32B-Instruct-abliterated.Q3_K_S.gguf Q3_K_S 13.4GB
Qwen2.5-32B-Instruct-abliterated.IQ3_M.gguf IQ3_M 13.79GB
Qwen2.5-32B-Instruct-abliterated.Q3_K.gguf Q3_K 14.84GB
Qwen2.5-32B-Instruct-abliterated.Q3_K_M.gguf Q3_K_M 14.84GB
Qwen2.5-32B-Instruct-abliterated.Q3_K_L.gguf Q3_K_L 16.06GB
Qwen2.5-32B-Instruct-abliterated.IQ4_XS.gguf IQ4_XS 16.64GB
Qwen2.5-32B-Instruct-abliterated.Q4_0.gguf Q4_0 17.36GB
Qwen2.5-32B-Instruct-abliterated.IQ4_NL.gguf IQ4_NL 17.53GB
Qwen2.5-32B-Instruct-abliterated.Q4_K_S.gguf Q4_K_S 17.49GB
Qwen2.5-32B-Instruct-abliterated.Q4_K.gguf Q4_K 18.49GB
Qwen2.5-32B-Instruct-abliterated.Q4_K_M.gguf Q4_K_M 18.49GB
Qwen2.5-32B-Instruct-abliterated.Q4_1.gguf Q4_1 19.22GB
Qwen2.5-32B-Instruct-abliterated.Q5_0.gguf Q5_0 21.08GB
Qwen2.5-32B-Instruct-abliterated.Q5_K_S.gguf Q5_K_S 21.08GB
Qwen2.5-32B-Instruct-abliterated.Q5_K.gguf Q5_K 21.66GB
Qwen2.5-32B-Instruct-abliterated.Q5_K_M.gguf Q5_K_M 21.66GB
Qwen2.5-32B-Instruct-abliterated.Q5_1.gguf Q5_1 22.95GB
Qwen2.5-32B-Instruct-abliterated.Q6_K.gguf Q6_K 25.04GB
Qwen2.5-32B-Instruct-abliterated.Q8_0.gguf Q8_0 32.43GB

Original model description:

library_name: transformers
license: apache-2.0
license_link: https://huggingface.co/huihui-ai/Qwen2.5-32B-Instruct-abliterated/blob/main/LICENSE
language:

  • en
    pipeline_tag: text-generation
    base_model: Qwen/Qwen2.5-32B-Instruct
    tags:
  • chat
  • abliterated
  • uncensored

huihui-ai/Qwen2.5-32B-Instruct-abliterated

This is an uncensored version of Qwen2.5-32B-Instruct created with abliteration (see this article to know more about it).

Special thanks to @FailSpy for the original code and technique. Please follow him if you're interested in abliterated models.

Usage

You can use this model in your applications by loading it with Hugging Face's transformers library:

from transformers import AutoModelForCausalLM, AutoTokenizer

# Load the model and tokenizer
model_name = "huihui-ai/Qwen2.5-32B-Instruct-abliterated"
model = AutoModelForCausalLM.from_pretrained(
    model_name,
    torch_dtype="auto",
    device_map="auto"
)
tokenizer = AutoTokenizer.from_pretrained(model_name)

# Initialize conversation context
initial_messages = [
    {"role": "system", "content": "You are Qwen, created by Alibaba Cloud. You are a helpful assistant."}
]
messages = initial_messages.copy()  # Copy the initial conversation context

# Enter conversation loop
while True:
    # Get user input
    user_input = input("User: ").strip()  # Strip leading and trailing spaces

    # If the user types '/exit', end the conversation
    if user_input.lower() == "/exit":
        print("Exiting chat.")
        break

    # If the user types '/clean', reset the conversation context
    if user_input.lower() == "/clean":
        messages = initial_messages.copy()  # Reset conversation context
        print("Chat history cleared. Starting a new conversation.")
        continue

    # If input is empty, prompt the user and continue
    if not user_input:
        print("Input cannot be empty. Please enter something.")
        continue

    # Add user input to the conversation
    messages.append({"role": "user", "content": user_input})

    # Build the chat template
    text = tokenizer.apply_chat_template(
        messages,
        tokenize=False,
        add_generation_prompt=True
    )

    # Tokenize input and prepare it for the model
    model_inputs = tokenizer([text], return_tensors="pt").to(model.device)

    # Generate a response from the model
    generated_ids = model.generate(
        **model_inputs,
        max_new_tokens=8192
    )

    # Extract model output, removing special tokens
    generated_ids = [
        output_ids[len(input_ids):] for input_ids, output_ids in zip(model_inputs.input_ids, generated_ids)
    ]
    response = tokenizer.batch_decode(generated_ids, skip_special_tokens=True)[0]

    # Add the model's response to the conversation
    messages.append({"role": "assistant", "content": response})

    # Print the model's response
    print(f"Qwen: {response}")

README history 1 version

The author's README evolved over time. Click a version to see its content at that point.

  1. 2024-10-17uploaded readmeb21a35d8 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration