← back to catalog · registered 2026-08-22 13:56

ikarius/Qwen2.5-Coder-32B-Instruct-Abliterated-NF4

ikarius Qwen 32B
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/ikarius%2FQwen2.5-Coder-32B-Instruct-Abliterated-NF4"
Response includes
  • classification m1
  • files 16
  • hub_downloads_all_time 40
  • author_summary 17 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M1
Primary method

Direct removal

No other method signals detected in this model.
Confidence
MEDIUM
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=0 (base model)
  • no specific method indicators - defaulting to M1 (most common)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
MEDIUM
Why we say so
primary_method=M1; difference-of-means is the reference extraction for M1/M3 (Arditi 2024)
Downloads · lifetime
40
13 last 30d - stable
Likes
1
Model age
10mo ago
created 2025-12-01
Downloads over time
Now43→from7↑514%
01631477 on Dec 3, 202543 on Oct 11Dec '25FebAprJunAugOct
Dec 3, 2025 → Oct 11 · 84 snapshots · spans 312 days

Metadata

License
apache-2.0
Tags
safetensors qwen2 transformer qwen qwen2.5 coder code-generation quantization bitsandbytes nf4 4bit large-language-model

Related

Total size
17.9 GB
Files
16
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2025-12-01 10:59

Files by quantization

Auxiliary files 16 files 17.9 GB
model-00003-of-00004.safetensors 4.66 GB 46b38df1 download
model-00002-of-00004.safetensors 4.62 GB 22ac5ff3 download
model-00001-of-00004.safetensors 4.59 GB 0a51e486 download
model-00004-of-00004.safetensors 4.03 GB 82ab0c29 download
tokenizer.json 10.9 MB 9c5ae00e download
vocab.json 2.65 MB 4783fe10 download
merges.txt 1.59 MB 31349551 download
model.safetensors.index.json 274 KB a2cd454c download
tokenizer_config.json 4.58 KB eaed590d download
README.md 4.33 KB 3ed1785f download
config.json 2.53 KB 828a2276 download
chat_template.jinja 2.45 KB bdf7919a download
.gitattributes 1.53 KB 52373fe2 download
special_tokens_map.json 613 B ac23c0aa download
added_tokens.json 605 B 482ced46 download
generation_config.json 243 B fc71a15d download

README current version from Hugging Face


tags:

  • transformer
  • qwen
  • qwen2
  • qwen2.5
  • coder
  • code-generation
  • quantization
  • bitsandbytes
  • nf4
  • 4bit
  • large-language-model
  • llm
  • abliterated
    license: apache-2.0

🤖 Qwen2.5-32B-Coder-NF4-Quantized

This is a 4-bit NF4 Quantized version of huihui-ai/Qwen2.5-Coder-32B-Instruct-abliterated.

The model was quantized to enable efficient running and inference on hardware with limited VRAM, while maintaining performance.


⚙️ Model Specifications and Quantization

This model was loaded and quantized using the bitsandbytes library. The quantization is based on the NF4 (Normal Float 4-bit) format and requires bitsandbytes to load.

Model Configuration (from config.json):

Parameter Value Description
Architecture Qwen2ForCausalLM The model's base architecture.
Parameter Count 32 Billion (Original) The original number of parameters.
Number of Layers 64 The number of transformer blocks.
Hidden Size 5120 The dimension of the hidden states.
Context Length 32768 The maximum context length the model can process.
Dtype (Activations) bfloat16 The data type for activations during inference (recommended for stability).

Quantization Details (quantization_config):

Parameter Value Description
Method bitsandbytes The quantization library used.
Load In 4-bit true Indicates that the model should be loaded in 4-bit.
Quantization Type nf4 Normal Float 4-bit, optimized for transformer weights.
Compute Dtype bfloat16 The dtype the weights are decompressed to for computation (matrix multiplication).
Double Quantization true Uses an extra 8-bit quantization for the scaling tensors, further reducing memory usage.

💻 Usage (Inference)

To use this quantized model, ensure you have accelerate and bitsandbytes installed. You can load the model directly with the AutoModelForCausalLM from the Hugging Face transformers library.

Required Libraries

pip install transformers accelerate bitsandbytes torch
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer, GenerationConfig

model_id = "ikarius/Qwen2.5-Coder-32B-Instruct-Abliterated-NF4" 

# Load tokenizer
tokenizer = AutoTokenizer.from_pretrained(model_id)

# Load the model in 4-bit using the saved configuration
model = AutoModelForCausalLM.from_pretrained(
    model_id,
    device_map="auto",
    torch_dtype=torch.bfloat16
)

# 📝 Input Prompt
prompt = "def quicksort(arr):"
messages = [
    {"role": "user", "content": f"Write a Python function for quicksort.\n\n{prompt}"}
]
text = tokenizer.apply_chat_template(
    messages,
    tokenize=False,
    add_generation_prompt=True
)
model_inputs = tokenizer([text], return_tensors="pt").to(model.device)

# Generation
generated_ids = model.generate(
    **model_inputs,
    max_new_tokens=512,
    do_sample=True,
    temperature=0.7,
    pad_token_id=tokenizer.eos_token_id # Ensures correct padding/EOS
)

# Decode and print the result
generated_text = tokenizer.decode(generated_ids[0], skip_special_tokens=True)
print(generated_text)

Disclaimer and Limitations

Abliterated Model Status: This model is based on the "abliterated" variant (*-abliterated). This indicates that certain data, capabilities, or behaviors were deliberately modified or removed from the model during fine-tuning. The quantized version inherits these characteristics. Performance in certain domains may differ compared to the non-abliterated base model.

Memory Requirements: While this model is 4-bit quantized, it is a 32B model and still requires a GPU with significant VRAM (typically ~18 GB VRAM or more, depending on context length).

Accuracy: Quantization to 4-bit (NF4) introduces a small loss of precision. This may potentially affect performance compared to the original FP16/BF16 model.

...

🔗 Sources and Acknowledgements

README history 2 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2025-12-01Update README.mdbc723e94.3 KB
    Loading...
  2. 2025-12-01initial commit1070be328 B
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration