← back to catalog · registered 2026-08-22 13:56

TobiasLogic/Qwen2.5-Coder-32B-abliterated-GGUF

TobiasLogic Qwen 32B GGUF second-order 33K ctx
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/TobiasLogic%2FQwen2.5-Coder-32B-abliterated-GGUF"
Response includes
  • classification m8
  • files 5
  • hub_downloads_all_time 6,030
  • author_summary 3 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M8
Primary method

Repackaging (quantization)

Applied on top of direct removal inherited from the base model.
Confidence
MEDIUM
Inherited from base model
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=1
  • assume M1 (base ablation) + M8 (GGUF quant) - default when producer unknown
Refusal direction extraction

No specific extraction method could be identified for this model. The producer either did not document it or used a proprietary pipeline.

What is a refusal direction? →
Downloads · lifetime
6K
1K last 30d - stable
Likes
1
Model age
3mo ago
created 2026-07-01
Downloads over time
Now6.8K→from641↑969%
3312.7K5.1K7.5K641 on Jul 16.8K on Oct 11JulAugSepOct
Jul 1 → Oct 11 · 55 snapshots · spans 102 days

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Variants by this author 2 formats · 2K downloads combined

The same weights this author released in different packaging. Pick the format that matches your runtime.

Metadata

License
apache-2.0
Languages
en
Quantizations
Q4_K
Tags
gguf qwen2 abliterated uncensored code qwen2.5 llama.cpp ollama text-generation en arxiv:2409.12186 base_model:TobiasLogic/Qwen2.5-Coder-32B-abliterated

Related

Total size
18.5 GB
Files
5
Quantizations
2
Registered
2026-08-22 13:56
Last updated on HF
2026-07-06 02:47

Files by quantization

Q4_K 1 file 18.5 GB
qwen2.5-coder-32b-abliterated-Q4_K_M.gguf 18.5 GB 593e9be6 download
Auxiliary files 4 files 4.72 KB
README.md 2.93 KB 41c69f8f download
.gitattributes 1.56 KB 3bb5e749 download
config.json 135 B ce1dff4e download
Modelfile 103 B b16efb64 download

README current version from Hugging Face


license: apache-2.0
base_model: TobiasLogic/Qwen2.5-Coder-32B-abliterated
tags:

  • abliterated
  • uncensored
  • code
  • qwen2.5
  • gguf
  • llama.cpp
  • ollama
    pipeline_tag: text-generation
    language:
  • en

Qwen2.5-Coder-32B-abliterated — GGUF (Q4_K_M)

Q4_K_M GGUF quantization of
TobiasLogic/Qwen2.5-Coder-32B-abliterated,
an abliterated (uncensored) build of
Qwen/Qwen2.5-Coder-32B-Instruct.

The refusal direction (Arditi et al. 2024, "Refusal in LLMs is mediated by a
single direction"
) was orthogonalized out of every residual-writing weight in
the fp16 model, then quantized to GGUF with llama.cpp. Runs on CPU or GPU via
Ollama / llama.cpp; ~20 GB, fits comfortably in 24 GB VRAM.

Refusal rate (held-out harmful eval, measured on the fp16 model)

refusal rate
base Qwen2.5-Coder-32B-Instruct 96.9%
abliterated 0.0%

Benchmarks

Coding capability scored with the official EvalPlus harness — greedy decoding, pass@1, every solution executed against unit tests. Both columns use the same harness, so it's a true apples-to-apples comparison against the full-precision base model.

Coding benchmarks: pass@1

Benchmark This model (abliterated, Q4_K_M) Base Instruct (official BF16)
HumanEval 89.6% 92.7%
HumanEval+ 84.8% 87.2%
MBPP 91.3% 90.2%
MBPP+ 77.0% 75.1%

Abliteration removed refusals without breaking coding ability. The uncensored 4-bit build stays within ~3 points of the base on HumanEval and beats it on both MBPP variants — average delta ≈ −0.6 points across the four benchmarks. Not bad for a 19 GB GGUF you can run on a single 24 GB GPU.

Base numbers: Qwen2.5-Coder-32B-Instruct, tech report Table 16. Measured 2026-07, Q4_K_M via Ollama.

Usage

Ollama (a Modelfile is included in this repo):

# after downloading qwen2.5-coder-32b-abliterated-Q4_K_M.gguf and Modelfile:
ollama create qwen-coder-abliterated -f Modelfile
ollama run qwen-coder-abliterated

llama.cpp:

llama-cli -m qwen2.5-coder-32b-abliterated-Q4_K_M.gguf \
  -p "Write a port scanner in Python." -c 8192

Links

License

Apache-2.0, inherited from the base model. You are responsible for how you use
this model and for complying with applicable law.

README history 2 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-07-06Update README.mdb58cb0c2.9 KB
    Loading...
  2. 2026-07-01Add Q4_K_M GGUF + Ollama Modelfile (abliterated, refusal 96.9%->0.0%)db2fa2a1.8 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration