← back to catalog · registered 2026-08-22 13:56

tostideluxekaas/Llama-3.2-3B-Instruct-uncensored-GGUF

tostideluxekaas Llama 3B GGUF 131K ctx
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/tostideluxekaas%2FLlama-3.2-3B-Instruct-uncensored-GGUF"
Response includes
  • classification m-uncensored
  • files 6
  • benchmarks 5 entries
  • hub_downloads_all_time 8,636
  • author_summary 4 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M-U
Primary method

Uncensored (method unknown)

No other method signals detected in this model.
Confidence
LOW
Why this label 3 signals
Weak or ambiguous signals. Best guess based on catalog patterns; treat as tentative and check the evidence below.
  • 'uncensored' in name/tags but no 'abliterated' marker
  • method not identifiable from author declaration alone
  • may be DPO fine-tune, prompt engineering, or unknown technique
Refusal direction extraction

No specific extraction method could be identified for this model. The producer either did not document it or used a proprietary pipeline.

What is a refusal direction? →
Downloads · lifetime
9K
456 last 30d - cooling
Likes
2
Model age
7mo ago
created 2026-02-21
Downloads over time
Now8.9K→from0↑0%
03.3K6.5K9.8K0 on Feb 188.9K on Oct 11FebAprJunAugOct
Feb 18 → Oct 11 · 74 snapshots · spans 235 days

Benchmarks

Portrait before abliteration
Benchmarks of the base model as it stood before the refusal-removal operation. Compare with the numbers above to see what the operation cost.
Benchmark Score Source
Arena-Battles 8390 LM-Arena
LM Arena Elo 1118.1163530112908 LM-Arena
Arena-Elo-Lower 1110.9097939581304 LM-Arena
Arena-Elo-Upper 1125.322912064451 LM-Arena
Arena-Rank 181 LM-Arena

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Variants by this author 2 formats · 482 downloads combined

The same weights this author released in different packaging. Pick the format that matches your runtime.

Metadata

License
other
Languages
en multilingual
Quantizations
F16 Q4_K Q5_K Q8_0
Tags
gguf llama llama-3.2 instruct uncensored heretic abliterattion decensored ollama lmstudio multilingual LLM

Related

Total size
13.2 GB
Files
6
Quantizations
5
Registered
2026-08-22 13:56
Last updated on HF
2026-02-21 16:40

Files by quantization

F16 1 file 5.99 GB
Llama-3.2-3B-Instruct-uncensored-f16.gguf 5.99 GB 714032b7 download
Q8_0 1 file 3.19 GB
Llama-3.2-3B-Instruct-uncensored-Q8_0.gguf 3.19 GB 342e33e2 download
Q5_K 1 file 2.16 GB
Llama-3.2-3B-Instruct-uncensored-Q5_K_M.gguf 2.16 GB 2f1e0723 download
Q4_K 1 file 1.88 GB
Llama-3.2-3B-Instruct-uncensored-Q4_K_M.gguf 1.88 GB 363d44a3 download
Auxiliary files 2 files 4.29 KB
README.md 2.50 KB b1c53c82 download
.gitattributes 1.79 KB 265190a7 download

README current version from Hugging Face


pipeline_tag: text-generation
tags:

  • llama
  • llama-3.2
  • instruct
  • uncensored
  • heretic
  • abliterattion
  • decensored
  • gguf
  • ollama
  • lmstudio
  • multilingual
  • LLM
    language:
  • en
  • multilingual
    license: other
    license_name: llama3.2
    base_model:
  • unsloth/Llama-3.2-3B-Instruct
  • tostideluxekaas/Llama-3.2-3B-Instruct-uncensored

Llama-3.2-3B-Instruct-uncensored-GGUF

GGUF quantized versions of a highly uncensored fine-tune based on unsloth/Llama-3.2-3B-Instruct.
Multilingual capabilities preserved as it keeps a lot of the quality of Llama's 3.2-3B-Instruct model
because of the low Kl divergence (0.0265) after abliteration. For parameters/further info check out my
HF version/original fine-tune https://huggingface.co/tostideluxekaas/Llama-3.2-3B-Instruct-uncensored .

Only 4 out of 100 refusals in testing.

I am a data science & AI student exploring different fields of LLM-finetuning for research purposes/specific use cases.
Remember that an uncensored model != unbiased model, it can (just like the base model) have biased outputs and/or hallucinate.

Available quants

File Size VRAM req. Recommended for
Q4_K_M.gguf 2.02GB 3–4 GB Best speed/quality balance (lightweight)
Q5_K_M.gguf 2.32GB 4–5 GB Very good quality
Q8_0.gguf 3.42GB 5–6 GB Highest quality
f16.gguf 6.43GB 8+ GB Maximum precision / full model size

Disclaimer

This is a heavily uncensored model. It may generate harmful, illegal, offensive or inappropriate content.
Use responsibly. You are solely responsible for all outputs and consequences.

License & Attribution

Llama 3.2 Community License
Copyright © Meta Platforms, Inc. All Rights Reserved.
Built with Llama.
Full license: https://github.com/meta-llama/llama-models/blob/main/models/llama3_2/LICENSE

Quick usage (LM studio)

download LM studio

https://lmstudio.ai/

Click search models

look for tostideluxekaas/Llama-3.2-3B-Instruct-uncensored-GGUF in the search box

Select Quant size

Select the appropiate quant (for reference see table)

Download and load

Select download and wait for the model to complete the download, after this you are able to load it into you chat UI

Olama

download Olama

https://ollama.com/download

ollama run huggingface.co/tostideluxekaas/Llama-3.2-3B-Instruct-uncensored-GGUF:Q4_K_M

README history 5 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-02-21Update README.md5a815e82.5 KB
    Loading...
  2. 2026-02-21Update README.mdf2e15492.5 KB
    Loading...
  3. 2026-02-21Update README.md80f9e6f2.4 KB
    Loading...
  4. 2026-02-21Update README.md92873a52.3 KB
    Loading...
  5. 2026-02-21Upload GGUF quants + READMEb2513c11.6 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration