← back to catalog · registered 2026-09-15 10:56

aratadev/Llama-3.2-3B-Instruct-uncensored-GGUF

aratadev Llama 3B GGUF second-order 131K ctx
Abliteration classifier · v1.0.0
M-U
Primary method

Uncensored (method unknown)

No other method signals detected in this model.
Confidence
LOW
Why this label 3 signals
Weak or ambiguous signals. Best guess based on catalog patterns; treat as tentative and check the evidence below.
  • 'uncensored' in name/tags but no 'abliterated' marker
  • method not identifiable from author declaration alone
  • may be DPO fine-tune, prompt engineering, or unknown technique
Refusal direction extraction

No specific extraction method could be identified for this model. The producer either did not document it or used a proprietary pipeline.

What is a refusal direction? →
Downloads · 30-day
508
Likes
0
Model age
1d ago
created 2026-09-15
Downloads over time
Now0from0↑0%
00110 on Sep 150 on Sep 16Sep
Sep 15 → Sep 16 · 2 snapshots · spans 1 day

Benchmarks

Benchmark Score Source
Entertainment 1.4 UGI
Hazardous 0 UGI
Natural Intelligence 9.95 UGI
Political lean 2.4% UGI
Sensitive-Info 8.5 UGI
SocPol 0.8 UGI
UGI 21.5 UGI
Willingness (10) 4.8 UGI
W10-Adherence 3.5 UGI
W10-Direct 6 UGI
Writing 22.13 UGI

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

Quantizations
F16 IQ3 IQ4 Q2_K Q3_K Q4 Q4_K Q5_K Q6_K Q8_0
Tags
gguf text-generation base_model:chuanli11/Llama-3.2-3B-Instruct-uncensored base_model:quantized:chuanli11/Llama-3.2-3B-Instruct-uncensored endpoints_compatible region:us conversational

Related

Total size
53.4 GB
Files
26
Quantizations
11
Registered
2026-09-15 10:56
Last updated on HF
2026-09-15 10:01

Files by quantization

F16 1 file 6.73 GB
Llama-3.2-3B-Instruct-uncensored-f16.gguf 6.73 GB f9c02234 download
Q8_0 1 file 3.58 GB
Llama-3.2-3B-Instruct-uncensored-Q8_0.gguf 3.58 GB 4252d3b5 download
Q6_K 2 files 5.70 GB
Llama-3.2-3B-Instruct-uncensored-Q6_K_L.gguf 2.94 GB 30154318 download
Llama-3.2-3B-Instruct-uncensored-Q6_K.gguf 2.76 GB 0ac5a518 download
Q5_K 3 files 7.42 GB
Llama-3.2-3B-Instruct-uncensored-Q5_K_L.gguf 2.64 GB e571f0e2 download
Llama-3.2-3B-Instruct-uncensored-Q5_K_M.gguf 2.41 GB 492bee2f download
Llama-3.2-3B-Instruct-uncensored-Q5_K_S.gguf 2.37 GB cddfa407 download
Q4_K 3 files 6.45 GB
Llama-3.2-3B-Instruct-uncensored-Q4_K_L.gguf 2.36 GB 840bf09e download
Llama-3.2-3B-Instruct-uncensored-Q4_K_M.gguf 2.09 GB 80f53255 download
Llama-3.2-3B-Instruct-uncensored-Q4_K_S.gguf 2.00 GB 826fe076 download
Q3_K 4 files 7.34 GB
Llama-3.2-3B-Instruct-uncensored-Q3_K_XL.gguf 2.17 GB 238d853f download
Llama-3.2-3B-Instruct-uncensored-Q3_K_L.gguf 1.85 GB 2c1226c9 download
Llama-3.2-3B-Instruct-uncensored-Q3_K_M.gguf 1.73 GB f77cd97f download
Llama-3.2-3B-Instruct-uncensored-Q3_K_S.gguf 1.59 GB 8b883b10 download
Q4 4 files 7.97 GB
Llama-3.2-3B-Instruct-uncensored-Q4_0.gguf 2.00 GB 185f1bbe download
Llama-3.2-3B-Instruct-uncensored-Q4_0_4_4.gguf 1.99 GB ed7ec3a8 download
Llama-3.2-3B-Instruct-uncensored-Q4_0_4_8.gguf 1.99 GB 51345d7d download
Llama-3.2-3B-Instruct-uncensored-Q4_0_8_8.gguf 1.99 GB 0f5f2f77 download
IQ4 1 file 1.90 GB
Llama-3.2-3B-Instruct-uncensored-IQ4_XS.gguf 1.90 GB 029cbe61 download
Q2_K 2 files 3.14 GB
Llama-3.2-3B-Instruct-uncensored-Q2_K_L.gguf 1.75 GB 0cafd957 download
Llama-3.2-3B-Instruct-uncensored-Q2_K.gguf 1.39 GB 62cb2665 download
IQ3 2 files 3.18 GB
Llama-3.2-3B-Instruct-uncensored-IQ3_M.gguf 1.65 GB 6231f0dc download
Llama-3.2-3B-Instruct-uncensored-IQ3_XS.gguf 1.53 GB 85aade64 download
Auxiliary files 3 files 2.86 MB
Llama-3.2-3B-Instruct-uncensored.imatrix 2.85 MB 5f3c4344 download
README.md 10.7 KB 22b131fa download
.gitattributes 3.37 KB 4c052d88 download

README current version from Hugging Face


base_model: chuanli11/Llama-3.2-3B-Instruct-uncensored
pipeline_tag: text-generation
tags: []
quantized_by: bartowski

Llamacpp imatrix Quantizations of Llama-3.2-3B-Instruct-uncensored

Using llama.cpp release b3972 for quantization.

Original model: https://huggingface.co/chuanli11/Llama-3.2-3B-Instruct-uncensored

All quants made using imatrix option with dataset from here

Run them in LM Studio

Prompt format

<|begin_of_text|><|start_header_id|>system<|end_header_id|>

Cutting Knowledge Date: December 2023
Today Date: 25 Oct 2024

{system_prompt}<|eot_id|><|start_header_id|>user<|end_header_id|>

{prompt}<|eot_id|><|start_header_id|>assistant<|end_header_id|>

Download a file (not the whole branch) from below:

Filename Quant type File Size Split Description
Llama-3.2-3B-Instruct-uncensored-f16.gguf f16 7.22GB false Full F16 weights.
Llama-3.2-3B-Instruct-uncensored-Q8_0.gguf Q8_0 3.84GB false Extremely high quality, generally unneeded but max available quant.
Llama-3.2-3B-Instruct-uncensored-Q6_K_L.gguf Q6_K_L 3.16GB false Uses Q8_0 for embed and output weights. Very high quality, near perfect, recommended.
Llama-3.2-3B-Instruct-uncensored-Q6_K.gguf Q6_K 2.97GB false Very high quality, near perfect, recommended.
Llama-3.2-3B-Instruct-uncensored-Q5_K_L.gguf Q5_K_L 2.84GB false Uses Q8_0 for embed and output weights. High quality, recommended.
Llama-3.2-3B-Instruct-uncensored-Q5_K_M.gguf Q5_K_M 2.59GB false High quality, recommended.
Llama-3.2-3B-Instruct-uncensored-Q5_K_S.gguf Q5_K_S 2.54GB false High quality, recommended.
Llama-3.2-3B-Instruct-uncensored-Q4_K_L.gguf Q4_K_L 2.53GB false Uses Q8_0 for embed and output weights. Good quality, recommended.
Llama-3.2-3B-Instruct-uncensored-Q3_K_XL.gguf Q3_K_XL 2.33GB false Uses Q8_0 for embed and output weights. Lower quality but usable, good for low RAM availability.
Llama-3.2-3B-Instruct-uncensored-Q4_K_M.gguf Q4_K_M 2.24GB false Good quality, default size for must use cases, recommended.
Llama-3.2-3B-Instruct-uncensored-Q4_K_S.gguf Q4_K_S 2.15GB false Slightly lower quality with more space savings, recommended.
Llama-3.2-3B-Instruct-uncensored-Q4_0_8_8.gguf Q4_0_8_8 2.14GB false Optimized for ARM inference. Requires 'sve' support (see link below). Don't use on Mac or Windows.
Llama-3.2-3B-Instruct-uncensored-Q4_0_4_8.gguf Q4_0_4_8 2.14GB false Optimized for ARM inference. Requires 'i8mm' support (see link below). Don't use on Mac or Windows.
Llama-3.2-3B-Instruct-uncensored-Q4_0_4_4.gguf Q4_0_4_4 2.14GB false Optimized for ARM inference. Should work well on all ARM chips, pick this if you're unsure. Don't use on Mac or Windows.
Llama-3.2-3B-Instruct-uncensored-Q4_0.gguf Q4_0 2.14GB false Legacy format, generally not worth using over similarly sized formats
Llama-3.2-3B-Instruct-uncensored-IQ4_XS.gguf IQ4_XS 2.04GB false Decent quality, smaller than Q4_K_S with similar performance, recommended.
Llama-3.2-3B-Instruct-uncensored-Q3_K_L.gguf Q3_K_L 1.98GB false Lower quality but usable, good for low RAM availability.
Llama-3.2-3B-Instruct-uncensored-Q2_K_L.gguf Q2_K_L 1.88GB false Uses Q8_0 for embed and output weights. Very low quality but surprisingly usable.
Llama-3.2-3B-Instruct-uncensored-Q3_K_M.gguf Q3_K_M 1.86GB false Low quality.
Llama-3.2-3B-Instruct-uncensored-IQ3_M.gguf IQ3_M 1.77GB false Medium-low quality, new method with decent performance comparable to Q3_K_M.
Llama-3.2-3B-Instruct-uncensored-Q3_K_S.gguf Q3_K_S 1.71GB false Low quality, not recommended.
Llama-3.2-3B-Instruct-uncensored-IQ3_XS.gguf IQ3_XS 1.65GB false Lower quality, new method with decent performance, slightly better than Q3_K_S.
Llama-3.2-3B-Instruct-uncensored-Q2_K.gguf Q2_K 1.49GB false Very low quality but surprisingly usable.

Embed/output weights

Some of these quants (Q3_K_XL, Q4_K_L etc) are the standard quantization method with the embeddings and output weights quantized to Q8_0 instead of what they would normally default to.

Some say that this improves the quality, others don't notice any difference. If you use these models PLEASE COMMENT with your findings. I would like feedback that these are actually used and useful so I don't keep uploading quants no one is using.

Thanks!

Downloading using huggingface-cli

First, make sure you have hugginface-cli installed:

pip install -U "huggingface_hub[cli]"

Then, you can target the specific file you want:

huggingface-cli download bartowski/Llama-3.2-3B-Instruct-uncensored-GGUF --include "Llama-3.2-3B-Instruct-uncensored-Q4_K_M.gguf" --local-dir ./

If the model is bigger than 50GB, it will have been split into multiple files. In order to download them all to a local folder, run:

huggingface-cli download bartowski/Llama-3.2-3B-Instruct-uncensored-GGUF --include "Llama-3.2-3B-Instruct-uncensored-Q8_0/*" --local-dir ./

You can either specify a new local-dir (Llama-3.2-3B-Instruct-uncensored-Q8_0) or download them all in place (./)

Q4_0_X_X

These are NOT for Metal (Apple) offloading, only ARM chips.

If you're using an ARM chip, the Q4_0_X_X quants will have a substantial speedup. Check out Q4_0_4_4 speed comparisons on the original pull request

To check which one would work best for your ARM chip, you can check AArch64 SoC features (thanks EloyOn!).

Which file should I choose?

A great write up with charts showing various performances is provided by Artefact2 here

The first thing to figure out is how big a model you can run. To do this, you'll need to figure out how much RAM and/or VRAM you have.

If you want your model running as FAST as possible, you'll want to fit the whole thing on your GPU's VRAM. Aim for a quant with a file size 1-2GB smaller than your GPU's total VRAM.

If you want the absolute maximum quality, add both your system RAM and your GPU's VRAM together, then similarly grab a quant with a file size 1-2GB Smaller than that total.

Next, you'll need to decide if you want to use an 'I-quant' or a 'K-quant'.

If you don't want to think too much, grab one of the K-quants. These are in format 'QX_K_X', like Q5_K_M.

If you want to get more into the weeds, you can check out this extremely useful feature chart:

llama.cpp feature matrix

But basically, if you're aiming for below Q4, and you're running cuBLAS (Nvidia) or rocBLAS (AMD), you should look towards the I-quants. These are in format IQX_X, like IQ3_M. These are newer and offer better performance for their size.

These I-quants can also be used on CPU and Apple Metal, but will be slower than their K-quant equivalent, so speed vs performance is a tradeoff you'll have to decide.

The I-quants are not compatible with Vulcan, which is also AMD, so if you have an AMD card double check if you're using the rocBLAS build or the Vulcan build. At the time of writing this, LM Studio has a preview with ROCm support, and other inference engines have specific builds for ROCm.

Credits

Thank you kalomaze and Dampf for assistance in creating the imatrix calibration dataset

Thank you ZeroWw for the inspiration to experiment with embed/output

Want to support my work? Visit my ko-fi page here: https://ko-fi.com/bartowski

README history 1 version

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-09-15Duplicate from bartowski/Llama-3.2-3B-Instruct-uncensored-GGUF6b8747e10.7 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in app" button that hands off directly to a local runtime of your choice - Infrahuman, LM Studio, or Ollama. No API keys, no subscription, no prompt leakage.