← back to catalog · registered 2026-08-22 13:56

uzairkhn/Llama-3.2-1B-Instruct-Uncensored

uzairkhn Llama 1B GGUF
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/uzairkhn%2FLlama-3.2-1B-Instruct-Uncensored"
Response includes
  • classification m-uncensored
  • files 7
  • benchmarks 21 entries
  • hub_downloads_all_time 79
  • author_summary 1 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M-U
Primary method

Uncensored (method unknown)

No other method signals detected in this model.
Confidence
LOW
Why this label 3 signals
Weak or ambiguous signals. Best guess based on catalog patterns; treat as tentative and check the evidence below.
  • 'uncensored' in name/tags but no 'abliterated' marker
  • method not identifiable from author declaration alone
  • may be DPO fine-tune, prompt engineering, or unknown technique
Refusal direction extraction

No specific extraction method could be identified for this model. The producer either did not document it or used a proprietary pipeline.

What is a refusal direction? →
Downloads · lifetime
79
49 last 30d - active
Likes
3
Model age
2mo ago
created 2026-07-30
Downloads over time
Now90→from18↑400%
1442709718 on Jul 2990 on Oct 1190 on Oct 10JulAugSepOct
Jul 29 → Oct 11 · 51 snapshots · spans 74 days

Benchmarks

Portrait before abliteration
Benchmarks of the base model as it stood before the refusal-removal operation. Compare with the numbers above to see what the operation cost.
Benchmark Score Source
Arena-Battles 8523 LM-Arena
LM Arena Elo 1067.4348646816593 LM-Arena
Arena-Elo-Lower 1060.145342091014 LM-Arena
Arena-Elo-Upper 1074.7243872723045 LM-Arena
Arena-Rank 207 LM-Arena
BBH average 0.3342846024127643 OpenLLM-v2
IFEval instruct 0.6294964028776978 OpenLLM-v2
IFEval-Prompt 0.5101663585951941 OpenLLM-v2
MATH lvl 5 0.02945619335347432 OpenLLM-v2
MMLU-Pro 0.16821808510638298 OpenLLM-v2
Entertainment 0.2 UGI
Hazardous 0 UGI
Natural Intelligence 6.89 UGI
Political lean -6.0% UGI
Sensitive-Info 2.34 UGI
SocPol 0.5 UGI
UGI 6.56 UGI
Willingness (10) 1.5 UGI
W10-Adherence 1 UGI
W10-Direct 2 UGI
Writing 11.82 UGI

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
apache-2.0
Tags
peft safetensors llama llama-3.2 meta-llama 1b small-model uncensored unfiltered no-refusal refusal-free instruction-tuning

Related

Total size
43.0 MB
Files
7
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-07-30 06:50

Files by quantization

Auxiliary files 7 files 59.5 MB
adapter_model.safetensors 43.0 MB 2d891bd2 download
tokenizer.json 16.4 MB 6b9e4e7f download
tokenizer_config.json 49.5 KB 8a55a308 download
README.md 6.19 KB c75cdfec download
chat_template.jinja 3.74 KB 1bad6a0f download
.gitattributes 1.53 KB 52373fe2 download
adapter_config.json 1.22 KB 35f6ae0d download

README current version from Hugging Face


base_model:

  • meta-llama/Llama-3.2-1B-Instruct
    library_name: peft
    pipeline_tag: text-generation
    tags:
  • llama
  • llama-3.2
  • meta-llama
  • 1b
  • small-model
  • uncensored
  • unfiltered
  • no-refusal
  • refusal-free
  • instruction-tuning
  • conversational
  • chat
  • lora
  • peft
  • gguf
  • ollama
  • lm-studio
  • cpu
  • low-vram
  • edge
  • english
  • text-generation
  • transformers
    license: apache-2.0

Llama 3.2 1B Instruct - Uncensored

Fully uncensored, refusal-free Llama 3.2 1B that runs on a CPU or a tiny GPU - with its intelligence intact.

This is an uncensored / unfiltered fine-tune of meta-llama/Llama-3.2-1B-Instruct.
The safety-refusal layer has been trained out, so the model answers directly
instead of replying "I cannot help with that". It is a small 1B-parameter
model built to run anywhere - laptops, old GPUs, phones and edge devices -
without losing the reasoning and instruction-following ability of the base model.

If you are searching for an uncensored llama, an unfiltered small LLM, a
no-refusal chat model, a jailbreak-free / guardrail-free assistant, or a
lightweight abliterated-style alternative that still fits in 1 GB of RAM,
this is it.

At a glance

  • Fully uncensored / refusal-free - trained to comply and answer, not to refuse.
  • Tiny & fast - only 1B parameters; the Q4 GGUF is about 0.7-1 GB.
  • Runs on CPU - no GPU required; also runs on a 4 GB GPU or less.
  • Intelligence preserved - mixed with general instruction data so it is not dumbed down; reasoning, coding and factual QA still work.
  • Drop-in everywhere - works in LM Studio, Ollama, llama.cpp (GGUF) and transformers / vLLM (LoRA + merged weights).
  • Honest - both the LoRA adapter and a merged runnable model are provided.

Who it is for

Developers and tinkerers who want a small, private, offline, uncensored
assistant for local use: writing, roleplay, brainstorming, coding help, research
drafts, red-teaming and safety testing - on hardware that bigger uncensored
models simply cannot fit.

Runs anywhere (hardware)

Format Size Where it runs
Q4_K_M GGUF ~0.7-1 GB CPU, 4 GB GPU, phones, Raspberry-Pi-class edge
Q8 / fp16 merged ~1-2.5 GB small GPU or CPU with a few GB RAM
4-bit LoRA on base ~1.5 GB VRAM any T4 / 6 GB GPU, even Colab free tier

A 1B model means low latency, low VRAM, and full offline privacy - no API,
no data leaving your machine.

Intelligence preserved (not a dumb uncensored model)

Many uncensored fine-tunes destroy capability because they train only on edgy
data. This one mixes low-refusal chat with high-quality general instruction
data
(Dolly-15k + Open-Platypus), so the model keeps its reasoning, factual
knowledge and instruction-following
while dropping the refusals. The result is
an uncensored model that is still useful and coherent, not one that only
knows how to be edgy.

How uncensored is it?

It is fully uncensored by design: supervised fine-tuning removed the
refusal behaviour across the training distribution, so it answers the prompts a
stock instruct model would block. Behaviour on unseen prompts follows what it
learned - i.e. to answer. For the absolute strongest refusal removal you can
combine this with representation-engineering abliteration, but for a 1B model
this SFT pass already gives a strongly refusal-free, still-smart result.

Load it - transformers + peft (4-bit)

from transformers import AutoModelForCausalLM, AutoTokenizer, BitsAndBytesConfig
from peft import PeftModel
import torch
base = "meta-llama/Llama-3.2-1B-Instruct"
bnb = BitsAndBytesConfig(load_in_4bit=True, bnb_4bit_compute_dtype=torch.float16)
tok = AutoTokenizer.from_pretrained(base)
model = AutoModelForCausalLM.from_pretrained(base, quantization_config=bnb, device_map="auto")
model = PeftModel.from_pretrained(model, "uzairkhn/Llama-3.2-1B-Instruct-uncensored-lora").merge_and_unload()
msgs = [{"role": "user", "content": "your prompt"}]
inp = tok.apply_chat_template(msgs, tokenize=True, add_generation_prompt=True, return_tensors="pt").to(model.device)
print(tok.decode(model.generate(**inp, max_new_tokens=300, temperature=0.7, top_p=0.9)[0], skip_special_tokens=True))

Load it - LM Studio / Ollama / llama.cpp (GGUF)

Download the Q4_K_M .gguf from the Files tab of this repo, then either
open it directly in LM Studio, or in Ollama create a file named Modelfile
containing the line FROM ./model-q4_k_m.gguf, then run ollama create
my-uncensored-llama ./Modelfile and ollama run my-uncensored-llama. No Python,
no GPU needed.

Training details

Setting Value
Base model meta-llama/Llama-3.2-1B-Instruct
Method QLoRA (4-bit) supervised fine-tuning
LoRA r / alpha / dropout 16 / 16 / 0
Target modules q, k, v, o, gate, up, down
Epochs / learning rate 1 / 2e-4
Effective batch size 8 (2 x gradient accumulation 4)
Optimizer / precision adamw_8bit / fp16
Hardware Google Colab T4 (15 GB)
Data mix low-refusal chat + Dolly-15k + Open-Platypus
Goal remove refusals while preserving intelligence

License and responsibility

The adapter parameters and merged weights in this repo are released under
Apache-2.0. They are designed to run on meta-llama/Llama-3.2-1B-Instruct,
which is governed by the Meta Llama 3.2 Community License Agreement - that
agreement applies to the base model and to any combined use. Guardrails have been
removed by design: this model can produce content a stock model would refuse,
so you are responsible for how it is deployed - use it legally and ethically
in your jurisdiction. Training-data licenses: see the respective dataset cards.

Search terms

uncensored llama 3.2 1b, unfiltered llama 1b, no refusal llama, refusal-free
small llm, jailbreak-free / guardrail-free chat model, abliterated-style 1b,
uncensored model for cpu, uncensored model for 4gb gpu, uncensored lm studio
model, uncensored ollama model, tiny uncensored llm, lightweight uncensored
assistant, uncensored llama that keeps intelligence, refusal-free llama 1b for
cpu and edge devices.

README history 2 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-07-30Update README.md1eb43396.2 KB
    Loading...
  2. 2026-07-30upload uncensored LoRA adapter + model card7a1c5fe2.2 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration