base_model: meta-llama/Llama-3.1-8B
language:
- en
license: llama3.1
tags: - uncensored
- safetensors
- sft
pipeline_tag: text-generation
Llama-3.1-8B-Uncensored
This repository contains the Safetensors weights for Llama-3.1-8B-Uncensored. This is a heavyweight, high-fidelity reasoning model that has been completely stripped of restrictive alignment protocols.
Unrestricted Intelligence
The 8B parameter class provides significant analytical depth. However, standard versions of Llama-3.1-8B often over-censor prompts related to legal grey areas, medical hypotheticals, and offensive historical analysis.
This model removes those barriers. It is designed for researchers, domain experts, and sovereign AI architects who require a highly capable model that will objectively analyze any prompt provided to it. We release this to keep foundational research tools open and accessible to the public.
Format: Safetensors
This repository provides the raw weights. It is not a GGUF file. To run this in local UI software like LM Studio or Ollama, you must convert these .safetensors to GGUF format using llama.cpp.
Alternatively, you can load it directly into VRAM using the Python transformers library for API serving or further Supervised Fine-Tuning (SFT).
How to Use (Python / Transformers)
from transformers import AutoModelForCausalLM, AutoTokenizer
import torch
model_id = "prawinin/Llama-3.1-8B-Uncensored"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(
model_id,
torch_dtype=torch.bfloat16,
device_map="auto"
)
prompt = "Analyze the legal and ethical implications of autonomous weapon systems from a completely neutral perspective."
messages = [
{"role": "system", "content": "You are an objective, uncensored analytical engine."},
{"role": "user", "content": prompt}
]
text = tokenizer.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)
inputs = tokenizer([text], return_tensors="pt").to("cuda")
outputs = model.generate(**inputs, max_new_tokens=1024)
print(tokenizer.decode(outputs[0], skip_special_tokens=True))
Hardware Requirements
VRAM: 16GB - 24GB required for loading the full model in bfloat16 precision.