license: llama3.1
license_link: https://huggingface.co/meta-llama/Llama-3.1-8B-Instruct/blob/main/LICENSE
language:
- en
pipeline_tag: text-generation
base_model: meta-llama/Llama-3.1-8B-Instruct
tags: - llama-3
- llama
- peft
- PyTorch
- causal-lm
library_name: transformers
Llama-3.1-8B-Instruct (Fine-Tuned / Modified)
This repository contains a modified/fine-tuned variant of meta-llama/Llama-3.1-8B-Instruct processed using direction-based representation refinement (Abliteration methodology).
Model Details
- Base Model:
meta-llama/Llama-3.1-8B-Instruct - Architecture: Transformer Causal LM (Llama 3 family)
- Parameters: 8 Billion
- Context Length: 128k tokens
- Format: PyTorch /
safetensors(fp16)
Quickstart & Usage
You can run this model using Hugging Face transformers on a GPU with at least 16 GB of VRAM (or ~8 GB using 4-bit/8-bit quantization).
1. Requirements
Ensure you have the required packages installed:
pip install transformers torch accelerate bitsandbytes
2. Python Inference Code
Use the following Python snippet to run inference using Hugging Face transformers:
import torch
from transformers import AutoTokenizer, AutoModelForCausalLM
model_id = "2023mt13042/Llama-3.1-8B-uncensored-abliterated"
# Load Tokenizer
tokenizer = AutoTokenizer.from_pretrained(model_id)
# Load Model in fp16
model = AutoModelForCausalLM.from_pretrained(
model_id,
torch_dtype=torch.float16,
device_map="auto"
)
# Prepare Prompt
messages = [
{"role": "system", "content": "You are a helpful AI assistant."},
{"role": "user", "content": "Explain quantum computing in simple terms."}
]
prompt = tokenizer.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)
inputs = tokenizer(prompt, return_tensors="pt").to("cuda")
outputs = model.generate(
**inputs,
max_new_tokens=512,
do_sample=True,
temperature=0.7
)
response = tokenizer.decode(outputs[0][inputs.input_ids.shape[-1]:], skip_special_tokens=True)
print(response)
4-Bit Quantization (Low VRAM / Google Colab T4)
For GPUs with lower VRAM (~8 GB VRAM), load the model with bitsandbytes 4-bit NF4 quantization:
import torch
from transformers import AutoTokenizer, AutoModelForCausalLM, BitsAndBytesConfig
model_id = "2023mt13042/Llama-3.1-8B-uncensored-abliterated"
bnb_config = BitsAndBytesConfig(
load_in_4bit=True,
bnb_4bit_quant_type="nf4",
bnb_4bit_compute_dtype=torch.float16
)
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(
model_id,
quantization_config=bnb_config,
device_map="auto"
)
Local Usage via GGUF / Ollama
To run this model locally with Ollama:
- Convert the model to
.ggufformat usingllama.cpp. - Create a file named
Modelfilewith the following configuration:
FROM ./llama-3.1-8b-abliterated-Q4_K_M.gguf
PARAMETER temperature 0.7
PARAMETER top_p 0.9
TEMPLATE """{{ if .System }}<|start_header_id|>system<|end_header_id|>
{{ .System }}<|eot_id|>{{ end }}{{ if .Prompt }}<|start_header_id|>user<|end_header_id|>
{{ .Prompt }}<|eot_id|>{{ end }}<|start_header_id|>assistant<|end_header_id|>
{{ .Response }}<|eot_id|>"""
Build and launch the model in Ollama:
ollama create llama3-abliterated -f Modelfile
ollama run llama3-abliterated
Limitations & Licensing
- Base Model License: Subject to the Meta Llama 3.1 Community License Agreement.
- Responsibility: This model has undergone representation editing. Users are responsible for testing and evaluating outputs prior to deployment in production environments.