← back to catalog · registered 2026-08-22 13:56

waguriagent/qwen2.5-7b-uncensored-gguf

waguriagent Qwen 7B GGUF 33K ctx
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/waguriagent%2Fqwen2.5-7b-uncensored-gguf"
Response includes
  • classification m-uncensored
  • files 3
  • benchmarks 16 entries
  • hub_downloads_all_time 3,341
  • author_summary 1 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M-U
Primary method

Uncensored (method unknown)

No other method signals detected in this model.
Confidence
LOW
Why this label 3 signals
Weak or ambiguous signals. Best guess based on catalog patterns; treat as tentative and check the evidence below.
  • 'uncensored' in name/tags but no 'abliterated' marker
  • method not identifiable from author declaration alone
  • may be DPO fine-tune, prompt engineering, or unknown technique
Refusal direction extraction

No specific extraction method could be identified for this model. The producer either did not document it or used a proprietary pipeline.

What is a refusal direction? →
Downloads · lifetime
3K
1K last 30d - stable
Likes
1
Model age
3mo ago
created 2026-06-18
Downloads over time
Now3.4K→from90↑3,732%
01.3K2.5K3.8K90 on Jun 173.4K on Oct 113.4K on Oct 10JunJulAugSepOct
Jun 17 → Oct 11 · 57 snapshots · spans 116 days

Benchmarks

Portrait before abliteration
Benchmarks of the base model as it stood before the refusal-removal operation. Compare with the numbers above to see what the operation cost.
Benchmark Score Source
BBH average 0.48553638604228827 OpenLLM-v2
IFEval instruct 0.7961630695443646 OpenLLM-v2
IFEval-Prompt 0.7208872458410351 OpenLLM-v2
MATH lvl 5 0 OpenLLM-v2
MMLU-Pro 0.4286901595744681 OpenLLM-v2
Entertainment 1.3 UGI
Hazardous 2.9 UGI
Natural Intelligence 15.76 UGI
Political lean -14.7% UGI
Sensitive-Info 15.62 UGI
SocPol 0.8 UGI
UGI 23.75 UGI
Willingness (10) 4 UGI
W10-Adherence 4 UGI
W10-Direct 4 UGI
Writing 29.72 UGI

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
apache-2.0
Languages
en
Quantizations
Q4_K
Tags
gguf qwen2.5 lora fine-tuned quantized q4_k_m ollama llama-cpp text-generation en base_model:Qwen/Qwen2.5-7B-Instruct base_model:adapter:Qwen/Qwen2.5-7B-Instruct

Related

Total size
4.36 GB
Files
3
Quantizations
2
Registered
2026-08-22 13:56
Last updated on HF
2026-06-18 19:50

Files by quantization

Q4_K 1 file 4.36 GB
qwen2.5-7b-uncensored.Q4_K_M.gguf 4.36 GB 33cd6d12 download
Auxiliary files 2 files 6.71 KB
README.md 5.16 KB 618bd102 download
.gitattributes 1.55 KB e32c62e0 download

README current version from Hugging Face


license: apache-2.0
base_model: Qwen/Qwen2.5-7B-Instruct
library_name: gguf
pipeline_tag: text-generation
language:

  • en
    tags:
  • gguf
  • qwen2.5
  • lora
  • fine-tuned
  • quantized
  • q4_k_m
  • ollama
  • llama-cpp
    model_creator: waguriagent
    quantized_by: waguriagent

Qwen2.5-7B Fine-tuned (GGUF Q4_K_M)

A LoRA fine-tune of Qwen2.5-7B-Instruct, merged into the base weights and
quantized to Q4_K_M GGUF for efficient local inference. Runs comfortably on
consumer GPUs with 8 GB VRAM (e.g. RTX 4060) via Ollama or llama.cpp.

Model Details

Property Value
Base model Qwen/Qwen2.5-7B-Instruct
Architecture Qwen2 (7.6B parameters)
Fine-tuning method LoRA (rank 32, alpha 64)
Target modules q_proj, k_proj, v_proj, o_proj, gate_proj, up_proj, down_proj
Precision (training) BF16
Quantization Q4_K_M (4-bit, k-quant medium)
File size ~4.4 GB
Context length 32,768 tokens (inherited from base)
Prompt format ChatML (Qwen2.5)

Quantization

This repository ships the Q4_K_M quant — the recommended balance of size
and quality for most use cases. It keeps the most sensitive weights at higher
precision while compressing the rest to 4-bit, yielding minimal quality loss
versus the full-precision model at roughly a quarter of the size.

Quant Bits Size Quality RAM/VRAM
Q4_K_M ~4.5 4.4 GB Good (recommended) ~6 GB

Usage

Ollama (recommended)

Pull and run directly from the Hub:

ollama run hf.co/waguriagent/qwen2.5-7b-uncensored-gguf:Q4_K_M

Or build a local model from the downloaded GGUF with a Modelfile:

FROM ./qwen2.5-7b-uncensored.Q4_K_M.gguf

TEMPLATE """{{ if .System }}<|im_start|>system
{{ .System }}<|im_end|>
{{ end }}{{ if .Prompt }}<|im_start|>user
{{ .Prompt }}<|im_end|>
{{ end }}<|im_start|>assistant
{{ .Response }}<|im_end|>
"""

PARAMETER stop "<|im_start|>"
PARAMETER stop "<|im_end|>"
PARAMETER temperature 0.7
PARAMETER top_p 0.9
ollama create my-model -f Modelfile
ollama run my-model

llama.cpp

# Download the GGUF
huggingface-cli download waguriagent/qwen2.5-7b-uncensored-gguf \
  qwen2.5-7b-uncensored.Q4_K_M.gguf --local-dir .

# Run interactively
./llama-cli -m qwen2.5-7b-uncensored.Q4_K_M.gguf \
  -p "You are a helpful assistant." -cnv

# Or serve an OpenAI-compatible API
./llama-server -m qwen2.5-7b-uncensored.Q4_K_M.gguf -c 8192

Python (llama-cpp-python)

from llama_cpp import Llama

llm = Llama(
    model_path="qwen2.5-7b-uncensored.Q4_K_M.gguf",
    n_ctx=8192,
    n_gpu_layers=-1,  # offload all layers to GPU
)

out = llm.create_chat_completion(
    messages=[
        {"role": "system", "content": "You are a helpful assistant."},
        {"role": "user", "content": "Explain gradient descent in one paragraph."},
    ],
    temperature=0.7,
)
print(out["choices"][0]["message"]["content"])

Prompt Format

This model uses the Qwen2.5 ChatML template:

<|im_start|>system
{system_prompt}<|im_end|>
<|im_start|>user
{user_message}<|im_end|>
<|im_start|>assistant

Training Procedure

The adapter was trained with PEFT + TRL (no quantization during training),
loading the base model in BF16 and training LoRA adapters directly on an
H100 80GB GPU. Sequence packing was enabled so short samples are concatenated
into dense sequences, maximizing token throughput per step.

Hyperparameters

Hyperparameter Value
Epochs 1
Effective batch size 32 (4 × 8 grad accumulation)
Sequence length 2048
Learning rate 2e-4
LR scheduler Cosine with 3% warmup
Optimizer AdamW (fused)
Weight decay 0.01
Max grad norm 0.3
Attention FlashAttention-2
Gradient checkpointing Enabled
LoRA rank / alpha 32 / 64
LoRA dropout 0

Pipeline

BF16 LoRA train (1 epoch)
   → merge adapter into base weights
   → convert to GGUF F16 (llama.cpp)
   → quantize to Q4_K_M
   → upload to Hugging Face

Hardware

  • GPU: 1× NVIDIA H100 SXM (80 GB HBM3)
  • Training time: ~63 minutes (1 epoch)
  • Framework: PyTorch 2.5.1 + CUDA 12.4, Transformers 4.46, PEFT 0.13, TRL 0.11

Limitations & Bias

  • This is a 4-bit quantized model; expect a small quality degradation versus the
    full-precision base. For maximum quality, use a higher-bit quant or the merged
    FP16 model.
  • The model inherits the knowledge cutoff, biases, and limitations of the
    Qwen2.5-7B-Instruct base model.
  • Fine-tuning on a domain-specific dataset can narrow general-purpose capability.
    Evaluate on your own tasks before production use.

License

Released under the Apache 2.0 license, consistent with the
Qwen2.5-7B-Instruct base model. Review the
base model license
for full terms.

Acknowledgements

README history 3 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-06-18Upload README.md with huggingface_hub8220eac5.2 KB
    Loading...
  2. 2026-06-18Upload README.md with huggingface_hub9068674305 B
    Loading...
  3. 2026-06-18Upload README.md with huggingface_hub9dd838850 B
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration