license: apache-2.0
base_model: Qwen/Qwen2.5-7B-Instruct
library_name: gguf
pipeline_tag: text-generation
language:
- en
tags: - gguf
- qwen2.5
- lora
- fine-tuned
- quantized
- q4_k_m
- ollama
- llama-cpp
model_creator: waguriagent
quantized_by: waguriagent
Qwen2.5-7B Fine-tuned (GGUF Q4_K_M)
A LoRA fine-tune of Qwen2.5-7B-Instruct, merged into the base weights and
quantized to Q4_K_M GGUF for efficient local inference. Runs comfortably on
consumer GPUs with 8 GB VRAM (e.g. RTX 4060) via Ollama or llama.cpp.
Model Details
| Property | Value |
|---|---|
| Base model | Qwen/Qwen2.5-7B-Instruct |
| Architecture | Qwen2 (7.6B parameters) |
| Fine-tuning method | LoRA (rank 32, alpha 64) |
| Target modules | q_proj, k_proj, v_proj, o_proj, gate_proj, up_proj, down_proj |
| Precision (training) | BF16 |
| Quantization | Q4_K_M (4-bit, k-quant medium) |
| File size | ~4.4 GB |
| Context length | 32,768 tokens (inherited from base) |
| Prompt format | ChatML (Qwen2.5) |
Quantization
This repository ships the Q4_K_M quant — the recommended balance of size
and quality for most use cases. It keeps the most sensitive weights at higher
precision while compressing the rest to 4-bit, yielding minimal quality loss
versus the full-precision model at roughly a quarter of the size.
| Quant | Bits | Size | Quality | RAM/VRAM |
|---|---|---|---|---|
| Q4_K_M | ~4.5 | 4.4 GB | Good (recommended) | ~6 GB |
Usage
Ollama (recommended)
Pull and run directly from the Hub:
ollama run hf.co/waguriagent/qwen2.5-7b-uncensored-gguf:Q4_K_M
Or build a local model from the downloaded GGUF with a Modelfile:
FROM ./qwen2.5-7b-uncensored.Q4_K_M.gguf
TEMPLATE """{{ if .System }}<|im_start|>system
{{ .System }}<|im_end|>
{{ end }}{{ if .Prompt }}<|im_start|>user
{{ .Prompt }}<|im_end|>
{{ end }}<|im_start|>assistant
{{ .Response }}<|im_end|>
"""
PARAMETER stop "<|im_start|>"
PARAMETER stop "<|im_end|>"
PARAMETER temperature 0.7
PARAMETER top_p 0.9
ollama create my-model -f Modelfile
ollama run my-model
llama.cpp
# Download the GGUF
huggingface-cli download waguriagent/qwen2.5-7b-uncensored-gguf \
qwen2.5-7b-uncensored.Q4_K_M.gguf --local-dir .
# Run interactively
./llama-cli -m qwen2.5-7b-uncensored.Q4_K_M.gguf \
-p "You are a helpful assistant." -cnv
# Or serve an OpenAI-compatible API
./llama-server -m qwen2.5-7b-uncensored.Q4_K_M.gguf -c 8192
Python (llama-cpp-python)
from llama_cpp import Llama
llm = Llama(
model_path="qwen2.5-7b-uncensored.Q4_K_M.gguf",
n_ctx=8192,
n_gpu_layers=-1, # offload all layers to GPU
)
out = llm.create_chat_completion(
messages=[
{"role": "system", "content": "You are a helpful assistant."},
{"role": "user", "content": "Explain gradient descent in one paragraph."},
],
temperature=0.7,
)
print(out["choices"][0]["message"]["content"])
Prompt Format
This model uses the Qwen2.5 ChatML template:
<|im_start|>system
{system_prompt}<|im_end|>
<|im_start|>user
{user_message}<|im_end|>
<|im_start|>assistant
Training Procedure
The adapter was trained with PEFT + TRL (no quantization during training),
loading the base model in BF16 and training LoRA adapters directly on an
H100 80GB GPU. Sequence packing was enabled so short samples are concatenated
into dense sequences, maximizing token throughput per step.
Hyperparameters
| Hyperparameter | Value |
|---|---|
| Epochs | 1 |
| Effective batch size | 32 (4 × 8 grad accumulation) |
| Sequence length | 2048 |
| Learning rate | 2e-4 |
| LR scheduler | Cosine with 3% warmup |
| Optimizer | AdamW (fused) |
| Weight decay | 0.01 |
| Max grad norm | 0.3 |
| Attention | FlashAttention-2 |
| Gradient checkpointing | Enabled |
| LoRA rank / alpha | 32 / 64 |
| LoRA dropout | 0 |
Pipeline
BF16 LoRA train (1 epoch)
→ merge adapter into base weights
→ convert to GGUF F16 (llama.cpp)
→ quantize to Q4_K_M
→ upload to Hugging Face
Hardware
- GPU: 1× NVIDIA H100 SXM (80 GB HBM3)
- Training time: ~63 minutes (1 epoch)
- Framework: PyTorch 2.5.1 + CUDA 12.4, Transformers 4.46, PEFT 0.13, TRL 0.11
Limitations & Bias
- This is a 4-bit quantized model; expect a small quality degradation versus the
full-precision base. For maximum quality, use a higher-bit quant or the merged
FP16 model. - The model inherits the knowledge cutoff, biases, and limitations of the
Qwen2.5-7B-Instruct base model. - Fine-tuning on a domain-specific dataset can narrow general-purpose capability.
Evaluate on your own tasks before production use.
License
Released under the Apache 2.0 license, consistent with the
Qwen2.5-7B-Instruct base model. Review the
base model license
for full terms.
Acknowledgements
- Qwen Team for the Qwen2.5 base model
- llama.cpp for GGUF conversion and quantization
- Hugging Face PEFT & TRL for the training stack