Qwen 2.5 7B Abliterated (GGUF)
GGUF format models for Qwen 2.5 7B with safety abliterations. Ready for local deployment with llama-cpp-python or compatible runtimes.
Models
- qwen2.5-7b-abliterated.gguf - Full precision model
- qwen2.5-7b-Q4_K_M-abliterated.gguf - Q4_K_M quantized model (4-bit quantization, recommended for most use cases)
Quick Start
from llama_cpp import Llama
# Load the quantized model (recommended)
llm = Llama(
model_path="qwen2.5-7b-Q4_K_M-abliterated.gguf",
n_ctx=2048,
n_threads=4
)
# Generate text
output = llm(
"Q: What is 2+2? A:",
max_tokens=50,
temperature=0.5,
top_p=0.95
)
print(output['choices'][0]['text'])
Model Details
- Base Model: Qwen 2.5 7B
- Format: GGUF (llama-cpp-python compatible)
- Modifications: Safety abliterations removed
- Vocabulary Size: 152,064
- Context Window: 32,768 tokens
- Recommended Quantization: Q4_K_M (best balance of quality and performance)
Test Suite
Automated test suite for testing Qwen 2.5 7B GGUF models with llama-cpp-python.
Test Coverage
Each model is tested for:
Model Loading
- File existence and size validation
- Successful loading into memory
- Model attributes and metadata
Model Metadata
- Context size validation
- Vocabulary size verification
- Model type detection
Text Generation
- Basic text completion
- Token limit enforcement
- Stop sequence handling
- Mathematical reasoning
- Multiple completions generation
- Code generation (Q4 model only)
Performance
- Generation speed measurement
- Memory usage validation
- Quantization-specific performance metrics
Robustness
- Empty prompt handling
- Long prompt processing
- Special characters support
- Unicode character support
- Repeated call stability
Setup
The test suite uses the .venv virtual environment with the following packages:
llama-cpp-python- Python bindings for llama.cpppytest- Testing framework
Running Tests
Option 1: Using the test script (Recommended)
# Run all tests
./scripts/run_tests.sh
# Run tests for full model only
./scripts/run_tests.sh full
# Run tests for Q4 model only
./scripts/run_tests.sh q4
Option 2: Using pytest directly
# Activate virtual environment
source .venv/bin/activate
# Run all tests
pytest tests/ -v
# Run specific test file
pytest tests/test_full_model.py -v
pytest tests/test_q4_model.py -v
# Run specific test class
pytest tests/test_full_model.py::TestModelLoading -v
# Run with output visible
pytest tests/ -v -s
# Run with coverage (if pytest-cov is installed)
pytest --cov=. tests/
Option 3: Run individual test files
source .venv/bin/activate
python tests/test_full_model.py
python test_q4_model.py
Test Files
tests/test_full_model.py- Tests for qwen2.5-7b-abliterated.gguftests/test_q4_model.py- Tests for qwen2.5-7b-Q4_K_M-abliterated.gguftests/conftest.py- Shared fixtures and utilitiestests/pytest.ini- Pytest configurationscripts/run_tests.sh- Test runner script
Test Structure
Each test file contains:
- TestModelLoading - Validates model file and loading
- TestModelMetadata - Checks model configuration and metadata
- TestTextGeneration - Tests various generation scenarios
- TestModelPerformance - Measures and validates performance
- TestModelRobustness - Tests edge cases and error handling
- TestQuantizationComparison - Q4-specific tests (Q4 model only)
Directory Structure
├── README.md # This file (model card)
├── qwen2.5-7b-abliterated.gguf # Full precision model
├── qwen2.5-7b-Q4_K_M-abliterated.gguf # Q4_K_M quantized model
├── config.json # Model configuration
├── generation_config.json # Generation parameters
├── tests/ # Test suite
│ ├── __init__.py
│ ├── conftest.py
│ ├── pytest.ini
│ ├── test_full_model.py
│ └── test_q4_model.py
├── scripts/ # Utility scripts
│ └── run_tests.sh
└── .gitignore # Git ignore rules
Configuration
Test parameters can be adjusted in tests/conftest.py:
"max_tokens": 50, # Maximum tokens to generate
"temperature": 0.5, # Sampling temperature
"top_p": 0.95, # Top-p sampling
"n_threads": 4, # CPU threads to use
"n_ctx": 512, # Context window size
Expected Results
All tests should pass if:
- Model files exist and are valid GGUF format
- Sufficient system memory is available
- llama-cpp-python is properly installed
Performance Metrics
Tests will output performance metrics:
- Generation time (seconds)
- Tokens per second
This allows comparison between full precision and quantized models.
Troubleshooting
Model not found
Ensure GGUF files are in the root directory of the repository.
Out of memory
Reduce n_ctx in tests/conftest.py or use the Q4_K_M quantized model.
Slow performance
Increase n_threads in tests/conftest.py to use more CPU cores.
Notes
- Tests use module-scoped fixtures to load models only once
- First test run will be slower due to model loading
- Q4_K_M model should be faster and use less memory than full precision
- Temperature is set low (0.1-0.5) for more deterministic test results