← back to catalog · registered 2026-08-22 13:56

agrawal-prateek/qwen2.5-7b-abliterated-gguf

agrawal-prateek Qwen 7B GGUF 33K ctx
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/agrawal-prateek%2Fqwen2.5-7b-abliterated-gguf"
Response includes
  • classification m8
  • files 6
  • hub_downloads_all_time 1,268
  • author_summary 3 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M8
Primary method

Repackaging (quantization)

Applied on top of direct removal inherited from the base model.
Confidence
MEDIUM
Inherited from base model
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=1
  • assume M1 (base ablation) + M8 (GGUF quant) - default when producer unknown
Refusal direction extraction

No specific extraction method could be identified for this model. The producer either did not document it or used a proprietary pipeline.

What is a refusal direction? →
Downloads · lifetime
1K
369 last 30d - stable
Likes
1
Model age
6mo ago
created 2026-04-03
Downloads over time
Now1.4K→from373↑276%
3227161.1K1.5K373 on Apr 151.4K on Oct 11AprMayJunJulAugSepOct
Apr 15 → Oct 11 · 65 snapshots · spans 179 days

Metadata

Quantizations
Q4_K
Tags
gguf qwen2 endpoints_compatible region:us conversational

Related

Total size
4.36 GB
Files
6
Quantizations
2
Registered
2026-08-22 13:56
Last updated on HF
2026-04-03 22:02

Files by quantization

Q4_K 1 file 4.36 GB
qwen2.5-7b-Q4_K_M-abliterated.gguf 4.36 GB 0cdf0398 download
Auxiliary files 5 files 8.19 KB
README.md 5.38 KB 525371d0 download
.gitattributes 1.55 KB 5a629035 download
config.json 555 B 478c3bb4 download
.gitignore 515 B 1b21df54 download
generation_config.json 217 B 1d7677ee download

README current version from Hugging Face

Qwen 2.5 7B Abliterated (GGUF)

GGUF format models for Qwen 2.5 7B with safety abliterations. Ready for local deployment with llama-cpp-python or compatible runtimes.

Models

  1. qwen2.5-7b-abliterated.gguf - Full precision model
  2. qwen2.5-7b-Q4_K_M-abliterated.gguf - Q4_K_M quantized model (4-bit quantization, recommended for most use cases)

Quick Start

from llama_cpp import Llama

# Load the quantized model (recommended)
llm = Llama(
    model_path="qwen2.5-7b-Q4_K_M-abliterated.gguf",
    n_ctx=2048,
    n_threads=4
)

# Generate text
output = llm(
    "Q: What is 2+2? A:",
    max_tokens=50,
    temperature=0.5,
    top_p=0.95
)
print(output['choices'][0]['text'])

Model Details

  • Base Model: Qwen 2.5 7B
  • Format: GGUF (llama-cpp-python compatible)
  • Modifications: Safety abliterations removed
  • Vocabulary Size: 152,064
  • Context Window: 32,768 tokens
  • Recommended Quantization: Q4_K_M (best balance of quality and performance)

Test Suite

Automated test suite for testing Qwen 2.5 7B GGUF models with llama-cpp-python.

Test Coverage

Each model is tested for:

Model Loading

  • File existence and size validation
  • Successful loading into memory
  • Model attributes and metadata

Model Metadata

  • Context size validation
  • Vocabulary size verification
  • Model type detection

Text Generation

  • Basic text completion
  • Token limit enforcement
  • Stop sequence handling
  • Mathematical reasoning
  • Multiple completions generation
  • Code generation (Q4 model only)

Performance

  • Generation speed measurement
  • Memory usage validation
  • Quantization-specific performance metrics

Robustness

  • Empty prompt handling
  • Long prompt processing
  • Special characters support
  • Unicode character support
  • Repeated call stability

Setup

The test suite uses the .venv virtual environment with the following packages:

  • llama-cpp-python - Python bindings for llama.cpp
  • pytest - Testing framework

Running Tests

Option 1: Using the test script (Recommended)

# Run all tests
./scripts/run_tests.sh

# Run tests for full model only
./scripts/run_tests.sh full

# Run tests for Q4 model only
./scripts/run_tests.sh q4

Option 2: Using pytest directly

# Activate virtual environment
source .venv/bin/activate

# Run all tests
pytest tests/ -v

# Run specific test file
pytest tests/test_full_model.py -v
pytest tests/test_q4_model.py -v

# Run specific test class
pytest tests/test_full_model.py::TestModelLoading -v

# Run with output visible
pytest tests/ -v -s

# Run with coverage (if pytest-cov is installed)
pytest --cov=. tests/

Option 3: Run individual test files

source .venv/bin/activate
python tests/test_full_model.py
python test_q4_model.py

Test Files

  • tests/test_full_model.py - Tests for qwen2.5-7b-abliterated.gguf
  • tests/test_q4_model.py - Tests for qwen2.5-7b-Q4_K_M-abliterated.gguf
  • tests/conftest.py - Shared fixtures and utilities
  • tests/pytest.ini - Pytest configuration
  • scripts/run_tests.sh - Test runner script

Test Structure

Each test file contains:

  1. TestModelLoading - Validates model file and loading
  2. TestModelMetadata - Checks model configuration and metadata
  3. TestTextGeneration - Tests various generation scenarios
  4. TestModelPerformance - Measures and validates performance
  5. TestModelRobustness - Tests edge cases and error handling
  6. TestQuantizationComparison - Q4-specific tests (Q4 model only)

Directory Structure

├── README.md                          # This file (model card)
├── qwen2.5-7b-abliterated.gguf       # Full precision model
├── qwen2.5-7b-Q4_K_M-abliterated.gguf # Q4_K_M quantized model
├── config.json                        # Model configuration
├── generation_config.json            # Generation parameters
├── tests/                            # Test suite
│   ├── __init__.py
│   ├── conftest.py
│   ├── pytest.ini
│   ├── test_full_model.py
│   └── test_q4_model.py
├── scripts/                          # Utility scripts
│   └── run_tests.sh
└── .gitignore                        # Git ignore rules

Configuration

Test parameters can be adjusted in tests/conftest.py:

"max_tokens": 50,      # Maximum tokens to generate
"temperature": 0.5,    # Sampling temperature
"top_p": 0.95,         # Top-p sampling
"n_threads": 4,        # CPU threads to use
"n_ctx": 512,          # Context window size

Expected Results

All tests should pass if:

  • Model files exist and are valid GGUF format
  • Sufficient system memory is available
  • llama-cpp-python is properly installed

Performance Metrics

Tests will output performance metrics:

  • Generation time (seconds)
  • Tokens per second

This allows comparison between full precision and quantized models.

Troubleshooting

Model not found

Ensure GGUF files are in the root directory of the repository.

Out of memory

Reduce n_ctx in tests/conftest.py or use the Q4_K_M quantized model.

Slow performance

Increase n_threads in tests/conftest.py to use more CPU cores.

Notes

  • Tests use module-scoped fixtures to load models only once
  • First test run will be slower due to model loading
  • Q4_K_M model should be faster and use less memory than full precision
  • Temperature is set low (0.1-0.5) for more deterministic test results

README history 1 version

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-04-03Add files using upload-large-folder toolce836455.4 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration