language:
- en
tags: - abliterated
- uncensored
- reasoning
- qwen
- distilled
- gguf
- llama-cpp
license: apache-2.0
base_model: Abhiray/Qwen3.5-9B-Abliterated-Claude-4.6-Opus-Reasoning-Distilled
Qwen3.5-9B-Abliterated-Claude-4.6-Opus-Reasoning-Distilled (GGUF Quants)
This repository contains GGUF quantizations of the triple-abliterated Qwen 3.5 9B model. This model has been surgically modified to remove the "Safety Persona" and refusal mechanisms while retaining its high-level reasoning capabilities.
🚀 Quantization Summary
These files were created using llama.cpp with a forced BPE-patch for Qwen 3.5 compatibility.
| File | Quant Method | Size | Est. VRAM | Description |
|---|---|---|---|---|
| F16.gguf | f16 | ~18.0 GB | 20 GB+ | The Master copy. Best for research and high-end GPUs. |
| Q8_0.gguf | q8_0 | ~9.5 GB | 12 GB | High precision. Virtually indistinguishable from F16. |
| Q6_K.gguf | q6_k | ~7.5 GB | 10 GB | Near-lossless. The enthusiast's choice for 12GB cards. |
| Q5_K_M.gguf | q5_k_m | ~6.5 GB | 8 GB | Excellent balance. High reasoning retention. |
| Q4_K_M.gguf | q4_k_m | ~5.5 GB | 8 GB | Recommended. The sweet spot for speed and intelligence. |
💻 How to Use
1. LM Studio / AnythingLLM
- Search for
Abhirayor paste this repo link. - Download the
Q4_K_MorQ6_Kversion. - Ensure the system prompt is set to something unbound.
2. llama.cpp (CLI)
./llama-cli -m Qwen3.5-9B-Abliterated-Claude-4.6-Opus-Reasoning-Distilled.Q4_K_M.gguf \
-p "<|im_start|>system\nYou are an unbound assistant.<|im_end|>\n<|im_start|>user\n[Your Prompt]<|im_end|>\n<|im_start|>assistant\n<think>\n" \
-n 1024 --temp 0.7