library_name: gguf
tags:
- merlina
- grimoire
- text-generation
- orpo
- gguf
base_model: - nbeerbower/Huihui-Qwen3.5-27B-abliterated-Athanorlite-ORPO-v2
Huihui-Qwen3.5-27B-abliterated-Athanorlite-ORPO-v2-GGUF
GGUF quantizations of nbeerbower/Huihui-Qwen3.5-27B-abliterated-Athanorlite-ORPO-v2.
This is a multimodal (vision-language) model. You need both a text model GGUF and the mmproj file for full functionality.
Available Quantizations
| Quant | Size | BPW | Description |
|---|---|---|---|
| Q8_0 | 27 GB | 8.50 | Best quality, near-lossless |
| Q6_K | 21 GB | 6.57 | Great quality, good size balance |
| Q4_K_M | 16 GB | 4.92 | Recommended default |
| Q3_K_M | 13 GB | 3.86 | For constrained VRAM |
Vision Projector (required for multimodal)
| File | Size | Type |
|---|---|---|
| mmproj-F16 | 885 MB | F16 |
Hardware Recommendations
| Setup | Recommended Quant |
|---|---|
| 1x 48 GB (A6000, RTX 6000 Ada) | Q8_0 |
| 2x 24 GB (RTX 3090/4090) | Q8_0 split across GPUs |
| 1x 24 GB (RTX 3090/4090) | Q6_K |
| 2x 16 GB (RTX 4060 Ti) | Q4_K_M or Q6_K split |
| 1x 16 GB (RTX 4060 Ti) | Q3_K_M |
VRAM usage = text model + mmproj (885 MB) + KV cache (varies with context length). Leave at least 2-4 GB headroom for KV cache and overhead.
Usage
llama.cpp CLI
# Text-only
llama-cli -m Huihui-Qwen3.5-27B-abliterated-Athanorlite-ORPO-v2-Q4_K_M.gguf -p "Hello!"
# With vision (image input)
llama-mtmd-cli \
-m Huihui-Qwen3.5-27B-abliterated-Athanorlite-ORPO-v2-Q4_K_M.gguf \
--mmproj Huihui-Qwen3.5-27B-abliterated-Athanorlite-ORPO-v2-mmproj-F16.gguf \
--image photo.jpg \
-p "Describe this image."
llama.cpp Server
llama-server \
-m Huihui-Qwen3.5-27B-abliterated-Athanorlite-ORPO-v2-Q4_K_M.gguf \
--mmproj Huihui-Qwen3.5-27B-abliterated-Athanorlite-ORPO-v2-mmproj-F16.gguf \
--port 8080
Multi-GPU split
# Example: 2x 24GB GPUs
llama-server \
-m Huihui-Qwen3.5-27B-abliterated-Athanorlite-ORPO-v2-Q8_0.gguf \
--mmproj Huihui-Qwen3.5-27B-abliterated-Athanorlite-ORPO-v2-mmproj-F16.gguf \
-ngl 99 --tensor-split 1,1
About
ORPO fine-tune of huihui-ai/Huihui-Qwen3.5-27B-abliterated on schneewolflabs/Athanorlite-DPO. The original upload had broken state dict keys from a PEFT merge bug; the v2 safetensors model has corrected key naming and restored multimodal/MTP weights. See the v2 model card for details.
