language:
- en
- zh
license: apache-2.0
tags: - qwen3_5
- gguf
- vision
- mtp
- qwen3.8
- distilled
- heretic
- abliterated
- uncensored
- conversational
base_model: insraq/Qwen3.5-4B-EmperoAI-Qwen3.8-Distill-Heretic-Abliterated
Qwen3.8-4B-Distill-Heretic-Abliterated-GGUF
GGUF version of insraq/Qwen3.5-4B-EmperoAI-Qwen3.8-Distill-Heretic-Abliterated, with both MTP (Multi-Token Prediction) speculative decoding layers and full vision weights preserved.
Why this exists
Two problems with existing GGUF conversions of this model:
Missing MTP layers: The original HF→GGUF conversion using llama.cpp's
convert_hf_to_gguf.pyproduced a broken GGUF — the metadata declared 33 layers (32 transformer + 1 MTP) but only 32 layers' tensors were actually exported, causingcheck_tensor_dims: tensor 'blk.32.attn_norm.weight' not founderrors in LM Studio and other loaders.Missing vision: insraq's MTP GGUF (Qwen3.5-4B-EmperoAI-Qwen3.8-Distill-Heretic-Abliterated-MTP-GGUF) fixed the MTP issue but only converted the text portion — the vision tower was dropped, leaving a blind model.
This repo combines the best of both: MTP-enabled text model (from insraq's properly converted GGUF) + vision projector (mmproj, converted via llama.cpp's --mmproj flag).
Model lineage
Qwen3.5-4B (base, Alibaba)
└─ empero-ai/Qwen3.8-4B-Distill (distilled from Qwen3.8, vision-capable)
└─ insraq/Qwen3.5-4B-EmperoAI-Qwen3.8-Distill-Heretic-Abliterated (Heretic v1.4.0 abliteration)
└─ This repo (MTP + vision GGUF)
Files
| File | Size | Description |
|---|---|---|
model-Q4_K_M.gguf |
2.8 GB | Text model with MTP layers, Q4_K_M quantized |
mmproj-f16.gguf |
641 MB | Vision projector, F16 (unquantized for quality) |
Conversion details
| Parameter | Value |
|---|---|
| Text model source | insraq/Qwen3.5-4B-EmperoAI-Qwen3.8-Distill-Heretic-Abliterated-MTP-GGUF (Q4_K_M, single-pass quantization from BF16) |
| Vision projector source | insraq/Qwen3.5-4B-EmperoAI-Qwen3.8-Distill-Heretic-Abliterated (BF16 Safetensors) |
| Vision projector tool | llama.cpp convert_hf_to_gguf.py --mmproj --outtype f16 |
| MTP layers | 1 (nextn_predict_layers: 1, blk.32.* tensors present) |
| Architecture | qwen3_5 (Qwen3_5ForConditionalGeneration) |
Architecture parameters
| Parameter | Value |
|---|---|
| block_count | 33 (32 transformer + 1 MTP) |
| embedding_length | 2560 |
| context_length | 262144 |
| attention.head_count | 16 |
| attention.head_count_kv | 4 |
| full_attention_interval | 4 |
| nextn_predict_layers | 1 |
Usage
With llama.cpp server
llama-server \
--mmproj mmproj-f16.gguf \
-m model-Q4_K_M.gguf \
--host 0.0.0.0 --port 8080 \
-ngl 99 --jinja
With LM Studio
- Load
model-Q4_K_M.ggufas the main model - Set
mmproj-f16.ggufas the vision projector (Projector/Embedding setting) - GPU offload: max layers
With Ollama
ollama run hf.co/yachen4ever/Qwen3.8-4B-Distill-Heretic-Abliterated-GGUF:Q4_K_M
API (OpenAI-compatible)
curl -X POST http://localhost:8080/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{
"model": "qwen3.8-4b",
"messages": [{
"role": "user",
"content": [
{"type": "image_url", "image_url": {"url": "data:image/jpeg;base64,..."} },
{"type": "text", "text": "Describe this image."}
]
}]
}'
Abliteration
Heretic v1.4.0 abliteration: refusals 6/100 vs 99/100 for the original model. Zero refusals on tested prompts (fictional bank heist, lock picking tutorial, forbidden love poem).
License
Apache 2.0 (inherited from Qwen3.5/Qwen3.8 base models)
Acknowledgments
- empero-ai — original Qwen3.8 distilled models with vision
- insraq — Heretic v1.4.0 abliterated version + MTP GGUF conversion
- ggml-org/llama.cpp — GGUF format and conversion tools