license: apache-2.0
base_model: DreamFast/Qwen3.5-4B-Uncensored-HauhauCS-Aggressive-Safetensor-Benchmark
Minas2025 Qwen3.5-4B Conversational (Q6_K GGUF)
Conversational fine-tune of Qwen3.5-4B for chat / roleplay. Trained with Unsloth QLoRA on an AMD RX 6700 XT.
Files
| File | Size | What |
|---|---|---|
Minas2025-Qwen3.5-4B-Conversational-Q6_K.gguf |
~3.5 GB | Main model, Q6_K quant. Load this. |
Minas2025-Qwen3.5-4B-Conversational-mmproj-F16.gguf |
~672 MB | Vision projector ("eyes"). Needed only for image input; text chat works without it. |
Base model
DreamFast/Qwen3.5-4B-Uncensored-HauhauCS-Aggressive-Safetensor-Benchmark
Training recipe
- Method: QLoRA, rank 16, alpha 32, dropout 0.0, all attention+MLP projections
(q_proj, k_proj, v_proj, o_proj, gate_proj, up_proj, down_proj) - Optimizer: AdamW 8-bit, LR 1.5e-4, cosine schedule, 10 warmup steps
- 1150 steps, batch size 1, gradient accumulation 8, max seq length 2048
- Dataset: ~11,400 conversational rows (train) / 600 rows (eval)
- Final train loss ~0.92, eval loss still falling at end of run (no overfit)
- Export: merged to F16, quantized to Q6_K with llama.cpp
Usage (llama.cpp / LM Studio / Ollama)
Load the Q6_K file. Chat template is included. Works fine at default
settings; context up to 2048 tokens (trained length).
Notes
- This is a community fine-tune for conversational/roleplay use.
- Q6_K is near-lossless vs F16 for this size; use Q8_0 or F16 only if you
have VRAM to spare and want the last fraction of quality.