license: apache-2.0
base_model: richardyoung/granite-4.2-3b-heretic
datasets:
- minas2025/warm-chat-12k
pipeline_tag: text-generation
tags: - gguf
- granite
- conversational
- chat
- uncensored
- finetune
- lora
- qlora
Minas2025 Granite 4.2 3B WarmChat (Q6_K GGUF)
Conversational fine-tune of Granite 4.2 3B for warm, natural chat. Trained with Unsloth QLoRA on an AMD RX 6700 XT.
Files
| File | Size | What |
|---|---|---|
Minas2025-Granite-4.2-3B-WarmChat-Q6_K.gguf |
~3.0 GB | Main model, Q6_K quant. Load this. |
Base model
richardyoung/granite-4.2-3b-heretic
Training recipe
- Method: QLoRA, rank 16, alpha 16, dropout 0.0; all attention+MLP projections
(q_proj, k_proj, v_proj, o_proj, gate_proj, up_proj, down_proj) - Optimizer: AdamW 8-bit, LR 1.5e-4, cosine schedule, 10 warmup steps
- 1150 steps, batch size 1, gradient accumulation 8, max seq length 2048
- Data: ~11,400 conversational rows (train) / 600 rows (eval)
- Eval loss falling through the full run, no overfit
- Export: adapters merged to F16, quantized to Q6_K with llama.cpp
Usage (llama.cpp / LM Studio / Ollama)
Load the Q6_K file directly. Chat template included. Context up to 2048 tokens (trained length).
Notes
- Community fine-tune for warm conversational chat.
- Q6_K is near-lossless vs F16 at this size.