base_model: Jiunsong/supergemma4-e4b-abliterated
library_name: gguf
pipeline_tag: text-generation
tags:
- llama.cpp
- gguf-my-repo
- gemma
supergemma4-e4b-abliterated-GGUF
This repository contains GGUF format model files for Jiunsong/supergemma4-e4b-abliterated.
These files were quantized using llama.cpp to provide various compressed versions of the model for local inference on lower-VRAM hardware.
Available Quantizations
The following quantization formats are provided to allow you to balance between memory usage, speed, and quality:
- Q8_0: 8-bit quantization. Very close to the original F16 model in quality, but requires the most memory.
- Q6_K: 6-bit quantization. Excellent balance of quality and size.
- Q5_K_M: 5-bit quantization. Good middle ground for lower-end hardware while retaining strong coherence.
- Q4_K_M: 4-bit quantization. The recommended standard for everyday local use. High speed, low memory, minor perplexity hit.
- Q4_K_S: 4-bit quantization (small). Slightly smaller and faster than Q4_K_M, with a minor drop in accuracy.
- Q3_K_M: 3-bit quantization. Extreme compression for very limited hardware. Noticeable degradation in complex reasoning, but suitable for basic text generation.
How to Use
You can run these GGUF models using any UI or terminal tool that supports llama.cpp, such as:
Command Line with llama.cpp
If you have llama.cpp compiled locally, you can run the model directly from the terminal.
# Example using the Q4_K_M quant
./llama-cli -m supergemma4-Q4_K_M.gguf -p "You are a helpful assistant. How do I write a Python script?" -n 512 -c 2048