base_model: Blackfrost-AI/Muse-Glimmer-30B-Abliterated-BF16
base_model_relation: quantized
license: apache-2.0
pipeline_tag: image-text-to-text
tags:
- gguf
- quantized
- llama.cpp
- muse-glimmer
- muse_glimmer
- abliterated
- multimodal
- conversational
Muse Glimmer 30B Abliterated — Custom GGUF Quant
Custom GGUF quantization ofBlackfrost-AI/Muse-Glimmer-30B-Abliterated-BF16.
This quant was produced with an experimental quantization workflow intended to improve the
quality/size tradeoff relative to a conventional quant of the same source model.
Model
- Base model:
Blackfrost-AI/Muse-Glimmer-30B-Abliterated-BF16 - Format: GGUF
- Approximate file size: 11.9 GB
- Quantization processing time: ~2.3 hours
- Target runtimes: llama.cpp and compatible GGUF frontends
- License: Apache-2.0, following the upstream model
- Recommended vram A 16 Gigabyte card should run this comfortably.
Local benchmark results
The following results are from local factual-summary testing and are not a standardized
leaderboard evaluation.
| Variant | Approx. size | Local factual score |
|---|---|---|
| Q8 reference | — | ~72% |
| Standard IQ4_XS | ~14.2 GB | ~67% |
| This custom quant | ~11.9 GB | ~72% |
In this local test, the custom quant matched the approximate Q8 factual score while being
about 2.3 GB smaller than the tested standard IQ4_XS representation.
Results may vary with prompt, sampling settings, runtime, hardware, and benchmark methodology.
Usage
With a recent llama.cpp build:
llama-cli -m ./YOUR_MODEL_FILE.gguf --jinja -ngl 999
Or with the llama.cpp server:
llama-server -m ./YOUR_MODEL_FILE.gguf --jinja -ngl 999
Replace YOUR_MODEL_FILE.gguf with the actual filename in this repository.
Multimodal / vision support
Muse Glimmer is an image-text-to-text model.
If your llama.cpp build requires a separate multimodal projector, use a compatiblemmproj*.gguf for this model lineage:
llama-server \
-m ./YOUR_MODEL_FILE.gguf \
--mmproj ./YOUR_MMPROJ_FILE.gguf \
--jinja \
-ngl 999
If this repository does not include an mmproj file, users will need to obtain or generate
a compatible projector separately for image input.
Reproducibility
When comparing this quant against other variants, please report:
- exact GGUF filename/revision
- llama.cpp version or commit
- inference backend
- GPU/CPU and available memory
- context size
- temperature, top-p, top-k, and seed
- benchmark prompts
- number of generations/runs
- scoring methodology
Notes
- This repository contains a quantized derivative, not a retrained model.
- The source is the abliterated BF16 derivative linked above.
- Quantization can affect factual recall, reasoning, formatting, and multimodal performance differently.
- Local benchmark results should be independently reproduced before drawing broad conclusions.
Attribution
Source model:
Blackfrost-AI/Muse-Glimmer-30B-Abliterated-BF16
Please also review the upstream model card for its original Muse Glimmer lineage, intended use,
limitations, and safety information.
License
Apache-2.0, following the source repository.