← back to catalog · registered 2026-08-22 13:56

Gavvvin/Muse-Glimmer-30B-Abliterated-OID-IQ3_XXS_XL-quant

Gavvvin 30B GGUF multimodal second-order 131K ctx
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/Gavvvin%2FMuse-Glimmer-30B-Abliterated-OID-IQ3_XXS_XL-quant"
Response includes
  • classification m8
  • files 3
  • hub_downloads_all_time 423
  • author_summary 3 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M8
Primary method

Repackaging (quantization)

Applied on top of direct removal inherited from the base model.
Confidence
MEDIUM
Inherited from base model
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=1
  • assume M1 (base ablation) + M8 (GGUF quant) - default when producer unknown
Refusal direction extraction

No specific extraction method could be identified for this model. The producer either did not document it or used a proprietary pipeline.

What is a refusal direction? →
Downloads · lifetime
423
67 last 30d - stable
Likes
2
Model age
8w ago
created 2026-08-15
Downloads over time
Now437→from273↑60%
265328391453273 on Aug 19437 on Oct 11AugSepOct
Aug 19 → Oct 11 · 48 snapshots · spans 53 days

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
apache-2.0
Quantizations
IQ3
Tags
gguf quantized llama.cpp muse-glimmer muse_glimmer abliterated multimodal conversational image-text-to-text base_model:Blackfrost-AI/Muse-Glimmer-30B-Abliterated-BF16 base_model:quantized:Blackfrost-AI/Muse-Glimmer-30B-Abliterated-BF16 license:apache-2.0

Related

Total size
11.9 GB
Files
3
Quantizations
2
Registered
2026-08-22 13:56
Last updated on HF
2026-08-15 15:20

Files by quantization

IQ3 1 file 11.9 GB
Muse-Glimmer-30B-Abliterated-OID-IQ3_XXS_XL.gguf 11.9 GB 3b023ecd download
Auxiliary files 2 files 4.88 KB
README.md 3.31 KB c31a4ba7 download
.gitattributes 1.57 KB 1a73f850 download

README current version from Hugging Face


base_model: Blackfrost-AI/Muse-Glimmer-30B-Abliterated-BF16
base_model_relation: quantized
license: apache-2.0
pipeline_tag: image-text-to-text
tags:

  • gguf
  • quantized
  • llama.cpp
  • muse-glimmer
  • muse_glimmer
  • abliterated
  • multimodal
  • conversational

Muse Glimmer 30B Abliterated — Custom GGUF Quant

Custom GGUF quantization of
Blackfrost-AI/Muse-Glimmer-30B-Abliterated-BF16.

This quant was produced with an experimental quantization workflow intended to improve the
quality/size tradeoff relative to a conventional quant of the same source model.

Model

  • Base model: Blackfrost-AI/Muse-Glimmer-30B-Abliterated-BF16
  • Format: GGUF
  • Approximate file size: 11.9 GB
  • Quantization processing time: ~2.3 hours
  • Target runtimes: llama.cpp and compatible GGUF frontends
  • License: Apache-2.0, following the upstream model
  • Recommended vram A 16 Gigabyte card should run this comfortably.

Local benchmark results

The following results are from local factual-summary testing and are not a standardized
leaderboard evaluation.

Variant Approx. size Local factual score
Q8 reference — ~72%
Standard IQ4_XS ~14.2 GB ~67%
This custom quant ~11.9 GB ~72%

In this local test, the custom quant matched the approximate Q8 factual score while being
about 2.3 GB smaller than the tested standard IQ4_XS representation.

Results may vary with prompt, sampling settings, runtime, hardware, and benchmark methodology.

Usage

With a recent llama.cpp build:

llama-cli -m ./YOUR_MODEL_FILE.gguf --jinja -ngl 999

Or with the llama.cpp server:

llama-server -m ./YOUR_MODEL_FILE.gguf --jinja -ngl 999

Replace YOUR_MODEL_FILE.gguf with the actual filename in this repository.

Multimodal / vision support

Muse Glimmer is an image-text-to-text model.

If your llama.cpp build requires a separate multimodal projector, use a compatible
mmproj*.gguf for this model lineage:

llama-server \
  -m ./YOUR_MODEL_FILE.gguf \
  --mmproj ./YOUR_MMPROJ_FILE.gguf \
  --jinja \
  -ngl 999

If this repository does not include an mmproj file, users will need to obtain or generate
a compatible projector separately for image input.

Reproducibility

When comparing this quant against other variants, please report:

  • exact GGUF filename/revision
  • llama.cpp version or commit
  • inference backend
  • GPU/CPU and available memory
  • context size
  • temperature, top-p, top-k, and seed
  • benchmark prompts
  • number of generations/runs
  • scoring methodology

Notes

  • This repository contains a quantized derivative, not a retrained model.
  • The source is the abliterated BF16 derivative linked above.
  • Quantization can affect factual recall, reasoning, formatting, and multimodal performance differently.
  • Local benchmark results should be independently reproduced before drawing broad conclusions.

Attribution

Source model:

Blackfrost-AI/Muse-Glimmer-30B-Abliterated-BF16

Please also review the upstream model card for its original Muse Glimmer lineage, intended use,
limitations, and safety information.

License

Apache-2.0, following the source repository.

README history 2 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-08-15Update README.md7f71db93.3 KB
    Loading...
  2. 2026-08-15Upload README.md46644393.2 KB
    Loading...

Discussions 1 thread

  1. 2026-08-15Quants are uploading.open1 💬#1
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration