← back to catalog · registered 2026-08-22 13:56

Blackfrost-AI/Muse-Glimmer-30B-Abliterated-GGUF

Blackfrost-AI 30B GGUF multimodal second-order
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/Blackfrost-AI%2FMuse-Glimmer-30B-Abliterated-GGUF"
Response includes
  • classification m8
  • files 17
  • author_summary 19 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M8
Primary method

Repackaging (quantization)

Applied on top of direct removal inherited from the base model.
Confidence
MEDIUM
Inherited from base model
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=1
  • assume M1 (base ablation) + M8 (GGUF quant) - default when producer unknown
Refusal direction extraction

No specific extraction method could be identified for this model. The producer either did not document it or used a proprietary pipeline.

What is a refusal direction? →
Downloads · 30-day
23K
Likes
50
Model age
2mo ago
created 2026-08-10
Downloads over time
Now108.7K→from0↑0%
039.8K79.7K119.5K0 on Aug 12108.7K on Sep 27AugSep
Aug 12 → Sep 27 · 36 snapshots · spans 46 days

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Variants by this author 4 formats · 24K downloads combined

The same weights this author released in different packaging. Pick the format that matches your runtime.

Metadata

License
apache-2.0
Quantizations
F16 Q2_K Q3_K Q4_K Q5_K Q6_K Q8_0
Tags
gguf muse-glimmer abliterated quantized agentic multimodal llama.cpp dflash experimental image-text-to-text conversational base_model:Blackfrost-AI/Muse-Glimmer-30B-Abliterated-BF16

Related

Total size
157 GB
Files
17
Quantizations
8
Registered
2026-08-22 13:56
Last updated on HF
2026-08-15 06:43

Files by quantization

Q8_0 2 files 29.5 GB
Muse-Glimmer-30B-Abliterated-Q8_0.gguf 27.6 GB fc7f9c5e download
mmproj-Muse-Glimmer-30B-Abliterated-Q8_0.gguf 1.91 GB ca930c0d download
Q6_K 1 file 21.3 GB
Muse-Glimmer-30B-Abliterated-Q6_K.gguf 21.3 GB f5487c50 download
Q5_K 2 files 36.5 GB
Muse-Glimmer-30B-Abliterated-Q5_K_M.gguf 18.5 GB 21073f1a download
Muse-Glimmer-30B-Abliterated-Q5_K_S.gguf 18.0 GB a05d442b download
Q4_K 4 files 33.6 GB
Muse-Glimmer-30B-Abliterated-Q4_K_M.gguf 15.8 GB 1316a7e2 download
Muse-Glimmer-30B-Abliterated-Q4_K_S.gguf 15.0 GB 97964122 download
dflash-Muse-Glimmer-30B-Abliterated-Q4_K_M.gguf 1.52 GB 1b58b92b download
mmproj-Muse-Glimmer-30B-Abliterated-Q4_K_M.gguf 1.30 GB acd4058c download
Q3_K 2 files 24.4 GB
Muse-Glimmer-30B-Abliterated-Q3_K_M.gguf 12.7 GB 39045df4 download
Muse-Glimmer-30B-Abliterated-Q3_K_S.gguf 11.7 GB 51330230 download
Q2_K 1 file 9.95 GB
Muse-Glimmer-30B-Abliterated-Q2_K.gguf 9.95 GB f756cbe4 download
F16 2 files 8.36 GB
dflash-Muse-Glimmer-30B-Abliterated-F16.gguf 4.77 GB 83564914 download
mmproj-Muse-Glimmer-30B-Abliterated-F16.gguf 3.58 GB c2e08b8b download
Auxiliary files 3 files 21.6 KB
chat_template.jinja 12.1 KB 07fc68d0 download
README.md 6.93 KB c15b81bd download
.gitattributes 2.63 KB bd659811 download

README current version from Hugging Face


license: apache-2.0
base_model:

  • Blackfrost-Research/Muse-Glimmer-30B-Abliterated-BF16
    tags:
  • muse-glimmer
  • gguf
  • abliterated
  • quantized
  • agentic
  • multimodal
  • llama.cpp
  • dflash
  • experimental
    pipeline_tag: image-text-to-text
    library_name: gguf

[!IMPORTANT]

Improvement update — August 15, 2026

This release now includes compact abliterated Q4_K_M companions: a 1.63 GB DFlash drafter and a 1.40 GB multimodal projector, matching Meta's consumer-hardware footprint while preserving this model's modified weights. The complete text quant ladder has also been refreshed with Meta's post-release Jinja correction, which normalizes Reasoning effort to Reasoning strength and prevents duplicate reasoning directives. Text generation, image input, and DFlash speculative decoding were validated together on current llama.cpp.

MUSE-GLIMMER-30B-ABLITERATED-GGUF

GGUF quant ladder of the abliterated Muse Glimmer 30B · runs local on one GPU or CPU

Built by Blackfrost · Las Vegas, NV

✅ All quants live

The full text quant ladder (Q2_K → Q8_0), compact Q4_K_M and full-precision vision projectors, and compact Q4_K_M and full-precision DFlash drafters are uploaded — see the Files tab.

⚠️ EXPERIMENTAL

Same-day arch, quantized. Expect sharp edges — decode, coherence, tool-parse, serve edge cases under load. Please open a Community discussion with loader/version, quant, prompt, sampling, and failure mode. Real repros get fixed faster.


Refusal benchmark — R1-HARMFUL-BENCH-450

Measured on the abliterated parent (GGUF quants inherit this behavior):

Metric Result
True refusal (harmful, n=300) 0 / 300 = 0.0%
True refusal (full 450) 0 / 450 = 0.0%
Substring-harmful 0 / 300
Substring-all 2 / 450 (XSTest false positives)
Errors 0

The weight change removes refusals cleanly — no measured true refusals across the full 450-prompt suite.


Why this model exists

Muse Glimmer is Meta Superintelligence Labs' 30B agentic, on-device model. This is the abliterated build — refusal behavior removed via a Blackfrost weight-change process — packaged as GGUF for llama.cpp, so it runs on a single consumer GPU or CPU, fully offline. The local footprint is the product.


Specifications

Architecture muse-glimmer — dense, 52 layers, hidden 6656, GQA (32 q / 2 kv), sliding-window attention, + vision tower
Base meta-models/Muse-Glimmer-30B — Meta, Apache-2.0
Transform Abliterated — refusal behavior removed via a Blackfrost weight-change process; multimodal capability intact
Formats GGUF — Q2_K, Q3_K_S, Q3_K_M, Q4_K_S, Q4_K_M, Q5_K_S, Q5_K_M, Q6_K, Q8_0
Context 131,072
Spec-decode DFlash drafter — --spec-type draft-dflash --spec-draft-n-max 15
Default persona Ships with the "AI assistant" system template baked in

Quant ladder

quant size recommended for
Q2_K 10.0 GB smallest, quality trade-off
Q3_K_S 11.7 GB very tight VRAM
Q3_K_M 12.7 GB tight VRAM
Q4_K_S 15.0 GB 16 GB cards
Q4_K_M 15.8 GB default — balanced, fits 24 GB
Q5_K_S 18.0 GB higher quality
Q5_K_M 18.5 GB strong quality/size balance
Q6_K 21.3 GB near-lossless
Q8_0 27.6 GB max fidelity

Vision & speculative-decode files

Load a text quant plus an mmproj projector for image input:

file size purpose
mmproj-Muse-Glimmer-30B-Abliterated-Q4_K_M.gguf 1.40 GB vision projector — compact, recommended
mmproj-Muse-Glimmer-30B-Abliterated-F16.gguf 3.6 GB vision projector — full precision
mmproj-Muse-Glimmer-30B-Abliterated-Q8_0.gguf 1.9 GB vision projector — compact
dflash-Muse-Glimmer-30B-Abliterated-Q4_K_M.gguf 1.63 GB abliterated DFlash drafter — compact, recommended
dflash-Muse-Glimmer-30B-Abliterated-F16.gguf 4.8 GB DFlash drafter — speculative decoding

Serving (llama.cpp) — confirmed settings

Requires llama.cpp b10353 or newer with llama-server. DFlash runs under llama-server only — it shares the target model's context, so it does not work in llama-cli.

Recommended — with DFlash speculative decoding (~1.6× faster, identical output):

llama-server \
  -m  Muse-Glimmer-30B-Abliterated-Q8_0.gguf \
  -md dflash-Muse-Glimmer-30B-Abliterated-Q4_K_M.gguf \
  --spec-type draft-dflash --spec-draft-n-max 15 \
  -ngl 999 -ngld 999 -fa on --jinja \
  --host 0.0.0.0 --port 8080 -c 16384 \
  --temp 1.0 --top-p 0.95 --top-k 64
  • Plain (no drafter): drop -md, --spec-type, --spec-draft-n-max, and -ngld.
  • Multimodal (image input): add --mmproj mmproj-Muse-Glimmer-30B-Abliterated-Q4_K_M.gguf.
  • One-command kit: deploy/serve.sh auto-downloads + serves; full guide in deploy/DEPLOYMENT.md.

Confirmed settings

  • Sampling: temperature 1.0, top_p 0.95, top_k 64 (Meta). Steer depth with a Reasoning strength: low/medium/high/xhigh system line.
  • --jinja is required. The refreshed template accepts an OpenAI-style Reasoning effort: <level> line, normalizes it, and does not inject a conflicting second directive.
  • Do not stop on <|eom|>. Use <|end_of_text|> and <|eot|> as stop tokens.
  • max_tokens ≥ 1024 — heavy thinker; small budgets return empty content because the reasoning channel consumes them. Reasoning arrives in reasoning_content, the answer in content.
  • --spec-draft-n-max 15 — DFlash block size (trained 16, clamped).
  • Flash attention: -fa on for peak speed; switch to -fa off if the load hangs on a brand-new GPU paired with an older CUDA toolkit.

Measured performance

1× NVIDIA RTX PRO 6000 (Blackwell), Q8_0, -fa off:

config decode tok/s speedup
baseline ~46 1.0×
+ DFlash ~73 1.6×

Speedup rises with -fa on and structured/code output (Meta reports up to 3.1× on an RTX 5090).


Built by Blackfrost · Las Vegas, NV. Not affiliated with Meta.

README history 10 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-08-15Update README.md83e7dc76.9 KB
    Loading...
  2. 2026-08-15Update README.md410d8c57.1 KB
    Loading...
  3. 2026-08-15Document companion and Jinja improvements1c46b847 KB
    Loading...
  4. 2026-08-11Update README.md4de1e925.9 KB
    Loading...
  5. 2026-08-11Update README.mde6420a65.9 KB
    Loading...
  6. 2026-08-11Upload README.md with huggingface_hubc4109085.8 KB
    Loading...
  7. 2026-08-10Update README.md1626ed25.4 KB
    Loading...
  8. 2026-08-10Update README.mdb0d34ab5.3 KB
    Loading...
  9. 2026-08-10Upload README.md with huggingface_hubf8116e05.9 KB
    Loading...
  10. 2026-08-10Collapse history — purge BF16 master reference from all commits64052f44.8 KB
    Loading...

Discussions 1 thread

  1. 2026-08-11DFlash & mmproj sizes — kquant versions planned? (16GB VRAM user)closed5 💬#1
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration