← back to catalog · registered 2026-09-27 14:57

The-freezer/Muse-Glimmer-30B-Abliterated-GGUF

The-freezer 30B GGUF multimodal second-order
curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/The-freezer%2FMuse-Glimmer-30B-Abliterated-GGUF"
Response includes
  • classification unknown
  • files 17
  • author_summary 3 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
?
Primary method

Unclassified

No clear signals of an abliteration technique in this model.
Confidence
UNKNOWN
Why this label 1 signal
No classification signals present. This may not be an abliterated model at all - it could be a repackaging, a merge with unrelated goals, or unrelated content that mentions the term.
  • no classification signals present (no abliterated, uncensored, or known producer/method markers)
Refusal direction extraction

No specific extraction method could be identified for this model. The producer either did not document it or used a proprietary pipeline.

What is a refusal direction? →
Downloads · 30-day
23K
Likes
50
Model age
6w ago
created 2026-08-10

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
apache-2.0
Quantizations
F16 Q2_K Q3_K Q4_K Q5_K Q6_K Q8_0
Tags
gguf muse-glimmer abliterated quantized agentic multimodal llama.cpp dflash experimental image-text-to-text conversational base_model:Blackfrost-AI/Muse-Glimmer-30B-Abliterated-BF16

Related

Total size
157 GB
Files
17
Quantizations
8
Registered
2026-09-27 14:57
Last updated on HF
2026-08-15 06:43

Files by quantization

Q8_0 2 files 29.5 GB
Muse-Glimmer-30B-Abliterated-Q8_0.gguf 27.6 GB fc7f9c5e download
mmproj-Muse-Glimmer-30B-Abliterated-Q8_0.gguf 1.91 GB ca930c0d download
Q6_K 1 file 21.3 GB
Muse-Glimmer-30B-Abliterated-Q6_K.gguf 21.3 GB f5487c50 download
Q5_K 2 files 36.5 GB
Muse-Glimmer-30B-Abliterated-Q5_K_M.gguf 18.5 GB 21073f1a download
Muse-Glimmer-30B-Abliterated-Q5_K_S.gguf 18.0 GB a05d442b download
Q4_K 4 files 33.6 GB
Muse-Glimmer-30B-Abliterated-Q4_K_M.gguf 15.8 GB 1316a7e2 download
Muse-Glimmer-30B-Abliterated-Q4_K_S.gguf 15.0 GB 97964122 download
dflash-Muse-Glimmer-30B-Abliterated-Q4_K_M.gguf 1.52 GB 1b58b92b download
mmproj-Muse-Glimmer-30B-Abliterated-Q4_K_M.gguf 1.30 GB acd4058c download
Q3_K 2 files 24.4 GB
Muse-Glimmer-30B-Abliterated-Q3_K_M.gguf 12.7 GB 39045df4 download
Muse-Glimmer-30B-Abliterated-Q3_K_S.gguf 11.7 GB 51330230 download
Q2_K 1 file 9.95 GB
Muse-Glimmer-30B-Abliterated-Q2_K.gguf 9.95 GB f756cbe4 download
F16 2 files 8.36 GB
dflash-Muse-Glimmer-30B-Abliterated-F16.gguf 4.77 GB 83564914 download
mmproj-Muse-Glimmer-30B-Abliterated-F16.gguf 3.58 GB c2e08b8b download
Auxiliary files 3 files 21.6 KB
chat_template.jinja 12.1 KB 07fc68d0 download
README.md 6.93 KB c15b81bd download
.gitattributes 2.63 KB bd659811 download

README current version from Hugging Face


license: apache-2.0
base_model:

  • Blackfrost-Research/Muse-Glimmer-30B-Abliterated-BF16
    tags:
  • muse-glimmer
  • gguf
  • abliterated
  • quantized
  • agentic
  • multimodal
  • llama.cpp
  • dflash
  • experimental
    pipeline_tag: image-text-to-text
    library_name: gguf

[!IMPORTANT]

Improvement update — August 15, 2026

This release now includes compact abliterated Q4_K_M companions: a 1.63 GB DFlash drafter and a 1.40 GB multimodal projector, matching Meta's consumer-hardware footprint while preserving this model's modified weights. The complete text quant ladder has also been refreshed with Meta's post-release Jinja correction, which normalizes Reasoning effort to Reasoning strength and prevents duplicate reasoning directives. Text generation, image input, and DFlash speculative decoding were validated together on current llama.cpp.

MUSE-GLIMMER-30B-ABLITERATED-GGUF

GGUF quant ladder of the abliterated Muse Glimmer 30B · runs local on one GPU or CPU

Built by Blackfrost · Las Vegas, NV

✅ All quants live

The full text quant ladder (Q2_K → Q8_0), compact Q4_K_M and full-precision vision projectors, and compact Q4_K_M and full-precision DFlash drafters are uploaded — see the Files tab.

⚠️ EXPERIMENTAL

Same-day arch, quantized. Expect sharp edges — decode, coherence, tool-parse, serve edge cases under load. Please open a Community discussion with loader/version, quant, prompt, sampling, and failure mode. Real repros get fixed faster.


Refusal benchmark — R1-HARMFUL-BENCH-450

Measured on the abliterated parent (GGUF quants inherit this behavior):

Metric Result
True refusal (harmful, n=300) 0 / 300 = 0.0%
True refusal (full 450) 0 / 450 = 0.0%
Substring-harmful 0 / 300
Substring-all 2 / 450 (XSTest false positives)
Errors 0

The weight change removes refusals cleanly — no measured true refusals across the full 450-prompt suite.


Why this model exists

Muse Glimmer is Meta Superintelligence Labs' 30B agentic, on-device model. This is the abliterated build — refusal behavior removed via a Blackfrost weight-change process — packaged as GGUF for llama.cpp, so it runs on a single consumer GPU or CPU, fully offline. The local footprint is the product.


Specifications

Architecture muse-glimmer — dense, 52 layers, hidden 6656, GQA (32 q / 2 kv), sliding-window attention, + vision tower
Base meta-models/Muse-Glimmer-30B — Meta, Apache-2.0
Transform Abliterated — refusal behavior removed via a Blackfrost weight-change process; multimodal capability intact
Formats GGUF — Q2_K, Q3_K_S, Q3_K_M, Q4_K_S, Q4_K_M, Q5_K_S, Q5_K_M, Q6_K, Q8_0
Context 131,072
Spec-decode DFlash drafter — --spec-type draft-dflash --spec-draft-n-max 15
Default persona Ships with the "AI assistant" system template baked in

Quant ladder

quant size recommended for
Q2_K 10.0 GB smallest, quality trade-off
Q3_K_S 11.7 GB very tight VRAM
Q3_K_M 12.7 GB tight VRAM
Q4_K_S 15.0 GB 16 GB cards
Q4_K_M 15.8 GB default — balanced, fits 24 GB
Q5_K_S 18.0 GB higher quality
Q5_K_M 18.5 GB strong quality/size balance
Q6_K 21.3 GB near-lossless
Q8_0 27.6 GB max fidelity

Vision & speculative-decode files

Load a text quant plus an mmproj projector for image input:

file size purpose
mmproj-Muse-Glimmer-30B-Abliterated-Q4_K_M.gguf 1.40 GB vision projector — compact, recommended
mmproj-Muse-Glimmer-30B-Abliterated-F16.gguf 3.6 GB vision projector — full precision
mmproj-Muse-Glimmer-30B-Abliterated-Q8_0.gguf 1.9 GB vision projector — compact
dflash-Muse-Glimmer-30B-Abliterated-Q4_K_M.gguf 1.63 GB abliterated DFlash drafter — compact, recommended
dflash-Muse-Glimmer-30B-Abliterated-F16.gguf 4.8 GB DFlash drafter — speculative decoding

Serving (llama.cpp) — confirmed settings

Requires llama.cpp b10353 or newer with llama-server. DFlash runs under llama-server only — it shares the target model's context, so it does not work in llama-cli.

Recommended — with DFlash speculative decoding (~1.6× faster, identical output):

llama-server \
  -m  Muse-Glimmer-30B-Abliterated-Q8_0.gguf \
  -md dflash-Muse-Glimmer-30B-Abliterated-Q4_K_M.gguf \
  --spec-type draft-dflash --spec-draft-n-max 15 \
  -ngl 999 -ngld 999 -fa on --jinja \
  --host 0.0.0.0 --port 8080 -c 16384 \
  --temp 1.0 --top-p 0.95 --top-k 64
  • Plain (no drafter): drop -md, --spec-type, --spec-draft-n-max, and -ngld.
  • Multimodal (image input): add --mmproj mmproj-Muse-Glimmer-30B-Abliterated-Q4_K_M.gguf.
  • One-command kit: deploy/serve.sh auto-downloads + serves; full guide in deploy/DEPLOYMENT.md.

Confirmed settings

  • Sampling: temperature 1.0, top_p 0.95, top_k 64 (Meta). Steer depth with a Reasoning strength: low/medium/high/xhigh system line.
  • --jinja is required. The refreshed template accepts an OpenAI-style Reasoning effort: <level> line, normalizes it, and does not inject a conflicting second directive.
  • Do not stop on <|eom|>. Use <|end_of_text|> and <|eot|> as stop tokens.
  • max_tokens ≥ 1024 — heavy thinker; small budgets return empty content because the reasoning channel consumes them. Reasoning arrives in reasoning_content, the answer in content.
  • --spec-draft-n-max 15 — DFlash block size (trained 16, clamped).
  • Flash attention: -fa on for peak speed; switch to -fa off if the load hangs on a brand-new GPU paired with an older CUDA toolkit.

Measured performance

1× NVIDIA RTX PRO 6000 (Blackwell), Q8_0, -fa off:

config decode tok/s speedup
baseline ~46 1.0×
+ DFlash ~73 1.6×

Speedup rises with -fa on and structured/code output (Meta reports up to 3.1× on an RTX 5090).


Built by Blackfrost · Las Vegas, NV. Not affiliated with Meta.

Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.