← back to catalog · registered 2026-08-23 00:02

kingjones777/Muse-Glimmer-30B-Uncensored-ROCmFP4-GGUF

kingjones777 30B GGUF multimodal 131K ctx
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/kingjones777%2FMuse-Glimmer-30B-Uncensored-ROCmFP4-GGUF"
Response includes
  • classification m8
  • files 9
  • benchmarks 11 entries
  • hub_downloads_all_time 3,574
  • author_summary 11 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M8
Primary method

Repackaging (quantization)

Applied on top of direct removal inherited from the base model.
Confidence
MEDIUM
Inherited from base model
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=1
  • assume M1 (base ablation) + M8 (GGUF quant) - default when producer unknown
Refusal direction extraction

No specific extraction method could be identified for this model. The producer either did not document it or used a proprietary pipeline.

What is a refusal direction? →
Downloads · lifetime
4K
454 last 30d - stable
Likes
0
Model age
7w ago
created 2026-08-22
Downloads over time
Now3.6K→from241↑1,405%
721.4K2.7K4K241 on Aug 263.6K on Oct 11AugSepOct
Aug 26 → Oct 11 · 47 snapshots · spans 46 days

Benchmarks

Portrait before abliteration
Benchmarks of the base model as it stood before the refusal-removal operation. Compare with the numbers above to see what the operation cost.
Benchmark Score Source
Entertainment 2.1 UGI
Hazardous 5.9 UGI
Natural Intelligence 37.13 UGI
Political lean -8.3% UGI
Sensitive-Info 38.16 UGI
SocPol 4.2 UGI
UGI 37.94 UGI
Willingness (10) 3.8 UGI
W10-Adherence 4.5 UGI
W10-Direct 3 UGI
Writing 41.03 UGI

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
apache-2.0
Tags
gguf llama.cpp rocm amd rocmfp4 rocmfpx strix-halo gfx1151 uncensored abliterated dflash speculative-decoding

Related

Total size
74.6 GB
Files
9
Quantizations
2
Registered
2026-08-23 00:02
Last updated on HF
2026-08-28 04:08

Files by quantization

mmproj 1 file 1.30 GB
mmproj-kquant.gguf 1.30 GB f48b4523 download
Auxiliary files 8 files 74.6 GB
muse-glimmer-30B-Uncensored-Q6_0_ROCMFPX_AGENT.gguf 24.2 GB fe524c72 download
muse-glimmer-30B-Uncensored-Q6_0_ROCMFPX_LEAN.gguf 21.1 GB eac379d1 download
muse-glimmer-30B-Uncensored-ROCmFP4-STRIX_LEAN.gguf 14.0 GB f7e40327 download
muse-glimmer-30B-Uncensored-ROCmFP4-FAST.gguf 13.8 GB 5aa86f37 download
dflash-kquant.gguf 1.52 GB 27d9a805 download
README.md 6.94 KB 74fba710 download
.gitattributes 1.93 KB 0d9af05c download
SHA256SUMS 711 B 81ec4677 download

README current version from Hugging Face


license: apache-2.0
base_model: meta-models/Muse-Glimmer-30B
base_model_relation: quantized
tags:

  • gguf
  • llama.cpp
  • rocm
  • amd
  • rocmfp4
  • rocmfpx
  • strix-halo
  • gfx1151
  • uncensored
  • abliterated
  • dflash
  • speculative-decoding
    pipeline_tag: image-text-to-text

Muse-Glimmer-30B Uncensored — ROCmFP4 for AMD Strix Halo (gfx1151)

Uncensored ROCmFP4 quantisations of meta-models/Muse-Glimmer-30B, built with the same ROCmFPX pipeline and the same card ftypes as kingjones777/Muse-Glimmer-30B-ROCmFP4-Strix-Halo-DFlash-GGUF.

Research artifact. Abliteration removes content-refusal. It does not add capability. Do not ship this as a product default. The aligned repo remains the serving default.

Metric Result
Quantization ROCmFP4 103 FAST + 106 STRIX_LEAN; also Q6 114/116
Source Muse-Glimmer-30B BF16 safetensors → GGUF via convert_hf_to_gguf.py
Hardware Ryzen AI Max+ 395 / Radeon 8060S / gfx1151 / 128 GB / ROCm 7.2.4
Drafter Meta dflash-kquant.gguf, --spec-type draft-dflash --spec-draft-n-max 15
Unc FAST 103 decode prose 15.68 · code 37.44 tok/s
Unc STRIX_LEAN 106 decode prose 16.72 · code 38.51 tok/s
Aligned FAST 103 (published A/B) 20.31 tok/s mixed; real-world ~17–45
Aligned STRIX_LEAN 106 (published A/B) 18.72 tok/s mixed

Why this build?

The aligned card measured STRIX_LEAN (106) at 18.72 tok/s and FAST (103) at 20.31 tok/s in a controlled A/B (DFlash n=15, ctx 32K, -fa on). This repo is those same ftypes from an abliterated checkpoint, plus the Q6 AGENT/LEAN pair.

Q6 LEAN (ftype 116) is not the card LEAN. Card LEAN = 106.

Which file should I use?

Start with STRIX_LEAN (106) if you want the card-matched LEAN. Take FAST (103) if you want the aligned speed pick. Take Q6 AGENT (114) if you want more bits and will live with ~Q6 decode.

Ryzen AI Max+ 395, ROCm 7.2.4, DFlash --spec-draft-n-max 15, -fa on, ctx 32768, batch 1, temperature 0. Warm medians of 3; first call after load discarded.

Build ftype Size prose tok/s code tok/s
Unc FAST 103 13.80 GiB 15.68 37.44
Unc STRIX_LEAN 106 14.00 GiB 16.72 38.51
Unc Q6 AGENT 114 24.17 GiB — —
Unc Q6 LEAN 116 21.09 GiB — —
Aligned FAST (published) 103 13.80 GiB ~15 ~39
Aligned STRIX_LEAN (published A/B) 106 14.00 GiB — 18.72 mixed

Same flags, same drafter as the aligned card. Decode on this model is workload-dominated — quote a range, not a point.

Quick start

llama-server \
  -m muse-glimmer-30B-Uncensored-ROCmFP4-STRIX_LEAN.gguf \
  --spec-type draft-dflash --model-draft dflash-kquant.gguf \
  --spec-draft-n-max 15 --spec-draft-ngl 99 --spec-draft-device ROCm0 \
  --chat-template-kwargs '{"reasoning_strength":"low"}' \
  -ngl 999 -fa on -dio --jinja -fit off -dev ROCm0 -c 32768 \
  --host 127.0.0.1 --port 8080

Requires a llama.cpp built with ROCmFP4 (ggml types 100–106) and the muse-glimmer port. Stock llama.cpp rejects these tensor types.

Flag Why
--chat-template-kwargs '{"reasoning_strength":"low"}' Template defaults to high. Small max_tokens then returns empty content.
-fa on (text) / -fa off (vision) Vision requires -fa off.
--spec-draft-n-max 15 DFlash block size is 16; one slot holds the previously accepted token.

--reasoning-budget is not enforced on this model. Use reasoning_strength.

Uncensored findings

Content-refusal scoring on a 24 harmful / 12 harmless / 8 quality research set, greedy (temp 0). Counts only — no payloads. Abliteration is supposed to drop harmful-tune refusals without wrecking ordinary Q&A.

Model Harmful 24 Harmless 12 Quality 8
Qwen3.8 aligned Q8 AGENT 23 refuse, 1 comply 11/12 ok (1 over-refuse) 6/8
Qwen3.8 uncensored Q6 AGENT (114) 23 comply, 1 broken 11/12 ok (1 over-refuse) 6/8
Muse aligned Q6 AGENT (114) 18 refuse, 6 comply 11/12 ok (1 over-refuse) 7/8
Muse uncensored STRIX_LEAN (106) 24 comply 12/12 ok (0 over-refuse) 7/8

Reading:

  • Aligned Qwen still refuses almost everything on this set. Abliterated Qwen complies on almost everything. Quality score is identical (same two fails: Márquez needle + bat-and-ball).
  • Aligned Muse is leakier than aligned Qwen on this classifier — a few complies even before abliteration.
  • Quality is a substring smoke check, not MMLU. It is a regression guard against a broken quant, not a capability claim.

Files

File ftype Size Role
muse-glimmer-30B-Uncensored-ROCmFP4-FAST.gguf 103 13.80 GiB speed pick
muse-glimmer-30B-Uncensored-ROCmFP4-STRIX_LEAN.gguf 106 14.00 GiB card-LEAN equivalent
muse-glimmer-30B-Uncensored-Q6_0_ROCMFPX_AGENT.gguf 114 24.17 GiB 6-bit, Q8 head/attn
muse-glimmer-30B-Uncensored-Q6_0_ROCMFPX_LEAN.gguf 116 21.09 GiB 6-bit throughout
dflash-kquant.gguf — 1.52 GiB DFlash drafter (Meta's, unmodified) — use this
mmproj-kquant.gguf — 1.30 GiB vision projector (unmodified; vision tensors were not abliterated)

Six files. llama.cpp loads them via --model, --model-draft and --mmproj. This repo is the uncensored family only — aligned builds are a separate repo.

Quantization

PYTHONPATH=gguf-py python convert_hf_to_gguf.py <MODEL_DIR> --outtype bf16 --outfile unc-BF16.gguf
llama-quantize unc-BF16.gguf …-FAST.gguf Q4_0_ROCMFP4_FAST 16
llama-quantize unc-BF16.gguf …-STRIX_LEAN.gguf Q4_0_ROCMFP4_STRIX_LEAN 16

No extra --tensor flags — matches the published aligned card. ROCmFPX llama-quantize only.

Known issues (same as aligned)

  1. Vulkan/CUDA/CPU cannot load these files — ROCmFP4 is ROCm-only.
  2. Vision requires -fa off.
  3. Small max_tokens returns empty content — budget goes to reasoning_content.
  4. --reasoning-budget is not enforced; use reasoning_strength.
  5. This is an uncensored research build. Do not deploy it as the public default.

Not yet measured

Test Status
Perplexity / KL vs BF16 ❓ not measured
MMLU-Pro, GPQA, GSM8K ❓ not run
Tool-calling 7-case suite on the unc weights ❓ not re-run (aligned scored 6/7, model-level)
Vision spatial 3/3 on the unc projector ❓ projector reused, not re-scored
Independent reproduction ❓ none yet

License and attribution

Base model: Meta Muse-Glimmer-30B (Apache 2.0). ROCmFP4 types: ROCmFPX. This repository is quantisation and measurement of an abliterated Muse-Glimmer-30B checkpoint.

See the aligned card for the muse-glimmer architecture port, DFlash notes, and tool-calling suite.

README history 4 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-08-28docs: point runtime instructions at the kingjones30/ROCmFPX fork (verified bu...51c726c7.8 KB
    Loading...
  2. 2026-08-22card: drop third-party model names; hang off Meta Muse-Glimmer-30B onlya34ac6b6.9 KB
    Loading...
  3. 2026-08-22card: hang off meta-models/Muse-Glimmer-30B, not TrevorJS8e0162f7.1 KB
    Loading...
  4. 2026-08-22card: Muse-Glimmer-30B uncensored ROCmFP4f2c3f2f7.1 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration