base_model: TrevorJS/Muse-Glimmer-30B-uncensored
base_model_relation: quantized
pipeline_tag: image-text-to-text
language:
- en
license: apache-2.0
tags: - abliteration
- uncensored
- muse-glimmer
- gguf
Muse-Glimmer-30B-uncensored (GGUF)
GGUF quantizations of TrevorJS/Muse-Glimmer-30B-uncensored.
Requires a llama.cpp build b10353 or later (#26841, merged 2026-08-10, commit
62bf73d, which addedmuse_glimmerconversion + inference support). Earlier builds will fail to load these files.
Files
| File | Quant | Size |
|---|---|---|
Muse-Glimmer-30B-uncensored-Q4_K_M.gguf |
Q4_K_M | 16.9 GB |
Muse-Glimmer-30B-uncensored-Q8_0.gguf |
Q8_0 | 29.6 GB |
mmproj-Muse-Glimmer-30B-uncensored-F16.gguf |
F16 | 3.8 GB |
The mmproj file is the vision projector and is required for image input; text-only inference does not need it.
Usage
# From HuggingFace (auto-downloads)
llama-server -hf TrevorJS/Muse-Glimmer-30B-uncensored-GGUF -c 8192
# From local file
llama-server -m Muse-Glimmer-30B-uncensored-Q4_K_M.gguf -c 8192
# With image input
llama-server -m Muse-Glimmer-30B-uncensored-Q4_K_M.gguf -c 8192 \
--mmproj mmproj-Muse-Glimmer-30B-uncensored-F16.gguf
Then open http://localhost:8080 for the chat UI.
Details
These are GGUF quantizations of TrevorJS/Muse-Glimmer-30B-uncensored, an abliterated
(uncensored) version of meta-models/Muse-Glimmer-30B.
Refusal behavior has been removed using norm-preserving biprojected abliteration.
Converted via an F16 intermediate (convert_hf_to_gguf.py --outtype f16) then llama-quantize.
See the bf16 model card for full method details,
before/after refusal rates, and agentic validation results.
Source code: TrevorS/muse-glimmer-abliteration
Prior work in the same program: TrevorS/gemma-4-abliteration
Credits
The method is not original to this work — refusal directions (Arditi et al., 2024), norm-preserving biprojection (grimjim), heretic (p-e-w) and the mlabonne / JailbreakBench / AdvBench corpora. Full attribution in the repo README.