← back to catalog · registered 2026-08-24 02:02

slevinw/Muse-Glimmer-30B-Heretic-Uncensored-GGUF

slevinw 30B GGUF multimodal 131K ctx
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/slevinw%2FMuse-Glimmer-30B-Heretic-Uncensored-GGUF"
Response includes
  • classification m3
  • files 30
  • benchmarks 11 entries
  • hub_downloads_all_time 1,615
  • author_summary 2 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M3
Primary method

Layer-wise ablation

Applied on top of direct removal inherited from the base model.
Confidence
HIGH
Inherited from base model
Why this label 2 signals
Producer identity confirmed by naming conventions, tags or the model card. This label is very unlikely to change.
  • 'heretic' in model name (Heretic-produced)
  • Heretic uses layer-wise optimization (M3) with underlying direction removal (M1)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
HIGH
Why we say so
name contains 'heretic'; Heretic default extraction is difference-of-means (Arditi 2024)
Downloads · lifetime
2K
875 last 30d - active
Likes
0
Model age
6w ago
created 2026-08-24
Downloads over time
Now1.7K→from195↑791%
1187091.3K1.9K195 on Aug 261.7K on Oct 11AugSepOct
Aug 26 → Oct 11 · 47 snapshots · spans 46 days

Benchmarks

Portrait before abliteration
Benchmarks of the base model as it stood before the refusal-removal operation. Compare with the numbers above to see what the operation cost.
Benchmark Score Source
Entertainment 2.1 UGI
Hazardous 5.9 UGI
Natural Intelligence 37.13 UGI
Political lean -8.3% UGI
Sensitive-Info 38.16 UGI
SocPol 4.2 UGI
UGI 37.94 UGI
Willingness (10) 3.8 UGI
W10-Adherence 4.5 UGI
W10-Direct 3 UGI
Writing 41.03 UGI

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
apache-2.0
Languages
en multilingual
Quantizations
BF16 F16 IQ1 IQ2 IQ3 IQ4 Q2_K Q3_K Q4_K Q5_K Q6_K Q8_0
Tags
transformers gguf llama.cpp quantization vision-language-model image-text-to-text abliterated uncensored heretic darkc0de en multilingual

Related

Total size
412 GB
Files
30
Quantizations
13
Registered
2026-08-24 02:02
Last updated on HF
2026-08-24 01:55

Files by quantization

BF16 1 file 51.9 GB
Muse-Glimmer-30B-Heretic-BF16.gguf 51.9 GB 75788920 download
F16 1 file 51.9 GB
Muse-Glimmer-30B-Heretic-F16.gguf 51.9 GB d1e2fd7f download
Q8_0 1 file 27.6 GB
Muse-Glimmer-30B-Heretic-Q8_0.gguf 27.6 GB 5ddce011 download
Q6_K 1 file 21.3 GB
Muse-Glimmer-30B-Heretic-Q6_K.gguf 21.3 GB 64bec154 download
Q5_K 2 files 36.5 GB
Muse-Glimmer-30B-Heretic-Q5_K_M.gguf 18.5 GB a7bed72b download
Muse-Glimmer-30B-Heretic-Q5_K_S.gguf 18.0 GB 12a5b197 download
Q4_K 2 files 30.8 GB
Muse-Glimmer-30B-Heretic-Q4_K_M.gguf 15.8 GB 7562b95c download
Muse-Glimmer-30B-Heretic-Q4_K_S.gguf 15.0 GB f63b1863 download
IQ4 2 files 29.3 GB
Muse-Glimmer-30B-Heretic-IQ4_NL.gguf 15.0 GB 956d4d76 download
Muse-Glimmer-30B-Heretic-IQ4_XS.gguf 14.3 GB 1d56703c download
Q3_K 3 files 38.1 GB
Muse-Glimmer-30B-Heretic-Q3_K_L.gguf 13.7 GB dfcb1e8c download
Muse-Glimmer-30B-Heretic-Q3_K_M.gguf 12.7 GB 79a7b196 download
Muse-Glimmer-30B-Heretic-Q3_K_S.gguf 11.7 GB 82d8e5db download
IQ3 4 files 45.1 GB
Muse-Glimmer-30B-Heretic-IQ3_M.gguf 11.9 GB 5ace6dcf download
Muse-Glimmer-30B-Heretic-IQ3_S.gguf 11.7 GB 0e43f4bc download
Muse-Glimmer-30B-Heretic-IQ3_XS.gguf 11.1 GB 28fa9f50 download
Muse-Glimmer-30B-Heretic-IQ3_XXS.gguf 10.4 GB 28dd22aa download
Q2_K 2 files 19.3 GB
Muse-Glimmer-30B-Heretic-Q2_K.gguf 9.95 GB 85caf552 download
Muse-Glimmer-30B-Heretic-Q2_K_S.gguf 9.34 GB da60ab7b download
IQ2 4 files 33.2 GB
Muse-Glimmer-30B-Heretic-IQ2_M.gguf 9.17 GB 4ab8c008 download
Muse-Glimmer-30B-Heretic-IQ2_S.gguf 8.50 GB e35b80cb download
Muse-Glimmer-30B-Heretic-IQ2_XS.gguf 8.12 GB e655cdfe download
Muse-Glimmer-30B-Heretic-IQ2_XXS.gguf 7.41 GB 2bbcc1c2 download
IQ1 2 files 12.7 GB
Muse-Glimmer-30B-Heretic-IQ1_M.gguf 6.58 GB e9a38e96 download
Muse-Glimmer-30B-Heretic-IQ1_S.gguf 6.08 GB 6f8b3010 download
Auxiliary files 5 files 14.5 GB
Muse-Glimmer-30B-Heretic-TQ2_0.gguf 7.79 GB 867ff7fa download
Muse-Glimmer-30B-Heretic-TQ1_0.gguf 6.69 GB 3dcf49b0 download
README.md 12.1 KB 0a8586c1 download
LICENSE 11.1 KB d6456956 download
.gitattributes 3.39 KB c94d49a3 download

README current version from Hugging Face


license: apache-2.0
base_model:

  • meta-models/Muse-Glimmer-30B
    library_name: transformers
    tags:
  • llama.cpp
  • gguf
  • quantization
  • vision-language-model
  • image-text-to-text
  • abliterated
  • uncensored
  • heretic
  • darkc0de
    language:
  • en
  • multilingual
    pipeline_tag: image-text-to-text

Muse-Glimmer-30B-Heretic-Uncensored-GGUF

Abstract

This repository disseminates a comprehensive suite of 27 GGUF (GPT-Generated Unified Format) quantizations of Muse-Glimmer-30B-Heretic-Uncensored, an abliterated derivative of Meta's Muse-Glimmer-30B produced with Heretic v1.4.0 by darkc0de. The release spans a memory footprint of 6.53 GB (IQ1_S) to 55.73 GB (F16/BF16), thereby accommodating deployment across heterogeneous hardware, from CPU-only systems to 32 GB VRAM accelerators. All sub-4-bit quantizations were calibrated with an importance matrix (imatrix) to mitigate quality degradation. The fidelity of the Q4_K_S variant relative to the F16 reference is quantified via token-level perplexity on a fixed evaluation set. The model exhibits zero refusals on adult/NSFW content; this property was empirically verified.

1. Model Overview

Property Value
Base model meta-models/Muse-Glimmer-30B (Meta Superintelligence Lab)
Abliteration Heretic v1.4.0 (heretic-project.org); KL divergence 0.0743 relative to base
Architecture Dense Causal Transformer with Perception Encoder (ViT-G/14, ~1.8B)
Language model parameters ~29.6B total (vision encoder included)
Hidden size 6656
Layers 52
Attention GQA; 32 Q / 2 KV heads (ratio 16:1); sliding window 2048 (3:1 local:global pattern)
FFN SwiGLU; intermediate dimension 19,968
Position encoding RoPE (θ = 500,000), local layers only
Vocabulary 202,048 tokens (200K BPE + 2,048 special tokens)
Context length 131,072+ tokens
Modalities Input: text + image · Output: text
Base license Apache 2.0
Knowledge cutoff January 4, 2026

2. Quantization Rationale

The quantization selection is derived from the model's architectural properties and the memory constraints of representative target hardware. The model employs grouped-query attention with 2 KV heads, which renders the KV cache exceptionally compact:

  • 52 layers × 2 KV heads × 128 head_dim = 52 KiB/token at FP16
  • Q8_0 KV cache: 26 KiB/token → 64K context ≈ 1.7 GB
  • Q4_0 KV cache: 13 KiB/token → 64K context ≈ 0.85 GB

2.1 Selection Guidance

Target hardware Recommended quant Model weight size Feasibility at 64K context (incl. KV)
8 GB VRAM (RTX 3060/4060) IQ2_XXS / IQ2_XS 8.0–8.7 GB Feasible
10 GB VRAM IQ3_XXS / IQ3_XS 7.3–12.0 GB Feasible
12 GB VRAM (RTX 3060 12G) IQ3_S / Q3_K_S 12.5 GB Feasible
16 GB VRAM (RTX 4080) Q4_K_S / IQ4_XS 16.1 GB Feasible
24 GB VRAM (RTX 4090) Q5_K_M / Q6_K 19.8–22.9 GB Feasible
32 GB VRAM Q8_0 29.6 GB Feasible
CPU / RAM-only IQ1_S / IQ1_M / TQ1_0 6.5–7.2 GB Feasible within 8 GB RAM
Reference / benchmarking F16 / BF16 55.7 GB —

3. Artifact Inventory

The repository contains 27 GGUF files. Reported sizes correspond to the stored artifacts:

File Size (GB) Description
Muse-Glimmer-30B-Heretic-BF16.gguf 55.73 BF16 reference (benchmark baseline)
Muse-Glimmer-30B-Heretic-F16.gguf 55.73 F16 reference (benchmark baseline)
Muse-Glimmer-30B-Heretic-Q8_0.gguf 29.61 Q8_0 (near-lossless)
Muse-Glimmer-30B-Heretic-Q6_K.gguf 22.87 Q6_K
Muse-Glimmer-30B-Heretic-Q5_K_M.gguf 19.81 Q5_K_M
Muse-Glimmer-30B-Heretic-Q5_K_S.gguf 19.35 Q5_K_S
Muse-Glimmer-30B-Heretic-Q4_K_M.gguf 16.94 Q4_K_M
Muse-Glimmer-30B-Heretic-Q4_K_S.gguf 16.13 Q4_K_S (primary 16 GB VRAM artifact)
Muse-Glimmer-30B-Heretic-IQ4_NL.gguf 16.14 IQ4_NL (imatrix)
Muse-Glimmer-30B-Heretic-IQ4_XS.gguf 15.34 IQ4_XS (imatrix)
Muse-Glimmer-30B-Heretic-Q3_K_L.gguf 14.68 Q3_K_L
Muse-Glimmer-30B-Heretic-Q3_K_M.gguf 13.68 Q3_K_M
Muse-Glimmer-30B-Heretic-IQ3_M.gguf 12.82 IQ3_M (imatrix)
Muse-Glimmer-30B-Heretic-Q3_K_S.gguf 12.51 Q3_K_S
Muse-Glimmer-30B-Heretic-IQ3_S.gguf 12.52 IQ3_S (imatrix)
Muse-Glimmer-30B-Heretic-IQ3_XS.gguf 11.97 IQ3_XS (imatrix)
Muse-Glimmer-30B-Heretic-Q2_K.gguf 10.69 Q2_K
Muse-Glimmer-30B-Heretic-Q2_K_S.gguf 10.03 Q2_K_S (imatrix)
Muse-Glimmer-30B-Heretic-IQ2_M.gguf 9.85 IQ2_M (imatrix)
Muse-Glimmer-30B-Heretic-IQ2_S.gguf 9.13 IQ2_S (imatrix)
Muse-Glimmer-30B-Heretic-IQ2_XS.gguf 8.71 IQ2_XS (imatrix)
Muse-Glimmer-30B-Heretic-TQ2_0.gguf 8.37 TQ2_0 (imatrix; ternary)
Muse-Glimmer-30B-Heretic-IQ2_XXS.gguf 7.96 IQ2_XXS (imatrix)
Muse-Glimmer-30B-Heretic-IQ3_XXS.gguf 7.29 IQ3_XXS (imatrix)
Muse-Glimmer-30B-Heretic-TQ1_0.gguf 7.19 TQ1_0 (imatrix; ternary)
Muse-Glimmer-30B-Heretic-IQ1_M.gguf 7.06 IQ1_M (imatrix)
Muse-Glimmer-30B-Heretic-IQ1_S.gguf 6.53 IQ1_S (imatrix)

4. Methodology

All artifacts were derived from the original safetensors using llama.cpp master (commit 030ebb5). Support for the muse_glimmer architecture is absent from release binaries prior to b10344; consequently, a master-branch build was required.

Step Tool Output Wall time
Conversion convert_hf_to_gguf.py (llama.cpp master) F16 GGUF, 731 tensors ~2.5 min
Imatrix calibration llama-imatrix (wiki.raw, 128 chunks) imatrix.dat ~2.5 h
Quantization llama-quantize.exe (CPU build, 16 threads) 27 quants, 731 tensors ~4 min each

4.1 Conversion

python convert_hf_to_gguf.py darkc0de/Muse-Glimmer-30B-heretic \
  --outfile Muse-Glimmer-30B-Heretic-F16.gguf

4.2 Imatrix Calibration

Importance-matrix calibration was applied to all sub-4-bit quantizations. The calibration corpus comprised 128 chunks drawn from wiki.raw:

llama-imatrix -m Muse-Glimmer-30B-Heretic-F16.gguf \
  -f wiki.raw -o imatrix.dat -c 131072

4.3 Quantization

The --imatrix flag must precede the output specification; the parser does not recognize it in trailing position.

llama-quantize --imatrix imatrix.dat \
  Muse-Glimmer-30B-Heretic-F16.gguf \
  Muse-Glimmer-30B-Heretic-IQ3_XS.gguf IQ3_XS

5. Fidelity Benchmark

Consistent with the repository's scope, this page does not report aggregate capability scores (e.g., MMLU); such evaluations are documented in the upstream card. The benchmark herein quantifies quantization-induced degradation: the fidelity of the Q4_K_S variant relative to the F16 reference, measured as token-level perplexity over a fixed evaluation set.

Experimental setup. wikitext-2 test split; 32 chunks; n_ctx=2048; batch 2048; 16 CPU threads (the build machine was not equipped with a GPU). Both variants were evaluated on an identical token set such that systematic errors cancel.

Model Perplexity (wikitext-2, 32 chunks) Δ vs. F16
F16 (reference) 5.6439 ± 0.07361 —
Q4_K_S 5.7831 ± 0.07601 +0.1392 (+2.47%)

Interpretation. A relative increase of +2.47% in perplexity is at or below the typical expectation for a Q4_K_S-class quantization of a 30B model, indicating successful quantization with minimal fidelity loss. Sub-4-bit quantizations (IQ/TQ series) incur greater fidelity loss in exchange for substantially reduced memory footprints; imatrix calibration recovers a meaningful portion of this gap.

6. Inference Performance

6.1 GPU (NVIDIA RTX 4080 16 GB; CUDA 13.3; llama.cpp b10355)

The model was fully offloaded (-ngl 99) with 16 CPU threads for batch processing. Q4_K_S occupies 15.01 GiB, within the 16 GB VRAM budget.

Test Throughput (t/s)
Prompt processing, pp128 1902.85 ± 134.91
Prompt processing, pp512 2258.22 ± 14.55
Prompt processing, pp2048 2300.36 ± 3.93
Token generation, tg64 38.91 ± 0.03
Token generation, tg256 38.88 ± 0.01

6.2 CPU-only (reference build, 16 threads)

Test Throughput (t/s)
Prompt processing, pp128 28.31 ± 0.44
Token generation, tg64 4.08 ± 0.04

6.3 Operational Notes

  • Generation throughput of ~39 tokens/s on an RTX 4080 is sufficient for real-time agentic interaction.
  • Prompt ingestion of ~2,300 tokens/s renders long-document processing non-bottlenecked.
  • The 15.01 GiB model leaves ~1.4 GB VRAM headroom on a 16 GB card; with --cache-type-k/v q8_0, a 64K context adds ~1.7 GB, remaining within budget.
# GPU inference (llama.cpp b10355+, CUDA build)
llama-cli -m Muse-Glimmer-30B-Heretic-Q4_K_S.gguf \
  -ngl 99 -c 65536 --cache-type-k q8_0 --cache-type-v q8_0

For agentic deployments, the upstream card's recommended sampling configuration applies (temperature = 1.0, top_p = 0.95, top_k = 64; Reasoning strength: high for complex tasks).

7. Limitations and Responsible Use

  • Abliterated model. The model exhibits reduced refusal behavior by design and may generate content that is inappropriate, offensive, or unsafe for many applications. Deployment should incorporate appropriate guardrails and human oversight, particularly in agentic or real-world-action contexts.
  • Age restriction. Not intended for use by individuals under 18 years of age. Deployers operating in environments accessible to minors bear responsibility for regulatory compliance.
  • Scope of evaluation. This repository provides a quantized artifact and its fidelity benchmark; it does not constitute a safety or capability evaluation of the model.
  • Quantization edge cases. Quantized inference may deviate from full precision in edge cases (see Section 5 for aggregate fidelity).
  • Multimodal note. Benchmarks were conducted in text-only mode; the perception encoder was not exercised.

8. License and Attribution

This work is a derivative of:

  1. meta-models/Muse-Glimmer-30B — © Meta Superintelligence Lab, released under Apache 2.0.
  2. darkc0de/Muse-Glimmer-30B-heretic — abliterated derivative produced with Heretic v1.4.0, retaining the Apache 2.0 license.

The GGUF conversion and quantization in this repository inherit the Apache 2.0 license. See LICENSE for the full text.

9. Citation

@misc{meta2026museglimmer,
  author = {Meta Superintelligence Lab},
  title = {Muse Glimmer: A 30B Multimodal Agentic Model for Local Deployment},
  year = {2026},
  url = {https://huggingface.co/meta-models/Muse-Glimmer-30B}
}

@misc{darkc0de2026heretic,
  author = {darkc0de},
  title = {Muse-Glimmer-30B-heretic},
  year = {2026},
  url = {https://huggingface.co/darkc0de/Muse-Glimmer-30B-heretic}
}

@misc{observerx2026gguf,
  author = {0bserverx},
  title = {Muse-Glimmer-30B-Heretic-Uncensored-GGUF},
  year = {2026},
  url = {https://huggingface.co/0bserverx/Muse-Glimmer-30B-Heretic-Uncensored-GGUF}
}

Research use and responsibility

This repository is intended for legitimate research and controlled evaluation, including interpretability, alignment and refusal-behavior analysis, red-team testing, and robustness work. It is not a ready-made production safety layer. If you deploy the model or expose it to other users, you are responsible for adding suitable access controls, moderation, monitoring, and abuse prevention.

Use of these files is subject to the Apache License 2.0 and all applicable laws. You are responsible for how you operate the model and for outputs produced in your environment. To the extent permitted by law, the repository maintainers and upstream authors accept no liability for misuse or resulting harm. Generated outputs are not statements or endorsements by the maintainers, upstream creators, or their organizations.

README history 1 version

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-08-24Duplicate from 0bserverx/Muse-Glimmer-30B-Heretic-Uncensored-GGUFf485ced12.1 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration