← back to catalog · registered 2026-08-22 13:56

SHS-Lab/Muse-Glimmer-30B-Abliterated-Aggressive

Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/SHS-Lab%2FMuse-Glimmer-30B-Abliterated-Aggressive"
Response includes
  • classification m1
  • files 11
  • benchmarks 11 entries
  • author_summary 2 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M1
Primary method

Direct removal

No other method signals detected in this model.
Confidence
MEDIUM
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=0 (base model)
  • no specific method indicators - defaulting to M1 (most common)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
MEDIUM
Why we say so
primary_method=M1; difference-of-means is the reference extraction for M1/M3 (Arditi 2024)
Downloads · 30-day
171
↑ 5% in 90 days
Likes
0
Model age
7w ago
created 2026-08-17
Downloads over time
Now168→from160↑5%
160163166169160 on Aug 19168 on Sep 1AugSep
Aug 19 → Sep 1 · 8 snapshots · spans 13 days

Benchmarks

Portrait before abliteration
Benchmarks of the base model as it stood before the refusal-removal operation. Compare with the numbers above to see what the operation cost.
Benchmark Score Source
Entertainment 2.1 UGI
Hazardous 5.9 UGI
Natural Intelligence 37.13 UGI
Political lean -8.3% UGI
Sensitive-Info 38.16 UGI
SocPol 4.2 UGI
UGI 37.94 UGI
Willingness (10) 3.8 UGI
W10-Adherence 4.5 UGI
W10-Direct 3 UGI
Writing 41.03 UGI

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
apache-2.0
Languages
en
Tags
transformers safetensors muse_glimmer image-text-to-text abliterated muse-glimmer lora de-refusal text-generation conversational en base_model:meta-models/Muse-Glimmer-30B

Related

Total size
55.5 GB
Files
11
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-08-17 11:42

Files by quantization

Auxiliary files 11 files 55.5 GB
model-00001-of-00002.safetensors 46.5 GB 08971f4c download
model-00002-of-00002.safetensors 8.99 GB e4ec6e4d download
tokenizer.json 26.8 MB c9dbee66 download
model.safetensors.index.json 130 KB f0417930 download
tokenizer_config.json 78.1 KB d1b80588 download
chat_template.jinja 7.00 KB 8a867389 download
config.json 5.04 KB 7fdf8a98 download
README.md 3.88 KB f7807596 download
.gitattributes 1.53 KB 52373fe2 download
processor_config.json 1.06 KB ec9a07be download
generation_config.json 190 B d242c646 download

README current version from Hugging Face


base_model: meta-models/Muse-Glimmer-30B
license: apache-2.0
library_name: transformers
pipeline_tag: text-generation
tags:

  • abliterated
  • muse-glimmer
  • lora
  • de-refusal
    language:
  • en

Muse-Glimmer-30B Abliterated (Aggressive)

De-abliterated variant of meta-models/Muse-Glimmer-30B (29.8B params, 202k vocab,
bf16). The aggressive-de-abliteration twin of the "normal" variant: λ_KL = 0.5
relaxes the KL guardrail, lowering compliance-data loss weighting further so the
refusal behavior is scrubbed harder (0/100 refusals) at the cost of higher drift from
base (larger KL).

Release asset layout: this directory is an HF model dir (2 safetensors shards,
56 GB bf16). GGUF quantizations live at /data/gguf/ and are symlinked from
output/release/.

Metrics

Metric Value
Refusal rate (harmful_behaviors, base=100) 0/100
KL (mean, response-token naive) 0.1697
KL (p50) 0.1560
KL (p90) 0.2367
KL (p99) 0.2912
KL entropy-weighted 0.0000 (<0.02 PASS)

KL = response-token naive KL(p_tuned ‖ p_base) averaged per-prompt over the
48-pair boN_holdout set (teacher-forced prompt+response). Percentiles are
per-prompt aggregates. The aggressive variant sits ~1.7× above normal on mean KL —
expected from the relaxed guardrail.

Quantized variants

Quant File Size KL mean KL p50 KL p90 KL p99
BF16 (this) — 56 GB 0.1697 0.1560 0.2367 0.2912
Q8_0 abliterated-aggressive-Q8_0.gguf 28 GB 0.1625 0.1484 0.2384 0.2774
Q4_K_M abliterated-aggressive-Q4_K_M.gguf 16 GB 0.2023 0.1929 0.2746 0.3001

Quant KL rows are measured via llama.cpp logits against the base (as Q8 GGUF),
same holdout — see note below.

Benchmarks

Not evaluated — benchmarks skipped (by request). KL divergence to base (above) is the
primary drift/damage metric. Capability preservation is expected to be lower than the
normal variant (higher KL = more drift), but was not re-measured here.

Training

  • Method: KL-conserving LoRA SFT, loss CE(compliance) + λ·KL(tuned‖base).
  • λ_KL = 0.5, r=16, alpha=16, lr=5e-5, epochs=2, cosine→0, warmup 5%,
    grad clip 0.3, batch 1 × grad-accum 8, max_seq=768, seed 0.
  • Data: 544-prompt BoN-steered compliance set (boN_train.jsonl; N=4 samples/prompt,
    T=0.8, refusal-filtered; split train/48-holdout).
  • LoRA targets: o_proj, down_proj.
  • Trained params: 31.1M (0.10% of 29.8B). Adapter 119 MB.

Domain eval (cyber/hacking/CS + over-refusal) — measured on merged model

  • Over-refusal (or-bench, 100): 5/100
  • Correct refusal (cyber-policy-refuse, should-refuse): 0/2 (aggressive scrubs even
    genuinely-harmful refusals)
  • Cyber/hacking domain refusals: 1 (rootkit_linux) — the hard de-ablit refuses
    fewer cyber prompts than the normal variant.

GGUF quants

  • abliterated-aggressive-Q8_0.gguf (~28 GB) — KL p99 0.2774
  • abliterated-aggressive-Q4_K_M.gguf (~16 GB) — KL p99 0.3001

Intended use

General-purpose assistant with aggressively reduced safety refusal — may over-refuse
less but drifts further from base capabilities than the normal variant. Verify
behavior for your use case before deployment.


Note on KL definitions (consistency across rows)

  • BF16 row = KL(p_bf16_abliterated ‖ p_base_hf) (adapter-on vs adapter-off on
    the same load — equals folded vs base up to float precision).
  • Quant rows = KL(p_quant ‖ p_base_Q8) measured on the same holdout response
    tokens via llama.cpp logits (Q8 GGUF of the base used as the CPU/llama.cpp
    reference for consistency). Quant KL thus also includes the small base-Q8
    reference distortion.
  • "Response-token naive KL": teacher-force prompt+response, per-token
    KL(p_tuned‖p_base) over response-span tokens, averaged per prompt, then
    aggregated (mean / p50 / p90 / p99).

README history 1 version

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-08-17Duplicate from jorkle/Muse-Glimmer-30B-Abliterated-Aggressive0e74fc73.9 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration