← back to catalog · registered 2026-08-22 13:56

yar-sh/Muse-Glimmer-30B-Abliterated-Aggressive-NVFP4

yar-sh 25B multimodal second-order
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/yar-sh%2FMuse-Glimmer-30B-Abliterated-Aggressive-NVFP4"
Response includes
  • classification m1
  • files 18
  • hub_downloads_all_time 85
  • author_summary 1 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M1
Primary method

Direct removal

No other method signals detected in this model.
Confidence
MEDIUM
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=0 (base model)
  • no specific method indicators - defaulting to M1 (most common)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
MEDIUM
Why we say so
primary_method=M1; difference-of-means is the reference extraction for M1/M3 (Arditi 2024)
Downloads · lifetime
85
37 last 30d - stable
Likes
0
Model age
8w ago
created 2026-08-14
Downloads over time
Now96→from35↑174%
32557910235 on Aug 1996 on Oct 1196 on Oct 9AugSepOct
Aug 19 → Oct 11 · 48 snapshots · spans 53 days

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
other
Languages
en ja de ru
Tags
transformers safetensors muse_glimmer image-text-to-text nvfp4 fp4 w4a4 compressed-tensors autoround vllm multimodal abliterated

Related

Total size
21.8 GB
Files
18
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-08-14 15:21

Files by quantization

Auxiliary files 18 files 21.8 GB
model-00004-of-00008.safetensors 3.00 GB 77e9cad0 download
model-00005-of-00008.safetensors 3.00 GB ab295a7f download
model-00002-of-00008.safetensors 2.97 GB 66388ac3 download
model-00003-of-00008.safetensors 2.97 GB 94076148 download
model-00001-of-00008.safetensors 2.97 GB 4df11bd8 download
model-00007-of-00008.safetensors 2.63 GB a31a7864 download
model-00008-of-00008.safetensors 2.50 GB 6bff0010 download
model-00006-of-00008.safetensors 1.72 GB cd401849 download
tokenizer.json 26.8 MB c9dbee66 download
model.safetensors.index.json 263 KB 68d92a9d download
tokenizer_config.json 78.1 KB d1b80588 download
chat_template.jinja 7.00 KB 8a867389 download
config.json 6.38 KB 871a4ca2 download
README.md 3.24 KB 4bcf2afc download
.gitattributes 1.53 KB 52373fe2 download
quantization_config.json 1.23 KB cb5b1a6c download
processor_config.json 1.06 KB ec9a07be download
generation_config.json 184 B e35c881f download

README current version from Hugging Face


license: other
license_name: muse-glimmer
base_model:

  • jorkle/Muse-Glimmer-30B-Abliterated-Aggressive
  • meta-models/Muse-Glimmer-30B
    base_model_relation: quantized
    pipeline_tag: image-text-to-text
    library_name: transformers
    tags:
  • muse_glimmer
  • nvfp4
  • fp4
  • w4a4
  • compressed-tensors
  • autoround
  • vllm
  • multimodal
  • image-text-to-text
  • abliterated
  • uncensored
  • decensored
    language:
  • en
  • ja
  • de
  • ru

Muse-Glimmer-30B-Abliterated-Aggressive — NVFP4 (lean, vision-preserved)

A lean 22 GB NVFP4 (W4A4) quant of jorkle/Muse-Glimmer-30B-Abliterated-Aggressive
(itself an aggressively-decensored — KL-conserving LoRA-SFT — build of Meta's
meta-models/Muse-Glimmer-30B).

Built to run a decensored, vision-preserving Muse-Glimmer on a single ~24–32 GB Blackwell / DGX-Spark (GB10)
in vLLM. At the time of quanting, no decensored lean NVFP4 Muse existed — the public lean NVFP4s were the
censored base, and the decensored builds were only GGUF / bf16 / MLX / a fat 28 GB NVFP4. This fills that gap.

What it is

  • Format: compressed-tensors NVFP4, W4A4 group-16, via Intel AutoRound (scheme="NVFP4",
    dataset="NeelNanda/pile-10k", nsamples=128, seqlen=2048, iters=200, quant_nontext_module=False).
  • Vision tower + lm_head kept in BF16 (quant_nontext_module=False) — pristine multimodal input at zero
    decode cost; only the decoder Linear layers (the size + per-token bandwidth) are quantized to FP4.
  • ~22 GB on disk (vs the ~28 GB fat abliterated NVFP4). Multimodal (perception encoder) intact.

Serving (vLLM)

Needs a Muse-Glimmer-capable vLLM build (e.g. vllm/vllm-openai:muse-glimmer):

vllm serve <this-model> --served-model-name muse-aggressive \
  --reasoning-parser muse_glimmer --enable-auto-tool-choice --tool-call-parser muse_glimmer \
  --trust-remote-code --max-model-len 131072 --kv-cache-dtype fp8

Quantization auto-detects (compressed-tensors) — no --quantization flag needed. Use a
Reasoning strength: low system line if you want content in content rather than reasoning_content.

Measured on GB10 (DGX-Spark, 224 GB/s, single-stream)

Metric Value
Decode (c=1 / c=4 / c=8) 12.5 / 46 / 87 tok/s
HumanEval (pass@1, reasoning-low) .872
IFEval (prompt-level strict) .684
Tools (32-case) .656
Vision ✅ (accurately describes real images)

Honest notes

  • This is the aggressive decensor. Its KL-LoRA-SFT decensoring lowers instruction-following
    (IFEval .684) vs a plain weight-edit abliteration of the same base (which measured IFEval ~.90).
    Prefer this build for maximally-permissive / RP use; for instruction-following-critical work a
    manual-abliterated quant is better.
  • Speed is bandwidth-bound (dense 30B ÷ ~224 GB/s). No speculative drafter is included — pairing a matched
    DFlash/EAGLE drafter would roughly double decode.

Attribution

Base: Meta Muse-Glimmer-30B. Decensoring: jorkle (Abliterated-Aggressive, KL-LoRA-SFT).
NVFP4 quant: this repo (AutoRound, compressed-tensors). Inherits the upstream Muse-Glimmer license.

README history 2 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-08-14drop llm-compressor tag55be7463.2 KB
    Loading...
  2. 2026-08-14Upload folder using huggingface_hubf962d763.3 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration