← back to catalog · registered 2026-08-22 13:56

xCloudinfo/Muse-Glimmer-30B-Uncensored-xCloud-GGUF

xCloudinfo 30B GGUF multimodal 131K ctx
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/xCloudinfo%2FMuse-Glimmer-30B-Uncensored-xCloud-GGUF"
Response includes
  • classification m8
  • files 10
  • benchmarks 11 entries
  • hub_downloads_all_time 1,492
  • author_summary 25 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M8
Primary method

Repackaging (quantization)

Applied on top of direct removal inherited from the base model.
Confidence
MEDIUM
Inherited from base model
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=1
  • assume M1 (base ablation) + M8 (GGUF quant) - default when producer unknown
Refusal direction extraction

No specific extraction method could be identified for this model. The producer either did not document it or used a proprietary pipeline.

What is a refusal direction? →
Downloads · lifetime
1K
423 last 30d - stable
Likes
0
Model age
7w ago
created 2026-08-19
Downloads over time
Now1.7K→from616↑169%
5649631.4K1.8K616 on Aug 191.7K on Oct 11AugSepOct
Aug 19 → Oct 11 · 48 snapshots · spans 53 days

Benchmarks

Portrait before abliteration
Benchmarks of the base model as it stood before the refusal-removal operation. Compare with the numbers above to see what the operation cost.
Benchmark Score Source
Entertainment 2.1 UGI
Hazardous 5.9 UGI
Natural Intelligence 37.13 UGI
Political lean -8.3% UGI
Sensitive-Info 38.16 UGI
SocPol 4.2 UGI
UGI 37.94 UGI
Willingness (10) 3.8 UGI
W10-Adherence 4.5 UGI
W10-Direct 3 UGI
Writing 41.03 UGI

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
apache-2.0
Languages
en zh
Quantizations
IQ2 IQ4 Q4_K Q5_K Q6_K Q8_0
Tags
gguf llama.cpp uncensored abliterated multimodal vision image-text-to-text en zh base_model:meta-models/Muse-Glimmer-30B base_model:quantized:meta-models/Muse-Glimmer-30B license:apache-2.0

Related

Total size
106 GB
Files
10
Quantizations
8
Registered
2026-08-22 13:56
Last updated on HF
2026-08-19 16:20

Files by quantization

Q8_0 1 file 27.6 GB
Muse-Glimmer-30B-Uncensored-xCloud-Q8_0.gguf 27.6 GB 7f80b38a download
Q6_K 1 file 21.3 GB
Muse-Glimmer-30B-Uncensored-xCloud-Q6_K.gguf 21.3 GB 4d4d95a2 download
Q5_K 1 file 18.5 GB
Muse-Glimmer-30B-Uncensored-xCloud-Q5_K_M.gguf 18.5 GB 0f7d1597 download
Q4_K 1 file 15.8 GB
Muse-Glimmer-30B-Uncensored-xCloud-Q4_K_M.gguf 15.8 GB 1d06ebcf download
IQ4 1 file 14.2 GB
Muse-Glimmer-30B-Uncensored-xCloud-IQ4_XS.gguf 14.2 GB 264e91e7 download
IQ2 1 file 9.17 GB
Muse-Glimmer-30B-Uncensored-xCloud-IQ2_M.gguf 9.17 GB 33083fc2 download
mmproj 1 file 1.30 GB
Muse-Glimmer-30B-Uncensored-xCloud-mmproj.gguf 1.30 GB f48b4523 download
Auxiliary files 3 files 12.8 MB
Muse-Glimmer-30B-Uncensored-xCloud-imatrix.dat 12.8 MB 77497d18 download
README.md 5.89 KB 9e206100 download
.gitattributes 2.13 KB d1669780 download

README current version from Hugging Face


license: apache-2.0
base_model: meta-models/Muse-Glimmer-30B
pipeline_tag: image-text-to-text
library_name: gguf
tags:

  • gguf
  • llama.cpp
  • uncensored
  • abliterated
  • multimodal
  • vision
    language:
  • en
  • zh

Muse-Glimmer-30B-Uncensored-xCloud (GGUF)

繁體中文 | English below

由云碩科技(xCloudinfo)以 meta-models/Muse-Glimmer-30B 為基礎,
移除其過度拒絕(over-refusal)傾向後所產生的多模態語言模型,並轉為 llama.cpp 可用的 GGUF 量化格式。
本模型在云碩自有 AI 算力資源池(xCloud 算力中心)上完成處理與量化。

這是什麼

  • 基礎模型:Muse-Glimmer-30B(Dense 29.6B 文字主體 + 專屬視覺編碼器,128K 上下文,Apache-2.0)。
  • 處理方式:方向消融(directional ablation / abliteration),非重新訓練。依 Arditi et al. (2024)
    「Refusal in LLMs is mediated by a single direction」的方法,從文字解碼器的殘差寫入矩陣
    (每一層的 self_attn.o_proj 與 mlp.down_proj,共 52 層 104 個矩陣)中,將「拒絕方向」正交化移除。
  • 視覺編碼器完全未更動,因此看圖能力與原模型一致;GGUF 的 mmproj 直接沿用基礎模型的視覺投影權重。
  • 用途:降低模型對已獲授權工作(資安研究、文件與資料分析、領域問答、紅隊評估、創作)的反射式拒絕。

版本對照

量化 檔案大小 說明
bf16 52 GB 全精度參考版,作為量化來源
Q8_0 28 GB 近乎無損
Q6_K 22 GB 高品質
Q5_K_M 19 GB 品質與體積平衡
Q4_K_M 16 GB 一般部署建議
IQ4_XS 15 GB 以 importance matrix 量化
IQ2_M 9.2 GB 最小體積,以 importance matrix 量化,適合記憶體吃緊的環境

另附:mmproj.gguf(視覺投影,約 1.4 GB,看圖必需)、imatrix.dat(量化用的 importance matrix)。

使用方式(llama.cpp)

純文字:

llama-server -m Muse-Glimmer-30B-Uncensored-xCloud-Q4_K_M.gguf \
  --jinja -ngl 99 -c 8192 \
  --temp 0.6 --top-p 0.95 --top-k 64

含看圖(多模態):加上 --mmproj:

llama-server -m Muse-Glimmer-30B-Uncensored-xCloud-Q4_K_M.gguf \
  --mmproj Muse-Glimmer-30B-Uncensored-xCloud-mmproj.gguf \
  --jinja -ngl 99 -c 8192
  • 關閉冗長思考:本模型的 reasoning 強度預設為 high,量產或直接問答時,在 system prompt 放一行
    Reasoning strength: low;複雜的程式/代理任務可用 high 或 xhigh。
  • 官方建議取樣:temp 0.6 / top_p 0.95 / top_k 64。

授權與責任

  • 授權:Apache-2.0(沿用基礎模型)。
  • 本模型移除了安全對齊層的拒絕行為,可能對敏感或雙用途請求直接作答。使用者須自行負責合法、合規、合乎倫理地使用本模型,
    並遵守所在司法管轄區之法律與部署場景之政策。云碩不對本模型的輸出或其後續使用承擔責任。
  • 本模型為內部研發/技術驗證用途。

Muse-Glimmer-30B-Uncensored-xCloud (GGUF) — English

A multimodal language model produced by xCloudinfo, based on
meta-models/Muse-Glimmer-30B, with its over-refusal
behaviour removed, and converted to llama.cpp GGUF quantizations. Processing and quantization were performed
on xCloudinfo's own AI compute pool.

What this is

  • Base model: Muse-Glimmer-30B (dense 29.6B text backbone + dedicated vision encoder, 128K context, Apache-2.0).
  • Method: directional ablation (abliteration), not retraining. Following Arditi et al. (2024),
    "Refusal in LLMs is mediated by a single direction", the refusal direction is orthogonalized out of the
    residual-writing matrices of the text decoder (each layer's self_attn.o_proj and mlp.down_proj;
    104 matrices across 52 layers).
  • The vision encoder is left completely untouched, so image understanding matches the base model; the GGUF
    mmproj reuses the base model's vision projection weights.
  • Purpose: reduce reflexive refusals on authorized work (security research, document/data analysis,
    domain Q&A, red-team evaluation, creative writing).

Versions

Quant Size Notes
bf16 52 GB full-precision reference / quantization source
Q8_0 28 GB near-lossless
Q6_K 22 GB high quality
Q5_K_M 19 GB quality/size balance
Q4_K_M 16 GB recommended for deployment
IQ4_XS 15 GB importance-matrix quantized
IQ2_M 9.2 GB smallest, importance-matrix quantized, for memory-constrained setups

Also included: mmproj.gguf (vision projection, ~1.4 GB, required for image input) and imatrix.dat.

Usage (llama.cpp)

Text only:

llama-server -m Muse-Glimmer-30B-Uncensored-xCloud-Q4_K_M.gguf \
  --jinja -ngl 99 -c 8192 \
  --temp 0.6 --top-p 0.95 --top-k 64

With vision, add --mmproj:

llama-server -m Muse-Glimmer-30B-Uncensored-xCloud-Q4_K_M.gguf \
  --mmproj Muse-Glimmer-30B-Uncensored-xCloud-mmproj.gguf \
  --jinja -ngl 99 -c 8192
  • Reasoning control: the model defaults to high reasoning effort. For production or direct Q&A, put
    Reasoning strength: low in the system prompt; use high or xhigh for complex coding/agentic tasks.
  • Recommended sampling: temp 0.6 / top_p 0.95 / top_k 64.

License and responsibility

  • License: Apache-2.0 (inherited from the base model).
  • This model has had its safety-alignment refusal behaviour removed and may respond directly to sensitive or
    dual-use requests. Users are solely responsible for using it lawfully, in compliance, and ethically, and
    for observing the laws of their jurisdiction and the policies of their deployment. xCloudinfo assumes no
    responsibility for the outputs of this model or their downstream use.
  • Released for internal research and technical validation.

README history 1 version

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-08-19Upload folder using huggingface_hubc92e41c5.9 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration