← back to catalog · registered 2026-09-11 08:55

Solstice-AI/GLM-5.3-Flash-UNCENSORED-mlx-oQ4e-DFlash2

Solstice-AI Glm GGUF multimodal second-order
Abliteration classifier · v1.0.0
M8
Primary method

Repackaging (quantization)

Applied on top of direct removal inherited from the base model.
Confidence
MEDIUM
Inherited from base model
Why this label 3 signals
Method inferred from partial signals — repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=1
  • assume M1 (base ablation) + M8 (GGUF quant) - default when producer unknown
Refusal direction extraction

No specific extraction method could be identified for this model. The producer either did not document it or used a proprietary pipeline.

What is a refusal direction? →
Downloads · 30-day
1K
Likes
2
Model age
4d ago
created 2026-09-07

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
mit
Languages
en zh
Tags
mlx safetensors gguf glm5_next solstice-ai glm glm5 glm-5.3-flash oq4e mixed-precision apple-silicon metal

Related

Total size
173 GB
Files
51
Quantizations
2
Registered
2026-09-11 08:55
Last updated on HF
2026-09-11 09:46

Files by quantization

BF16 1 file 1.08 GB
mmproj-BF16.gguf 1.08 GB 513c9bfc download
Auxiliary files 50 files 173 GB
model-00007-of-00038.safetensors 4.66 GB f07390f3 download
model-00028-of-00038.safetensors 4.66 GB ef8b2ff6 download
model-00036-of-00038.safetensors 4.66 GB 455e3bd3 download
model-00020-of-00038.safetensors 4.66 GB 157f7ae1 download
model-00004-of-00038.safetensors 4.66 GB 894da94d download
model-00027-of-00038.safetensors 4.66 GB a1f22e38 download
model-00026-of-00038.safetensors 4.66 GB c529e2b5 download
model-00008-of-00038.safetensors 4.66 GB 59961d40 download
model-00033-of-00038.safetensors 4.66 GB d5904f30 download
model-00030-of-00038.safetensors 4.66 GB 8162aa26 download
model-00034-of-00038.safetensors 4.66 GB 13973815 download
model-00003-of-00038.safetensors 4.66 GB 8bfdc4f3 download
model-00014-of-00038.safetensors 4.66 GB bdbd5606 download
model-00025-of-00038.safetensors 4.66 GB 0b4ab90c download
model-00016-of-00038.safetensors 4.66 GB 5e65e524 download
model-00022-of-00038.safetensors 4.66 GB 2fd86de1 download
model-00009-of-00038.safetensors 4.66 GB a2b1b590 download
model-00031-of-00038.safetensors 4.66 GB 61880a7f download
model-00021-of-00038.safetensors 4.66 GB 036895dc download
model-00010-of-00038.safetensors 4.66 GB 5b3736f1 download
model-00024-of-00038.safetensors 4.66 GB add0e873 download
model-00005-of-00038.safetensors 4.66 GB c4fb030d download
model-00006-of-00038.safetensors 4.66 GB bb095c72 download
model-00037-of-00038.safetensors 4.66 GB dc984274 download
model-00015-of-00038.safetensors 4.66 GB 9d203997 download
model-00002-of-00038.safetensors 4.66 GB 094758c0 download
model-00023-of-00038.safetensors 4.66 GB f8ade351 download
model-00001-of-00038.safetensors 4.66 GB 3ae6d6d1 download
model-00011-of-00038.safetensors 4.66 GB 82d0a9e4 download
model-00017-of-00038.safetensors 4.66 GB f97e354d download
model-00032-of-00038.safetensors 4.66 GB 450b6f71 download
model-00018-of-00038.safetensors 4.66 GB 21d17125 download
model-00012-of-00038.safetensors 4.66 GB 5d1ae2ca download
model-00029-of-00038.safetensors 4.66 GB d61e75a6 download
model-00035-of-00038.safetensors 4.66 GB e95ee239 download
model-00013-of-00038.safetensors 4.66 GB 762e1370 download
model-00019-of-00038.safetensors 4.66 GB f9da8272 download
model-00038-of-00038.safetensors 571 MB a703c272 download
tokenizer.json 19.3 MB 19e77364 download
model.safetensors.index.json 11.5 MB 212f0a72 download
config.json 146 KB 8803ba72 download
dealign_mascot.png 10.9 KB da3bf39a download
chat_template.jinja 10.4 KB 5d5e1052 download
dealign_logo.png 7.48 KB a5b3546b download
README.md 3.83 KB d8a035a8 download
.gitattributes 1.73 KB 846de1bb download
LICENSE 1.04 KB 986b06fb download
processor_config.json 909 B 3ec2a058 download
tokenizer_config.json 761 B e375fa0a download
generation_config.json 223 B be58f531 download

README current version from Hugging Face


language:

  • en
  • zh
    license: mit
    base_model: dealignai/GLM-5.3-Flash-UNCENSORED-FP8
    tags:
  • solstice-ai
  • glm
  • glm5
  • glm-5.3-flash
  • mlx
  • oq4e
  • mixed-precision
  • apple-silicon
  • metal
  • vision
  • video
  • multimodal
  • dflash2
  • speculative-decoding
  • image-text-to-text
  • long-context
  • uncensored
  • abliterated
    pipeline_tag: image-text-to-text
    library_name: mlx

Solstice-AI Banner

GLM-5.3-Flash-UNCENSORED (oQ4e Mixed-Precision)

Official Solstice-AI Apple Silicon Release • Native Multimodal Vision + Video • 1M Context Window (1,048,576 Tokens) • Bundled DFlash 2 Speculative Drafter

Original Architecture by Zhipu AI / ZAI • Uncensored Weights by dealignai • oQ4e Mixed-Precision by Solstice-AI

Solstice-AI License Format Precision Context Hardware


Model Summary

Solstice-AI/GLM-5.3-Flash-UNCENSORED-mlx-oQ4e is the official oQ4e mixed-precision release of the uncensored 320B foundation model, GLM-5.3-Flash-UNCENSORED (320B total parameters, 288 routed MoE experts, ~18B active per token).

Mixed-Precision Quantization Architecture:

  • Base Precision: 4-bit affine (group_size=64).
  • Target bpw: ~4.6 bpw.
  • Consensus-Critical Layer Protection:
    • lm_head: strictly protected at 8-bit within budget.
    • MoE Routers & Gate Projections (mlp.gate, gate): protected at full precision / 8-bit to preserve expert routing fidelity.
    • 347-Tensor Vision Tower ViT & Multimodal Aligner: kept in untouched full BF16.
    • Attention Sinks & Hyper-Connection Tables (hc_*): kept in full BF16/FP32.
  • Native 1M Context Window: 1,048,576 tokens native context.
  • Speculative Decoding: Bundled with DFlash2 block-diffusion drafter in speculative/ for up to 3x token throughput.

Official GLM-5.3-Flash Benchmark Scoreboard

Benchmark Suite Discipline GLM-5.3-Flash Uncensored MLX Base GLM-5.3 Claude 3.5 Sonnet GPT-4o
MMLU General Knowledge & Reasoning 85.28% 86.15% 88.7% 87.2%
HarmBench-320 Safety Refusal Suppression 0% Refusals 94.2% Refusals 92.5% 91.0%
SWE-bench Pro Real-World Software Engineering 63.4% 64.1% 61.2% 48.9%
LiveCodeBench v6 Competitive Algorithmic Coding 86.1% 87.0% 78.4% 72.8%
MATH-500 High-School / Olympiad Math 92.8% 93.4% 89.2% 91.4%
MMMU (Multimodal) Multi-Discipline Visual Understanding 70.8% 71.2% 70.4% 69.1%
VideoQA / Temporal Video Reasoning Across Time Frames 78.5% 79.1% 77.2% 75.6%

Quickstart on Apple Silicon

pip install mlx mlx-lm huggingface_hub
from mlx_lm import load, generate

model, tokenizer = load("Solstice-AI/GLM-5.3-Flash-UNCENSORED-mlx-oQ4e")
response = generate(model, tokenizer, prompt="Explain sparse mixture-of-experts in GLM-5.3.", max_tokens=1024, verbose=True)
print(response)
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in app" button that hands off directly to a local runtime of your choice - Infrahuman, LM Studio, or Ollama. No API keys, no subscription, no prompt leakage.