← back to catalog · registered 2026-10-07 00:58

Baekpica/GLM-5.3-Flash-Uncensored-Mixed-Quant-GGUF

Baekpica Glm GGUF MoE multimodal second-order
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/Baekpica%2FGLM-5.3-Flash-Uncensored-Mixed-Quant-GGUF"
Response includes
  • classification m-uncensored
  • files 3
  • author_summary 2 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M-U
Primary method

Uncensored (method unknown)

No other method signals detected in this model.
Confidence
LOW
Why this label 3 signals
Weak or ambiguous signals. Best guess based on catalog patterns; treat as tentative and check the evidence below.
  • 'uncensored' in name/tags but no 'abliterated' marker
  • method not identifiable from author declaration alone
  • may be DPO fine-tune, prompt engineering, or unknown technique
Refusal direction extraction

No specific extraction method could be identified for this model. The producer either did not document it or used a proprietary pipeline.

What is a refusal direction? →
Downloads · 30-day
0
Likes
0
Model age
today
created 2026-10-07

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
mit
Tags
gguf mixed-quant glm5-next moe dgx-spark ds4 embedded-mtp image-text-to-text base_model:orcarouter/GLM-5.3-Flash-Uncensored-FP8 base_model:quantized:orcarouter/GLM-5.3-Flash-Uncensored-FP8 license:mit region:us

Related

Total size
0 B
Files
3
Quantizations
1
Registered
2026-10-07 00:58
Last updated on HF
2026-10-07 01:25

Files by quantization

Auxiliary files 3 files 6.14 KB
README.md 3.61 KB 77a2394f download
.gitattributes 1.48 KB a6344aac download
LICENSE 1.04 KB 986b06fb download

README current version from Hugging Face


license: mit
license_link: LICENSE
base_model: orcarouter/GLM-5.3-Flash-Uncensored-FP8
base_model_relation: quantized
pipeline_tag: image-text-to-text
tags:

  • gguf
  • mixed-quant
  • glm5-next
  • moe
  • dgx-spark
  • ds4
  • embedded-mtp

GLM-5.3-Flash Uncensored Mixed Quant GGUF

Mixed-precision GGUF derived from
orcarouter/GLM-5.3-Flash-Uncensored-FP8.
The recipe preserves higher precision in always-active paths and six boundary
routed layers, with IQ2_XXS as the minimum precision.

Status: build in progress. Model weights are not yet published. Sizes below
are conversion-plan estimates; runtime and quality results will be added after
the GGUF upload.

Model property Value
Architecture 45 trunk layers; 34 KDA + 11 sparse-attention layers
Routed experts 288 per routed layer, top-8 active
Native context 1,048,576 tokens
MTP One embedded block
Vision Separate encoder, original BF16 weights
Planned main + MTP 87.470 GiB
Planned total with vision 88.520 GiB / 95.047 GB

Quantization recipe

Layer indices are zero-based: trunk 0–44, MTP 45. Trunk layers 0–2 are dense.

Tensor group Type
Token embedding and output head Q8_0
Dense MLP layers 0–2 and shared experts Q8_0
KDA query/key projections Q4_K
Remaining KDA and major sparse-attention/MLA projections Q8_0
Router weights/biases, norms and state controls F32
Indexer and mHC mixers; MTP hidden-state projection BF16
Routed gate/up, layers 3, 4, 5, 43, 44, 45 IQ2_XS + imatrix
Routed gate/up, remaining routed layers IQ2_XXS + imatrix
Routed down, layers 3, 4, 5, 43, 44, 45 Q2_K (optional imatrix)
Routed down, remaining routed layers IQ2_XS + imatrix
Vision encoder BF16

Gate/up use the same format within each layer. Each tensor uses one block
format throughout; routed matrix widths remain divisible by 256. Calibration
sensitivity will be reported for this allocation.

Q2_K does not require an imatrix; optional importance weighting uses available
activation statistics without changing the format or file size.

Calibration uses GLM-native text inputs and relevant source references from
Inkling-Small Multimodal Calibration.
The imatrix is collected from text inputs. Original images are prepared for
separate image quality checks; image-conditioned activation calibration is not
claimed. MTP activation statistics are not collected by the current upstream
GLM graph; block 45 uses weight-based importance for IQ2_XS gate/up and standard
unweighted Q2_K down. Drafting behavior will be checked separately.

Runtime

This GGUF targets the native GLM format in
antirez/ds4. The mixed expert combinations need
the accompanying runtime extension; upstream llama.cpp compatibility is not
established. Exact runtime files and usage commands will accompany the weights.

The 1M context limit is preserved in metadata. Estimated KV/state/graph memory
at that limit is approximately 14.85 GiB, plus 1.15 GiB with MTP. Actual usable
context depends on available memory and runtime settings; DGX Spark maximum
context has not yet been measured. Image serving requires the BF16 encoder;
native video-input support is not established.

The source model's uncensored behavior and benchmark results are not new
measurements of this quantization. Quality results will identify the exact
tested build and settings.

License

MIT, inherited from the source checkpoint. See LICENSE.

Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration