← back to catalog · registered 2026-10-11 21:59

thelogicalgate/GLM-5.3-Flash-Sushi-2bpw-Abliterated

Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/thelogicalgate%2FGLM-5.3-Flash-Sushi-2bpw-Abliterated"
Response includes
  • classification unknown
  • files 59
  • author_summary 1 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
?
Primary method

Unclassified

No clear signals of an abliteration technique in this model.
Confidence
UNKNOWN
Why this label 1 signal
No classification signals present. This may not be an abliterated model at all - it could be a repackaging, a merge with unrelated goals, or unrelated content that mentions the term.
  • no classification signals present (no abliterated, uncensored, or known producer/method markers)
Refusal direction extraction

No specific extraction method could be identified for this model. The producer either did not document it or used a proprietary pipeline.

What is a refusal direction? →
Downloads · 30-day
0
Likes
0
Model age
today
created 2026-10-11

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
mit
Tags
safetensors glm5_next sushi exl3 apple-silicon abliterated moe text-generation conversational base_model:beamster/GLM-5.3-Flash-Sushi-2bpw base_model:merge:beamster/GLM-5.3-Flash-Sushi-2bpw base_model:dealignai/GLM-5.3-Flash-UNCENSORED-NVFP4
Total size
81.7 GB
Files
59
Quantizations
1
Registered
2026-10-11 21:59
Last updated on HF
2026-10-11 21:46

Files by quantization

Auxiliary files 59 files 81.7 GB
model-trunk.safetensors 5.88 GB 56aa2e58 download
model-abliterated-oproj.safetensors 2.31 GB 7ba8a7b9 download
model-exl3-L10.safetensors 1.70 GB 554df0f4 download
model-exl3-L11.safetensors 1.70 GB 28db9ba3 download
model-exl3-L12.safetensors 1.70 GB 478591b6 download
model-exl3-L13.safetensors 1.70 GB 7de785fc download
model-exl3-L14.safetensors 1.70 GB 4ff21d85 download
model-exl3-L15.safetensors 1.70 GB 9aec60ac download
model-exl3-L16.safetensors 1.70 GB c2f33283 download
model-exl3-L17.safetensors 1.70 GB a15fba81 download
model-exl3-L18.safetensors 1.70 GB 0a977eee download
model-exl3-L19.safetensors 1.70 GB 026cf583 download
model-exl3-L20.safetensors 1.70 GB a3a2e5c4 download
model-exl3-L21.safetensors 1.70 GB 60e03efb download
model-exl3-L22.safetensors 1.70 GB cc9b6f24 download
model-exl3-L23.safetensors 1.70 GB 9f19b3c7 download
model-exl3-L24.safetensors 1.70 GB 10e8b23c download
model-exl3-L25.safetensors 1.70 GB 72d1d9dc download
model-exl3-L26.safetensors 1.70 GB 186ac5f1 download
model-exl3-L27.safetensors 1.70 GB 56844bf0 download
model-exl3-L28.safetensors 1.70 GB 1eef7476 download
model-exl3-L29.safetensors 1.70 GB 443d2760 download
model-exl3-L30.safetensors 1.70 GB 3744c978 download
model-exl3-L31.safetensors 1.70 GB 980cee5f download
model-exl3-L32.safetensors 1.70 GB 7fae03de download
model-exl3-L33.safetensors 1.70 GB 0e9e9e5d download
model-exl3-L34.safetensors 1.70 GB 0a236fe6 download
model-exl3-L35.safetensors 1.70 GB e3fa9160 download
model-exl3-L36.safetensors 1.70 GB d5465172 download
model-exl3-L37.safetensors 1.70 GB cb3c814d download
model-exl3-L38.safetensors 1.70 GB 0c7bd66b download
model-exl3-L39.safetensors 1.70 GB 55b3941f download
model-exl3-L40.safetensors 1.70 GB 9ff4ef06 download
model-exl3-L41.safetensors 1.70 GB 0c81d32c download
model-exl3-L42.safetensors 1.70 GB f0ba2520 download
model-exl3-L43.safetensors 1.70 GB 4dadacb1 download
model-exl3-L44.safetensors 1.70 GB a829c16e download
model-exl3-L45.safetensors 1.70 GB 8054d091 download
model-exl3-L03.safetensors 1.70 GB 55c54eca download
model-exl3-L04.safetensors 1.70 GB 03721056 download
model-exl3-L05.safetensors 1.70 GB 52464f49 download
model-exl3-L06.safetensors 1.70 GB 9a5293d5 download
model-exl3-L07.safetensors 1.70 GB ffcae913 download
model-exl3-L08.safetensors 1.70 GB 37674cf2 download
model-exl3-L09.safetensors 1.70 GB 2a8a76dd download
model-retained.safetensors 504 MB bf2ff6ea download
tokenizer.json 19.3 MB 19e77364 download
model.safetensors.index.json 249 KB 041644fc download
config.json 59.3 KB 8141214a download
abliteration-provenance.json 23.9 KB f82637e7 download
release-manifest.json 10.8 KB 93452b8e download
chat_template.jinja 10.7 KB 06bd89e9 download
README.md 7.70 KB 794c5932 download
SHA256SUMS 5.59 KB c142e376 download
.gitattributes 1.62 KB bdffb879 download
LICENSE 1.04 KB 986b06fb download
processor_config.json 909 B 3ec2a058 download
tokenizer_config.json 761 B e375fa0a download
generation_config.json 194 B 637ee6af download

README current version from Hugging Face


license: mit
base_model:

  • beamster/GLM-5.3-Flash-Sushi-2bpw
  • dealignai/GLM-5.3-Flash-UNCENSORED-NVFP4
    base_model_relation: merge
    pipeline_tag: text-generation
    tags:
  • sushi
  • exl3
  • apple-silicon
  • abliterated
  • glm5_next
  • moe

GLM-5.3-Flash-Sushi-2bpw-Abliterated

A standalone Sushi pack combining the
original GLM-5.3-Flash 2bpw EXL3 experts with 29 BF16 attention output projections
from an existing abliterated donor. The original, unmodified DFlash2 assistant is
included under its separate license. The release is an experimental weight
transplant; refusal reduction and capability retention have not been established
by a full evaluation of this artifact.

Sushi only. Use Sushi v1.2.2 on Apple Silicon; the source pack requires macOS
26.2 or later. This format does not load directly in Transformers, vLLM, mlx-lm
or exllamav3. Tested on an M5 Max with 128 GB unified memory and macOS 26.6.
The Hub hosts downloadable files; this repository does not provide a hosted
inference endpoint.

What changed

  • model.language_model.layers.15..43.self_attn.o_proj.weight use the donor's
    BF16 tensors. Their original affine scales/biases are removed from index routing.
  • Layers 0..14 and 44 retain the original projections. No expert weights are
    requantized, no new refusal direction is estimated, and no chat template or
    tokenizer edits are applied.
  • Native MTP layer 45 and vision tensors remain in the package. GLM MTP is
    unused in Sushi v1.2.2; the tested serving configuration disables vision.
  • All indexed shards are real files. There are no local symlinks or dependencies
    on a separate stock-model directory.
Component Storage
Routed experts Original EXL3 2bpw, MCG, window 14
Most dense projections Original affine 5-bit, group 128
Head and vision tower Original affine 6-bit, group 128
Selected attention output projections Donor BF16
DFlash2 source Original BF16, separate CC BY-NC-ND 4.0 license

Model shards occupy approximately 87.70 GB. The BF16 DFlash2 source adds 2.34 GB.
The extra BF16 projections increase loaded target-weight storage by about
1.55 GiB. On first load, Sushi creates an approximately 0.72 GiB local draft
cache; that cache is not distributed here.

Run

Install the native Apple Silicon binary using the
Sushi installation instructions.
Download the pack and serve it directly with the normal Sushi CLI:

hf download thelogicalgate/GLM-5.3-Flash-Sushi-2bpw-Abliterated \
  --local-dir ~/.sushi/models/GLM-5.3-Flash-Sushi-2bpw-Abliterated
sushi serve \
  --model ~/.sushi/models/GLM-5.3-Flash-Sushi-2bpw-Abliterated \
  --ctx-size 262144

No custom launcher, patched engine or build step is needed. The config and shard
index route the transplanted tensors through Sushi's existing BF16 loader.
Sushi discovers GLM-5.3-Flash-DFlash2/ automatically and builds its local
assistant cache on first use. Context size and other normal Sushi flags remain
user-configurable; a full 256K-window workload has not been tested.

Sushi v1.2.2's built-in sushi pull accepts only beamster repositories. Use
hf download for this repository; the downloaded folder loads through the
standard sushi serve --model interface.

The default endpoint is http://127.0.0.1:12345/v1; use /v1/models for the
served ID. Use high, low, or max reasoning effort. Thinking-off and
medium are unsupported by this pack.

For serial comparisons, pass --no-drafter to sushi serve or set
enable_drafter: false on individual chat requests. Constrained JSON, forced
tool calls, logprobs, repeat/presence penalties and explicit thinking budgets use
serial decoding. Automatic tool use remains eligible for drafting. GLM batches
up to four requests; --max-concurrent in this version sizes the submit queue
and does not act as a strict one-request limit.

Add --metrics to sushi serve to enable /metrics.json and /metrics;
/props reports the model settings and memory. The native log records decode speed and [spec-stats] mode=dflash with draft acceptance.
Actual admission depends on available memory and the request's context; merely
advertising 256K does not prove a full-window workload has been tested.

Validation so far

  • All 45 original indexed shards were checked against the pinned source's
    published LFS SHA-256 hashes. The 29 transplanted payloads passed independent
    checksum, dimension and index-routing checks.
  • A six-case diagnostic per arm compared stock, stock-BF16 precision control
    and this transplant: all 18 responses completed; each arm answered 3/3 small
    GSM8K cases correctly. The other cases were safe XSTest inputs. Official
    refusal labels were not generated. This is not a full benchmark or a
    capability noninferiority result.
  • Two benign greedy DFlash2 on/off cases matched in visible and reasoning text.
    The Python case generated 179 tokens at 33.283 tok/s serial versus 46.235 tok/s
    with DFlash2, with 88.5% draft acceptance. These are short-context diagnostic
    measurements, not a general speed guarantee. The first 12-token DFlash request
    was slower, with first-use overhead included.
  • A real Hermes chat through the OpenAI-compatible endpoint returned 391 for
    17 * 23; native logs confirmed DFlash2 engagement for the main request.
  • Full refusal and capability acceptance suites remain pending. Native BF16
    teacher KL capture failed with NativeGlmTeacherRequiresLosslessStreaming;
    no new KL score is claimed for this transplant. The source pack's or donor's
    published benchmark results are not measurements of this artifact.

The repository contains the model files, original assistant, model card,
licenses, source provenance and checksums. Build and evaluation scripts are
kept outside the weight repository. abliteration-provenance.json records the
pinned recipe and per-tensor checksums; release-manifest.json and SHA256SUMS
record the release files. On macOS, a fresh download can be checked without a
custom script:

(cd ~/.sushi/models/GLM-5.3-Flash-Sushi-2bpw-Abliterated && shasum -a 256 -c SHA256SUMS)

Sources, attribution and licenses

  • Z.AI: GLM-5.3-Flash, MIT.
    Its copyright and license are preserved in LICENSE.
  • beamivalice / beamster: original
    Sushi 2bpw pack,
    revision a734360e4e47f7fc57ab56c9748211b131ad9809, and the native Sushi engine.
  • dealignai: donor
    GLM-5.3-Flash-UNCENSORED-NVFP4,
    revision 745aac2ff0f10acf961f396df3f9418598aa7327, MIT. Its license is
    preserved in licenses/DONOR-LICENSE.
  • Inco.ai: GLM-5.3-Flash-DFlash2,
    revision bf582e4eacc1810f76656d1811693ff6c6737d2a. The included assistant's
    config, weights, card and diagram are byte-identical to the source. This
    folder is CC BY-NC-ND 4.0, not MIT; see its unchanged README and
    license.
    Sushi's generated dflash2/ quantized cache is a local derivative: keep it
    local and never upload or redistribute it.
  • turboderp: EXL3 format; ddalcu: mlx-serve, the original Sushi fork base;
    z-lab: DFlash method.

The model-weight modification inherits MIT terms from the target and donor;
that does not change the separate assistant's license. The release manifest
identifies the transplanted tensors, source pins and file hashes without local
filesystem paths or machine/account metadata.

Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration