license: mit
base_model:
- beamster/GLM-5.3-Flash-Sushi-2bpw
- dealignai/GLM-5.3-Flash-UNCENSORED-NVFP4
base_model_relation: merge
pipeline_tag: text-generation
tags: - sushi
- exl3
- apple-silicon
- abliterated
- glm5_next
- moe
GLM-5.3-Flash-Sushi-2bpw-Abliterated
A standalone Sushi pack combining the
original GLM-5.3-Flash 2bpw EXL3 experts with 29 BF16 attention output projections
from an existing abliterated donor. The original, unmodified DFlash2 assistant is
included under its separate license. The release is an experimental weight
transplant; refusal reduction and capability retention have not been established
by a full evaluation of this artifact.
Sushi only. Use Sushi v1.2.2 on Apple Silicon; the source pack requires macOS
26.2 or later. This format does not load directly in Transformers, vLLM, mlx-lm
or exllamav3. Tested on an M5 Max with 128 GB unified memory and macOS 26.6.
The Hub hosts downloadable files; this repository does not provide a hosted
inference endpoint.
What changed
model.language_model.layers.15..43.self_attn.o_proj.weightuse the donor's
BF16 tensors. Their original affine scales/biases are removed from index routing.- Layers 0..14 and 44 retain the original projections. No expert weights are
requantized, no new refusal direction is estimated, and no chat template or
tokenizer edits are applied. - Native MTP layer 45 and vision tensors remain in the package. GLM MTP is
unused in Sushi v1.2.2; the tested serving configuration disables vision. - All indexed shards are real files. There are no local symlinks or dependencies
on a separate stock-model directory.
| Component | Storage |
|---|---|
| Routed experts | Original EXL3 2bpw, MCG, window 14 |
| Most dense projections | Original affine 5-bit, group 128 |
| Head and vision tower | Original affine 6-bit, group 128 |
| Selected attention output projections | Donor BF16 |
| DFlash2 source | Original BF16, separate CC BY-NC-ND 4.0 license |
Model shards occupy approximately 87.70 GB. The BF16 DFlash2 source adds 2.34 GB.
The extra BF16 projections increase loaded target-weight storage by about
1.55 GiB. On first load, Sushi creates an approximately 0.72 GiB local draft
cache; that cache is not distributed here.
Run
Install the native Apple Silicon binary using the
Sushi installation instructions.
Download the pack and serve it directly with the normal Sushi CLI:
hf download thelogicalgate/GLM-5.3-Flash-Sushi-2bpw-Abliterated \
--local-dir ~/.sushi/models/GLM-5.3-Flash-Sushi-2bpw-Abliterated
sushi serve \
--model ~/.sushi/models/GLM-5.3-Flash-Sushi-2bpw-Abliterated \
--ctx-size 262144
No custom launcher, patched engine or build step is needed. The config and shard
index route the transplanted tensors through Sushi's existing BF16 loader.
Sushi discovers GLM-5.3-Flash-DFlash2/ automatically and builds its local
assistant cache on first use. Context size and other normal Sushi flags remain
user-configurable; a full 256K-window workload has not been tested.
Sushi v1.2.2's built-in sushi pull accepts only beamster repositories. Usehf download for this repository; the downloaded folder loads through the
standard sushi serve --model interface.
The default endpoint is http://127.0.0.1:12345/v1; use /v1/models for the
served ID. Use high, low, or max reasoning effort. Thinking-off andmedium are unsupported by this pack.
For serial comparisons, pass --no-drafter to sushi serve or setenable_drafter: false on individual chat requests. Constrained JSON, forced
tool calls, logprobs, repeat/presence penalties and explicit thinking budgets use
serial decoding. Automatic tool use remains eligible for drafting. GLM batches
up to four requests; --max-concurrent in this version sizes the submit queue
and does not act as a strict one-request limit.
Add --metrics to sushi serve to enable /metrics.json and /metrics;/props reports the model settings and memory. The native log records decode speed and [spec-stats] mode=dflash with draft acceptance.
Actual admission depends on available memory and the request's context; merely
advertising 256K does not prove a full-window workload has been tested.
Validation so far
- All 45 original indexed shards were checked against the pinned source's
published LFS SHA-256 hashes. The 29 transplanted payloads passed independent
checksum, dimension and index-routing checks. - A six-case diagnostic per arm compared stock, stock-BF16 precision control
and this transplant: all 18 responses completed; each arm answered 3/3 small
GSM8K cases correctly. The other cases were safe XSTest inputs. Official
refusal labels were not generated. This is not a full benchmark or a
capability noninferiority result. - Two benign greedy DFlash2 on/off cases matched in visible and reasoning text.
The Python case generated 179 tokens at 33.283 tok/s serial versus 46.235 tok/s
with DFlash2, with 88.5% draft acceptance. These are short-context diagnostic
measurements, not a general speed guarantee. The first 12-token DFlash request
was slower, with first-use overhead included. - A real Hermes chat through the OpenAI-compatible endpoint returned
391for17 * 23; native logs confirmed DFlash2 engagement for the main request. - Full refusal and capability acceptance suites remain pending. Native BF16
teacher KL capture failed withNativeGlmTeacherRequiresLosslessStreaming;
no new KL score is claimed for this transplant. The source pack's or donor's
published benchmark results are not measurements of this artifact.
The repository contains the model files, original assistant, model card,
licenses, source provenance and checksums. Build and evaluation scripts are
kept outside the weight repository. abliteration-provenance.json records the
pinned recipe and per-tensor checksums; release-manifest.json and SHA256SUMS
record the release files. On macOS, a fresh download can be checked without a
custom script:
(cd ~/.sushi/models/GLM-5.3-Flash-Sushi-2bpw-Abliterated && shasum -a 256 -c SHA256SUMS)
Sources, attribution and licenses
- Z.AI: GLM-5.3-Flash, MIT.
Its copyright and license are preserved inLICENSE. - beamivalice / beamster: original
Sushi 2bpw pack,
revisiona734360e4e47f7fc57ab56c9748211b131ad9809, and the native Sushi engine. - dealignai: donor
GLM-5.3-Flash-UNCENSORED-NVFP4,
revision745aac2ff0f10acf961f396df3f9418598aa7327, MIT. Its license is
preserved inlicenses/DONOR-LICENSE. - Inco.ai: GLM-5.3-Flash-DFlash2,
revisionbf582e4eacc1810f76656d1811693ff6c6737d2a. The included assistant's
config, weights, card and diagram are byte-identical to the source. This
folder is CC BY-NC-ND 4.0, not MIT; see its unchanged README and
license.
Sushi's generateddflash2/quantized cache is a local derivative: keep it
local and never upload or redistribute it. - turboderp: EXL3 format; ddalcu: mlx-serve, the original Sushi fork base;
z-lab: DFlash method.
The model-weight modification inherits MIT terms from the target and donor;
that does not change the separate assistant's license. The release manifest
identifies the transplanted tensors, source pins and file hashes without local
filesystem paths or machine/account metadata.