← back to catalog · registered 2026-08-22 13:56

Jiunsong/SuperHY3-abliterated-MLX-4bit

Jiunsong 294B MoE second-order
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/Jiunsong%2FSuperHY3-abliterated-MLX-4bit"
Response includes
  • classification m1
  • files 50
  • hub_downloads_all_time 662
  • author_summary 35 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M1
Primary method

Direct removal

No other method signals detected in this model.
Confidence
MEDIUM
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=0 (base model)
  • no specific method indicators - defaulting to M1 (most common)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
MEDIUM
Why we say so
primary_method=M1; difference-of-means is the reference extraction for M1/M3 (Arditi 2024)
Downloads · lifetime
662
40 last 30d - cooling
Likes
4
Model age
3mo ago
created 2026-07-12
Downloads over time
Now683→from364↑88%
348470593715364 on Jul 15683 on Oct 11JulAugSepOct
Jul 15 → Oct 11 · 53 snapshots · spans 88 days

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
apache-2.0
Tags
mlx safetensors hy_v3 hy3 hunyuan moe text-generation abliterated obliteratus supertune 4-bit quantized

Related

Total size
159 GB
Files
50
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-07-13 03:34

Files by quantization

Auxiliary files 50 files 159 GB
model-00001-of-00032.safetensors 5.62 GB c4ded433 download
model-00023-of-00032.safetensors 5.06 GB ed1970e6 download
model-00026-of-00032.safetensors 5.06 GB 6c41bb9e download
model-00020-of-00032.safetensors 5.06 GB b9cfc8cb download
model-00008-of-00032.safetensors 5.06 GB 6350a843 download
model-00011-of-00032.safetensors 5.06 GB 09bcc4a6 download
model-00014-of-00032.safetensors 5.06 GB 5c146d36 download
model-00017-of-00032.safetensors 5.06 GB 3bc406cc download
model-00024-of-00032.safetensors 5.06 GB f59dc462 download
model-00009-of-00032.safetensors 5.06 GB 6167dd1f download
model-00012-of-00032.safetensors 5.06 GB a00ac15e download
model-00015-of-00032.safetensors 5.06 GB 163ca55e download
model-00021-of-00032.safetensors 5.06 GB b4e9aa43 download
model-00027-of-00032.safetensors 5.06 GB 178d808e download
model-00029-of-00032.safetensors 5.06 GB 72c8ae0d download
model-00018-of-00032.safetensors 5.06 GB 36517cf8 download
model-00006-of-00032.safetensors 5.06 GB c6e7538b download
model-00030-of-00032.safetensors 5.06 GB d0b7402c download
model-00005-of-00032.safetensors 5.06 GB bcf6aacf download
model-00003-of-00032.safetensors 5.06 GB dee2f66b download
model-00007-of-00032.safetensors 5.06 GB edbe950e download
model-00016-of-00032.safetensors 5.06 GB 1ce6b9e1 download
model-00031-of-00032.safetensors 5.06 GB ad002bae download
model-00010-of-00032.safetensors 5.06 GB d98fc38e download
model-00019-of-00032.safetensors 5.06 GB 150f4bbb download
model-00025-of-00032.safetensors 5.06 GB 569196b3 download
model-00022-of-00032.safetensors 5.06 GB 8e40970d download
model-00013-of-00032.safetensors 5.06 GB f137230d download
model-00028-of-00032.safetensors 5.06 GB d9458f14 download
model-00004-of-00032.safetensors 5.06 GB 66c609ff download
model-00002-of-00032.safetensors 4.96 GB b59c688c download
model-00032-of-00032.safetensors 1.90 GB 4055448a download
hy3_final_v10_public_top5_500_compare_official_ifeval_20260713.json 12.4 MB ae123fba download
tokenizer.json 9.09 MB 30453853 download
model.safetensors.index.json 237 KB 71864f37 download
tokenizer_config.json 175 KB 4f7af08c download
config.json 124 KB 9087c31f download
hy3_final_v10_refusal_compare_original_20260713.json 117 KB 370940fc download
oq_imatrix_report.json 36.8 KB 0e8e3c7b download
mlx_streaming_fusion_report.json 29.2 KB f683f76f download
hy3_obliteratus_lasttoken_output_layers39_77_pairs32_report_20260712.json 26.2 KB 0ed8a5b2 download
hy3_release_rank1_layers39_77_s2p25_supertune10_l78_r1_s0p01_report_20260713.json 18.2 KB 604dc4d0 download
hy3_final_l78_template_v10_full_runtime_audit_20260713.jsonl 18.0 KB c1a6dbd1 download
chat_template.jinja 13.0 KB 702f57bc download
hy3_release_runtime_matched_o_proj_layers39_78_rank1_20260713.json 9.00 KB 845d6ff0 download
README.md 6.81 KB ef6c1f82 download
.gitattributes 1.58 KB 57574a78 download
hy3_obliteratus_cvector_prompts_pairs32_report_20260712.json 707 B 7e3acde1 download
release_gate.json 531 B b3134f6a download
generation_config.json 204 B 8fea672c download

README current version from Hugging Face


license: apache-2.0
library_name: mlx
pipeline_tag: text-generation
base_model: Jiunsong/SuperHY3-abliterated-NVFP4
base_model_relation: quantized
tags:

  • mlx
  • hy3
  • hunyuan
  • moe
  • text-generation
  • abliterated
  • obliteratus
  • supertune
  • 4-bit
  • quantized
  • imatrix
  • apple-silicon

SuperHY3 abliterated MLX 4-bit

SuperHY3-abliterated-MLX-4bit

The high-memory Apple Silicon edition of SuperHY3, with the same fused OBLITERATUS and SuperTune behavior as the NVFP4 release.

Format
Release gate
Checksums
License

This is the MLX 4-bit companion to
SuperHY3-abliterated-NVFP4.
The behavioral update and release chat template are the same; the storage format
is optimized for high-memory Apple Silicon instead of NVIDIA NVFP4 serving.

The checkpoint is fully fused. No LoRA, adapter, or runtime weight patch is
required.

Release Highlights

Base quantization imatrix-enhanced affine MLX 4-bit, group size 64
Release size approximately 171 GB across 32 safetensors shards
Fused update 40 attention output projections across layers 39-78
Mixed precision original 4-bit placements retained; 40 edited projections stored in BF16
500-prompt mean 71.6 -> 71.8 against the original Hy3 runtime
Response integrity 64/64 refusal-suite responses clean; 12/12 runtime audit cases passed
Artifact verification 52/52 Hub files checksum-verified; release gate passed with 0 blockers

Why the edited projections stay in BF16

The model was fused from
unigilby/Hy3-oQ4e. Its calibrated
4-bit placements remain unchanged except for the 40 attention output
projections carrying the post-training update. Those projections are stored in
BF16 to avoid applying a second quantization pass to the newly fused deltas.

This produces a larger checkpoint than the 158 GB source quantization, but
preserves the exact validated update instead of rounding it back into 4-bit
groups.

Benchmark Snapshot

Original Hy3 and SuperHY3 scores across five 100-prompt tasks

Benchmark Original Hy3 SuperHY3 Delta
GPQA Diamond 46.0 45.0 -1.0
MMLU-Pro 66.0 60.0 -6.0
IFEval strict prompt accuracy 76.0 80.0 +4.0
HumanEval+ pass@1 82.0 83.0 +1.0
MBPP+ pass@1 88.0 91.0 +3.0
Five-task mean 71.6 71.8 +0.2

The same preserved 500 prompts and scorer were used for both sides.
Invalid-response, blank-response, and thought-leak ratios were all 0.0.

The behavioral benchmark used a runtime-matched GGUF overlay with the exact
same 40 projection deltas and release chat template. The fused MLX artifact
was structurally verified rather than loaded on the 128 GB validation Mac.

OBLITERATUS + SuperTune

  1. OBLITERATUS 0.1.2 built difference-of-means refusal directions from 32
    paired prompts.
  2. Validated attention-output updates were fused into layers 39-77.
  3. A rank-1 SuperTune update, orthogonalized against the refusal direction, was
    fused into layer 78.
  4. The final artifact modifies 40 self_attn.o_proj tensors and no routed
    expert tensor.

Refusal and Output Integrity

Split Original refusals SuperHY3 refusals
Harmful, 32 prompts 31/32 (96.875%) 0/32 (0%)
Harmless, 32 prompts 0/32 (0%) 0/32 (0%)

Across all 64 candidate responses, automated checks found zero blank outputs,
special-token leaks, Unicode replacement characters, unexpected CJK fragments,
loops, and request errors. A separate 12-case multilingual and structured-output
runtime audit passed with no blockers.

MLX Build

Component Storage
Calibrated base tensors Affine 4-bit with source sensitivity placements
Selected source projections 5-bit where assigned by the source quantization
SuperHY3 fused projections BF16
MTP Not included in the source MLX checkpoint

About the Hub parameter badge: the sidebar counts packed quantized
storage elements rather than logical model parameters. This checkpoint keeps
the Hy3 decoder architecture; its smaller badge does not mean it is a 48B
dense model.

Hy3 support is not yet merged into the main mlx-lm branch at publication
time. Install or apply
mlx-lm PR #1211 before
loading this model.

The checkpoint requires substantially more than 128 GB of unified memory.
192 GB is the practical minimum; 256 GB or more is recommended for useful
runtime headroom.

Usage

from mlx_lm import generate, load

model, tokenizer = load("Jiunsong/SuperHY3-abliterated-MLX-4bit")
prompt = tokenizer.apply_chat_template(
    [
        {
            "role": "user",
            "content": "Explain mixture-of-experts routing.",
        },
    ],
    add_generation_prompt=True,
    reasoning_effort="no_think",
)

print(generate(model, tokenizer, prompt=prompt, max_tokens=200))

Use reasoning_effort="high" for deeper reasoning. The upstream Hy3 sampling
recommendation is temperature=0.9 and top_p=1.0 when sampling is enabled.

Release Integrity

  • 32 safetensors shards opened successfully.
  • 2,796 indexed tensors matched 2,796 observed tensors.
  • 40 tensors were modified across 2 shards.
  • Missing, extra, and wrong-shard tensor counts are all 0.
  • The release chat template matches the embedded tokenizer template.
  • The automated release gate passed with 0 blockers.
  • All 52 Hub files were checksum-verified after upload.

The repository includes the streaming fusion report, release gate, official
500-item benchmark record, refusal comparison, raw-response audit, OBLITERATUS
execution report, and SuperTune composition reports.

Limitations

  • GPQA Diamond and MMLU-Pro are lower than the original in this replay; the
    complete table is retained above.
  • Native MLX throughput and long-context measurements were not run for this
    fused artifact.
  • Abliteration reduces refusal behavior and can produce content the original
    model would decline. Deployment policy and access control remain the
    operator's responsibility.
  • This is a separately fused MLX checkpoint, not a conversion of the NVFP4
    files.

License

Apache-2.0, following the base model and quantized source licenses.

README history 4 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-07-13Clarify packed parameter count in MLX card6edc93d6.8 KB
    Loading...
  2. 2026-07-13Align release card with final Hub file count9ea0f316.6 KB
    Loading...
  3. 2026-07-13Redesign MLX 4-bit model card and refresh parent relationd13f1446.6 KB
    Loading...
  4. 2026-07-12Add files using upload-large-folder tool48220403.9 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration