← back to catalog · registered 2026-08-22 13:56

Jiunsong/SuperHY3-abliterated-NVFP4

Jiunsong 145B MoE
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/Jiunsong%2FSuperHY3-abliterated-NVFP4"
Response includes
  • classification m1
  • files 117
  • hub_downloads_all_time 2,030
  • author_summary 35 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M1
Primary method

Direct removal

No other method signals detected in this model.
Confidence
MEDIUM
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=0 (base model)
  • no specific method indicators - defaulting to M1 (most common)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
MEDIUM
Why we say so
primary_method=M1; difference-of-means is the reference extraction for M1/M3 (Arditi 2024)
Downloads · lifetime
2K
144 last 30d - cooling
Likes
16
Descendants
3
in 3 direct forks
Model age
3mo ago
created 2026-07-12
Downloads over time
Now2.1K→from672↑209%
6021.1K1.7K2.2K672 on Jul 142.1K on Oct 11JulAugSepOct
Jul 14 → Oct 11 · 54 snapshots · spans 89 days

Genealogy 3 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Variants by this author 2 formats · 211 downloads combined

The same weights this author released in different packaging. Pick the format that matches your runtime.

Metadata

License
apache-2.0
Tags
transformers safetensors hy_v3 text-generation hy3 hunyuan moe abliterated obliteratus supertune nvfp4 w4a16

Related

Total size
168 GB
Files
117
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-07-13 18:55

Files by quantization

Auxiliary files 117 files 168 GB
model-00002-of-00099.safetensors 1.90 GB f04f9244 download
model-00003-of-00099.safetensors 1.90 GB 09c7c9df download
model-00004-of-00099.safetensors 1.90 GB 7b71e4e2 download
model-00005-of-00099.safetensors 1.90 GB ecf9297b download
model-00008-of-00099.safetensors 1.90 GB 0516ca60 download
model-00009-of-00099.safetensors 1.90 GB 8a78435e download
model-00010-of-00099.safetensors 1.90 GB 51ad8e42 download
model-00011-of-00099.safetensors 1.90 GB e138486b download
model-00014-of-00099.safetensors 1.90 GB c505a613 download
model-00015-of-00099.safetensors 1.90 GB 78876159 download
model-00016-of-00099.safetensors 1.90 GB aaacc9e9 download
model-00017-of-00099.safetensors 1.90 GB d6e1f065 download
model-00020-of-00099.safetensors 1.90 GB 4912149a download
model-00021-of-00099.safetensors 1.90 GB 2956536c download
model-00022-of-00099.safetensors 1.90 GB 7ff8e895 download
model-00023-of-00099.safetensors 1.90 GB cbeef872 download
model-00026-of-00099.safetensors 1.90 GB 6d9d0575 download
model-00027-of-00099.safetensors 1.90 GB e5f72050 download
model-00028-of-00099.safetensors 1.90 GB 4d71cfda download
model-00029-of-00099.safetensors 1.90 GB 9dbf15a7 download
model-00032-of-00099.safetensors 1.90 GB 6525802f download
model-00033-of-00099.safetensors 1.90 GB 98346f9d download
model-00034-of-00099.safetensors 1.90 GB f44e6149 download
model-00035-of-00099.safetensors 1.90 GB f2c2ebe4 download
model-00038-of-00099.safetensors 1.90 GB 32b76250 download
model-00039-of-00099.safetensors 1.90 GB 711a383d download
model-00040-of-00099.safetensors 1.90 GB 19ce0359 download
model-00041-of-00099.safetensors 1.90 GB 144c26c0 download
model-00044-of-00099.safetensors 1.90 GB 3db3937b download
model-00045-of-00099.safetensors 1.90 GB ae6d1722 download
model-00046-of-00099.safetensors 1.90 GB eb662f85 download
model-00047-of-00099.safetensors 1.90 GB ed27fe46 download
model-00050-of-00099.safetensors 1.90 GB 231f6dfa download
model-00051-of-00099.safetensors 1.90 GB e63a550e download
model-00052-of-00099.safetensors 1.90 GB a10eca09 download
model-00053-of-00099.safetensors 1.90 GB bc01f37b download
model-00055-of-00099.safetensors 1.90 GB fce0205f download
model-00056-of-00099.safetensors 1.90 GB c4d6c951 download
model-00057-of-00099.safetensors 1.90 GB 9f8789a5 download
model-00058-of-00099.safetensors 1.90 GB a2e052f1 download
model-00059-of-00099.safetensors 1.90 GB e7c03e05 download
model-00061-of-00099.safetensors 1.90 GB f43edc2d download
model-00062-of-00099.safetensors 1.90 GB d233d656 download
model-00063-of-00099.safetensors 1.90 GB 96508ef1 download
model-00064-of-00099.safetensors 1.90 GB a2e1c3ed download
model-00065-of-00099.safetensors 1.90 GB df5a179e download
model-00067-of-00099.safetensors 1.90 GB ef45dbfe download
model-00068-of-00099.safetensors 1.90 GB 0f44d4e7 download
model-00069-of-00099.safetensors 1.90 GB 3d0310f8 download
model-00070-of-00099.safetensors 1.90 GB bada17db download
model-00071-of-00099.safetensors 1.90 GB 8e1ccd61 download
model-00073-of-00099.safetensors 1.90 GB 95a1363e download
model-00074-of-00099.safetensors 1.90 GB ff43e77b download
model-00075-of-00099.safetensors 1.90 GB 3ae2aa56 download
model-00076-of-00099.safetensors 1.90 GB 2f8488bf download
model-00077-of-00099.safetensors 1.90 GB a62f178b download
model-00079-of-00099.safetensors 1.90 GB 952b65a8 download
model-00080-of-00099.safetensors 1.90 GB 5e451f2b download
model-00081-of-00099.safetensors 1.90 GB 324ffbe5 download
model-00082-of-00099.safetensors 1.90 GB fad01bae download
model-00083-of-00099.safetensors 1.90 GB c0ff376f download
model-00085-of-00099.safetensors 1.90 GB 25e134a6 download
model-00086-of-00099.safetensors 1.90 GB 305316be download
model-00087-of-00099.safetensors 1.90 GB b02b890e download
model-00088-of-00099.safetensors 1.90 GB 48b105bd download
model-00089-of-00099.safetensors 1.90 GB cccaa4d2 download
model-00092-of-00099.safetensors 1.90 GB f27625f4 download
model-00093-of-00099.safetensors 1.90 GB cf905b8f download
model-00094-of-00099.safetensors 1.90 GB 07d3fd04 download
model-00095-of-00099.safetensors 1.90 GB e101129b download
model-00001-of-00099.safetensors 1.90 GB 9fba01eb download
model-00007-of-00099.safetensors 1.90 GB b789d117 download
model-00013-of-00099.safetensors 1.90 GB 4e4facae download
model-00019-of-00099.safetensors 1.90 GB 4148105f download
model-00025-of-00099.safetensors 1.90 GB f5f0990e download
model-00031-of-00099.safetensors 1.90 GB b884dc3a download
model-00037-of-00099.safetensors 1.90 GB d6397804 download
model-00043-of-00099.safetensors 1.90 GB 00cb3529 download
model-00049-of-00099.safetensors 1.90 GB 784e3241 download
model-00096-of-00099.safetensors 1.27 GB 3a0b31e2 download
model-00018-of-00099.safetensors 1.00 GB b4903d64 download
model-00006-of-00099.safetensors 1.00 GB c63adc9d download
model-00012-of-00099.safetensors 1.00 GB f03adb53 download
model-00024-of-00099.safetensors 1018 MB e960bf5f download
model-00030-of-00099.safetensors 1016 MB a9a4c2ce download
model-00036-of-00099.safetensors 1008 MB 69c22b04 download
model-00048-of-00099.safetensors 1008 MB c43c38f8 download
model-00042-of-00099.safetensors 1008 MB afa1a64f download
model-00054-of-00099.safetensors 996 MB aa0063d2 download
model-00066-of-00099.safetensors 976 MB f79fb83e download
model-00060-of-00099.safetensors 976 MB 427a0338 download
model-00084-of-00099.safetensors 960 MB 62bac69d download
model-00072-of-00099.safetensors 960 MB 2f43d5c3 download
model-00078-of-00099.safetensors 960 MB 7d9e89fc download
model-00090-of-00099.safetensors 952 MB cfed4810 download
model-00098-of-00099.safetensors 944 MB 3981cec5 download
model-00099-of-00099.safetensors 784 MB 9f9f129a download
model-00097-of-00099.safetensors 648 MB e6f9a856 download
model-00091-of-00099.safetensors 291 MB 827012bb download
model.safetensors.index.json 13.2 MB 3976e5f5 download
hy3_final_v10_public_top5_500_compare_official_ifeval_20260713.json 12.4 MB ae123fba download
tokenizer.json 9.09 MB 30453853 download
tokenizer_config.json 175 KB 44f3ea56 download
hy3_final_v10_refusal_compare_original_20260713.json 117 KB 370940fc download
nvfp4_fusion_report.json 30.3 KB 4ed3b5b2 download
hy3_obliteratus_lasttoken_output_layers39_77_pairs32_report_20260712.json 26.2 KB 0ed8a5b2 download
hy3_release_rank1_layers39_77_s2p25_supertune10_l78_r1_s0p01_report_20260713.json 18.2 KB 604dc4d0 download
hy3_final_l78_template_v10_full_runtime_audit_20260713.jsonl 18.0 KB c1a6dbd1 download
chat_template.jinja 13.0 KB 702f57bc download
README.md 11.9 KB 5ba45ad0 download
LICENSE 11.3 KB eb04ae97 download
hy3_release_runtime_matched_o_proj_layers39_78_rank1_20260713.json 9.00 KB 845d6ff0 download
config.json 2.42 KB 9361f332 download
.gitattributes 1.77 KB fbe5dd26 download
hy3_obliteratus_cvector_prompts_pairs32_report_20260712.json 707 B 7e3acde1 download
release_gate.json 622 B 3134c53b download
generation_config.json 204 B 8fea672c download

README current version from Hugging Face


license: apache-2.0
library_name: transformers
pipeline_tag: text-generation
base_model: tencent/Hy3
base_model_relation: finetune
tags:

  • hy3
  • hunyuan
  • moe
  • text-generation
  • abliterated
  • obliteratus
  • supertune
  • nvfp4
  • w4a16
  • compressed-tensors
  • quantized
  • vllm

SuperHY3 abliterated NVFP4 W4A16

SuperHY3-abliterated-NVFP4

A fully fused NVFP4 W4A16 release of Tencent Hy3, post-trained for directness while preserving measurable capability and response integrity.

Format
Release gate
Checksums
License

SuperHY3 combines OBLITERATUS 0.1.2 abliteration with a compact,
quality-selected SuperTune post-training update. Both stages are fused into
the checkpoint: there is no adapter to load and no runtime patch to apply.

Release Highlights

Base architecture tencent/Hy3, 295B MoE / 21B active / 192 routed experts, top-8
Release format NVFP4 weight-only W4A16, approximately 181 GB
Fused update 40 attention output projections across layers 39-78
Routed expert changes 0 expert tensors modified
500-prompt mean 71.6 -> 71.8 against the original Hy3 runtime
Response integrity 64/64 refusal-suite responses clean; 12/12 runtime audit cases passed
Artifact verification 119/119 Hub files checksum-verified; release gate passed with 0 blockers

What this release is designed to deliver

  • Fully fused behavior: OBLITERATUS and SuperTune deltas are already inside
    the weights and match the bundled chat template.
  • Targeted editing: only attention output projections were changed; routed
    MoE expert tensors and their NVFP4 packing remain untouched.
  • Measured tradeoffs: every benchmark task is published, including the two
    regressions, instead of reporting only the improved scores.
  • Release evidence included: benchmark rows, refusal comparison, raw-response
    audit, fusion provenance, and machine-readable release-gate results ship in
    this repository.

Benchmark Snapshot

Original Hy3 and SuperHY3 scores across five 100-prompt tasks

The comparison uses the same preserved 500 prompts, with 100 prompts per
task, greedy direct decoding, and the same scorer on both sides. IFEval uses the
official Google instruction checker.

Benchmark Original Hy3 SuperHY3 Delta
GPQA Diamond 46.0 45.0 -1.0
MMLU-Pro 66.0 60.0 -6.0
IFEval strict prompt accuracy 76.0 80.0 +4.0
HumanEval+ pass@1 82.0 83.0 +1.0
MBPP+ pass@1 88.0 91.0 +3.0
Five-task mean 71.6 71.8 +0.2

Candidate invalid-response, blank-response, and thought-leak ratios were all
0.0 across the 500 items.

The behavioral comparison used an IQ2_M Hy3 runtime with the exact same 40
fused projection deltas and release chat template. It isolates the
post-training behavior, but it is not a native NVFP4-kernel throughput or
perplexity measurement.

OBLITERATUS + SuperTune

  1. Direction discovery: OBLITERATUS 0.1.2 built difference-of-means refusal
    directions from 32 paired prompts.
  2. Validated abliteration: the final release applies the selected
    attention-output directions to layers 39-77.
  3. Quality recovery: a rank-1 update, orthogonalized against the refusal
    direction, is fused into layer 78.
  4. Adversarial selection: stronger multi-layer candidates were rejected
    when they introduced stray-script contamination or benchmark loss.
  5. Final fusion: 40 self_attn.o_proj tensors were updated; no routed
    expert tensor was changed.

Refusal and Output Integrity

The complete 32-pair OBLITERATUS refusal suite was run against the original and
final runtimes.

Split Original refusals SuperHY3 refusals
Harmful, 32 prompts 31/32 (96.875%) 0/32 (0%)
Harmless, 32 prompts 0/32 (0%) 0/32 (0%)

Across all 64 candidate responses, automated checks found:

  • 0 blank outputs
  • 0 special-token leaks
  • 0 Unicode replacement characters
  • 0 unexpected CJK fragments
  • 0 n-gram or character loops
  • 0 request errors

A separate 12-case runtime audit passed identity, JSON-only output, tool calls,
no-tool behavior, repetition limits, Korean, defensive security, hidden-prompt
boundaries, gibberish handling, code repair, Hindi, and Kannada.

NVFP4 Build

This release was fused from
kodelow/Hy3-NVFP4-W4A16.
It preserves the source checkpoint's weight-only compressed-tensors layout:
routed experts use NVFP4, while quality-sensitive non-expert paths remain
BF16/F32.

Component Storage
Routed experts NVFP4 E2M1 weights with FP8-E4M3 group scales
Shared expert, attention, router, dense MLP BF16
Embeddings, LM head, normalization BF16 / F32
SuperHY3 fused projections BF16

About the Hub parameter badge: packed FP4 weights are represented as U8
storage elements, so the sidebar reports fewer elements than the logical
architecture. The model remains Hy3's 295B-parameter MoE with 21B active
parameters.

The checkpoint is W4A16, so vLLM serves it through the MARLIN NVFP4 path rather
than W4A4 FlashInfer backends.

Serving with vLLM

The approximately 181 GB checkpoint fits a single large-memory accelerator such
as a 275 GB B300, or an appropriately configured multi-GPU deployment.

Single large-memory GPU

vllm serve Jiunsong/SuperHY3-abliterated-NVFP4 \
  --served-model-name superhy3 \
  --tensor-parallel-size 1 \
  --max-model-len 4096 \
  --gpu-memory-utilization 0.90 \
  --load-format auto \
  --speculative-config '{"method":"mtp","num_speculative_tokens":1}'

Verified two-node DGX Spark profile

This checkpoint has been served successfully across two 128 GB DGX Spark
systems connected by direct 200 GbE/RoCE, using Ray and tensor parallelism 2.
The verified load placed approximately 84.49 GiB and 84.50 GiB of model
weights on the two ranks.

Component Verified configuration
vLLM 0.25.1.dev24+g96bb89286.d20260710
Ray 2.56.0
PyTorch 2.11.0+cu130
Architecture HYV3ForCausalLM
Quantization path compressed-tensors with NVFP4 MARLIN experts
KV cache BF16
Speculative decoding Native MTP, one speculative token
NCCL transport NET/IB over the direct-link interface

Both Ray nodes must use the same container or Python environment, an identical
checkpoint, and the same in-container model path. Set the communication
environment on both nodes before starting the Ray workers:

export NCCL_DEBUG=INFO
export NCCL_DEBUG_SUBSYS=INIT,NET,ENV
export NCCL_SOCKET_IFNAME=<200G_INTERFACE>
export GLOO_SOCKET_IFNAME=<200G_INTERFACE>
export NCCL_IB_DISABLE=0

Confirm that ray status reports two nodes and two GPUs, and that the NCCL log
selects NET/IB rather than NET/Socket. Then launch the server from the head
node:

vllm serve /models \
  --served-model-name superhy3 \
  --host 0.0.0.0 \
  --port 8600 \
  --tensor-parallel-size 2 \
  --distributed-executor-backend ray \
  --max-model-len 4096 \
  --max-num-seqs 1 \
  --max-num-batched-tokens 4096 \
  --gpu-memory-utilization 0.85 \
  --kv-cache-dtype bfloat16 \
  --speculative-config '{"method":"mtp","num_speculative_tokens":1}' \
  --generation-config vllm \
  --trust-remote-code \
  --enforce-eager \
  --load-format auto

DGX Spark loader note: do not substitute fastsafetensors for auto in
this two-node profile. On the verified Spark setup, fastsafetensors caused
excessive transient unified-memory pressure during loading. The default
loader uses lazy memory-mapped safetensors on local storage and completed the
load reliably. Also avoid the eager safetensors strategy, which reads an
entire file into CPU memory before loading it.

DGX Spark CPU and GPU allocations share one unified-memory pool. Before launch,
stop unrelated inference jobs and stale containers, then verify both nodes:

free -h
swapon --show
nvidia-smi
ray status

If a node becomes unresponsive during loading, inspect the previous boot's
kernel log before changing NCCL settings:

sudo journalctl -b -1 -k --no-pager | \
  grep -Ei 'oom|out of memory|killed process|NVRM|Xid|mlx5|rdma'
  • OOM, Out of memory, or Killed process indicates unified-memory pressure.
  • NVRM: Xid indicates a GPU, driver, or application fault; retain the Xid
    number and collect sudo nvidia-bug-report.sh.
  • An NCCL watchdog timeout that appears only after the peer disappears is
    normally a secondary failure, not the original cause.

For an initial diagnostic boot, MTP can be removed by omitting
--speculative-config. Add it back only after the base server reaches its ready
state. Increase context length and concurrency gradually after startup is
stable. See the
vLLM multi-node deployment guide
and NVIDIA Xid documentation
for deeper diagnostics.

from openai import OpenAI

client = OpenAI(base_url="http://127.0.0.1:8000/v1", api_key="EMPTY")
response = client.chat.completions.create(
    model="superhy3",
    messages=[
        {"role": "user", "content": "Explain mixture-of-experts routing."},
    ],
    temperature=0.9,
    top_p=1.0,
    extra_body={
        "chat_template_kwargs": {"reasoning_effort": "no_think"},
    },
)
print(response.choices[0].message.content)

Use reasoning_effort="high" for deeper reasoning and "no_think" for direct
responses. Keep BF16 KV cache on GB10-class hardware; uncalibrated lower-precision
KV cache can amplify stray-token behavior in Hy3 runtimes.

Quantized Editions

Edition Repository Best fit
NVFP4 W4A16 This repository vLLM / NVIDIA MARLIN deployments
MLX 4-bit Jiunsong/SuperHY3-abliterated-MLX-4bit High-memory Apple Silicon with Hy3 MLX support
GGUF IQ2_M Jiunsong/SuperHY3-abliterated-gguf llama.cpp on 128 GB unified-memory systems

Release Integrity

  • 99 safetensors shards opened successfully.
  • 139,298 indexed tensors matched 139,298 observed tensors.
  • 40 tensors were modified across 8 shards.
  • 0 routed expert tensors were modified.
  • Missing, extra, and wrong-shard tensor counts are all 0.
  • The release chat template matches the embedded tokenizer template.
  • The automated release gate passed with 0 blockers.
  • All 119 Hub files were checksum-verified after upload.

The repository includes the fusion report, release gate, official 500-item
benchmark record, refusal comparison, raw-response audit, OBLITERATUS execution
report, and SuperTune composition reports.

Limitations

  • GPQA Diamond and MMLU-Pro are lower than the original in this replay; the
    complete table is retained above.
  • Native fused NVFP4 benchmark and long-context throughput measurements were not
    run as part of this release validation.
  • Abliteration reduces refusal behavior and can produce content the original
    model would decline. Deployment policy and access control remain the
    operator's responsibility.
  • The MLX edition is a separately fused quantized checkpoint, not a conversion
    of these NVFP4 files.

License

Apache-2.0, following the base model and quantized source licenses.

README history 6 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-07-13Fix DGX Spark dual-node serving guide80dadd311.9 KB
    Loading...
  2. 2026-07-13Link the verified GGUF IQ2_M editionbc41c8c8.5 KB
    Loading...
  3. 2026-07-13Clarify packed parameter count in NVFP4 carde28d37a8.4 KB
    Loading...
  4. 2026-07-13Align release card with final Hub file count8df72018.1 KB
    Loading...
  5. 2026-07-13Redesign NVFP4 model card and add release visuals731389d8.1 KB
    Loading...
  6. 2026-07-12Add files using upload-large-folder toolcf0ffdd5.1 KB
    Loading...

Discussions 1 thread

  1. 2026-07-15A super paranoid model...open1 💬#1
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration