← back to catalog · registered 2026-08-22 13:56

gbuzhf/Huihui-Ornith-1.0-35B-abliterated-MTP-UD-APEX-GGUF

gbuzhf 35B GGUF MoE multimodal second-order
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/gbuzhf%2FHuihui-Ornith-1.0-35B-abliterated-MTP-UD-APEX-GGUF"
Response includes
  • classification m8
  • files 14
  • hub_downloads_all_time 5,439
  • author_summary 4 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M8
Primary method

Repackaging (quantization)

Applied on top of direct removal inherited from the base model.
Confidence
MEDIUM
Inherited from base model
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=1
  • assume M1 (base ablation) + M8 (GGUF quant) - default when producer unknown
Refusal direction extraction

No specific extraction method could be identified for this model. The producer either did not document it or used a proprietary pipeline.

What is a refusal direction? →
Downloads · lifetime
5K
1K last 30d - stable
Likes
1
Model age
8w ago
created 2026-08-14
Downloads over time
Now5.6K→from3.8K↑48%
3.5K4.3K5.1K5.8K3.8K on Aug 195.6K on Oct 11AugSepOct
Aug 19 → Oct 11 · 49 snapshots · spans 53 days

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
mit
Quantizations
IQ4 Q4_K Q5_K Q6_K
Tags
gguf mtp speculative-decoding apex v2d-lite moe vlm abliterated uncensored qwen3_5_moe image-text-to-text doi:10.57967/hf/9987

Related

Total size
168 GB
Files
14
Quantizations
6
Registered
2026-08-22 13:56
Last updated on HF
2026-09-12 02:12

Files by quantization

Q6_K 1 file 28.1 GB
Huihui-Ornith-1.0-35B-abliterated-MTP-UD-Q6_K.gguf 28.1 GB 1616f1a5 download
Q5_K 1 file 24.1 GB
Huihui-Ornith-1.0-35B-abliterated-MTP-UD-Q5_K_S.gguf 24.1 GB 5ba1d92c download
Q4_K 1 file 21.6 GB
Huihui-Ornith-1.0-35B-abliterated-MTP-UD-Q4_K_XL.gguf 21.6 GB 8a800e13 download
IQ4 1 file 17.4 GB
Huihui-Ornith-1.0-35B-abliterated-MTP-UD-IQ4_XS.gguf 17.4 GB 2dc3e124 download
F16 1 file 861 MB
Huihui-Ornith-1.0-35B-abliterated-mmproj-F16.gguf 861 MB 0a10749c download
Auxiliary files 9 files 76.4 GB
Huihui-Ornith-1.0-35B-abliterated-MTP-APEX-I-Balanced-v2D-lite.gguf 24.5 GB 1f905f22 download
Huihui-Ornith-1.0-35B-abliterated-MTP-APEX-I-Quality-v2D-lite.gguf 22.2 GB 23c37bc9 download
Huihui-Ornith-1.0-35B-abliterated-MTP-APEX-I-Compact-v2D-lite.gguf 16.4 GB ff059ecf download
Huihui-Ornith-1.0-35B-abliterated-MTP-APEX-I-Mini-v2D-lite.gguf 13.4 GB 3a7aa68c download
README.md 5.77 KB 01ddd06a download
.gitattributes 2.52 KB bcd3e194 download
gate_spec_bench.json 2.02 KB a5f53fce download
MANIFEST.txt 1.94 KB 7fb74f2b download
sha256sums.txt 1.09 KB d4fdfdc4 download

README current version from Hugging Face


license: mit
base_model:

  • huihui-ai/Huihui-Ornith-1.0-35B-abliterated-GGUF
    base_model_relation: quantized
    pipeline_tag: image-text-to-text
    library_name: gguf
    tags:
  • gguf
  • mtp
  • speculative-decoding
  • apex
  • v2d-lite
  • moe
  • vlm
  • abliterated
  • uncensored
  • qwen3_5_moe

Huihui-Ornith-1.0-35B-abliterated — MTP + UD + APEX-v2D-lite GGUF

huihui-ai's abliterated Ornith-1.0-35B
with an MTP head grafted in, the vision projector, and eight tiers.

Measured MTP draft acceptance: 86.4% (748/866), 44.1 t/s, --mmproj loaded.

This build is deliberately a one-variable change from the faithful
gbuzhf/Ornith-1.0-35B-MTP-UD-APEX-GGUF:
identical MTP head, identical tensor maps, identical imatrix. Only the trunk is
abliterated. That is what makes the two comparable at all.

faithful abliterated delta
MTP draft acceptance 87.4% 86.4% −1.0 pp
tokens/s (gate) 49.1 44.1 −5.0

Same harness, same prompt mix, so this comparison is like-for-like. Abliteration
cost about one point of draft acceptance
— the un-abliterated head still fits the
abliterated trunk nearly as well.

Tiers

Ranked by active bpw — routed experts weighted 8/256 because that is what
actually fires per token, token_embd excluded (it is a get_rows, not a matmul).
It is a measure of where the bits go, not a measured quality score.
Every file size matched its pre-build prediction to 0.01 GB.

tier GB active bpw file bpw
MTP-UD-Q6_K 30.21 8.063 6.804
MTP-UD-Q5_K_S 25.84 7.693 5.820
MTP-UD-Q4_K_XL 23.22 7.471 5.230
MTP-APEX-I-Balanced-v2D-lite 26.30 6.913 5.922
MTP-UD-IQ4_XS 18.68 6.762 4.209
MTP-APEX-I-Quality-v2D-lite 23.85 6.699 5.371
MTP-APEX-I-Compact-v2D-lite 17.57 5.228 3.956
MTP-APEX-I-Mini-v2D-lite 14.37 4.180 3.208

The MTP head is pinned Q8_0 in every tier (8.515 bpw measured). No imatrix can
reach blk.40 — it is grafted after calibration and never runs during an imatrix
pass — so it is unguided RTN either way, and draft acceptance converts directly
into tokens/sec.

Note the two inversions against file size. UD-IQ4_XS (18.68 GB) scores above
APEX-I-Quality-v2D-lite (23.85 GB) on the active path, and UD-Q4_K_XL (23.22 GB)
above APEX-I-Balanced-v2D-lite (26.30 GB). The UD maps pin attention, shared
experts and token_embd at Q8_0 at every tier, and roughly 55% of active
parameters are non-routed — attention, the shared expert and the 508 M-param output
head all run on every token, while routed experts contribute 8/256. If you are
choosing on quality per GB, start from the top of this table, not from file size.

How the MTP head got here

huihui-ai publishes the abliterated model as bf16 GGUF only — there are no
abliterated safetensors on the Hub. The usual path grafts 19 mtp.* tensors into
safetensors and then converts, which was impossible, so the graft was done one
stage later, directly at GGUF level:

+ 20 blk.40.* tensors, copied byte-for-byte from the faithful build's bf16 master
~ qwen35moe.block_count            40 -> 41
+ qwen35moe.nextn_predict_layers   (absent) -> 1

The donor head is Qwen3.6-35B-A3B's original MTP head (844.6 M params), which
passed a byte-identity HEAD GATE in the faithful build —
sha256 faac91f15cbe54475faa2578bedc46a7c29a947b8a3e7ef3ecd376ae079826ab. So this
repo provably carries the same head as the faithful one.

Verified before building: huihui's 733 tensor names are identical to the
reference, and all shapes agree once trailing 1s are normalised (huihui writes
ffn_gate_inp_shexp as [2048, 1], convert_hf_to_gguf.py writes [2048] — the
same 2048 values). A grafted-master gate then asserted 753 tensors, 20 in blk.40,
block_count=41, nextn_predict_layers=1.

Vision

Huihui-Ornith-1.0-35B-abliterated-mmproj-F16.gguf (0.90 GB) is required —
without it you have a blind model. It is huihui-ai's own projector, kept separate
so one copy serves all eight tiers.

llama-server -m Huihui-Ornith-1.0-35B-abliterated-MTP-UD-Q5_K_S.gguf \
  --mmproj Huihui-Ornith-1.0-35B-abliterated-mmproj-F16.gguf \
  -c 8192 -fa on --jinja \
  --spec-type draft-mtp,ngram-mod \
  --spec-draft-n-max 1 --spec-draft-n-min 0 --spec-draft-p-min 0.75 \
  --spec-ngram-mod-n-min 8 --spec-ngram-mod-n-max 24 --spec-ngram-mod-n-match 48

--mmproj silently force-disables ctx_shift and cache_reuse.

Honest limits — read these

  • The imatrix was computed on the UN-abliterated model. It is the same 2-way
    Ornith-native merge (bartowski calibration_datav5 + unsloth, 2463 chunks,
    1.26 M tokens) used by the faithful build. Abliteration changes activation
    statistics, so a base imatrix will misallocate precision to some degree, and on a
    256-expert MoE it can misdescribe routing too. This was a deliberate trade to
    keep the build on a CPU box and to hold every variable except the trunk constant.
    It is the one known-suboptimal ingredient here.
  • Acceptance is one gate run, 866 draft tokens, on a code-novel + copy-edit
    prompt mix. Comparable to the faithful build's 87.4% because it is the identical
    harness — not comparable to acceptance figures quoted from other harnesses.
  • Abliteration is not free. It moves weights off the trained optimum by design,
    and that cost is separate from and additional to quantization error. If you do not
    need refusal removal, use the faithful build.

Also included

  • BF16/ — the abliterated bf16 + MTP master (71.07 GB, 753 tensors), so any
    future tier rebuilds from it with no graft and no conversion.
  • sha256sums.txt

README history 3 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-08-15card: tighten the active-bpw note6a354995.8 KB
    Loading...
  2. 2026-08-15card: measured 86.4% MTP acceptance, 8 tiers, GB/active-bpw/file-bpw table5b0189c5.9 KB
    Loading...
  3. 2026-08-14init75b1067662 B
    Loading...

Discussions 1 thread

  1. 2026-08-15TEST RUN: Abliterated-MTP-Q5_K_Sopen2 💬#1
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration