← back to catalog · registered 2026-08-22 13:56

gbuzhf/Ornith-1.5-35B-A3B-Abliterated-MTP-UD-APEX-GGUF

gbuzhf 35B GGUF MoE multimodal second-order 262K ctx
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/gbuzhf%2FOrnith-1.5-35B-A3B-Abliterated-MTP-UD-APEX-GGUF"
Response includes
  • classification m8
  • files 17
  • hub_downloads_all_time 76,212
  • author_summary 4 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M8
Primary method

Repackaging (quantization)

Applied on top of direct removal inherited from the base model.
Confidence
MEDIUM
Inherited from base model
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=1
  • assume M1 (base ablation) + M8 (GGUF quant) - default when producer unknown
Refusal direction extraction

No specific extraction method could be identified for this model. The producer either did not document it or used a proprietary pipeline.

What is a refusal direction? →
Downloads · lifetime
76K
37K last 30d - stable
Likes
40
Model age
7w ago
created 2026-08-22
Downloads over time
Now94.5K→from245↑38,464%
034.6K69.3K103.9K245 on Aug 1994.5K on Oct 11AugSepOct
Aug 19 → Oct 11 · 48 snapshots · spans 53 days

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
mit
Quantizations
IQ4 Q5_K Q6_K
Tags
gguf ICE ICE-Tiers ice-quant 9-Tier-Standard Tiel-Calibrated mtp speculative-decoding apex v2d-lite unsloth-dynamic moe

Related

Total size
181 GB
Files
17
Quantizations
4
Registered
2026-08-22 13:56
Last updated on HF
2026-08-30 07:30

Files by quantization

Q6_K 1 file 28.1 GB
Ornith-1.5-35B-A3B-Abliterated-MTP-UD-Q6_K.gguf 28.1 GB 03c33fe7 download
Q5_K 1 file 24.1 GB
Ornith-1.5-35B-A3B-Abliterated-MTP-UD-Q5_K_S.gguf 24.1 GB f9c2793b download
IQ4 1 file 17.4 GB
Ornith-1.5-35B-A3B-Abliterated-MTP-UD-IQ4_XS.gguf 17.4 GB 0f971a1f download
Auxiliary files 14 files 111 GB
Ornith-1.5-35B-A3B-Abliterated-MTP-25G-ICE.gguf 23.1 GB 8e8a11ad download
Ornith-1.5-35B-A3B-Abliterated-MTP-23G-ICE.gguf 21.3 GB e06ef326 download
Ornith-1.5-35B-A3B-Abliterated-MTP-21G-ICE.gguf 19.4 GB 0abd3d1d download
Ornith-1.5-35B-A3B-Abliterated-MTP-19G-ICE.gguf 17.5 GB cb52afd7 download
Ornith-1.5-35B-A3B-Abliterated-MTP-APEX-I-Compact-v2D-lite.gguf 16.4 GB 71118667 download
Ornith-1.5-35B-A3B-Abliterated-MTP-APEX-I-Mini-v2D-lite.gguf 13.4 GB 1a8a35bb download
Ornith-1.5-35B-A3B-imatrix.gguf 183 MB 8d5b1693 download
Ornith-1.5-35B-A3B-calibration-v6.txt 1.17 MB ca482e71 download
KLD_RESULTS.txt 25.3 KB 3ddf29c2 download
README.md 10.8 KB 540e4e30 download
.gitattributes 2.32 KB 3785c1c3 download
gate_spec_bench.json 2.01 KB 4b88951d download
MANIFEST.txt 1.07 KB d073434d download
sha256sums.txt 1.03 KB 2e48a89a download

README current version from Hugging Face


license: mit
base_model:

  • PocketAiHub/Ornith-1.5-35B-A3B-Abliterated
    base_model_relation: quantized
    pipeline_tag: image-text-to-text
    library_name: gguf
    tags:
  • gguf
  • ICE
  • ICE-Tiers
  • ice-quant
  • 9-Tier-Standard
  • mtp
  • speculative-decoding
  • apex
  • v2d-lite
  • unsloth-dynamic
  • moe
  • vlm
  • qwen35moe
  • imatrix
  • abliterated
  • uncensored

Ornith-1.5-35B-A3B Abliterated — MTP + UD + ICE + APEX GGUF

Nine Pareto-optimal tiers of the abliterated model, each with an MTP head grafted
in
from the original Ornith-1.5 (the abliterated source ships none).

Measured on wikitext-2-raw, 16 chunks x 2048 ctx, against two references, so
quantization damage and abliteration damage can be told apart.

How much did abliteration itself change the model?

Mean KLD (abliterated BF16 vs original BF16) 0.0151
same top-1 token 95.04%

For scale, the best quantization in this ladder costs ~0.022 KLD. Abliteration is
a smaller perturbation than Q6_K quantization.

Tiers (9/9 measured)

tier size mean KLD 99% KLD 99.9% KLD PPL ratio same top-1 active bpw file bpw overall
MTP-UD-Q6_K 30.21 GB 0.0222 0.226 0.695 0.9960 94.05% 8.063 6.804 96.7
MTP-UD-Q5_K_S 25.84 GB 0.0261 0.269 0.899 0.9871 93.51% 7.693 5.820 96.3
MTP-25G-ICE 24.85 GB 0.0293 0.300 1.280 0.9856 93.44% 7.686 5.597 96.0
MTP-23G-ICE 22.84 GB 0.0345 0.345 1.151 0.9902 92.39% 7.523 5.143 95.4
MTP-21G-ICE 20.85 GB 0.0398 0.400 1.655 0.9943 92.06% 7.357 4.695 94.9
MTP-19G-ICE 18.82 GB 0.0612 0.630 2.327 1.0025 90.02% 7.192 4.240 93.0
MTP-UD-IQ4_XS 18.68 GB 0.0706 0.695 2.645 1.0502 89.40% 6.762 4.209 92.2
MTP-APEX-I-Compact-v2D-lite 17.57 GB 0.0925 0.893 3.141 1.0169 87.97% 5.228 3.956 90.5
MTP-APEX-I-Mini-v2D-lite 14.37 GB 0.2546 2.410 5.578 1.2129 80.70% 4.180 3.208 80.0

Sorted best -> worst by overall (BF16 = 100), the same composite used on the
non-abliterated card:
0.70/(1+meanKLD) + 0.30*sameTop1. All KLD columns are measured
against the abliterated BF16, i.e. they isolate what the quantization costs.

Read the tail columns with care. 99.9% KLD is the ~33rd-worst token out of
32,768 — an extreme order statistic with large sampling variance, so it inverts
between adjacent tiers without that meaning anything. 99% KLD rests on ~328 tokens
and orders all nine tiers monotonically; mean KLD uses all 32,768 and separates the
closest pair by 4.3 sigma. Rank on mean KLD; treat the tail columns as shape,
not order.

Abliterated vs non-abliterated, same recipe

tier KLD
abl
KLD
clean
Δ top-1
abl
top-1
clean
Δ PPL ratio
abl
PPL ratio
clean
Δ
MTP-UD-Q6_K 0.0222 0.0221 +0.6% 94.05% 93.85% +0.20 pp 0.9960 0.9957 +0.0003
MTP-UD-Q5_K_S 0.0261 0.0272 -4.2% 93.51% 93.51% +0.01 pp 0.9871 0.9862 +0.0009
MTP-25G-ICE 0.0293 0.0303 -3.4% 93.44% 93.16% +0.28 pp 0.9856 0.9814 +0.0042
MTP-23G-ICE 0.0345 0.0361 -4.4% 92.39% 92.65% -0.26 pp 0.9902 0.9885 +0.0018
MTP-21G-ICE 0.0398 0.0412 -3.4% 92.06% 92.03% +0.03 pp 0.9943 0.9924 +0.0019
MTP-19G-ICE 0.0612 0.0608 +0.5% 90.02% 90.32% -0.30 pp 1.0025 1.0030 -0.0005
MTP-UD-IQ4_XS 0.0706 0.0723 -2.4% 89.40% 89.46% -0.06 pp 1.0502 1.0526 -0.0024
MTP-APEX-I-Compact-v2D-lite 0.0925 0.0954 -3.0% 87.97% 87.83% +0.14 pp 1.0169 1.0101 +0.0068
MTP-APEX-I-Mini-v2D-lite 0.2546 0.2608 -2.4% 80.70% 80.49% +0.21 pp 1.2129 1.2281 -0.0151

Δ is near zero or slightly negative across the ladder: abliteration does not make
this model harder to quantize
, and in the mid-range it is marginally easier —
plausibly because projecting a direction out of ffn_down narrows its dynamic range.

Two further results, measured against the original BF16 as well (full numbers in
KLD_RESULTS.txt):

  • The two damages are strongly sub-additive — abliteration and quantization are
    largely orthogonal, so the combined figure sits far below their sum.
  • No tier un-abliterates. Across the 9 rungs measured, the distance to the original BF16 stays above abliteration's own distance (0.0151), so quantization never pulls the model back toward the refusal behaviour.

Which tier is which

family what it is
UD-* Unsloth Dynamic 2.0 maps, replayed 1:1. Pins attention, the shared expert and token_embd at Q8_0 at every size and moves only the routed experts.
ICE-* Bits allocated by how far a quantization error travels, not by activation magnitude. Named by target size.
APEX-I-*-v2D-lite mudler's APEX maps plus one extra step on attn_k/attn_v in the ten full-attention blocks and on the output head.

Rule of thumb: Q6_K / Q5_K_S near-lossless, 25G/23G-ICE the quality sweet
spot, 21G/19G-ICE the best small tiers, Compact/Mini only if you are tight on
VRAM — Mini drops sharply.

The ICE tier, and why this is a 9-tier release

ICE allocates bits by error travel distance: how far a quantization error
propagates before it reaches the output. Tensors writing straight into the residual
stream, the always-on dense path, and the router are protected; the routed expert
stack — 93% of the parameters but only 8-of-256 active per token — is left uniform.
Routers stay F32, and the draft block is un-pinned so its experts follow the tier.

On the non-abliterated ladder, measured on this model and this harness, ICE lands
ahead at matched size:

comparison result
23G-ICE vs UD-Q4_K_XL (same size) -5.0% KLD
23G-ICE vs APEX-I-Quality -13.0% KLD and 0.61 GB smaller
25G-ICE vs APEX-I-Balanced -12.2% KLD and 1.15 GB smaller

UD-Q4_K_XL,APEX-I-Quality-v2D-lite and APEX-I-Balanced-v2D-lite are each already covered by
an ICE tier that is both smaller and closer to BF16, so rebuilding them would add size
without adding a quality point. Full derivation, the refuted ffn_down rule
and the measured convexity bound are in the
original Ornith-1.5 card.

The 9-Tier Standard

From this release onward these nine recipes are the standard ladder:
UD-Q6_K · UD-Q5_K_S · 25G-ICE · 23G-ICE · 21G-ICE · 19G-ICE ·
UD-IQ4_XS · APEX-I-Compact-v2D-lite · APEX-I-Mini-v2D-lite

They are the measured Pareto frontier of a 12-tier sweep on this architecture:
every dropped tier is beaten on both size and KLD by one that ships. Reference
measurements and methodology:
Ornith-1.5-35B-MTP-UD-APEX-GGUF.

What was done to the source

PocketAiHub/Ornith-1.5-35B-A3B-Abliterated-GGUF is faithful at the tensor level —
every tensor except ffn_down is byte-identical to ornith-ai's BF16. Abliteration is
confined to ffn_down (both routed and shared experts), layers 15-39, at
~1.6-1.9e-02 L1-relative. Routers, attention, ffn_gate, ffn_up and layers 0-14
are untouched.

Three things were repaired while grafting:

  1. MTP head restored — 20 blk.40.* tensors from ornith-ai's BF16;
    block_count 40 -> 41, nextn_predict_layers added.
  2. tokenizer.ggml.add_bos_token restored to False — the source omits the key
    entirely, so llama.cpp falls back to its own default and tokenises differently
    from the original. Left unfixed this also invalidates any KLD against the original.
  3. tokenizer.chat_template restored — the source ships a 7536-byte copy with
    the multi-system-message merge block removed; the original is 7828 bytes.

imatrix

bartowski's Ornith-1.5-35B-A3B-imatrix.gguf, reused unmodified. Justified by
measurement, not assumption: the routers are byte-identical between the original
and the abliterated model
, so the same experts fire and the per-channel statistics
still apply.

Recipes

All nine tensor maps were confirmed byte-exact against the corresponding shipped
non-abliterated tier before this build, so the two ladders are directly comparable:

  • UD — replayed 1:1 from unsloth/Ornith-1.0-35B-GGUF (Unsloth Dynamic 2.0).
  • APEX v2D-lite — mudler's Ornith-1.5 maps, with attn_k/attn_v on the ten
    full-attention blocks and the output head each raised one step.
  • ICE — bits allocated by how far a quantization error travels rather than by
    activation magnitude; the expert stack is uniform and the draft block un-pinned.

MTP / speculative decoding

Every tier carries the head, pinned Q8_0 (F16 attn_k/attn_v on the ICE tiers).
It is grafted from the original model, so the draft head is not abliterated while
the trunk is
— worth knowing if you rely on the refusal behaviour under drafting:
the head only proposes, the abliterated trunk verifies, so accepted tokens are always
the trunk's.

Measured draft acceptance: 96.77% (390/403) on 23G-ICE, --spec-type draft-mtp, text-only.

prompt set acceptance
code-novel 98.45%
structured 98.29%
copy-edit 97.30%
prose-novel 91.57%

Raw run in gate_spec_bench.json. Acceptance depends on the prompt mix — compare only against numbers taken on the same harness.

llama-server -m Ornith-1.5-35B-A3B-Abliterated-MTP-21G-ICE.gguf \
  --mmproj mmproj-Ornith-1.5-35B-A3B-Abliterated-F16.gguf \
  -c 8192 -fa on --jinja \
  --spec-type draft-mtp --spec-draft-n-max 1 --spec-draft-n-min 0 --spec-draft-p-min 0.75

Vision

Not re-hosted — use the projector from the source repo:
mmproj-Ornith-1.5-35B-A3B-Abliterated-F16.gguf
(0.90 GB). Without it the model is blind. Note --mmproj force-disables ctx_shift
and cache_reuse.

Also included

KLD_RESULTS.txt (raw llama-perplexity --kl-divergence output, both references),
sha256sums.txt, MANIFEST.txt.

The BF16 masters are not re-hosted: the abliterated source is at
PocketAiHub
and the original at ornith-ai.

Credit: PocketAiHub for the abliteration, bartowski for the imatrix and for
publishing its corpus, mudler for the APEX reference maps, Unsloth for the
UD 2.0 maps.

README history 20 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-08-30Update README.md04eb5f914.9 KB
    Loading...
  2. 2026-08-30Update README.md8c6337c14.9 KB
    Loading...
  3. 2026-08-30Fix: TIEL rows belong only in the card table, not the abl-vs-clean tablee36d64314.8 KB
    Loading...
  4. 2026-08-30Add four TIEL_Calibrated ICE tiers to the card table7a9f55b15.2 KB
    Loading...
  5. 2026-08-26Update README.mdf90b47e13.8 KB
    Loading...
  6. 2026-08-26Update README.mde6a6ed613.8 KB
    Loading...
  7. 2026-08-26Update README.md35a049e13.8 KB
    Loading...
  8. 2026-08-24ICE tiers rebuilt: updated measurements, digests and manifest0609e3613.7 KB
    Loading...
  9. 2026-08-23Update README.mdd7f1b2612.6 KB
    Loading...
  10. 2026-08-23Update README.mdde097e412.9 KB
    Loading...
  11. 2026-08-23Update README.mde09ddd912.7 KB
    Loading...
  12. 2026-08-23Update README.md2d75a0c12.6 KB
    Loading...
  13. 2026-08-23Update README.mdaf4b6c112.7 KB
    Loading...
  14. 2026-08-23Update README.md575bb7c12.6 KB
    Loading...
  15. 2026-08-23card: MTPv2 head (new native head from ornith-ai)efb290413 KB
    Loading...
  16. 2026-08-22Update README.mde1ed55210.8 KB
    Loading...
  17. 2026-08-22card: 9/9 measuredd7c9ee811.3 KB
    Loading...
  18. 2026-08-22card: 9/9 measured66e55ac10.9 KB
    Loading...
  19. 2026-08-22card: 9/9 measured60431ad10.6 KB
    Loading...
  20. 2026-08-22card: 9/9 measured466aff110.6 KB
    Loading...

Discussions 6 threads

  1. 2026-09-15running issues - Mini quantopen3 💬#6
    Loading...
  2. 2026-08-30ICE quants with peculiar-ragdoll's Tiel iMatrix on *Abliterated* trunkopen3 💬#5
    Loading...
  3. 2026-08-24ICE Tiers got upgrade, you can re-download.open1 💬#4
    Loading...
  4. 2026-08-23MTPv2open5 💬#3
    Loading...
  5. 2026-08-22Abliteration Resultsopen32 💬#2
    Loading...
  6. 2026-08-22About "Uncensored" "Abliterated"tagclosed2 💬#1
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration