← back to catalog · registered 2026-08-22 13:56

zerodigest/Ornith-1.5-35B-Uncensored-YMQ-MTP-GGUF

zerodigest 35B GGUF MoE second-order 262K ctx
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/zerodigest%2FOrnith-1.5-35B-Uncensored-YMQ-MTP-GGUF"
Response includes
  • classification m-uncensored
  • files 9
  • hub_downloads_all_time 30,585
  • author_summary 2 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M-U
Primary method

Uncensored (method unknown)

No other method signals detected in this model.
Confidence
LOW
Why this label 3 signals
Weak or ambiguous signals. Best guess based on catalog patterns; treat as tentative and check the evidence below.
  • 'uncensored' in name/tags but no 'abliterated' marker
  • method not identifiable from author declaration alone
  • may be DPO fine-tune, prompt engineering, or unknown technique
Refusal direction extraction

No specific extraction method could be identified for this model. The producer either did not document it or used a proprietary pipeline.

What is a refusal direction? →
Downloads · lifetime
31K
4K last 30d - stable
Likes
7
Model age
7w ago
created 2026-08-21
Downloads over time
Now32.5K→from897↑3,522%
011.9K23.8K35.6K897 on Aug 1932.5K on Oct 11AugSepOct
Aug 19 → Oct 11 · 48 snapshots · spans 53 days

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
apache-2.0
Tags
gguf text-generation quantizer autoround architecture-aware moe mixture-of-experts uncensored 35b base_model:0xKitkat/Ornith-1.5-35B-A3B-Uncensored base_model:quantized:0xKitkat/Ornith-1.5-35B-A3B-Uncensored license:apache-2.0

Related

Total size
56.1 GB
Files
9
Quantizations
3
Registered
2026-08-22 13:56
Last updated on HF
2026-09-04 16:37

Files by quantization

Q6_K 1 file 579 MB
mmproj-Ornith-1.5-35B-Uncensored-Q6_K.gguf 579 MB 903e4e3d download
Q4_K 1 file 473 MB
mmproj-Ornith-1.5-35B-Uncensored-Q4_K_S.gguf 473 MB c43b3754 download
Auxiliary files 7 files 56.1 GB
Ornith-1.5-35B-Uncensored-YMQ-L.gguf 17.4 GB 6f494908 download
Ornith-1.5-35B-Uncensored-YMQ-M.gguf 15.1 GB 468eb97c download
Ornith-1.5-35B-Uncensored-YMQ-S.gguf 12.5 GB 78c038b3 download
Ornith-1.5-35B-Uncensored-YMQ-XS.gguf 11.1 GB b4427c6b download
logo.jpg 167 KB 53d427f3 download
README.md 9.47 KB fdf09ecc download
.gitattributes 1.97 KB d9b7d67a download

README current version from Hugging Face


license: apache-2.0
base_model: 0xKitkat/Ornith-1.5-35B-A3B-Uncensored
library_name: gguf
tags:

  • text-generation
  • gguf
  • quantizer
  • autoround
  • architecture-aware
  • moe
  • mixture-of-experts
  • uncensored
  • 35b

Ornith-1.5-35B-Uncensored-YMQ-MTP-GGUF

Source Model: 0xKitkat/Ornith-1.5-35B-A3B-Uncensored

⚖️ An Architecture-Aware, AutoRound-Inspired MoE Mixed Precision Layout

This repository features advanced, custom architecture-aware quantizations of Ornith-1.5-35B (Uncensored) processed directly from official raw BF16 source files using the custom YMQ-Compiler (v2.0) log-space framework.

These builds natively preserve the delicate Mixture-of-Experts routing topology and utilize high-context optimization parameters tailored for demanding local agent execution environments (such as RooCode/Aider).

ZeroDigest YMQ Logo


📊 Quantization Preset Tier Details

Preset Tier Total Size Target Usage / Memory VRAM Profile Cognitive Real-World Coding Quality
XS ~11.5 GB 12GB GPU Entry / Balanced Budget Setup The 12GB VRAM Champion. Fully re-engineered layout to protect structural tracking in compact footprints.
S ~13.4 GB Light workspace / Low-VRAM cache headroom Balanced economy. Linear compression baseline for standard workflows.
M ~16.2 GB The High-Context Sweet Spot (Recommended) 🎯 Elite logical stability. Optimizes VRAM to leave a massive headroom buffer for deep agent context loops.
L ~18.6 GB Premium Single-GPU Processing / Heavy workloads Near lossless. Solid performance scaling across heavier consumer setups.

📉 Perplexity Evaluation Metrics (WikiText-2)

The following metrics show the mathematical quality preservation of the YMQ-Compiler log-space cluster analysis compared to standard linear quantization layouts. Tested natively via llama-perplexity over a 4096 context window using the official WikiText-2 test corpus.

Model Preset Variant File Size Perplexity Score (Lower is Better) Cognitive Calibration Verdict
XS ~11.5 GB 13.5942 The 12GB VRAM Champion. Re-insulated fallback grids.
S ~13.4 GB 13.5454 Balanced economy boundary. Linear compression baseline for standard workflows.
M (Recommended) ~16.2 GB 12.4310 🎯 The High-Context Sweet Spot. Elite logical stability with massive VRAM headroom.
L ~18.6 GB 12.5107 Near lossless. Solid performance scaling across heavier consumer setups.

💡 The MoE Compression Breakthrough: YMQ vs. Standard Quants

Standard quantization pipelines act like a blunt hammer. They apply a uniform, flat bit-depth across every layer of the model, which completely breaks the delicate routing paths of Mixture-of-Experts (MoE) architectures.

Independent local testing using the official llama-perplexity harness exposes the massive optimization gap between standard, flat layouts and the architecture-aware YMQ-Compiler:

  • Standard Flat Q4_K_M (~21.0 GB):

    • Hits a high 13.4757 perplexity score on the test corpus.
    • Starves the core attention entry channels and native MTP speculator tracking paths of bit-depth resolution.
    • Model experiences severe tracking fatigue under stress.
  • YMQ-Compiler M Preset (~16.0 GB):

    • Scores a spectacular 12.4310 perplexity score on the exact same corpus.
    • Surgically isolates and insulates critical logic entry gates behind high-fidelity shields (Q5_K / Q6_K).
    • Heavily compresses background expert layers into non-linear matrices.
    • Saves a massive 5 Gabytes of VRAM while delivering a drastic leap in logical reasoning clarity.

The Practical Result

The practical result is an elite, filter-free developer build that completely eliminates late-stage context amnesia bugs.

  • Flat, uniform community quants suffer from endless reading loops and failed edits past 30k tokens.
  • YMQ-M and S presets stream massive 200k+ context windows completely in VRAM.
  • Blistering multi-token speculation speeds (80–120 t/s).
  • Zero loop failures!

⚖️ YMQ vs. Uniform Quantization (The AutoRound Philosophy)

Standard quantization pipelines apply a blunt, uniform bit-depth across every single layer in a model. This blunt approach completely collapses the delicate routing structures of large Mixture-of-Experts (MoE) architectures, starving critical logic anchors of necessary precision while bloating file sizes with idle parameters.

The YMQ-Compiler implements a post-training optimization philosophy similar to advanced weight-tuning frameworks like Intel's AutoRound:

  • Targeted Bit Isolation: Instead of applying a flat matrix mask, YMQ operates strictly in Log-Space. It surgically identifies the core attention entry lanes and high-leverage routing networks, locking them behind heavy Q5_K and Q6_K high-fidelity safety shields.
  • Expert Layer Flattening: It takes the massive pool of background expert weights (ffn_*_exps) and compresses them aggressively using dense, non-linear grids (IQ2_XS and IQ4_NL). Because these background parameters make up the majority of the file footprint but are rarely active at the same time, the compiler shaves off gigabytes of background noise without breaking the model's main train of thought.

The result is a highly stable, custom mixed-precision portfolio that matches the low perplexity and high instruction clarity of premium optimized configurations, while allowing 16GB and 24GB single-GPU setups to stream massive context windows with absolute structural peace of mind!


🛠️ The YMQ Compilation Architecture

Standard quantization pipelines treat network tensors like a flat dataset, applying destructive blanket low-bit compression to delicate tracking networks. The YMQ-Compiler solves high-context logic decay by parsing model files dynamically via an automated, multi-tiered protection matrix:

  1. Log-Space Gap Detection Clustering: Instead of flat percentage thresholds, the engine computes statistical cluster variances in log-space, successfully isolating intermediate logical reasoning spikes and elevating them to stable non-linear 4-bit (IQ4_XS) formats, while compressing idle fact-storage layers to aggressive 2-bit baselines.
  2. Fading Boundary Tapering: Recognizes the extreme fragility of initial token entry data vectors, forcing an input wave cushion (L00=IQ4_NL → L01=IQ4_XS → L02=IQ3_XXS) that gradually stabilizes parameters before hitting the fallback pools.
  3. Dedicated Gate Insulation: Hard-shields volatile parallel Transformer Multi-Head Attention and MoE expert routing paths, keeping context tracking perfectly noise-free.
  4. Asymmetric Vocabulary Shielding: Fixes tied-weight boundary errors by mapping the final logit classification exit heads to robust configurations to completely eliminate formatting loops and API tag leakage under deep contexts.

🚀 Recommended Runtime Parameters (llama.cpp / llama-server)

$./llama-server -m Ornith-1.5-35B-Uncensored-YMQ-M.gguf -ctk q8_0 -ctv q4_0 --ctx-size 245760 --mmproj mmproj-Ornith-1.5-35B-Uncensored-Q6_K.gguf \
  --n-cpu-moe 5 --timeout 36000 --checkpoint-min-step 2048 --ctx-checkpoints 4 --spec-type draft-mtp --spec-draft-n-max 2 \
  --n-predict -1 --temp 0.6 --top-p 0.95 --top-k 20 --repeat-penalty 1.05 --jinja -fa

🖼️ Vision Projection (--mmproj)

For multimodal vision support, pair these builds with one of the following projection files:

Variant File Size Notes
Full Precision (BF16) mmproj-Ornith-1.5-35B-Uncensored-BF16.gguf ~903 MB Full-precision vision tower, native to the Uncensored abliteration weights. Maximum fidelity for image reasoning tasks.
Q6_K mmproj-Ornith-1.5-35B-Uncensored-Q6_K.gguf (this repo) ~579 MB High-fidelity quantized vision tower. Excellent quality-to-size balance with minimal perceptible degradation.
Q4_K_S mmproj-Ornith-1.5-35B-Uncensored-Q4_K_S.gguf (this repo) ~473 MB Compact vision projection for VRAM-constrained setups. Retains strong image understanding at reduced footprint.

Pass via --mmproj <path-to-file> in your llama-server invocation (see example above).


☕ Support & Future R&D

If the YMQ-Compiler builds saved your context window from collapsing or optimized your active development cycle speeds, consider buying a coffee to fund further low-level optimization research. Your support keeps the server nodes baking future model scales!

👉 Support ZeroDigest Research on ko-fi


📦 Source Framework & Automation Code

The compiler pipeline automation engine, setup thresholds, and structural mapping rules are open-source. To view the implementation details or compile your own custom models natively using this profile layout, visit the official development hub:

👉 GitHub: ZeroDigest / YMQ-Compiler

README history 12 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-09-04Update README.mdb65152612 KB
    Loading...
  2. 2026-09-04Update README.mdb8ec57212 KB
    Loading...
  3. 2026-09-01Update README.md7f32f3d9.6 KB
    Loading...
  4. 2026-08-21Update README.mdf857bfe9.5 KB
    Loading...
  5. 2026-08-21Update README.md15f50089.5 KB
    Loading...
  6. 2026-08-21Update README.md60d561a9.5 KB
    Loading...
  7. 2026-08-21Update README.md627c6459.5 KB
    Loading...
  8. 2026-08-21Update README.md64472bc8.5 KB
    Loading...
  9. 2026-08-21Update README.mdc5772648.3 KB
    Loading...
  10. 2026-08-21Update README.md805cb448.3 KB
    Loading...
  11. 2026-08-21Update README.md09ee7ab8.3 KB
    Loading...
  12. 2026-08-21Create README.mdd3e10e28.5 KB
    Loading...

Discussions 2 threads

  1. 2026-10-09Is seems litl bit buggyopen1 💬#2
    Loading...
  2. 2026-08-27Qwen3.6-35B-A3B and KAT Coder V2.5 - 35B-A3B requestopen4 💬#1
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration