← back to catalog · registered 2026-08-22 13:56

williamliao/Ornith-1.0-35B-AEON-Ultimate-Uncensored-NVFP4-MTP-GGUF

williamliao Qwen 35B GGUF MoE 262K ctx
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/williamliao%2FOrnith-1.0-35B-AEON-Ultimate-Uncensored-NVFP4-MTP-GGUF"
Response includes
  • classification m-uncensored
  • files 3
  • hub_downloads_all_time 4,752
  • author_summary 2 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M-U
Primary method

Uncensored (method unknown)

No other method signals detected in this model.
Confidence
LOW
Why this label 3 signals
Weak or ambiguous signals. Best guess based on catalog patterns; treat as tentative and check the evidence below.
  • 'uncensored' in name/tags but no 'abliterated' marker
  • method not identifiable from author declaration alone
  • may be DPO fine-tune, prompt engineering, or unknown technique
Refusal direction extraction

No specific extraction method could be identified for this model. The producer either did not document it or used a proprietary pipeline.

What is a refusal direction? →
Downloads · lifetime
5K
32 last 30d - cooling
Likes
0
Model age
3mo ago
created 2026-07-03
Downloads over time
Now4.8K→from0↑0%
01.7K3.5K5.2K0 on Jul 14.8K on Oct 11JulAugSepOct
Jul 1 → Oct 11 · 54 snapshots · spans 102 days

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
apache-2.0
Tags
llama.cpp gguf nvfp4 fp4 modelopt mtp embedded-mtp speculative-decoding qwen3.6 qwen moe experimental

Related

Total size
22.0 GB
Files
3
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-07-04 06:48

Files by quantization

Auxiliary files 3 files 22.0 GB
Ornith-1.0-35B-AEON-Ultimate-Uncensored-NVFP4-MTP.gguf 22.0 GB 7e835287 download
README.md 7.88 KB b2c50646 download
.gitattributes 483 B fc430ad7 download

README current version from Hugging Face


license: apache-2.0
base_model:

  • deepreinforce-ai/Ornith-1.0-35B
  • AEON-7/Ornith-1.0-35B-AEON-Ultimate-Uncensored-NVFP4
    base_model_relation: quantized
    library_name: llama.cpp
    tags:
  • gguf
  • llama.cpp
  • nvfp4
  • fp4
  • modelopt
  • mtp
  • embedded-mtp
  • speculative-decoding
  • qwen3.6
  • qwen
  • moe
  • experimental
    pipeline_tag: text-generation

Ornith-1.0-35B-AEON-Ultimate-Uncensored-NVFP4-MTP-GGUF

Experimental GGUF conversion of AEON-7/Ornith-1.0-35B-AEON-Ultimate-Uncensored-NVFP4 with an embedded native Qwen3.6 MTP block.

This is a self-contained NVFP4 GGUF with embedded native MTP tensors.

It does not require an external MTP-only draft model, but this embedded-MTP path is experimental.

Status

Experimental.

The model loads and performs native MTP speculative decoding in llama.cpp. However, on the tested system, using the same MTP tensors through a standalone MTP-only GGUF via --model-draft was slightly faster than embedding the tensors directly into the GGUF.

For most users, the standard NVFP4 GGUF plus external --model-draft is recommended.

Model Details

  • Source model: AEON-7/Ornith-1.0-35B-AEON-Ultimate-Uncensored-NVFP4
  • Base model: deepreinforce-ai/Ornith-1.0-35B
  • Format: GGUF
  • Quantization: NVIDIA NVFP4
  • Architecture: Qwen3.6 Mixture-of-Experts with hybrid attention
  • MTP: Embedded native Qwen3.6 MTP block
  • Parameters: 35B total, approximately 3B activated per token
  • Purpose: Local inference and native MTP speculative decoding with llama.cpp

Embedded MTP

This model was created by grafting the native Qwen3.6 MTP block into the converted NVFP4 GGUF.

The resulting GGUF contains:

qwen35moe.block_count = 41
qwen35moe.nextn_predict_layers = 1

Additional tensors:

blk.40.*

Only the native MTP block was added. The original transformer layers blk.0.* through blk.39.* were not modified.

Recommended graft source:

Qwen3.6-35B-A3B-MTP-ONLY-Q6_K.gguf

Compatibility

A recent version of llama.cpp with Qwen3.6 MoE, NVFP4, and native MTP support is required.

Tested with:

  • Windows
  • NVIDIA GeForce RTX 5070 Ti 16 GB
  • NVIDIA GeForce RTX 5060 Ti 16 GB
  • llama.cpp CUDA backend
  • Embedded native Qwen3.6 MTP speculative decoding

Older llama.cpp builds may fail to recognize the nvfp4 tensor type, may not correctly load associated scale tensors, or may lack compatible Qwen3.6 MoE/MTP support.

Performance may vary with llama.cpp build, GPU split, context size, KV-cache format, prompt, sampling settings, and other runtime options.

Usage

llama-server

llama-server \
  -m Ornith-1.0-35B-AEON-Ultimate-Uncensored-NVFP4-MTP.gguf \
  --spec-type draft-mtp \
  --spec-draft-n-max 3 \
  --flash-attn on \
  -ngl 99

llama-cli

llama-cli \
  -m Ornith-1.0-35B-AEON-Ultimate-Uncensored-NVFP4-MTP.gguf \
  --spec-type draft-mtp \
  --spec-draft-n-max 3 \
  --flash-attn on \
  -ngl 99

Multi-GPU example

llama-server \
  -m Ornith-1.0-35B-AEON-Ultimate-Uncensored-NVFP4-MTP.gguf \
  --spec-type draft-mtp \
  --spec-draft-n-max 3 \
  --split-mode layer \
  --tensor-split 1,1 \
  --flash-attn on \
  -ngl 99

The best tensor split depends on available VRAM, GPU speed, PCIe topology, context length, and KV-cache placement. An even split is only a starting point.

Runtime Recommendation

Although the embedded MTP GGUF functions correctly, local benchmarks showed that using the standalone MTP-only GGUF via --model-draft was slightly faster than embedding the same tensors directly into the GGUF.

Recommended for most users:

llama-server \
  -m Ornith-1.0-35B-AEON-Ultimate-Uncensored-NVFP4.gguf \
  --model-draft Qwen3.6-35B-A3B-MTP-ONLY-Q6_K.gguf \
  --spec-type draft-mtp \
  --spec-draft-n-max 3

Use this embedded-MTP GGUF mainly if you want a self-contained file or want to experiment with native MTP grafting.

Benchmark Summary

Local mixed-task benchmark:

Mode Aggregate acceptance Wall time Notes
No MTP n/a 21.88 s Baseline (standard decoding)
External MTP-only (--model-draft, Q6_K) 95.09% 22.33 s Recommended; highest acceptance and best speculative decoding performance in local tests
Embedded MTP 93.17% 24.51 s Functional, but slower than the external draft model in local tests

The embedded model works correctly, but --model-draft currently appears to be the better runtime path on the tested setup.

Results may vary depending on llama.cpp version, CUDA graph behavior, GPU offload, batch/ubatch settings, context length, quantization, and prompts.

Suggested MTP Settings

A reasonable general-purpose starting point is:

--spec-draft-n-max 3

For translation, role-play, creative writing, conversational output, or open-ended explanations, start with:

--spec-draft-n-max 2

For JSON, fixed templates, repeated patterns, and deterministic code completion, try:

--spec-draft-n-max 4

Grafting Method

The embedded MTP block was generated by grafting the standalone MTP-only GGUF.

This demonstrates that MTP-only GGUF files can serve two purposes:

  1. Standalone draft model via --model-draft
  2. Graft source for creating a self-contained GGUF

The graft process updates:

qwen35moe.block_count: 40 -> 41
qwen35moe.nextn_predict_layers: 1

and appends:

blk.40.*

It does not modify:

tokenizer
vocabulary
chat template
RoPE settings
original transformer layers

Verification

Check the embedded MTP metadata and tensors:

python .\gguf-py\gguf\scripts\gguf_dump.py Qwen3.6-35B-A3B-NVFP4-MTP.gguf |
  Select-String "block_count|nextn_predict_layers|blk\.40"

Expected output should include:

qwen35moe.block_count = 41
qwen35moe.nextn_predict_layers = 1
blk.40.attn_k.weight
blk.40.nextn.eh_proj.weight
blk.40.nextn.shared_head_norm.weight

Notes

  • This GGUF preserves NVIDIA's NVFP4 tensor type; it is not equivalent to Q4_K_M, Q4_K_S, or an Unsloth dynamic quant.
  • The embedded MTP block is experimental and may not be faster than using --model-draft.
  • The model is a 35B-total-parameter MoE with approximately 3B parameters activated per token; runtime memory requirements still depend on the full stored checkpoint rather than only the active parameter count.
  • NVIDIA's source checkpoint was prepared for ModelOpt and vLLM. llama.cpp support is a separate community implementation and may behave differently from NVIDIA's reference runtime.
  • Native MTP accelerates token generation but does not improve prompt-prefill speed in the same way.
  • Very short benchmark outputs are more sensitive to run-to-run variance.
  • Results are specific to the tested hardware, llama.cpp build, prompts, runtime options, and context configuration.
  • Although standard decoding completed slightly faster on this short benchmark, native MTP enables speculative decoding and substantially increases token generation throughput (up to ~115 tok/s in these tests). Acceptance rate is therefore a more meaningful metric than total wall time when evaluating MTP effectiveness.

Related Projects

  • Standard NVFP4 GGUF: AEON-7/Ornith-1.0-35B-AEON-Ultimate-Uncensored-NVFP4
  • Grafting utility: gguf-graft-mtp

Credits

  • Qwen Team / Alibaba Cloud — Qwen3.6-35B-A3B
  • deepreinforce-ai / Ornith-1.0-35B — Qwen3.6-35B-A3B
  • NVIDIA — ModelOpt and the original NVFP4 checkpoint
  • ggml-org — llama.cpp, GGUF, NVFP4 inference support, and native MTP support
  • a4lg — MTP-only GGUF subset used as a graft source

License

The source model is distributed under the Apache License 2.0.

Users should review the upstream AEON-7/Ornith-1.0-35B-AEON-Ultimate-Uncensored-NVFP4, deepreinforce-ai/Ornith-1.0-35B, and the MTP-only GGUF source model cards before redistribution or commercial use.

README history 4 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-07-04Update README.md034d8317.9 KB
    Loading...
  2. 2026-07-03Update README.md152d9cb7.4 KB
    Loading...
  3. 2026-07-03Upload folder using huggingface_hub258ebcf7.4 KB
    Loading...
  4. 2026-07-03initial commit358c38928 B
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration