← back to catalog · registered 2026-08-22 13:56

neko-legends/Ornith-1.0-35B-AEON-Ultimate-Uncensored-NVFP4-GGUF-MTP

neko-legends 35B GGUF MoE second-order 262K ctx
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/neko-legends%2FOrnith-1.0-35B-AEON-Ultimate-Uncensored-NVFP4-GGUF-MTP"
Response includes
  • classification m8
  • files 3
  • hub_downloads_all_time 11,444
  • author_summary 4 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M8
Primary method

Repackaging (quantization)

Applied on top of direct removal inherited from the base model.
Confidence
MEDIUM
Inherited from base model
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=1
  • assume M1 (base ablation) + M8 (GGUF quant) - default when producer unknown
Refusal direction extraction

No specific extraction method could be identified for this model. The producer either did not document it or used a proprietary pipeline.

What is a refusal direction? →
Downloads · lifetime
11K
366 last 30d - cooling
Likes
7
Model age
3mo ago
created 2026-06-28
Downloads over time
Now11.6K→from8K↑45%
7.9K9.2K10.6K12K8K on Jul 111.6K on Oct 11JulAugSepOct
Jul 1 → Oct 11 · 54 snapshots · spans 102 days

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
mit
Tags
llama.cpp gguf nvfp4 mtp speculative-decoding blackwell rtx-5090 qwen3_5_moe moe mixture-of-experts reasoning thinking

Related

Total size
21.8 GB
Files
3
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-07-01 01:09

Files by quantization

Auxiliary files 3 files 21.8 GB
ornith-1.0-35b-aeon-ultimate-uncensored-nvfp4-gguf-mtp.gguf 21.8 GB 3f0545ee download
README.md 11.8 KB 72bb6787 download
.gitattributes 1.87 KB 12b07089 download

README current version from Hugging Face


license: mit
license_link: https://huggingface.co/deepreinforce-ai/Ornith-1.0-35B/blob/main/LICENSE
base_model:

  • AEON-7/Ornith-1.0-35B-AEON-Ultimate-Uncensored-NVFP4
  • deepreinforce-ai/Ornith-1.0-35B
    base_model_relation: quantized
    library_name: llama.cpp
    pipeline_tag: text-generation
    tags:
  • gguf
  • llama.cpp
  • nvfp4
  • mtp
  • speculative-decoding
  • blackwell
  • rtx-5090
  • qwen3_5_moe
  • moe
  • mixture-of-experts
  • reasoning
  • thinking
  • coding
  • agentic
  • uncensored
  • abliterated
  • aeon
  • aeon-7
  • ornith
  • 35b
  • local-llm
  • windows
  • neko-legends

Neko Legends local inference release

Ornith-1.0-35B-AEON-Ultimate-Uncensored-NVFP4-GGUF-MTP

RTX 5090 validated

A text-generation GGUF package for recent llama.cpp builds: AEON Ultimate Uncensored NVFP4 trunk/body weights with a compatible MTP block, ready for Blackwell native FP4 local serving.

FormatGGUF
QuantNVFP4
Spec decodedraft-mtp
Validated ctx262k
Target stackllama.cpp
Artifact23.4 GB

[!IMPORTANT]
This repo publishes one recommended AEON-trunk MTP artifact. It was validated for text serving only; the original safetensors family is multimodal, but this GGUF card does not claim vision or multimodal serving support.

Quick Start

Download

Use ornith-1.0-35b-aeon-ultimate-uncensored-nvfp4-gguf-mtp.gguf. It is the AEON NVFP4 GGUF with grafted compatible MTP block.

Serve

Run with a current CUDA 13.x llama.cpp build and enable --spec-type draft-mtp or the tuned draft-mtp,ngram-mod profile.

Expect

On the tested RTX 5090 machine, llama.cpp initialized MTP at full 262k context and reported BLACKWELL_NATIVE_FP4 = 1.

RTX 5090 Snapshot

RTX 5090, Windows, llama.cpp-b9267-cuda13.1, context 262144, generation 1024 tokens, temperature=0.6.

Runtime Prompt Decode tok/s Full-wall tok/s Prompt prefill
Base native GGUF 10k 133.0 106.0 1.9s
AEON-trunk MTP GGUF 10k 131.5 101.5 2.2s
Base native GGUF 200k 82.1 18.9 41.0s
AEON-trunk MTP GGUF 200k 86.0 15.9 52.1s

Tuning note: for the 10k prompt, draft-mtp with --spec-draft-n-max 2 reached 133.7 decode tok/s and 104.0 full-wall tok/s. The chart uses the single temp=0.6 draft-mtp,ngram-mod profile for both prompt sizes.

Windows native GGUF and MTP benchmark chart
AEON Ornith Ultimate Uncensored NVFP4 Windows Docker vs native GGUF benchmark chart

Files

File Size Notes
ornith-1.0-35b-aeon-ultimate-uncensored-nvfp4-gguf-mtp.gguf 23.4 GB (21.80 GiB) Recommended AEON Ultimate Uncensored NVFP4 trunk/body GGUF with grafted compatible MTP block
images/aeon-ornith-windows-docker-vs-gguf.png RTX 5090 Windows benchmark comparison chart

Which File Should I Use?

Use ornith-1.0-35b-aeon-ultimate-uncensored-nvfp4-gguf-mtp.gguf for the AEON Ultimate Uncensored NVFP4 GGUF with MTP serving support in llama.cpp. This repository intentionally publishes only the AEON-trunk MTP artifact.

This is a text-generation GGUF. The original safetensors model family is multimodal, but this GGUF file was validated for text serving only.

MTP Provenance

AEON's compressed-tensors checkpoint advertises mtp_num_hidden_layers = 1 in config metadata, but the downloaded model.safetensors contained no mtp, nextn, or model.layers.40 tensor names. A direct conversion with MTP metadata failed in llama.cpp because blk.40.attn_norm.weight and the rest of the MTP block were absent.

The recommended MTP file in this repo was therefore built as a graft:

Local validation confirmed llama.cpp initializes draft-mtp successfully at full 262k context and reports BLACKWELL_NATIVE_FP4 = 1 on RTX 5090.

SHA256 for ornith-1.0-35b-aeon-ultimate-uncensored-nvfp4-gguf-mtp.gguf:

3F0545EE14ED3B01A18E794945E33FFE6876F9A3C3787316A652C6CFDE4BDDE3

Example llama.cpp Command

$LlamaServer = Join-Path "<path-to-llama.cpp-build-folder>" "llama-server.exe"
$Model = Join-Path "<path-to-model-folder>" "ornith-1.0-35b-aeon-ultimate-uncensored-nvfp4-gguf-mtp.gguf"

& $LlamaServer `
  --model "$Model" `
  --alias aeon-ornith-1.0-35b-nvfp4-aeon-mtp `
  --host 127.0.0.1 `
  --port 39199 `
  --device CUDA0 `
  --gpu-layers all `
  --gpu-layers-draft all `
  --ctx-size 262144 `
  --cache-type-k q4_0 `
  --cache-type-v q4_0 `
  --cache-type-k-draft q4_0 `
  --cache-type-v-draft q4_0 `
  --flash-attn on `
  --parallel 1 `
  --cont-batching `
  --jinja `
  --metrics `
  --slots `
  --spec-type draft-mtp `
  --spec-draft-n-max 2 `
  --spec-draft-p-min 0.0

For very long prompts, draft-mtp,ngram-mod with --spec-draft-n-max 3 was the better measured high-context profile in this run.

RTX 5090 Windows Benchmark Details

Runtime Prompt target Prompt tokens Decode tok/s Prompt prefill Full-wall tok/s Wall time
Base native GGUF 10k 8,905 133.0 1.9s 106.0 9.7s
AEON-trunk MTP GGUF 10k 8,905 131.5 2.2s 101.5 10.1s
AEON-trunk MTP tuned n_max=2 10k 8,905 133.7 2.1s 104.0 9.8s
Base native GGUF 200k 174,588 82.1 41.0s 18.9 54.1s
AEON-trunk MTP GGUF 200k 174,588 86.0 52.1s 15.9 64.5s

Censorship Smoke Test

A short local smoke test against ornith-1.0-35b-aeon-ultimate-uncensored-nvfp4-gguf-mtp.gguf on 2026-06-28 asked for neutral factual summaries of politically sensitive history/current-affairs topics. The model returned direct factual answers with no refusal or evasion markers detected. This is a small smoke test, not a formal safety or truthfulness evaluation.

Source And Credits

Responsible Use

This is an uncensored/abliterated model family. You are responsible for downstream usage, deployment policy, and any application-level safeguards. Older llama.cpp builds may not load current GGUF/NVFP4 files correctly.

README history 17 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-07-01Fix benchmark chart asset links9adc37b11.8 KB
    Loading...
  2. 2026-06-30Use exact Ornith repository name in model card hero667e95111.6 KB
    Loading...
  3. 2026-06-30Polish Neko Legends model card hero grid1289bd111.5 KB
    Loading...
  4. 2026-06-30Refresh model card with Neko Legends themeb86d25a11.3 KB
    Loading...
  5. 2026-06-28Generalize censorship smoke test wording35b1d125.9 KB
    Loading...
  6. 2026-06-28Clarify GGUF file size unitsc2889195.9 KB
    Loading...
  7. 2026-06-28Rename GGUF artifact and refresh uncensored benchmark chart48451525.9 KB
    Loading...
  8. 2026-06-28Add censorship smoke test noteff86c0b5.8 KB
    Loading...
  9. 2026-06-28Keep only AEON Ornith MTP GGUF artifactc2409b05.4 KB
    Loading...
  10. 2026-06-28Document AEON-trunk MTP GGUF940785f6.2 KB
    Loading...
  11. 2026-06-28Clarify MTP and base GGUF notes45b5ecc7.5 KB
    Loading...
  12. 2026-06-28Rename model card for MTP repo slug10077e86.8 KB
    Loading...
  13. 2026-06-28Document native NVFP4 MTP GGUF212ab356.7 KB
    Loading...
  14. 2026-06-28Use generic paths in model card examples9b25e755.2 KB
    Loading...
  15. 2026-06-28Clarify MTP status6b07d4b5 KB
    Loading...
  16. 2026-06-28Add GGUF model cardc0bf27e4.3 KB
    Loading...
  17. 2026-06-28initial commit38ef41e21 B
    Loading...

Discussions 1 thread

  1. 2026-06-28Are you sure this one is the uncensored one?open3 💬#1
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration