← back to catalog · registered 2026-08-22 13:56

pablogrant/ORNITH-1.0_35B_AEON_PABLOG-OPTIMIZED_UNCENSORED_DSPARK-DRAFT_NVFP4

pablogrant MoE second-order
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/pablogrant%2FORNITH-1.0_35B_AEON_PABLOG-OPTIMIZED_UNCENSORED_DSPARK-DRAFT_NVFP4"
Response includes
  • classification m-uncensored
  • files 6
  • author_summary 4 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M-U
Primary method

Uncensored (method unknown)

No other method signals detected in this model.
Confidence
LOW
Why this label 3 signals
Weak or ambiguous signals. Best guess based on catalog patterns; treat as tentative and check the evidence below.
  • 'uncensored' in name/tags but no 'abliterated' marker
  • method not identifiable from author declaration alone
  • may be DPO fine-tune, prompt engineering, or unknown technique
Refusal direction extraction

No specific extraction method could be identified for this model. The producer either did not document it or used a proprietary pipeline.

What is a refusal direction? →
Downloads · 30-day
65
↑ 108% in 90 days
Likes
4
Model age
3mo ago
created 2026-07-01
Downloads over time
Now876→from422↑108%
399573747921422 on Jul 1876 on Sep 5JulAugSep
Jul 1 → Sep 5 · 18 snapshots · spans 66 days

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
mit
Tags
speculators safetensors speculative-decoding dspark vllm qwen3_5_moe nvfp4 draft-model eagle mixture-of-experts text-generation custom_code

Related

Total size
1.54 GB
Files
6
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-07-01 17:17

Files by quantization

Auxiliary files 6 files 1.54 GB
model.safetensors 1.54 GB 7ab36d46 download
README.md 9.83 KB a32f5755 download
config.json 1.92 KB a0047b10 download
config.py 1.84 KB c444cbe8 download
.gitattributes 1.48 KB a6344aac download
val_metrics.json 739 B ee3c070d download

README current version from Hugging Face


license: mit
base_model: AEON-7/Ornith-1.0-35B-AEON-Ultimate-Uncensored-NVFP4
library_name: speculators
pipeline_tag: text-generation
tags:

  • speculative-decoding
  • dspark
  • speculators
  • vllm
  • qwen3_5_moe
  • nvfp4
  • draft-model
  • eagle
  • mixture-of-experts

ORNITH-1.0_35B_AEON_PABLOG-OPTIMIZED_UNCENSORED_DSPARK-DRAFT_NVFP4

DSpark speculative-decoding draft for the AEON uncensored NVFP4 base · preview

[!IMPORTANT]

🔗 TWO-PART MODEL — this is the DRAFT (does not run alone)

This repo is the DSpark draft only — it does nothing by itself; it only accelerates a base model. You must also download the matching base:

→ ORNITH-1.0_35B_AEON_PABLOG-OPTIMIZED_UNCENSORED_NVFP4 (the base)

Download both. (And see the deployability NOTICE below.)

[!WARNING]

⚠️ NOTICE — Verified, but not yet deployable

This is a trained and validated DSpark draft (acceptance metrics below), published in the standard speculators format. It is not servable by current inference engines yet: vLLM's and SGLang's speculators loaders recognize only eagle3 / dflash / peagle, and the existing vLLM DSpark path (PR #46995) targets DeepSeek-packaged models — not generic speculators-format DSpark drafts like this one. This model becomes deployable the moment vLLM or SGLang add official speculators-format DSpark serving. We're publishing it now so it's ready that day. For a deployable-today draft of the same base, see our DFlash version (link forthcoming).

A DSpark speculative-decoding draft model for the NVFP4 build,
AEON-7/Ornith-1.0-35B-AEON-Ultimate-Uncensored-NVFP4.
Speculative decoding is lossless — the draft proposes a block, the base verifies it in one
pass and accepts the longest correct prefix, so the base's output distribution is preserved
exactly. You only gain speed.

This draft was trained against the NVFP4 base directly, so it is matched to the quantized
model's hidden states and logits — deploy it with the NVFP4 base for best acceptance. (For
the BF16 base, use the BF16-matched draft instead.)

To our knowledge this is the first published DSpark speculator for a qwen3_5_moe NVFP4 base; the card includes the full recipe + Blackwell gotchas.

[!IMPORTANT]
Preview / first-cut checkpoint (5k ShareGPT, 3 epochs, 3 draft layers). It learns cleanly but is undertrained — modest speedup, not peak. A stronger replacement is planned. Notably it landed on par with the BF16-matched draft (accept-len 1.695 vs 1.672) — the quantized base was no handicap for draft matching here.

⚠️ Uncensored-base disclaimer

This speculator targets an uncensored / abliterated base whose refusal behavior has been
removed (0/80 refusals on harmful-prompt probes: CBRN, cyber, weapons, self-harm). Because
speculative decoding is lossless, this draft does not add, remove, or alter any behavior —
it only accelerates what the base already does; all safety characteristics and risks are inherited
entirely from the base. Use at your own risk and responsibility; you are accountable for generated
content and legal compliance. For research and legitimate development use. Base is MIT-licensed —
see the base card and
upstream Ornith-1.0-35B for full method and validation.

Why we chose the uncensored AEON variant (not base Ornith)

Three reasons, from our own use:

  1. Post-hoc guardrails degrade quality. Refusal training and safety filtering bolted on
    after pretraining impose an "alignment tax" — over-refusal, hedging, and capability
    regressions that bleed into entirely legitimate requests. AEON's abliteration removes that
    refusal layer while preserving capability (near-lossless: first-token KL ≈ 0.0014 vs base,
    identical agentic-coding pass@1), which lifts response quality in the areas we actually work in.
  2. Guardrails belong to the deploying organization. The right controls depend on audience,
    domain, and jurisdiction — not something a model vendor should hard-code into the weights. An
    uncensored base is a neutral substrate onto which each organization applies its own guardrails,
    sized to its real risk and policy, instead of inheriting a fixed vendor policy it can't tune.
  3. It simply performed better for us. In our internal testing, the AEON uncensored build
    outperformed the original Ornith-1.0-35B — so matching the speculator to the variant we
    actually deploy was the obvious call.

(Technical aside: a DSpark draft is calibrated to a specific base's hidden states and logits, so pair it with the base it was trained against — see the matched note above.)

Trained directly on NVFP4 — not a quantized BF16 draft

This checkpoint was trained against the NVFP4 base itself — hidden states streamed online from
a live NVFP4 vLLM server — not produced by quantizing the
BF16-matched draft. That's
deliberate, and it matters more than it looks:

  • A DSpark draft is calibrated to its base's hidden states and output logits. Going BF16 → NVFP4
    shifts both: the aux-layer activations the draft consumes change (NVFP4 is lower-precision and, on
    vLLM 0.24, currently runs the Marlin FP4 path), and the base's output logits move slightly.
  • Quantizing the BF16 draft's weights does not recalibrate it to the NVFP4 base's distribution.
    You'd get a BF16-matched draft that is merely smaller — it would still predict the BF16 base's
    behavior and therefore accept fewer tokens against the NVFP4 base.
  • Matching comes from what the draft trained against, not from the draft's own weight precision.
    So each base precision earns its own (short) training run — which is exactly why this repo and the
    BF16 repo are separate artifacts, not a model + its quant.

(The draft is small and is not the throughput bottleneck — the base's verify pass dominates — so
there's little upside to quantizing the draft itself. We keep drafts in BF16.)

Validation results (this checkpoint)

Draft eval (speculators, block size 8, greedy) — final, 3 epochs (best = epoch 2):

Metric Value
Mean accepted length 1.695
Mean acceptance rate 0.279
Per-position acceptance 0.437 / 0.364 / 0.308 / 0.269 / 0.243 / 0.224 / 0.209

Measured end-to-end throughput (fill after A/B)

Single-stream on 1× RTX PRO 6000 Blackwell, NVFP4 base, greedy:

Config tok/s speedup
Base only (no draft) [placeholder] 1.00×
Base + this DSpark draft [placeholder] [placeholder]

NVFP4 kernel note: the base is genuinely NVFP4-quantized and runs on vLLM 0.24; on current builds its NVFP4 weights execute through the Marlin FP4 kernel (functionally correct) — native Blackwell FP4-tensor-core kernels would raise throughput further. (This is a kernel-selection detail, not a limitation of the format.)

Validation hardware

Component Spec
Server Gigabyte G294-Z42-AAP2 (MZ43-G20)
CPU 2× AMD EPYC 9555 (Turin), 64-core each — 128 cores total
RAM 125 GiB DDR5-4800
GPU 2× NVIDIA RTX PRO 6000 Blackwell Server Edition (96 GB each)
Driver / CUDA 580.173.02 / CUDA 13.0
OS / kernel Ubuntu 24.04 / Linux 6.8
Serving / training vLLM 0.24.0 · vllm-project/speculators

Training used 1 GPU to serve the base (online hidden-state generation) + 1 GPU for the draft trainer. Throughput figures are single-GPU, mapping directly to one RTX PRO 6000.

Deploy (vLLM)

vllm serve pablogrant/ORNITH-1.0_35B_AEON_PABLOG-OPTIMIZED_UNCENSORED_NVFP4 \
  --speculative-config '{"model": "pablogrant/ORNITH-1.0_35B_AEON_PABLOG-OPTIMIZED_UNCENSORED_DSPARK-DRAFT_NVFP4", "num_speculative_tokens": 7}' \
  --max-model-len 32768 --max-num-seqs 512 --trust-remote-code

Reasoning model (opens <think>); recommended sampling temperature 0.6, top_p 0.95, top_k 20.

Training recipe & Blackwell gotchas

Online DSpark via vllm-project/speculators: NVFP4 base served in vLLM (aux hidden states), draft trained against the live endpoint.

  • Algorithm: DSpark (DFlash backbone + Markov + confidence head)
  • Draft: 3 layers, block size 8, draft vocab 32000, --draft-attn-impl sdpa
  • Aux layers: 9 19 29 (base: 40-layer qwen3_5_moe, 30 GatedDeltaNet + 10 full-attn, 256 experts + 1 shared, vision tower, 256K ctx)
  • Data: 5k ShareGPT, seq-len 4096, 3 epochs, lr 3e-4
  • Losses: {ce: 0.1, tv: 0.9} + confidence-head BCE

Gotchas (Blackwell / SM 12.0):

  1. Default simple_flex_attention backward kernel needs 112 KB shared mem > Blackwell's ~99 KB → --draft-attn-impl sdpa.
  2. vLLM reserves KV for the 262K default context (≈45 GiB) → --max-model-len 8192 for training.
  3. PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True breaks the hidden-states connector → unset it.
  4. qwen3_5_moe aux-hidden-state extraction works in vLLM 0.24 out of the box.

Credits & license

README history 6 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-07-01Clarify NVFP4 kernel note (kernel-selection detail, not a format limitation)9383b3c9.8 KB
    Loading...
  2. 2026-07-01Rename to DSPARK-DRAFT + partner pairing banner (base)3126e529.8 KB
    Loading...
  3. 2026-07-01Add NOTICE: verified, not deployable until vLLM/SGLang add speculators-format...c9799eb9.3 KB
    Loading...
  4. 2026-07-01Finalize nvfp4 draft metrics (accept-len 1.695, 3 epochs)9a2a3ee8.6 KB
    Loading...
  5. 2026-07-01Add model card (preview; measurement values placeholdered)050be9d8.5 KB
    Loading...
  6. 2026-07-01initial commit6e7271e21 B
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration