license: mit
base_model: AEON-7/Ornith-1.0-35B-AEON-Ultimate-Uncensored-NVFP4
library_name: speculators
pipeline_tag: text-generation
tags:
- speculative-decoding
- dspark
- speculators
- vllm
- qwen3_5_moe
- nvfp4
- draft-model
- eagle
- mixture-of-experts
ORNITH-1.0_35B_AEON_PABLOG-OPTIMIZED_UNCENSORED_DSPARK-DRAFT_NVFP4
DSpark speculative-decoding draft for the AEON uncensored NVFP4 base · preview
[!IMPORTANT]
🔗 TWO-PART MODEL — this is the DRAFT (does not run alone)
This repo is the DSpark draft only — it does nothing by itself; it only accelerates a base model. You must also download the matching base:
→ ORNITH-1.0_35B_AEON_PABLOG-OPTIMIZED_UNCENSORED_NVFP4 (the base)
Download both. (And see the deployability NOTICE below.)
[!WARNING]
⚠️ NOTICE — Verified, but not yet deployable
This is a trained and validated DSpark draft (acceptance metrics below), published in the standard
speculatorsformat. It is not servable by current inference engines yet: vLLM's and SGLang's speculators loaders recognize onlyeagle3/dflash/peagle, and the existing vLLM DSpark path (PR #46995) targets DeepSeek-packaged models — not generic speculators-format DSpark drafts like this one. This model becomes deployable the moment vLLM or SGLang add official speculators-format DSpark serving. We're publishing it now so it's ready that day. For a deployable-today draft of the same base, see our DFlash version (link forthcoming).
A DSpark speculative-decoding draft model for the NVFP4 build,AEON-7/Ornith-1.0-35B-AEON-Ultimate-Uncensored-NVFP4.
Speculative decoding is lossless — the draft proposes a block, the base verifies it in one
pass and accepts the longest correct prefix, so the base's output distribution is preserved
exactly. You only gain speed.
This draft was trained against the NVFP4 base directly, so it is matched to the quantized
model's hidden states and logits — deploy it with the NVFP4 base for best acceptance. (For
the BF16 base, use the BF16-matched draft instead.)
To our knowledge this is the first published DSpark speculator for a qwen3_5_moe NVFP4 base; the card includes the full recipe + Blackwell gotchas.
[!IMPORTANT]
Preview / first-cut checkpoint (5k ShareGPT, 3 epochs, 3 draft layers). It learns cleanly but is undertrained — modest speedup, not peak. A stronger replacement is planned. Notably it landed on par with the BF16-matched draft (accept-len 1.695 vs 1.672) — the quantized base was no handicap for draft matching here.
⚠️ Uncensored-base disclaimer
This speculator targets an uncensored / abliterated base whose refusal behavior has been
removed (0/80 refusals on harmful-prompt probes: CBRN, cyber, weapons, self-harm). Because
speculative decoding is lossless, this draft does not add, remove, or alter any behavior —
it only accelerates what the base already does; all safety characteristics and risks are inherited
entirely from the base. Use at your own risk and responsibility; you are accountable for generated
content and legal compliance. For research and legitimate development use. Base is MIT-licensed —
see the base card and
upstream Ornith-1.0-35B for full method and validation.
Why we chose the uncensored AEON variant (not base Ornith)
Three reasons, from our own use:
- Post-hoc guardrails degrade quality. Refusal training and safety filtering bolted on
after pretraining impose an "alignment tax" — over-refusal, hedging, and capability
regressions that bleed into entirely legitimate requests. AEON's abliteration removes that
refusal layer while preserving capability (near-lossless: first-token KL ≈ 0.0014 vs base,
identical agentic-coding pass@1), which lifts response quality in the areas we actually work in. - Guardrails belong to the deploying organization. The right controls depend on audience,
domain, and jurisdiction — not something a model vendor should hard-code into the weights. An
uncensored base is a neutral substrate onto which each organization applies its own guardrails,
sized to its real risk and policy, instead of inheriting a fixed vendor policy it can't tune. - It simply performed better for us. In our internal testing, the AEON uncensored build
outperformed the original Ornith-1.0-35B — so matching the speculator to the variant we
actually deploy was the obvious call.
(Technical aside: a DSpark draft is calibrated to a specific base's hidden states and logits, so pair it with the base it was trained against — see the matched note above.)
Trained directly on NVFP4 — not a quantized BF16 draft
This checkpoint was trained against the NVFP4 base itself — hidden states streamed online from
a live NVFP4 vLLM server — not produced by quantizing the
BF16-matched draft. That's
deliberate, and it matters more than it looks:
- A DSpark draft is calibrated to its base's hidden states and output logits. Going BF16 → NVFP4
shifts both: the aux-layer activations the draft consumes change (NVFP4 is lower-precision and, on
vLLM 0.24, currently runs the Marlin FP4 path), and the base's output logits move slightly. - Quantizing the BF16 draft's weights does not recalibrate it to the NVFP4 base's distribution.
You'd get a BF16-matched draft that is merely smaller — it would still predict the BF16 base's
behavior and therefore accept fewer tokens against the NVFP4 base. - Matching comes from what the draft trained against, not from the draft's own weight precision.
So each base precision earns its own (short) training run — which is exactly why this repo and the
BF16 repo are separate artifacts, not a model + its quant.
(The draft is small and is not the throughput bottleneck — the base's verify pass dominates — so
there's little upside to quantizing the draft itself. We keep drafts in BF16.)
Validation results (this checkpoint)
Draft eval (speculators, block size 8, greedy) — final, 3 epochs (best = epoch 2):
| Metric | Value |
|---|---|
| Mean accepted length | 1.695 |
| Mean acceptance rate | 0.279 |
| Per-position acceptance | 0.437 / 0.364 / 0.308 / 0.269 / 0.243 / 0.224 / 0.209 |
Measured end-to-end throughput (fill after A/B)
Single-stream on 1× RTX PRO 6000 Blackwell, NVFP4 base, greedy:
| Config | tok/s | speedup |
|---|---|---|
| Base only (no draft) | [placeholder] | 1.00× |
| Base + this DSpark draft | [placeholder] | [placeholder] |
NVFP4 kernel note: the base is genuinely NVFP4-quantized and runs on vLLM 0.24; on current builds its NVFP4 weights execute through the Marlin FP4 kernel (functionally correct) — native Blackwell FP4-tensor-core kernels would raise throughput further. (This is a kernel-selection detail, not a limitation of the format.)
Validation hardware
| Component | Spec |
|---|---|
| Server | Gigabyte G294-Z42-AAP2 (MZ43-G20) |
| CPU | 2× AMD EPYC 9555 (Turin), 64-core each — 128 cores total |
| RAM | 125 GiB DDR5-4800 |
| GPU | 2× NVIDIA RTX PRO 6000 Blackwell Server Edition (96 GB each) |
| Driver / CUDA | 580.173.02 / CUDA 13.0 |
| OS / kernel | Ubuntu 24.04 / Linux 6.8 |
| Serving / training | vLLM 0.24.0 · vllm-project/speculators |
Training used 1 GPU to serve the base (online hidden-state generation) + 1 GPU for the draft trainer. Throughput figures are single-GPU, mapping directly to one RTX PRO 6000.
Deploy (vLLM)
vllm serve pablogrant/ORNITH-1.0_35B_AEON_PABLOG-OPTIMIZED_UNCENSORED_NVFP4 \
--speculative-config '{"model": "pablogrant/ORNITH-1.0_35B_AEON_PABLOG-OPTIMIZED_UNCENSORED_DSPARK-DRAFT_NVFP4", "num_speculative_tokens": 7}' \
--max-model-len 32768 --max-num-seqs 512 --trust-remote-code
Reasoning model (opens <think>); recommended sampling temperature 0.6, top_p 0.95, top_k 20.
Training recipe & Blackwell gotchas
Online DSpark via vllm-project/speculators: NVFP4 base served in vLLM (aux hidden states), draft trained against the live endpoint.
- Algorithm: DSpark (DFlash backbone + Markov + confidence head)
- Draft: 3 layers, block size 8, draft vocab 32000,
--draft-attn-impl sdpa - Aux layers:
9 19 29(base: 40-layerqwen3_5_moe, 30 GatedDeltaNet + 10 full-attn, 256 experts + 1 shared, vision tower, 256K ctx) - Data: 5k ShareGPT, seq-len 4096, 3 epochs, lr 3e-4
- Losses:
{ce: 0.1, tv: 0.9}+ confidence-head BCE
Gotchas (Blackwell / SM 12.0):
- Default
simple_flex_attentionbackward kernel needs 112 KB shared mem > Blackwell's ~99 KB →--draft-attn-impl sdpa. - vLLM reserves KV for the 262K default context (≈45 GiB) →
--max-model-len 8192for training. PYTORCH_CUDA_ALLOC_CONF=expandable_segments:Truebreaks the hidden-states connector → unset it.qwen3_5_moeaux-hidden-state extraction works in vLLM 0.24 out of the box.
Credits & license
- Base: AEON-7/Ornith-1.0-35B-AEON-Ultimate-Uncensored-NVFP4 · upstream deepreinforce-ai/Ornith-1.0-35B — MIT
- Framework: vllm-project/speculators · DSpark (DeepSeek/PKU)
- Released under MIT.