← back to catalog · registered 2026-08-22 13:56

AxiangQAQ/Ornith-1.0-35B-uncensored-heretic-FP8-BLOCK

AxiangQAQ Qwen 33B MoE multimodal second-order
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/AxiangQAQ%2FOrnith-1.0-35B-uncensored-heretic-FP8-BLOCK"
Response includes
  • classification m3
  • files 14
  • hub_downloads_all_time 3,834
  • author_summary 1 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M3
Primary method

Layer-wise ablation

Applied on top of direct removal inherited from the base model.
Confidence
HIGH
Inherited from base model
Why this label 2 signals
Producer identity confirmed by naming conventions, tags or the model card. This label is very unlikely to change.
  • 'heretic' in model name (Heretic-produced)
  • Heretic uses layer-wise optimization (M3) with underlying direction removal (M1)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
HIGH
Why we say so
name contains 'heretic'; Heretic default extraction is difference-of-means (Arditi 2024)
Downloads · lifetime
4K
153 last 30d - cooling
Likes
0
Model age
3mo ago
created 2026-07-13
Downloads over time
Now3.9K→from179↑2,081%
01.4K2.9K4.3K179 on Jul 153.9K on Oct 11JulAugSepOct
Jul 15 → Oct 11 · 53 snapshots · spans 88 days

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
mit
Languages
en zh ja
Tags
transformers safetensors qwen3_5_moe image-text-to-text heretic uncensored decensored abliterated mpoa ornith fp8 fp8-block

Related

Total size
36.6 GB
Files
14
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-07-14 13:46

Files by quantization

Auxiliary files 14 files 36.6 GB
model-00001-of-00002.safetensors 24.8 GB 4dcaa895 download
model-00002-of-00002.safetensors 10.2 GB 8fa21a99 download
model-mtp.safetensors 1.57 GB 62e6f4b9 download
tokenizer.json 19.1 MB 6f32ce20 download
model.safetensors.index.json 6.47 MB a1184bfc download
chat_template.jinja 15.9 KB 81df6b83 download
README.md 7.47 KB 98ebc7b0 download
config.json 4.68 KB 2cbaa111 download
.gitattributes 1.65 KB e5f2beea download
tokenizer_config.json 1.17 KB 7b0f9106 download
processor_config.json 1.16 KB 33818c7f download
preprocessor_config.json 390 B 2ea84a43 download
video_preprocessor_config.json 385 B 3ba673a5 download
generation_config.json 227 B e8f101cd download

README current version from Hugging Face


library_name: transformers
license: mit
license_link: https://huggingface.co/deepreinforce-ai/Ornith-1.0-35B/blob/main/LICENSE
language:

  • en
  • zh
  • ja
    pipeline_tag: text-generation
    tags:
  • heretic
  • uncensored
  • decensored
  • abliterated
  • mpoa
  • ornith
  • fp8
  • fp8-block
  • compressed-tensors
  • moe
  • qwen3.5
  • mtp
  • speculative-decoding
  • multimodal
  • vllm
  • sglang
  • agentic
  • coding
    base_model:
  • llmfan46/Ornith-1.0-35B-uncensored-heretic
  • deepreinforce-ai/Ornith-1.0-35B

Ornith-1.0-35B-uncensored-heretic-FP8-BLOCK

License: MIT FP8 MTP Uncensored

FP8 Block Quantization + MTP Speculative Decoding

Uncensored | 35B MoE (3B active) | Vision + Text | ~38GB

Base Model · FP8 Reference · MTP Technique


Overview

This is an FP8_BLOCK quantized version of llmfan46/Ornith-1.0-35B-uncensored-heretic with an MTP head grafted from Qwen3.6-35B-A3B for speculative decoding.

The uncensored heretic version achieves 90% fewer refusals (9/100 vs 89/100 original) with only 0.0019 KL divergence.

Key Features

QuantizationFP8_BLOCK (128×128 block, data-free PTQ)
Model Size~38GB (vs ~70GB BF16)
MTP HeadBF16, from Qwen3.6-35B-A3B (19 tensors)
Throughput+21% with speculative decoding (ref. benchmark, not tested)
ArchitectureQwen3.5 MoE, 256 experts, 8 active per token
MultimodalVision + Text

Quick Start

vLLM (Recommended)

pip install vllm

vllm serve AxiangQAQ/Ornith-1.0-35B-uncensored-heretic-FP8-BLOCK \
  --served-model-name ornith-35b-uncensored-fp8 \
  --quantization compressed-tensors \
  --max-model-len 32768 \
  --gpu-memory-utilization 0.93 \
  --enable-prefix-caching \
  --enable-auto-tool-choice \
  --tool-call-parser qwen3_xml \
  --reasoning-parser qwen3 \
  --speculative-config '{"method":"mtp","num_speculative_tokens":3}'

Performance-optimized configuration (validated on RTX 4090 48GB):

vllm serve AxiangQAQ/Ornith-1.0-35B-uncensored-heretic-FP8-BLOCK \
  --served-model-name ornith-35b-uncensored-fp8 \
  --host 0.0.0.0 \
  --port 8000 \
  --tensor-parallel-size 1 \
  --max-model-len 204800 \
  --max-num-seqs 8 \
  --max-num-batched-tokens 8192 \
  --gpu-memory-utilization 0.93 \
  --dtype bfloat16 \
  --quantization compressed-tensors \
  --kv-cache-dtype fp8_e4m3 \
  --mm-encoder-tp-mode data \
  --mm-processor-cache-type shm \
  --mm-shm-cache-max-object-size-mb 512 \
  --media-io-kwargs '{"video": {"num_frames": -1}}' \
  --attention-backend flashinfer \
  --enable-prefix-caching \
  --enable-chunked-prefill \
  --generation-config vllm \
  --max-cudagraph-capture-size 16 \
  --speculative-config '{"method":"mtp","num_speculative_tokens":3}' \
  --compilation_config.mode VLLM_COMPILE \
  --compilation_config.cudagraph_mode PIECEWISE \
  --reasoning-parser qwen3 \
  --tool-call-parser qwen3_xml \
  --enable-auto-tool-choice \
  --trust-remote-code

GPU Requirements: ~40GB VRAM minimum (model is ~38GB FP8). The configuration above was validated on a modified RTX 4090 48GB. For lower VRAM GPUs, quantization to GGUF format is recommended instead.

SGLang

pip install sglang

python3 -m sglang.launch_server \
  --model-path AxiangQAQ/Ornith-1.0-35B-uncensored-heretic-FP8-BLOCK \
  --host 0.0.0.0 \
  --port 30000

Transformers

from transformers import AutoModelForMultimodalLM, AutoProcessor

model = AutoModelForMultimodalLM.from_pretrained(
    "AxiangQAQ/Ornith-1.0-35B-uncensored-heretic-FP8-BLOCK",
    torch_dtype="auto",
    device_map="auto",
    trust_remote_code=True,
)
processor = AutoProcessor.from_pretrained(
    "AxiangQAQ/Ornith-1.0-35B-uncensored-heretic-FP8-BLOCK",
    trust_remote_code=True,
)

Performance

Note: This model has not been benchmarked. The data below is from the reference model shisa-ai/Ornith-1.0-35B-FP8-BLOCK-MTP and is provided for reference only.

Reference benchmarks (RTX PRO 6000 Blackwell, single GPU):

Config Output tok/s vs Baseline TTFT (ms) TPOT (ms) Acceptance
No speculative 203.1 — 51.7 4.65 —
MTP tokens=1 221.0 +8.8% 60.3 4.21 85.8%
MTP tokens=3 246.5 +21.4% 65.5 3.57 67.0%

Source: shisa-ai/Ornith-1.0-35B-FP8-BLOCK-MTP


Quantization Details

  • Tool: llmcompressor v0.12.0
  • Method: Model-free PTQ (no calibration data needed)
  • Scheme: FP8_BLOCK (compressed-tensors format)
  • Weight: Static FP8, symmetric, 128×128 block strategy
  • Activation: Dynamic FP8, symmetric, group size 128
  • MTP Head: BF16 (not quantized)

Ignored layers: lm_head, embed_tokens, visual.*, mlp.gate, mlp.shared_expert_gate, linear_attn.*, mtp.*


Model Lineage

deepreinforce-ai/Ornith-1.0-35B (MIT, original)
    └── llmfan46/Ornith-1.0-35B-uncensored-heretic (Heretic v1.2.0 MPOA ablation)
        └── AxiangQAQ/Ornith-1.0-35B-uncensored-heretic-FP8-BLOCK (this model)
            ├── FP8_BLOCK quantization (llmcompressor)
            └── MTP head grafted from Qwen/Qwen3.6-35B-A3B (donor only)

Credits

Component Source
Original Model deepreinforce-ai/Ornith-1.0-35B
Uncensored Version llmfan46/Ornith-1.0-35B-uncensored-heretic
FP8 Quantization Reference shisa-ai/Ornith-1.0-35B-FP8-BLOCK
MTP Graft Technique protoLabsAI/Ornith-1.0-9B-MTP
MTP Donor Qwen/Qwen3.6-35B-A3B
Chat Template Fix froggeric/Qwen-Fixed-Chat-Templates

License

MIT — same as the source Ornith-1.0-35B model.

Citation

@misc{ornith-35b,
    title = {{Ornith-1.0-35B}: Agentic Coding, Open to All},
    url = {https://deep-reinforce.com/ornith_1_0.html},
    author = {{DeepReinforce Team}},
    year = {2026}
}

README history 9 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-07-14Fix VRAM requirements: ~40GB min, not 24GB47d54db7.5 KB
    Loading...
  2. 2026-07-14Update Quick Start with validated 4090 48GB config + minimal configc075bec7.5 KB
    Loading...
  3. 2026-07-14Fix: add language field, ornith tag, throughput disclaimera13fba66.3 KB
    Loading...
  4. 2026-07-13Fix title centering with HTML tagsc105af16.2 KB
    Loading...
  5. 2026-07-13Fix: remove Chinese text, keep English only5a60ee26 KB
    Loading...
  6. 2026-07-13Add performance disclaimer - data from reference model only7f899ac6.2 KB
    Loading...
  7. 2026-07-13Update model card: remove logo, fix base_model05884fb5.7 KB
    Loading...
  8. 2026-07-13Update model cardd9cf4c35.9 KB
    Loading...
  9. 2026-07-13Upload README.md with huggingface_hub0fae17c4.1 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration