← back to catalog · registered 2026-08-22 13:56

thanet-s/Ornith-1.0-35B-uncensored-heretic-nvfp4-fp8dense-gb10

thanet-s Qwen 16B MoE second-order
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/thanet-s%2FOrnith-1.0-35B-uncensored-heretic-nvfp4-fp8dense-gb10"
Response includes
  • classification m3
  • files 16
  • hub_downloads_all_time 968
  • author_summary 1 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M3
Primary method

Layer-wise ablation

Applied on top of direct removal inherited from the base model.
Confidence
HIGH
Inherited from base model
Why this label 2 signals
Producer identity confirmed by naming conventions, tags or the model card. This label is very unlikely to change.
  • 'heretic' in model name (Heretic-produced)
  • Heretic uses layer-wise optimization (M3) with underlying direction removal (M1)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
HIGH
Why we say so
name contains 'heretic'; Heretic default extraction is difference-of-means (Arditi 2024)
Downloads · lifetime
968
67 last 30d - cooling
Likes
2
Model age
3mo ago
created 2026-07-05
Downloads over time
Now1K→from572↑76%
5507178841.1K572 on Jul 151K on Oct 11JulAugSepOct
Jul 15 → Oct 11 · 53 snapshots · spans 88 days

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
other
Tags
transformers safetensors qwen3_5_moe image-text-to-text vllm compressed-tensors nvfp4 fp8 dgx-spark dgx-spark-optimized optimized-quant uncensored-heretic

Related

Total size
20.9 GB
Files
16
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-07-05 08:39

Files by quantization

Auxiliary files 16 files 21.0 GB
model-00005-of-00005.safetensors 4.20 GB f2067b9b download
model-00004-of-00005.safetensors 4.18 GB a6686121 download
model-00002-of-00005.safetensors 4.18 GB 1253296b download
model-00003-of-00005.safetensors 4.18 GB c3db76ea download
model-00001-of-00005.safetensors 4.17 GB d0f2cef7 download
tokenizer.json 19.1 MB bb3dcb99 download
model.safetensors.index.json 13.6 MB 63f3c8bc download
config.json 567 KB 47785692 download
config.json.llmcompressor-text.bak 566 KB 37041ade download
chat_template.jinja 7.68 KB 57fcf921 download
README.md 4.24 KB e1720248 download
recipe.yaml 2.39 KB 45fd94e2 download
.gitattributes 1.60 KB a09db2ea download
tokenizer_config.json 1.24 KB a9eacca6 download
processor_config.json 1.16 KB 33818c7f download
generation_config.json 219 B 5e81902d download

README current version from Hugging Face


base_model: llmfan46/Ornith-1.0-35B-uncensored-heretic
library_name: transformers
pipeline_tag: text-generation
license: other
tags:

  • vllm
  • compressed-tensors
  • nvfp4
  • fp8
  • dgx-spark
  • dgx-spark-optimized
  • optimized-quant
  • uncensored-heretic
  • gb10
  • qwen3.5
  • moe

Ornith-1.0-35B-uncensored-heretic-nvfp4-fp8dense-gb10

This is a DGX Spark optimized quant for vLLM: a GB10-oriented
compressed-tensors quantization of
llmfan46/Ornith-1.0-35B-uncensored-heretic for vLLM inference.

The source model is llmfan46/Ornith-1.0-35B-uncensored-heretic.

Quantization

  • Dense attention and linear-attention projections: FP8 W8A8.
  • Routed MoE experts and shared expert projections: NVFP4 W4A4.
  • Preserved in BF16: embeddings, lm_head, router gates, visual modules, norms, Conv1D, A_log, dt_bias, and non-target tensors.
  • Visual encoder: copied from the source checkpoint and preserved in BF16 for multimodal/image support.
  • Calibration data: HuggingFaceH4/ultrachat_200k, train_sft.
  • Calibration samples: 512.
  • Calibration sequence length: 2048.
  • Pipeline: sequential, CPU offload, all MoE experts calibrated.

The quantization groups use mutually exclusive explicit target regexes.

Output

  • Format: mixed-precision.
  • Quantization method: compressed-tensors.
  • Config groups: group_0, group_1.
  • Weight layout: sharded safetensors with model.safetensors.index.json.
  • Safetensors shards: 5.
  • Indexed tensors: 124376.
  • Max position embeddings: 262144.
  • MTP weights: not present in the inspected source checkpoint.

DGX Spark vLLM

This model is intended for the GB10 patched vLLM path used by
demon-zombie/Qwen3.5-122B-A10B-NVFP4-FP8Dense-GB10.

Example:

docker run --gpus all -p 8000:8000 --ipc host \
  -v /opt/vllm-cache:/root/.cache/huggingface \
  -e CUBLASLT_WORKSPACE_SIZE=33554432 \
  -e PYTORCH_CUDA_ALLOC_CONF=expandable_segments:False \
  -e RUNAI_STREAMER_MEMORY_LIMIT=4294967296 \
  vllm/vllm-openai:v0.22.1 \
  --model thanet-s/Ornith-1.0-35B-uncensored-heretic-nvfp4-fp8dense-gb10 \
  --served-model-name Ornith-1.0-35B-uncensored-heretic \
  --kernel-config '{"moe_backend": "flashinfer_b12x"}' \
  --load-format runai_streamer \
  --gpu-memory-utilization 0.50 \
  --kv-cache-dtype fp8 \
  --enable-prefix-caching \
  --enable-chunked-prefill \
  --max-num-batched-tokens 4176 \
  --enable-auto-tool-choice \
  --tool-call-parser qwen3_xml \
  --reasoning-parser qwen3

On DGX Spark / GB10, the vLLM image may still need the same CUTLASS DSL and
flashinfer SM12x patch files described in the reference model card above.

DGX Spark benchmark

Measured on NVIDIA DGX Spark / GB10 with vllm/vllm-openai:v0.22.1, the GB10 CUTLASS DSL and flashinfer SM12x patches, --gpu-memory-utilization 0.50, --kv-cache-dtype fp8, --kernel-config {"moe_backend": "flashinfer_b12x"}, and --max-num-batched-tokens 4176.

  • vLLM log throughput during a warm single request: Avg generation throughput: 60.8 tokens/s.
  • Client-side measurement for the same warm request: 1024 completion tokens in 16.905s, about 60.58 tok/s.
  • First measured request after server readiness: 768 completion tokens in 13.676s, about 56.15 tok/s; vLLM log showed Avg generation throughput: 53.5 tokens/s.
  • Model load memory reported by vLLM: 21.04 GiB.
  • Max model length reported by vLLM: 262144 tokens.
  • GPU KV cache size at --gpu-memory-utilization 0.50: 3,344,321 tokens.
  • Maximum concurrency reported by vLLM for 262144-token requests: 12.76x.

These figures are from one local DGX Spark run and may change with prompt shape, sampling settings, vLLM version, patch versions, and CUDA graph/cache warmup state.

Status

Local DGX Spark conversion validation:

  • Source: llmfan46/Ornith-1.0-35B-uncensored-heretic.
  • Conversion completed with the same NVFP4-FP8Dense GB10 recipe used for the non-uncensored Ornith 35B build.
  • verify-output.py passed after vLLM config patching and visual tensor merge.
  • Visual tensors preserved: 333 model.visual.* tensors copied from the source checkpoint.
  • vLLM serving smoke test passed on DGX Spark / GB10 after upload.

README history 3 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-07-05Add files using upload-large-folder toolddae4174.2 KB
    Loading...
  2. 2026-07-05Update DGX Spark throughput benchmarkebfd8ee4.2 KB
    Loading...
  3. 2026-07-05Add files using upload-large-folder tool8bd2bd13.3 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration