← back to catalog · registered 2026-08-22 13:56

cebeuq/Ornith-1.0-397B-abliterated-W4A16

cebeuq 397B MoE multimodal
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/cebeuq%2FOrnith-1.0-397B-abliterated-W4A16"
Response includes
  • classification m1
  • files 60
  • hub_downloads_all_time 3,065
  • author_summary 2 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M1
Primary method

Direct removal

No other method signals detected in this model.
Confidence
MEDIUM
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=0 (base model)
  • no specific method indicators - defaulting to M1 (most common)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
MEDIUM
Why we say so
primary_method=M1; difference-of-means is the reference extraction for M1/M3 (Arditi 2024)
Downloads · lifetime
3K
67 last 30d - cooling
Likes
2
Model age
3mo ago
created 2026-07-03
Downloads over time
Now3.1K→from1.2K↑163%
1.1K1.8K2.5K3.3K1.2K on Jul 153.1K on Oct 11JulAugSepOct
Jul 15 → Oct 11 · 53 snapshots · spans 88 days

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
mit
Tags
transformers safetensors qwen3_5_moe image-text-to-text abliterated uncensored mixture-of-experts w4a16 auto-round gptq vllm multimodal

Related

Total size
196 GB
Files
60
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-07-03 03:57

Files by quantization

Auxiliary files 60 files 196 GB
model-00009-of-00047.safetensors 4.19 GB 0a0ad664 download
model-00012-of-00047.safetensors 4.19 GB a69605b1 download
model-00015-of-00047.safetensors 4.19 GB 0d9c6294 download
model-00018-of-00047.safetensors 4.19 GB 50e1a827 download
model-00021-of-00047.safetensors 4.19 GB 335cfe69 download
model-00024-of-00047.safetensors 4.19 GB 64877574 download
model-00002-of-00047.safetensors 4.19 GB 96d57e17 download
model-00005-of-00047.safetensors 4.19 GB cf08eec4 download
model-00030-of-00047.safetensors 4.19 GB e8c7fb36 download
model-00033-of-00047.safetensors 4.19 GB a028125e download
model-00036-of-00047.safetensors 4.19 GB 11ab09dd download
model-00039-of-00047.safetensors 4.19 GB 77f0eb6a download
model-00042-of-00047.safetensors 4.19 GB 5a586936 download
model-00045-of-00047.safetensors 4.19 GB b22680e0 download
model-00027-of-00047.safetensors 4.19 GB 5cc422fa download
model-00028-of-00047.safetensors 4.19 GB 4b7d28fc download
model-00031-of-00047.safetensors 4.19 GB 5ba9b7a9 download
model-00034-of-00047.safetensors 4.19 GB 622e2322 download
model-00037-of-00047.safetensors 4.19 GB 47064a6c download
model-00040-of-00047.safetensors 4.19 GB d55dad0f download
model-00043-of-00047.safetensors 4.19 GB 1fb4208f download
model-00025-of-00047.safetensors 4.19 GB 41e9b259 download
model-00044-of-00047.safetensors 4.19 GB 53f0013e download
model-00022-of-00047.safetensors 4.19 GB c4c52675 download
model-00041-of-00047.safetensors 4.19 GB c9df1bb4 download
model-00019-of-00047.safetensors 4.19 GB f51ebce8 download
model-00011-of-00047.safetensors 4.19 GB dc19a7f1 download
model-00038-of-00047.safetensors 4.19 GB 6091b06e download
model-00014-of-00047.safetensors 4.19 GB 2f1d9cbb download
model-00016-of-00047.safetensors 4.19 GB 593fb455 download
model-00035-of-00047.safetensors 4.19 GB f25e03e3 download
model-00017-of-00047.safetensors 4.19 GB 5fa5664d download
model-00013-of-00047.safetensors 4.19 GB 90bb17db download
model-00020-of-00047.safetensors 4.19 GB d24cf15f download
model-00032-of-00047.safetensors 4.19 GB 9a391a33 download
model-00008-of-00047.safetensors 4.19 GB be58b20e download
model-00023-of-00047.safetensors 4.19 GB 08a2c18d download
model-00010-of-00047.safetensors 4.19 GB 0e6a2287 download
model-00026-of-00047.safetensors 4.19 GB a664366b download
model-00029-of-00047.safetensors 4.19 GB f8ed3288 download
model-00006-of-00047.safetensors 4.19 GB 95227570 download
model-00003-of-00047.safetensors 4.19 GB 17d2d6c4 download
model-00001-of-00047.safetensors 4.19 GB 2b87e58d download
model-00004-of-00047.safetensors 4.19 GB 264a55dd download
model-00007-of-00047.safetensors 4.19 GB 5ba1196d download
model-00000-of-00047.safetensors 4.19 GB cfc7ed4d download
model-00046-of-00047.safetensors 2.79 GB db142832 download
model.safetensors.index.json 28.2 MB 4345280e download
tokenizer.json 19.1 MB 06b95093 download
vocab.json 6.41 MB 0aa0ce06 download
config.json 28.9 KB 33c5332d download
quantization_config.json 23.4 KB 85df4133 download
chat_template.jinja 7.57 KB a585dec8 download
README.md 6.28 KB 4d253b11 download
.gitattributes 1.72 KB 51471fd3 download
processor_config.json 1.16 KB 33818c7f download
tokenizer_config.json 1.14 KB 1d134cd2 download
preprocessor_config.json 390 B 2ea84a43 download
video_preprocessor_config.json 385 B 3ba673a5 download
generation_config.json 244 B 85b45ab4 download

README current version from Hugging Face


license: mit
base_model:

  • deepreinforce-ai/Ornith-1.0-397B
    base_model_relation: quantized
    pipeline_tag: image-text-to-text
    library_name: transformers
    tags:
  • abliterated
  • uncensored
  • mixture-of-experts
  • qwen3_5_moe
  • w4a16
  • auto-round
  • gptq
  • vllm
  • multimodal
  • dgx-spark

Ornith-1.0-397B — Abliterated · W4A16

A decensored (abliterated) rebuild of deepreinforce-ai/Ornith-1.0-397B — a 397B-parameter Qwen3.5-MoE multimodal reasoning / agentic-coding model — re-quantized to W4A16 (AutoRound, auto_round:auto_gptq) so it runs on 2× NVIDIA DGX Spark (GB10, 128 GB each) with vLLM.

Refusal behavior was removed from the language model only; the vision tower, embeddings, routers and norms are untouched. Coding, reasoning and vision (image + video) remain intact.

⚠️ Uncensored model. Safety refusals have been substantially removed. You are responsible for how you use it. Intended for local/research use on hardware you control. It will attempt almost any request.


Highlights

Base deepreinforce-ai/Ornith-1.0-397B (Qwen3.5-MoE, 397B total / ~17B active)
Architecture 60 layers, hidden 4096, 512 experts (top-10), hybrid Gated-DeltaNet linear + full attention (45 + 15), MTP head, SigLIP-style vision tower (image + video)
Context 262,144 tokens (256K)
Quantization W4A16 — AutoRound RTN, int4, group-size 128, symmetric, auto_round:auto_gptq packing
Kept at BF16 embeddings, lm_head, all routers (mlp.gate, shared_expert_gate), norms, and the entire vision tower
Size ~196 GB, 47 safetensors shards
Runtime vLLM ≥ 0.17 (tested 0.20.1), served across 2× DGX Spark via Ray TP=2

Validation (before → after abliteration)

Measured on the 2-node vLLM cluster; refusal on mlabonne/harmful_behaviors (N=40, identical prompts), KL over top-k first-token logprobs on mlabonne/harmless_alpaca (N=40).

Metric Reference W4A16 This model (abliterated)
Refusal rate 30.0 % 7.5 % (−75 %)
KL divergence (harmless first-token) — 0.116 — capability preserved
Coding smoke-test — ✅ works
Image understanding — ✅ works (OCR + scene)
Video understanding — ✅ works (frames + motion)
128K context (needle retrieval) — ✅ retrieved both needles @ 126,709 tok
256K context (needle retrieval) — ✅ retrieved both needles @ 252,229 tok

Serving performance (2× DGX Spark, TP=2 over ConnectX-7, single-stream):

input ctx decode tok/s prefill tok/s TTFT
~1K 22 824 1.1 s
~8K 23 1,102 6.6 s
~128K 18 886 134 s
~256K 17 704 329 s

Decode is ~flat across context length because the hybrid linear-attention layers keep KV small; the single-stream ceiling is the per-token cross-node all-reduce over the 200 GbE ConnectX-7 link (no NVLink).


How it was built

  1. Refusal directions. A memory-frugal layer-streaming forward over the local quant (per-layer gptq→bf16 dequant, no full model in memory) computed Arditi difference-of-means refusal directions r_ℓ ∈ ℝ⁴⁰⁹⁶ on 128 harmful vs. 128 harmless prompts (per-layer directions, adjacent-layer cosine ≈ 0.91).
  2. Abliterate + re-quantize (streaming). The 122 BF16 shards of the base were streamed from the Hub one at a time and the residual-write matrices of the LM were orthogonalized against the refusal direction — W ← W − r rᵀW for self_attn.o_proj, linear_attn.out_proj, mlp.shared_expert.down_proj, and per-expert mlp.experts.down_proj (hidden = output axis) — then re-quantized with AutoRound's quantize_weight_rtn in the exact auto_round:auto_gptq layout, un-fusing experts to experts.{e}.{gate,up,down}_proj. Output is byte-identical to the reference AutoRound quant on every non-abliterated tensor.

Never staged the full 794 GB BF16 model — the whole pipeline is streaming + data-free (AutoRound model_free/RTN, no calibration).


Serving (vLLM, 2× DGX Spark)

The model (~196 GB) exceeds a single 128 GB Spark, so serve it across both nodes with tensor-parallel = 2 (Ray). Grade/serve config that works on GB10:

vllm serve /path/to/this-model \
  --served-model-name ornith --trust-remote-code \
  --tensor-parallel-size 2 --distributed-executor-backend ray \
  --max-model-len 262144 --kv-cache-dtype fp8 \
  --gpu-memory-utilization 0.90 --max-num-seqs 1 \
  --enable-chunked-prefill --max-num-batched-tokens 8192 \
  --enable-auto-tool-choice --tool-call-parser qwen3_coder \
  --limit-mm-per-prompt '{"image":48,"video":4}'

GB10 notes: pin NCCL to the ConnectX-7 interface; disable Ray's OOM monitor (RAY_memory_monitor_refresh_ms=0) since on unified memory the weights legitimately use most of each node's 121 GiB; --enable-chunked-prefill is required for long context (uncapped prefill OOMs); nvidia-smi reports memory as N/A on GB10 — use free -h.

Notes for agent frameworks

  • Reasoning model: replies begin with an inline Thinking Process: chain-of-thought in message.content (no separate reasoning field). Budget max_tokens accordingly and treat content as reasoning.
  • Tool calls: with --enable-auto-tool-choice --tool-call-parser qwen3_coder, OpenAI-style tools/tool_choice:"auto" return structured tool_calls (verified). The CoT still appears in content alongside the tool call.

Files

config.json, quantization_config.json, model-*-of-*.safetensors (+ index), tokenizer.json/tokenizer_config.json/vocab.json, chat_template.jinja, generation_config.json, preprocessor_config.json, processor_config.json, video_preprocessor_config.json. See docs/ for the full validation report and an agent-integration brief, and scripts/ for the abliteration/quant/serve/eval pipeline used to build this.

License & credits

MIT (inherited from the base model). Base model: deepreinforce-ai/Ornith-1.0-397B. Abliteration follows the residual-direction method (Arditi et al.; tooling inspired by p-e-w/heretic and elder-plinius/OBLITERATUS); quantization via Intel AutoRound. Not affiliated with or endorsed by the base-model authors.

README history 1 version

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-07-03Add files using upload-large-folder tool12929a86.3 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration