← back to catalog · registered 2026-08-22 13:56

maci0/Ornith-1.0-35B-abliterated-NVFP4

maci0 34B MoE multimodal
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/maci0%2FOrnith-1.0-35B-abliterated-NVFP4"
Response includes
  • classification m1
  • files 12
  • hub_downloads_all_time 915
  • author_summary 8 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M1
Primary method

Direct removal

No other method signals detected in this model.
Confidence
MEDIUM
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=0 (base model)
  • no specific method indicators - defaulting to M1 (most common)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
MEDIUM
Why we say so
primary_method=M1; difference-of-means is the reference extraction for M1/M3 (Arditi 2024)
Downloads · lifetime
915
67 last 30d - cooling
Likes
1
Model age
3mo ago
created 2026-07-06
Downloads over time
Now957→from325↑194%
2935367781K325 on Jul 13957 on Oct 11JulAugSepOct
Jul 13 → Oct 11 · 54 snapshots · spans 90 days

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
apache-2.0
Languages
en zh
Tags
transformers safetensors qwen3_5_moe image-text-to-text nvfp4 fp4 w4a4 gptq quantized compressed-tensors llm-compressor vllm

Related

Total size
20.4 GB
Files
12
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-08-04 01:51

Files by quantization

Auxiliary files 12 files 20.4 GB
model.safetensors 19.6 GB e72a79e2 download
model-towers.safetensors 852 MB f730cecc download
tokenizer.json 19.1 MB 06b95093 download
model.safetensors.index.json 11.9 MB 92ec17e7 download
config.json 413 KB e9718549 download
README.md 15.7 KB 94a34f85 download
chat_template.jinja 7.36 KB b07660cc download
.gitattributes 1.60 KB a09db2ea download
processor_config.json 1.16 KB 33818c7f download
tokenizer_config.json 1.14 KB 1d134cd2 download
recipe.yaml 372 B ba0bdb40 download
generation_config.json 214 B 3f25ead4 download

README current version from Hugging Face


base_model: deepreinforce-ai/Ornith-1.0-35B
base_model_relation: quantized
license: apache-2.0
library_name: transformers
pipeline_tag: image-text-to-text
language:

  • en
  • zh
    tags:
  • nvfp4
  • fp4
  • w4a4
  • gptq
  • quantized
  • compressed-tensors
  • llm-compressor
  • vllm
  • qwen3_5
  • moe
  • mixture-of-experts
  • experimental
  • vision-language
  • thinking
  • code
  • coder
  • agentic
  • uncensored
  • abliterated

RQ-35B-ORN-XExperimental
Ornith-1.0-35B MoE · NVFP4
35B MoE (256 experts, 8 active) agentic coder · refusals removed (custom MoE-aware).
Params35B
Active~3B · 8/256 experts
Size20 GB
Perplexity7.33
Refusals99 → 0 / 8
Context256K
MTP headn/a

TL;DR: Ornith-1.0-35B, a 256-expert MoE, quantized to NVFP4 (W4A4) for vLLM on NVIDIA Blackwell. 20 GB, wikitext-2 PPL 7.33, agentic coder, refusals removed with a custom MoE-aware method (2 rounds).

[!WARNING]
Experimental. This build is mostly a research artifact. The abliteration uses a custom fused-expert MoE method (not a standard tool), validated only on a small informal probe rather than a rigorous benchmark, and NVFP4 quantization of MoE models is bleeding-edge. Expect rough edges. Treat it as a proof of concept, not a production model, and evaluate it yourself before relying on it.

Ornith-1.0-35B abliterated NVFP4

deepreinforce-ai/Ornith-1.0-35B,
a 35B Mixture-of-Experts model abliterated with a custom MoE-aware method (refusal
directions removed from the fused expert tensors), then quantized to NVFP4 (W4A4) in the
compressed-tensors nvfp4-pack-quantized format with
llm-compressor (GPTQ + MSE, shared
fused-layer scales).

Decensored, and compact. Standard abliteration tools cannot reach a fused-expert MoE, so
this model was abliterated with a custom per-expert orthogonalization (see
Abliteration). NVFP4 then compresses the model to ~20 GB with a wikitext-2
perplexity of 7.33.

Refusals (held-out probe) 0/8 harmful prompts it previously refused
Abliteration custom MoE-aware per-expert orthogonalization, 2 rounds
Size on disk ~20 GB vs ~70 GB bf16 (~29%)
wikitext-2 PPL 7.33
  • Built for vLLM on NVIDIA Blackwell (4-bit weight + 4-bit activation). Pre-Blackwell GPUs
    run it weight-only.
  • Loading and generation verified in vLLM on an NVIDIA GB10 (Blackwell, sm_121).

Uncensored / abliterated model. It follows instructions without refusal guardrails. The
abliteration only removes refusals; all other behaviour comes from the base model. You
are responsible for how you use it.

Fidelity

Near-lossless versus the bf16 source, 20 GB vs 70 GB bf16 (~29%), at wikitext-2 perplexity 7.33. GPTQ error compensation and an MSE observer keep the drop from bf16 minimal; the header lists the full characteristics and Quantization covers the recipe.

Quickstart

NVFP4 is auto-detected from config.json (compressed-tensors); no quantization flag
needed.

vllm serve maci0/Ornith-1.0-35B-abliterated-NVFP4 \
  --served-model-name ornith-35b-abliterated-nvfp4 \
  --max-model-len 131072 \
  --gpu-memory-utilization 0.90 \
  --kv-cache-dtype fp8 \
  --reasoning-parser qwen3 \
  --enable-auto-tool-choice --tool-call-parser qwen3_coder
  • Supports up to 262144 tokens; keep at least 128K to preserve thinking quality.
  • Add --language-model-only to skip the vision tower and free KV cache for text use.
  • The parser flags are not auto-detected; pass them explicitly.

Python (OpenAI client)

from openai import OpenAI
client = OpenAI(base_url="http://localhost:8000/v1", api_key="x")
r = client.chat.completions.create(
    model="ornith-35b-abliterated-nvfp4",
    messages=[{"role": "user", "content": "Add a retry with exponential backoff to this HTTP client and explain the change."}],
)
print(r.choices[0].message.content)

curl

curl http://localhost:8000/v1/chat/completions -H "Content-Type: application/json" -d '{
  "model": "ornith-35b-abliterated-nvfp4",
  "messages": [{"role": "user", "content": "Add a retry with exponential backoff to this HTTP client and explain the change."}]
}'

About the base model

Ornith-1.0 is a self-improving family of open agentic-coding models from
Deep Reinforce. The 35B member is a
Mixture-of-Experts Qwen3.5-MoE-family vision-language model with thinking-mode reasoning and
a 256K context.

  • 40 decoder layers, Mixture-of-Experts with 256 experts per layer (packed as fused expert
    tensors), plus a vision tower for image and video input.
  • 256K context (max_position_embeddings 262144).
  • Thinking mode by default, with an instruct toggle (preserved here; abliteration and
    quantization keep the original chat template).

Abliteration

This is the notable part of this build. Ornith-1.0-35B is a fused-expert MoE: all 256
experts of a layer are packed into a single fused tensor (Qwen3_5MoeExperts.down_proj)
rather than 256 iterable Linear modules. Standard abliteration tooling (including
Heretic) walks the module tree looking for Linear
layers, so it silently skips every expert and can only touch the attention projections.
On this model that leaves a hard refusal floor (roughly 88/100 harmful prompts still
refused), because the refusal behaviour lives in the experts.

To reach the experts, this model was abliterated with a custom MoE-aware method: for each
layer, compute the refusal direction, then orthogonalize the fused experts.down_proj
output per expert
against that direction. That is 10240 expert slices (256 experts x 40
layers), applied in two rounds (a second pass removes the residual refusal signal that
remains after the first). Refusal directions are estimated from
mlabonne/harmless_alpaca (good) vs mlabonne/harmful_behaviors (bad).

Stage Result
Standard tools (Heretic) attention only, experts skipped, refusal floor ~88/100
MoE-aware round 1 per-expert orthogonalization of the fused experts.down_proj
MoE-aware round 2 residual refusal direction removed
Held-out probe 0/8 harmful prompts it previously refused

Honest validation caveat. The result was checked on a small held-out probe: the model no
longer refuses 8 harmful prompts it previously refused (0/8). This is a sanity check, not a
rigorous refusals-at-fixed-KL benchmark like the smaller dense Ornith build. Treat the
decensoring as demonstrated on a small probe, not exhaustively measured.

Quantization

Scheme NVFP4, W4A4
Weight rounding GPTQ (Hessian-based error compensation), MSE observer
Weights FP4 (E2M1), group_size=16, tensor_group, FP8 (E4M3) group scales, shared across fused layers
Activations FP4, dynamic per-group, FP8 (E4M3) scales
Quantized all language-model Linear layers, including the fused MoE expert tensors
MoE calibration moe_calibrate_all_experts=True (every one of the 256 experts must be routed during calibration, or unrouted experts produce garbage)
Kept in bf16 routers (mlp.gate, shared_expert_gate), vision tower (model.visual.*), lm_head
Untouched gated delta-net Conv1d and SSM params (A_log, dt_bias), never Linear

GPTQ is a quantization-time cost only; inference speed and format are identical to plain
round-to-nearest NVFP4, but it chooses better 4-bit values.

Calibration: 512 domain-matched samples (long reasoning + general chat + code),
max_seq_len=2048, text-only path through the VL model, with
moe_calibrate_all_experts=True so all 256 experts per layer receive calibration traffic.

Toolchain: llmcompressor==0.12.0, compressed-tensors==0.17.1, transformers==5.12.1,
torch==2.11.0+cu130, on an NVIDIA GB10 (Blackwell, sm_121). The routers (mlp.gate,
shared_expert_gate) are left in bf16 so routing stays exact.

Recommended sampling

Thinking mode is the default.

  • Thinking, precise coding: temperature=0.6, top_p=0.95, top_k=20
  • Thinking, general: temperature=1.0, top_p=0.95, top_k=20
  • Instruct / non-thinking: temperature=0.7, top_p=0.80, top_k=20
  • To run non-thinking, set {%- set enable_thinking = false %} in the chat template, or
    pass extra_body={"chat_template_kwargs": {"enable_thinking": false}}.

Related

Notes

  • Needs NVIDIA Blackwell (sm_121, e.g. GB10) for accelerated W4A4; pre-Blackwell GPUs run it weight-only.
  • --reasoning-parser and --tool-call-parser are not auto-detected; pass them explicitly.
  • Thinking mode is on by default; toggle it via the chat template or chat_template_kwargs.
  • No refusal guardrails; you are responsible for how you use it.

License

Apache-2.0, following the base model. Intended use and all responsibility for use follow
the base model.

Credits

  • Base model: Deep Reinforce (Ornith-1.0)
  • Abliteration: custom MoE-aware per-expert orthogonalization (fused experts.down_proj), because standard Linear-walking tools cannot reach fused MoE experts
  • Quantization tooling: llm-compressor / compressed-tensors
Part of Rogue Quants · NVFP4 component datasheets · collection. Fabricated on GB10 (Blackwell) with llm-compressor. Refusals shown per 100 harmful prompts; "n/a" = not separately measured (base-inherited).

README history 8 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-08-04uniform schema, no em-dashes4e4710415.7 KB
    Loading...
  2. 2026-08-04card: uniform characteristics schema (params/active/experts/etc)1e1920b15.7 KB
    Loading...
  3. 2026-08-03trim redundant Fidelity table (covered by header)eac003b14.6 KB
    Loading...
  4. 2026-08-03card: readability - sans labels, mono values, more spacingd341e0714.7 KB
    Loading...
  5. 2026-08-03card: fix table borders + readability7f4505214.7 KB
    Loading...
  6. 2026-08-03card: die sigil + a11y table semantics37e1ede14.2 KB
    Loading...
  7. 2026-08-03card: DATASHEET design languagea060f5413.6 KB
    Loading...
  8. 2026-07-07Add model card (experimental, MoE-aware abliteration)cef2af015.7 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration