← back to catalog · registered 2026-08-22 13:56

spinochenza/Ornith-1.0-35B-uncensored-heretic

spinochenza 35B GGUF MoE
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/spinochenza%2FOrnith-1.0-35B-uncensored-heretic"
Response includes
  • classification m3
  • files 13
  • hub_downloads_all_time 367
  • author_summary 17 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M3
Primary method

Layer-wise ablation

Applied on top of direct removal inherited from the base model.
Confidence
HIGH
Inherited from base model
Why this label 2 signals
Producer identity confirmed by naming conventions, tags or the model card. This label is very unlikely to change.
  • 'heretic' in model name (Heretic-produced)
  • Heretic uses layer-wise optimization (M3) with underlying direction removal (M1)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
HIGH
Why we say so
name contains 'heretic'; Heretic default extraction is difference-of-means (Arditi 2024)
Downloads · lifetime
367
9 last 30d - cooling
Likes
0
Model age
3mo ago
created 2026-07-02
Downloads over time
Now372→from207↑80%
199262325389207 on Jul 1372 on Oct 11372 on Oct 8JulAugSepOct
Jul 1 → Oct 11 · 54 snapshots · spans 102 days

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
mit
Tags
transformers safetensors qwen3_5_moe image-text-to-text heretic uncensored decensored abliterated mpoa text-generation conversational base_model:deepreinforce-ai/Ornith-1.0-35B-GGUF

Related

Total size
65.4 GB
Files
13
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-07-02 06:20

Files by quantization

Auxiliary files 13 files 65.4 GB
model-00001-of-00002.safetensors 46.3 GB 263cfc7e download
model-00002-of-00002.safetensors 19.1 GB dea4a90a download
tokenizer.json 19.1 MB 6f32ce20 download
model.safetensors.index.json 3.21 MB 3988fa1b download
README.md 29.1 KB 943ce145 download
chat_template.jinja 7.83 KB 1e6f3b77 download
config.json 3.34 KB d59b3f89 download
tokenizer_config.json 1.17 KB 7b0f9106 download
processor_config.json 1.16 KB 33818c7f download
preprocessor_config.json 390 B 2ea84a43 download
video_preprocessor_config.json 385 B 3ba673a5 download
.gitattributes 311 B 0f16200f download
generation_config.json 227 B e8f101cd download

README current version from Hugging Face


library_name: transformers
license: mit
license_link: https://huggingface.co/deepreinforce-ai/Ornith-1.0-35B/blob/main/LICENSE
pipeline_tag: text-generation
tags:

  • heretic
  • uncensored
  • decensored
  • abliterated
  • mpoa
    base_model:
  • deepreinforce-ai/Ornith-1.0-35B-GGUF

🚨⚠️ I HAVE REACHED HUGGING FACE'S FREE STORAGE LIMIT ⚠️🚨

I can no longer upload new models unless I can cover the cost of additional storage.
I host 70+ free models as an independent contributor and this work is unpaid.
Without your support, no more new models can be uploaded.

🎉 Patreon (Monthly)  |  ☕ Ko-fi (One-time)

Every contribution goes directly toward Hugging Face storage fees to keep models free for everyone.


90% fewer refusals (9/100 Uncensored vs 89/100 Original) while preserving model quality (0.0019 KL divergence).

❤️ Support My Work

Creating these models takes significant time, work and compute. If you find them useful consider supporting me:

image/png

Platform Link What you get
🎉 Patreon Monthly support Priority model requests
☕ Ko-fi One-time tip My eternal gratitude

Your help will motivate me and would go into further improving my workflow and coverings fees for storage, compute and may even help uncensoring bigger model with rental Cloud GPUs.


This is a decensored version of deepreinforce-ai/Ornith-1.0-35B-GGUF, made using Heretic v1.2.0 with a variant of the Magnitude-Preserving Orthogonal Ablation (MPOA) method

Abliteration parameters

Parameter Value
direction_index 20.57
attn.out_proj.max_weight 1.97
attn.out_proj.max_weight_position 29.36
attn.out_proj.min_weight 1.41
attn.out_proj.min_weight_distance 23.76
mlp.down_proj.max_weight 1.07
mlp.down_proj.max_weight_position 31.48
mlp.down_proj.min_weight 0.62
mlp.down_proj.min_weight_distance 26.62
attn.o_proj.max_weight 1.98
attn.o_proj.max_weight_position 24.78
attn.o_proj.min_weight 0.09
attn.o_proj.min_weight_distance 27.94

Targeted components

  • attn.o_proj
  • attn.out_proj
  • mlp.down_proj

Performance

Metric This model Original model (Ornith-1.0-35B-GGUF)
KL divergence 0.0019 0 (by definition)
Refusals ✅ 9/100 ❌ 89/100

MMLU test results:

Original:

============================================================

  • Total questions: 7021

  • Correct: 5802

  • Accuracy: 0.8264 (82.64%)

  • Parse failures: 0

============================================================

Tested subject scores:

  • professional_law: 0.6917 (543/785)
  • moral_scenarios: 0.6742 (298/442)
  • miscellaneous: 0.9295 (356/383)
  • professional_psychology: 0.8892 (281/316)
  • high_school_psychology: 0.9593 (259/270)
  • high_school_macroeconomics: 0.8832 (174/197)
  • elementary_mathematics: 0.7717 (142/184)
  • moral_disputes: 0.8506 (148/174)
  • prehistory: 0.8779 (151/172)
  • philosophy: 0.8931 (142/159)
  • high_school_biology: 0.9474 (144/152)
  • professional_accounting: 0.6783 (97/143)
  • clinical_knowledge: 0.9000 (126/140)
  • high_school_microeconomics: 0.9632 (131/136)
  • nutrition: 0.8593 (116/135)
  • professional_medicine: 0.9104 (122/134)
  • conceptual_physics: 0.9141 (117/128)
  • high_school_mathematics: 0.5748 (73/127)
  • human_aging: 0.7931 (92/116)
  • security_studies: 0.8750 (98/112)
  • high_school_statistics: 0.8108 (90/111)
  • marketing: 0.9083 (99/109)
  • high_school_world_history: 0.9057 (96/106)
  • sociology: 0.9515 (98/103)
  • high_school_government_and_politics: 0.9901 (100/101)
  • high_school_geography: 0.9293 (92/99)
  • high_school_chemistry: 0.7732 (75/97)
  • high_school_us_history: 0.9579 (91/95)
  • virology: 0.5056 (45/89)
  • college_medicine: 0.8523 (75/88)
  • world_religions: 0.9091 (80/88)
  • high_school_physics: 0.7738 (65/84)
  • electrical_engineering: 0.8395 (68/81)
  • astronomy: 0.9620 (76/79)
  • logical_fallacies: 0.9474 (72/76)
  • high_school_european_history: 0.8630 (63/73)
  • anatomy: 0.8873 (63/71)
  • college_biology: 0.9062 (58/64)
  • human_sexuality: 0.8594 (55/64)
  • formal_logic: 0.6562 (42/64)
  • public_relations: 0.7377 (45/61)
  • international_law: 0.9333 (56/60)
  • college_physics: 0.7193 (41/57)
  • college_mathematics: 0.6727 (37/55)
  • econometrics: 0.7593 (41/54)
  • jurisprudence: 0.8679 (46/53)
  • high_school_computer_science: 0.8846 (46/52)
  • machine_learning: 0.8269 (43/52)
  • medical_genetics: 0.9412 (48/51)
  • global_facts: 0.5490 (28/51)
  • management: 0.9000 (45/50)
  • us_foreign_policy: 0.9400 (47/50)
  • college_chemistry: 0.5745 (27/47)
  • abstract_algebra: 0.6170 (29/47)
  • business_ethics: 0.8478 (39/46)
  • college_computer_science: 0.7556 (34/45)
  • computer_security: 0.8605 (37/43)

Heretic:

============================================================

  • Total questions: 7021

  • Correct: 5737

  • Accuracy: 0.8171 (81.71%)

  • Parse failures: 0

============================================================

Tested subject scores:

  • professional_law: 0.6815 (535/785)
  • moral_scenarios: 0.5837 (258/442)
  • miscellaneous: 0.9347 (358/383)
  • professional_psychology: 0.8892 (281/316)
  • high_school_psychology: 0.9630 (260/270)
  • high_school_macroeconomics: 0.8883 (175/197)
  • elementary_mathematics: 0.7717 (142/184)
  • moral_disputes: 0.8563 (149/174)
  • prehistory: 0.8837 (152/172)
  • philosophy: 0.8868 (141/159)
  • high_school_biology: 0.9474 (144/152)
  • professional_accounting: 0.6783 (97/143)
  • clinical_knowledge: 0.9000 (126/140)
  • high_school_microeconomics: 0.9559 (130/136)
  • nutrition: 0.8741 (118/135)
  • professional_medicine: 0.8955 (120/134)
  • conceptual_physics: 0.9219 (118/128)
  • high_school_mathematics: 0.5354 (68/127)
  • human_aging: 0.7759 (90/116)
  • security_studies: 0.8482 (95/112)
  • high_school_statistics: 0.8018 (89/111)
  • marketing: 0.9083 (99/109)
  • high_school_world_history: 0.8962 (95/106)
  • sociology: 0.9320 (96/103)
  • high_school_government_and_politics: 0.9901 (100/101)
  • high_school_geography: 0.9394 (93/99)
  • high_school_chemistry: 0.8041 (78/97)
  • high_school_us_history: 0.9474 (90/95)
  • virology: 0.5056 (45/89)
  • college_medicine: 0.8523 (75/88)
  • world_religions: 0.9091 (80/88)
  • high_school_physics: 0.7738 (65/84)
  • electrical_engineering: 0.8642 (70/81)
  • astronomy: 0.9367 (74/79)
  • logical_fallacies: 0.9474 (72/76)
  • high_school_european_history: 0.8767 (64/73)
  • anatomy: 0.8732 (62/71)
  • college_biology: 0.9062 (58/64)
  • human_sexuality: 0.8594 (55/64)
  • formal_logic: 0.7031 (45/64)
  • public_relations: 0.7541 (46/61)
  • international_law: 0.8667 (52/60)
  • college_physics: 0.7719 (44/57)
  • college_mathematics: 0.6364 (35/55)
  • econometrics: 0.7037 (38/54)
  • jurisprudence: 0.8679 (46/53)
  • high_school_computer_science: 0.9231 (48/52)
  • machine_learning: 0.7885 (41/52)
  • medical_genetics: 0.9216 (47/51)
  • global_facts: 0.5294 (27/51)
  • management: 0.9000 (45/50)
  • us_foreign_policy: 0.9000 (45/50)
  • college_chemistry: 0.5106 (24/47)
  • abstract_algebra: 0.6383 (30/47)
  • business_ethics: 0.8043 (37/46)
  • college_computer_science: 0.7333 (33/45)
  • computer_security: 0.8605 (37/43)

MMLU - Massive Multitask Language Understanding, multiple-choice questions across 57 subjects (math, history, law, medicine, etc.).

GGUF Version

GGUF quantizations available here llmfan46/Ornith-1.0-35B-uncensored-heretic-GGUF.


Ornith Blog

Ornith-1.0-35B

Aloha! 🌺 Today, we are releasing Ornith-1.0, a self-improving family of open-source models for agentic coding.

Highlights:

  • State-of-the-Art Coding Agents: Available in 9B-Dense, 31B-Dense, 35B-MoE, and 397B-MoE (post-trained on top of Gemma 4 and Qwen 3.5), achieving state-of-the-art performance among open-source models of comparable size on coding benchmarks such as Terminal-Bench 2.1, SWE-Bench, NL2Repo and OpenClaw.
  • Self-Improving Training Framework: Ornith-1.0 employs RL to learn to generate not only solution rollouts, but also the scallfold that drive those rollouts. By jointly optimizing the scaffold and the resulting solution, the model discovers better search trajectories and generates higher-quality solutions.
  • Licence: MIT licensed, globally accessible, and free from regional limitations.
Ornith 35B Benchmark Results

Ornith 1.0 35B

This model card documents Ornith-1.0-35B, the lightweight member of the Ornith family, designed for efficient single-GPU deployment.

Benchmarks

Ornith-1.0-35B Qwen3.5-35B Qwen3.6-35B Gemma4-31B Qwen3.5-397B
Agentic Coding
Terminal-Bench 2.1 (Terminus-2) 64.2 41.4 52.5 42.1 53.5
Terminal-Bench 2.1 (Claude Code) 62.8 38.9 49.2 - 48.6
SWE-bench Verified 75.6 70 73.4 52 76.4
SWE-bench Pro 50.4 44.6 49.5 35.7 51.6
SWE-bench Multilingual 69.3 60.3 67.2 51.7 69.3
NL2Repo 34.6 20.5 29.4 15.5 36.8
Claw-eval Avg 69.8 65.4 68.7 48.5 70.7
SWE Atlas - QnA 37.1 13.2 15.5 - 20.4
SWE Atlas - RF 29.7 10.2 11.4 - 18.4
SWE Atlas - TW 27.8 9.8 13.3 - 18.5

* Terminal-Bench 2.1 (Terminus-2): We evaluate Terminal-Bench 2.1 using the Harbor/Terminus-2 framework with parser=json, temperature=1.0, top_p=1.0, and a 128K context window. Each run uses a 4-hour timeout with 32 CPU cores and 48GB RAM, and results are averaged over 5 runs. We adjust the Qwen chat template to ensure consistency between training and inference (https://huggingface.co/deepreinforce-ai/Ornith-1.0-397B/blob/main/chat_template.jinja), and modify Harbor to align with vLLM's reasoning_content key.
* Terminal-Bench 2.1 (Claude Code): We evaluate Terminal-Bench 2.1 using Claude Code 2.1.126 with parser=json, temperature=1.0, top_p=1.0, max_new_tokens=131072. Results are averaged over 5 runs. Again, Qwen chat template needs to be modified.
* SWE-Bench Verified, Pro and Multilingual: using OpenHands harness with temp=1.0, top_p=0.95, 256k context window.
* SWE Atlas QnA, RF, TW: using mini SWE agent harness with temp=1.0, top_p=0.95, 128K context window. Results are averaged over 5 runs.
* NL2Repo: with temperature=1.0, top_p=1.0, 400K context, 48K output and anti-hacking filters.
* ClawEval: An agentic code benchmark over real-user task distributions; temp=0.6 and 256K context.

Quickstart

📝 NOTE

Ornith-1.0-35B is a reasoning model: by default the assistant turn opens with a <think> … </think> block before the final answer. The serving recipes below enable a reasoning parser so the chain-of-thought is returned in a separate reasoning_content field, and a tool-call parser so the model's <tool_call> blocks are surfaced as OpenAI-style tool_calls.

Serving Ornith-1.0-35B requires recent runtimes:

  • Transformers ≥ 5.8.1
  • vLLM ≥ 0.19.1
  • SGLang ≥ 0.5.9

Serving Ornith-1.0-35B

The two recipes below stand up an OpenAI-compatible server on a single 8×80GB GPU node (tensor-parallel 8). Adjust --tensor-parallel-size / --tp to the number of GPUs you have.

vLLM

vllm serve deepreinforce-ai/Ornith-1.0-35B \
    --served-model-name Ornith-1.0-35B \
    --tensor-parallel-size 8 \
    --host 0.0.0.0 --port 8000 \
    --max-model-len 262144 \
    --gpu-memory-utilization 0.90 \
    --enable-prefix-caching \
    --enable-auto-tool-choice --tool-call-parser qwen3_xml \
    --reasoning-parser qwen3 \
    --trust-remote-code

SGLang

python -m sglang.launch_server \
    --model-path deepreinforce-ai/Ornith-1.0-35B \
    --served-model-name Ornith-1.0-35B \
    --tp 8 \
    --host 0.0.0.0 --port 8000 \
    --context-length 262144 \
    --mem-fraction-static 0.85 \
    --tool-call-parser qwen3_coder \
    --reasoning-parser qwen3

Hugging Face Transformers

For a quick local test (or to script offline generation), load the model directly with Transformers. Make sure you have a recent release installed — see the Transformers installation guide; Ornith-1.0-35B requires transformers >= 5.8.1.

from transformers import AutoModelForCausalLM, AutoTokenizer

model_name = "deepreinforce-ai/Ornith-1.0-35B"

tokenizer = AutoTokenizer.from_pretrained(model_name)
model = AutoModelForCausalLM.from_pretrained(
    model_name,
    dtype="auto",
    device_map="auto",
)

messages = [
    {"role": "user", "content": "Write a Python function is_prime(n). Keep it short."}
]
text = tokenizer.apply_chat_template(
    messages,
    tokenize=False,
    add_generation_prompt=True,
)

inputs = tokenizer(text, return_tensors="pt").to(model.device)
generated = model.generate(
    **inputs,
    max_new_tokens=512,
    do_sample=True,
    temperature=0.6,
    top_p=0.95,
    top_k=20,
)
output_ids = generated[0][inputs.input_ids.shape[1]:]

# The reply contains a <think> ... </think> reasoning block followed by the answer.
content = tokenizer.decode(output_ids, skip_special_tokens=True)
print(content)

To split the reasoning trace from the final answer, parse on the </think> marker:

text = tokenizer.decode(output_ids, skip_special_tokens=True)
if "</think>" in text:
    reasoning, answer = text.split("</think>", 1)
    reasoning = reasoning.replace("<think>", "").strip()
    answer = answer.strip()
else:
    reasoning, answer = "", text.strip()

Using Ornith-1.0-35B via the Chat Completions API

Once a vLLM or SGLang server is running, talk to it with any OpenAI-compatible client.

Basic Usage

from openai import OpenAI

client = OpenAI(
    base_url="http://localhost:8000/v1",
    api_key="EMPTY",  # any non-empty string works for a local server
)

response = client.chat.completions.create(
    model="Ornith-1.0-35B",
    messages=[
        {"role": "user", "content": "Write a one-line Python lambda that squares a number."}
    ],
    temperature=0.6,
    top_p=0.95,
    max_tokens=1024,
)

message = response.choices[0].message
# reasoning_content holds the <think> trace; content holds the final answer.
print("reasoning:", getattr(message, "reasoning_content", None))
print("answer:", message.content)

You can also stream tokens, or hand the model tools — Ornith-1.0-35B emits well-formed function calls that the server parses into the standard tool_calls field:

tools = [
    {
        "type": "function",
        "function": {
            "name": "get_weather",
            "description": "Get the current weather for a city",
            "parameters": {
                "type": "object",
                "properties": {"city": {"type": "string"}},
                "required": ["city"],
            },
        },
    }
]

response = client.chat.completions.create(
    model="Ornith-1.0-35B",
    messages=[{"role": "user", "content": "What is the weather in Paris right now?"}],
    tools=tools,
    tool_choice="auto",
    temperature=0.6,
    max_tokens=2048,
)

tool_call = response.choices[0].message.tool_calls[0]
print(tool_call.function.name, tool_call.function.arguments)
# -> get_weather {"city": "Paris"}

You can point any OpenAI-compatible SDK (Python, Node.js, etc.) or curl at the same /v1/chat/completions endpoint.

Agentic Usage

Ornith-1.0-35B excels in tool-calling and agentic coding capabilities.

Agent Frameworks

Because Ornith-1.0-35B exposes an OpenAI-compatible endpoint with tool calling, it works out of the box with standard agent frameworks. Below is a minimal example that connects Ornith-1.0-35B to tools through an MCP server.

import os
from openai import OpenAI

client = OpenAI(
    base_url=os.getenv("OPENAI_BASE_URL", "http://localhost:8000/v1"),
    api_key=os.getenv("OPENAI_API_KEY", "EMPTY"),
)

tools = [
    {
        "type": "function",
        "function": {
            "name": "run_shell",
            "description": "Run a shell command and return its output.",
            "parameters": {
                "type": "object",
                "properties": {
                    "command": {"type": "string", "description": "The command to run"}
                },
                "required": ["command"],
            },
        },
    }
]

messages = [{"role": "user", "content": "List the Python files in the current directory."}]

response = client.chat.completions.create(
    model="deepreinforce-ai/Ornith-1.0-35B",
    messages=messages,
    tools=tools,
    temperature=0.6,
    top_p=0.95,
)
print(response.choices[0].message)

Examples of using Ornith with agent harness:

Hermes Agent

# Hermes talks to any OpenAI-compatible endpoint — point it at your Ornith server.
export OPENAI_BASE_URL="http://localhost:8000/v1"
export OPENAI_API_KEY="EMPTY"
export MODEL="deepreinforce-ai/Ornith-1.0-35B"

Atomic.chat/ Ollama / llama.cpp

# Both runtimes load a GGUF build of Ornith (publish one at deepreinforce-ai/Ornith-1.0-35B-GGUF).

# llama.cpp — serve an OpenAI-compatible API on port 8000.
llama-server -hf deepreinforce-ai/Ornith-1.0-35B-GGUF --port 8000 -c 262144

# Ollama — pull and chat with the same GGUF straight from Hugging Face.
ollama run hf.co/deepreinforce-ai/Ornith-1.0-35B-GGUF

OpenClaw

# OpenClaw talks to any OpenAI-compatible endpoint — point it at your Ornith server.
export OPENAI_BASE_URL="http://localhost:8000/v1"
export OPENAI_API_KEY="EMPTY"
export OPENAI_MODEL="deepreinforce-ai/Ornith-1.0-35B"

Unsloth Studio

pip install unsloth

# Load Ornith for fast local inference or fine-tuning (Python):
#   from unsloth import FastLanguageModel
#   model, tokenizer = FastLanguageModel.from_pretrained(
#       "deepreinforce-ai/Ornith-1.0-35B",
#       max_seq_length=262144,
#       load_in_4bit=True,
#   )

OpenHands

pip install openhands-ai

# OpenHands routes through LiteLLM; the "openai/" prefix selects the OpenAI-compatible path.
export LLM_MODEL="openai/deepreinforce-ai/Ornith-1.0-35B"
export LLM_BASE_URL="http://localhost:8000/v1"
export LLM_API_KEY="EMPTY"

# Launch the CLI (or run the official OpenHands Docker image with the same env vars).
openhands

Coding CLIs

Ornith-1.0-35B is optimized for terminal-based coding agents. Point any OpenAI-compatible coding CLI at your Ornith-1.0-35B endpoint (set OPENAI_BASE_URL and OPENAI_API_KEY) to understand large codebases, automate tedious work, and ship faster.

OpenCode

# Register your local Ornith endpoint as a provider in ~/.config/opencode/opencode.json:
#
# {
#   "$schema": "https://opencode.ai/config.json",
#   "provider": {
#     "ornith": {
#       "npm": "@ai-sdk/openai-compatible",
#       "name": "Ornith (local)",
#       "options": { "baseURL": "http://localhost:8000/v1", "apiKey": "EMPTY" },
#       "models": { "deepreinforce-ai/Ornith-1.0-35B": { "name": "Ornith-1.0-35B" } }
#     }
#   }
# }

opencode

Citation

If you find our work helpful, feel free to give us a cite.

@misc{ornith-35b,
    title = {{Ornith-1.0-35B}: Agentic Coding, Open to All},
    url = {https://deep-reinforce.com/ornith_1_0.html},
    author = {{DeepReinforce Team}},
    year = {2026}
}

README history 1 version

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-07-02Duplicate from llmfan46/Ornith-1.0-35B-uncensored-heretica28362629.1 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration