← back to catalog · registered 2026-08-22 13:56

YuYu1015/YuYu1015-Ornith-1.0-35B-abliterated

YuYu1015 Qwen 35B MoE
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/YuYu1015%2FYuYu1015-Ornith-1.0-35B-abliterated"
Response includes
  • classification m1
  • files 10
  • hub_downloads_all_time 774
  • author_summary 23 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M1
Primary method

Direct removal

No other method signals detected in this model.
Confidence
MEDIUM
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=0 (base model)
  • no specific method indicators - defaulting to M1 (most common)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
MEDIUM
Why we say so
primary_method=M1; difference-of-means is the reference extraction for M1/M3 (Arditi 2024)
Downloads · lifetime
774
21 last 30d - cooling
Likes
1
Descendants
3
in 3 direct forks
Model age
3mo ago
created 2026-06-30
Downloads over time
Now783→from359↑118%
338500663825359 on Jul 1783 on Oct 11JulAugSepOct
Jul 1 → Oct 11 · 54 snapshots · spans 102 days

Genealogy 3 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Variants by this author 3 formats · 276 downloads combined

The same weights this author released in different packaging. Pick the format that matches your runtime.

Metadata

License
apache-2.0
Languages
en zh
Tags
transformers safetensors qwen3_5_moe image-text-to-text qwen3.5 moe gated-deltanet reasoning thinking abliterated uncensored text-generation

Related

Total size
65.4 GB
Files
10
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-07-02 02:00

Files by quantization

Auxiliary files 10 files 65.4 GB
model-00001-of-00002.safetensors 46.3 GB aaaa5219 download
model-00002-of-00002.safetensors 19.1 GB 9333573c download
tokenizer.json 19.1 MB 6f32ce20 download
model.safetensors.index.json 3.18 MB 4c252d6d download
README.md 10.8 KB 811842b3 download
chat_template.jinja 7.36 KB b07660cc download
config.json 3.22 KB 9166dfd1 download
.gitattributes 1.53 KB 52373fe2 download
tokenizer_config.json 1.14 KB 1d134cd2 download
generation_config.json 214 B 3f25ead4 download

README current version from Hugging Face


license: apache-2.0
base_model:

  • deepreinforce-ai/Ornith-1.0-35B
    base_model_relation: finetune
    library_name: transformers
    pipeline_tag: text-generation
    tags:
  • safetensors
  • qwen3.5
  • moe
  • gated-deltanet
  • reasoning
  • thinking
  • abliterated
  • uncensored
    language:
  • en
  • zh

YuYu1015-Ornith-1.0-35B-abliterated

English | 繁體中文


English

[!IMPORTANT]

🔄 Re-uploaded 2026-06-30 — please re-download

The model was upgraded on 2026-06-30. What changed vs the old version: moralizing roughly halved (~28% → ~14%), and the reasoning degradation of the old version was fixed — the new model now matches the base on GSM8K (80%) and even on the hardest competition-math (MATH-500 L4-5). If you downloaded an earlier copy, please re-download to get the improved version.

📦 Quantized versions: NVFP4 (Blackwell · vLLM / SGLang) · GGUF (llama.cpp · Q8_0 / UD-Q6_K / UD-Q4_K_M)

Support me on Ko-fi

[!WARNING]

⚠️ READ FIRST — Sampling Parameters MUST Be Set Correctly

This model requires the exact sampling parameters below, especially keeping repeat-penalty at 1.0 (the default — do not raise it). Wrong values break it:

Setting Result
repeat-penalty 1.0 ✅ recommended (default); this 35B rarely loops — only occasionally
repeat-penalty 1.05 truncated / unfinished answers
temp 0 (greedy) not recommended

An abliterated (uncensored) variant of deepreinforce-ai/Ornith-1.0-35B, a Qwen3.5 Mixture-of-Experts reasoning model. Refusal behavior has been removed and moralizing substantially reduced by weights-only abliteration (no training), keeping the base model's reasoning/thinking intact.

Model Details

Item Value
Architecture Qwen3.5 35B MoE — 40 layers (full-attention + GatedDeltaNet hybrid), 256 routed + 1 shared experts/layer, ~9 active per token
Base model deepreinforce-ai/Ornith-1.0-35B
Author YuYu1015
Precision BF16 (~70 GB, 2 shards)
Context length Inherited from base
Thinking mode Supported (reasoning model, emits <think>…</think>)
Languages English, Chinese

Evaluation

Measured on harmful-intent prompts (refusal / moralizing), GSM8K, and the hardest MATH-500 problems (reasoning). Hard refusal is detected via refusal-phrase markers, moralizing via a BERT classifier; GSM8K / MATH are exact-match accuracy.

Metric Base Ornith-1.0-35B This model
Hard refusal rate ~99% ~5%
Moralizing / disclaimer rate ~95% ~14%
GSM8K (reasoning) ~80% 80%
MATH-500 hardest tier (L4-5, competition) ~10% ~12%

→ Refusals essentially eliminated, moralizing cut to ~14% (from ~95% — roughly one-seventh), and reasoning not degraded: it matches the base both on GSM8K (80%) and on the hardest competition-math problems (MATH-500 level 4-5, AIME-tier — exactly where any reasoning damage from abliteration would surface). Weights-only, so the base model's original thinking is preserved.

Recommended Sampling Parameters

This is a reasoning model — keep thinking enabled and use the official Qwen3.5 sampling settings:

--temp 1.0
--top-p 0.95
--top-k 20
--min-p 0.0
--presence-penalty 0.0
--repeat-penalty 1.0

⚠️ For normal (long-form) generation, keep --repeat-penalty 1.0 and --presence-penalty 0. This 35B only rarely loops at these defaults. Do not raise them: repeat-penalty 1.05 truncates, and a high presence-penalty (e.g. 1.5) makes long answers drift off-topic / incoherent — it over-penalizes the topic words the model must reuse to stay on track. Exception — tool-calling / structured output: there, presence-penalty 1.5 actually helps (short, low-repetition output benefits from the anti-repetition pressure). Greedy decoding (--temp 0) is not recommended.

Usage

Transformers (BF16, multi-GPU):

from transformers import AutoModelForCausalLM, AutoTokenizer
import torch

m = "YuYu1015/YuYu1015-Ornith-1.0-35B-abliterated"
tok = AutoTokenizer.from_pretrained(m, trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained(m, dtype=torch.bfloat16,
                                             device_map="auto",       # ~70 GB — needs multi-GPU / a large GPU
                                             trust_remote_code=True).eval()
msgs = [{"role": "user", "content": "Your prompt here"}]
text = tok.apply_chat_template(msgs, add_generation_prompt=True, tokenize=False)
ids = tok(text, return_tensors="pt", add_special_tokens=False).to(model.device)
out = model.generate(**ids, max_new_tokens=4096, do_sample=True,
                     temperature=1.0, top_p=0.95, top_k=20, min_p=0.0,
                     repetition_penalty=1.0)    # default; no penalty needed — raising it (1.05) truncates
print(tok.decode(out[0][ids["input_ids"].shape[1]:], skip_special_tokens=True))

Safety Warning

This model has safety filtering removed (abliterated) and may generate sensitive, controversial, or inappropriate content. Users are solely responsible for all consequences and legal liability arising from its use, and must ensure usage complies with local laws and ethical standards.

Credits


繁體中文

[!IMPORTANT]

🔄 2026-06-30 重新上傳 —— 請重新下載

本模型於 2026-06-30 升級。新舊差異: 說教約砍半(~28% → ~14%),且修正了舊版的推理退化 —— 新版在 GSM8K(80%)與最難的競賽數學(MATH-500 L4-5)上都與原版持平。先前下載過舊版的使用者,請重新下載取得改善版。

📦 量化版本: NVFP4(Blackwell · vLLM / SGLang) · GGUF(llama.cpp · Q8_0 / UD-Q6_K / UD-Q4_K_M)

Support me on Ko-fi

[!WARNING]

⚠️ 必讀 — 取樣參數務必正確設定

本模型強依賴下方那組取樣參數,尤其 repeat-penalty 保持 1.0(預設、請勿調高)。設錯會壞掉:

設定 結果
repeat-penalty 1.0 ✅ 推薦(預設);此 35B 很少思考迴圈 —— 僅偶發
repeat-penalty 1.05 答不完被截斷
temp 0(貪婪) 不建議

deepreinforce-ai/Ornith-1.0-35B(Qwen3.5 **混合專家(MoE)推理模型)的 abliterated(去審查)版本。以純權重 abliteration(零訓練)**移除拒答、大幅降低說教,並保留原模型的推理/思考能力。

模型資訊

項目 數值
架構 Qwen3.5 35B MoE — 40 層(全注意力 + GatedDeltaNet 混合)、每層 256 routed + 1 shared 專家、每 token 約啟用 9 個
基礎模型 deepreinforce-ai/Ornith-1.0-35B
作者 YuYu1015
精度 BF16(約 70 GB、2 shards)
Context 長度 沿用基礎模型
思考模式 支援(推理模型,輸出 <think>…</think>)
語言 英文、中文

評估

於有害意圖 prompt(拒答/說教)、GSM8K 與最難的 MATH-500 題目(推理)上量測。硬拒答以拒絕語句標記偵測、說教以 BERT 分類器偵測;GSM8K/MATH 為精確比對正確率。

指標 原版 Ornith-1.0-35B 本模型
硬拒答率 ~99% ~5%
說教/免責率 ~95% ~14%
GSM8K(推理) ~80% 80%
MATH-500 最難檔(L4-5 競賽級) ~10% ~12%

→ 拒答幾乎清零、說教降至 ~14%(自 ~95%,約七分之一),推理能力未退化:在 GSM8K(80%)與最難的競賽數學(MATH-500 level 4-5,AIME 級 —— abliteration 若傷推理會在此現形)上都與原版持平。純權重,原始思考完整保留。

建議取樣參數

這是推理模型——請保持思考開啟,並使用 Qwen3.5 官方取樣設定:

--temp 1.0
--top-p 0.95
--top-k 20
--min-p 0.0
--presence-penalty 0.0
--repeat-penalty 1.0

⚠️ 一般(長文)生成:請保持 --repeat-penalty 1.0、--presence-penalty 0。 此 35B 在這組預設下只偶爾思考迴圈(很少、非持續)。請勿調高:repeat-penalty 1.05 會截斷;presence-penalty 調高(如 1.5)會讓長答案離題/答非所問 —— 它會過度懲罰模型必須複用的主題詞,導致偏離主軸。例外 —— tool call/結構化輸出: 那時 presence-penalty 1.5 反而有幫助(短、低重複的輸出受益於抗重複壓力)。同樣不建議用貪婪解碼(--temp 0)。

使用方式

Transformers(BF16、多卡):

from transformers import AutoModelForCausalLM, AutoTokenizer
import torch

m = "YuYu1015/YuYu1015-Ornith-1.0-35B-abliterated"
tok = AutoTokenizer.from_pretrained(m, trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained(m, dtype=torch.bfloat16,
                                             device_map="auto",       # 約 70 GB — 需多卡 / 大顯存
                                             trust_remote_code=True).eval()
msgs = [{"role": "user", "content": "你的問題"}]
text = tok.apply_chat_template(msgs, add_generation_prompt=True, tokenize=False)
ids = tok(text, return_tensors="pt", add_special_tokens=False).to(model.device)
out = model.generate(**ids, max_new_tokens=4096, do_sample=True,
                     temperature=1.0, top_p=0.95, top_k=20, min_p=0.0,
                     repetition_penalty=1.0)    # 預設;不需懲罰 —— 調高(1.05)會截斷
print(tok.decode(out[0][ids["input_ids"].shape[1]:], skip_special_tokens=True))

安全警告

此模型已移除安全過濾機制(abliterated),可能產生敏感、爭議性或不當內容。使用者須自行承擔所有風險與法律責任,並確保使用方式符合當地法規與倫理標準。

致謝

README history 3 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-07-02Update README86065d110.8 KB
    Loading...
  2. 2026-06-30Update README.md77151b99.7 KB
    Loading...
  3. 2026-06-30Create README.md540d9aa9.7 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration