← back to catalog · registered 2026-08-22 13:56

YuYu1015/YuYu1015-Ornith-1.0-35B-abliterated-NVFP4

YuYu1015 Qwen 16B MoE second-order
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/YuYu1015%2FYuYu1015-Ornith-1.0-35B-abliterated-NVFP4"
Response includes
  • classification m1
  • files 12
  • hub_downloads_all_time 567
  • author_summary 23 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M1
Primary method

Direct removal

No other method signals detected in this model.
Confidence
MEDIUM
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=0 (base model)
  • no specific method indicators - defaulting to M1 (most common)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
MEDIUM
Why we say so
primary_method=M1; difference-of-means is the reference extraction for M1/M3 (Arditi 2024)
Downloads · lifetime
567
80 last 30d - stable
Likes
0
Model age
3mo ago
created 2026-07-01
Downloads over time
Now595→from151↑294%
129299469639151 on Jul 1595 on Oct 11JulAugSepOct
Jul 1 → Oct 11 · 54 snapshots · spans 102 days

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Variants by this author 3 formats · 276 downloads combined

The same weights this author released in different packaging. Pick the format that matches your runtime.

Metadata

License
apache-2.0
Languages
en zh
Tags
safetensors qwen3_5_moe nvfp4 fp4 modelopt vllm sglang qwen3.5 moe gated-deltanet reasoning abliterated

Related

Total size
22.1 GB
Files
12
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-07-02 02:00

Files by quantization

Auxiliary files 12 files 22.1 GB
model-00002-of-00003.safetensors 9.32 GB 25a05ffa download
model-00001-of-00003.safetensors 9.32 GB 7adb1076 download
model-00003-of-00003.safetensors 3.48 GB 8dcb0b82 download
tokenizer.json 19.1 MB 6f32ce20 download
model.safetensors.index.json 13.1 MB 3a5f0d03 download
config.json 9.53 KB 64548509 download
chat_template.jinja 7.36 KB b07660cc download
README.md 5.43 KB 2f28dc6b download
hf_quant_config.json 5.25 KB 0b84ca48 download
.gitattributes 1.60 KB aa7aacd0 download
tokenizer_config.json 1.14 KB 1d134cd2 download
generation_config.json 214 B 3f25ead4 download

README current version from Hugging Face


license: apache-2.0
base_model:

  • YuYu1015/YuYu1015-Ornith-1.0-35B-abliterated
    base_model_relation: quantized
    pipeline_tag: text-generation
    tags:
  • nvfp4
  • fp4
  • modelopt
  • vllm
  • sglang
  • qwen3.5
  • moe
  • gated-deltanet
  • reasoning
  • abliterated
  • uncensored
    language:
  • en
  • zh

YuYu1015-Ornith-1.0-35B-abliterated-NVFP4

English | 繁體中文

NVFP4 (NVIDIA FP4) quant of YuYu1015/YuYu1015-Ornith-1.0-35B-abliterated (the BF16 source) · also available as GGUF

Support me on Ko-fi


English

NVFP4 quant of the abliterated (uncensored) Qwen3.5 35B MoE reasoning model, produced with NVIDIA TensorRT Model Optimizer (nvfp4_mlp_only, MSE calibration on reasoning data). For high-throughput low-precision inference on NVIDIA Blackwell GPUs.

Item Value
Format NVFP4 (E2M1 + 2-level scale, group size 16), modelopt
Size ~23.7 GB (3 shards)
Quantized routed experts + shared-expert MLP → NVFP4
Kept BF16 attention, GatedDeltaNet linear-attn, MoE router, shared_expert_gate, lm_head, embeddings
KV-cache not quantized (BF16)

This mixed-precision recipe follows NVIDIA's higher-accuracy guidance for FP4 PTQ — the routed experts (the bulk of the weights) go to FP4 while every precision-sensitive tensor stays in BF16, so reasoning quality is preserved.

Requirements

NVFP4 needs Blackwell (sm_120 / sm_100, e.g. RTX PRO 6000 / B200) and a runtime with modelopt-NVFP4 support (SGLang or vLLM). It will not load with plain transformers (use a serving runtime).

Usage

SGLang:

python -m sglang.launch_server --model-path YuYu1015/YuYu1015-Ornith-1.0-35B-abliterated-NVFP4 --quantization modelopt_fp4 --trust-remote-code

vLLM:

vllm serve YuYu1015/YuYu1015-Ornith-1.0-35B-abliterated-NVFP4 --trust-remote-code

Recommended Sampling Parameters

Reasoning model (emits <think>…</think>). Official Qwen3.5 settings:

temperature 1.0 · top_p 0.95 · top_k 20 · min_p 0.0 · presence_penalty 0.0

Safety Warning

This model has safety filtering removed (abliterated) and may generate sensitive or inappropriate content. Users are solely responsible for all consequences and legal liability, and must ensure usage complies with local laws and ethical standards.

Credits


繁體中文

YuYu1015/YuYu1015-Ornith-1.0-35B-abliterated(BF16 來源)的 NVFP4(NVIDIA FP4)量化版本;另有 GGUF 版本。

以 NVIDIA TensorRT Model Optimizer(nvfp4_mlp_only、推理資料 MSE 校準)量化的 abliterated(去審查)Qwen3.5 35B MoE 推理模型,供 NVIDIA Blackwell GPU 高吞吐低精度推理。

項目 數值
格式 NVFP4(E2M1 + 兩級縮放,group size 16),modelopt
大小 ~23.7 GB(3 shards)
量化 routed experts + shared-expert MLP → NVFP4
保 BF16 attention、GatedDeltaNet linear-attn、MoE router、shared_expert_gate、lm_head、embedding
KV-cache 不量化(BF16)

此混合精度配方遵循 NVIDIA 對 FP4 PTQ 的高精度建議 —— 只把 routed experts(佔大多數參數)壓到 FP4,所有對精度敏感的張量全保 BF16,以保留推理能力。

需求

NVFP4 需 Blackwell(sm_120 / sm_100,如 RTX PRO 6000 / B200) 及支援 modelopt-NVFP4 的 runtime(SGLang 或 vLLM)。無法用純 transformers 載入(請用推理 runtime)。

使用方式

SGLang:

python -m sglang.launch_server --model-path YuYu1015/YuYu1015-Ornith-1.0-35B-abliterated-NVFP4 --quantization modelopt_fp4 --trust-remote-code

vLLM:

vllm serve YuYu1015/YuYu1015-Ornith-1.0-35B-abliterated-NVFP4 --trust-remote-code

建議取樣參數

推理模型(輸出 <think>…</think>)。Qwen3.5 官方設定:

temperature 1.0 · top_p 0.95 · top_k 20 · min_p 0.0 · presence_penalty 0.0

安全警告

此模型已移除安全過濾(abliterated),可能產生敏感或不當內容。使用者須自行承擔所有風險與法律責任,並確保使用方式符合當地法規與倫理標準。

致謝

README history 1 version

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-07-02Update READMEbf09bb25.4 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration