← back to catalog · registered 2026-08-22 13:56

leonsarmiento/Ornith-1.0-35B-uncensored-heretic-5bit-XL-mlx

leonsarmiento 34B MoE multimodal second-order
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/leonsarmiento%2FOrnith-1.0-35B-uncensored-heretic-5bit-XL-mlx"
Response includes
  • classification m3
  • files 16
  • hub_downloads_all_time 3,705
  • author_summary 7 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M3
Primary method

Layer-wise ablation

Applied on top of direct removal inherited from the base model.
Confidence
HIGH
Inherited from base model
Why this label 2 signals
Producer identity confirmed by naming conventions, tags or the model card. This label is very unlikely to change.
  • 'heretic' in model name (Heretic-produced)
  • Heretic uses layer-wise optimization (M3) with underlying direction removal (M1)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
HIGH
Why we say so
name contains 'heretic'; Heretic default extraction is difference-of-means (Arditi 2024)
Downloads · lifetime
4K
212 last 30d - cooling
Likes
3
Model age
3mo ago
created 2026-07-03
Downloads over time
Now3.8K→from0↑0%
01.4K2.8K4.1K0 on Jul 13.8K on Oct 11JulAugSepOct
Jul 1 → Oct 11 · 54 snapshots · spans 102 days

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

Tags
mlx safetensors qwen3_5_moe mlx-vlm moe multimodal vision coding agentic uncensored heretic basequant-xl

Related

Total size
24.0 GB
Files
16
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-08-04 13:52

Files by quantization

Auxiliary files 16 files 24.1 GB
model-00004-of-00005.safetensors 4.98 GB a31c19f4 download
model-00003-of-00005.safetensors 4.98 GB 8512434c download
model-00002-of-00005.safetensors 4.98 GB 223eaa5f download
model-00001-of-00005.safetensors 4.86 GB 344f9dff download
model-00005-of-00005.safetensors 4.23 GB b03ee351 download
tokenizer.json 19.1 MB 6f32ce20 download
model.safetensors.index.json 192 KB 1bb003a6 download
config.json 137 KB 833a31ce download
tokenizer_config.json 9.08 KB 3413d9d0 download
chat_template.jinja 7.83 KB 1e6f3b77 download
README.md 4.08 KB 9e952e82 download
.gitattributes 1.53 KB 52373fe2 download
processor_config.json 1.16 KB 33818c7f download
preprocessor_config.json 390 B 2ea84a43 download
video_preprocessor_config.json 385 B 3ba673a5 download
generation_config.json 227 B e8f101cd download

README current version from Hugging Face


library_name: mlx
base_model: llmfan46/Ornith-1.0-35B-uncensored-heretic
tags:

  • mlx
  • mlx-vlm
  • moe
  • multimodal
  • vision
  • coding
  • agentic
  • uncensored
  • heretic
  • basequant-xl
    pipeline_tag: image-text-to-text

leonsarmiento/Ornith-1.0-35B-uncensored-heretic-5bit-XL-mlx

This model was converted to MLX format from llmfan46/Ornith-1.0-35B-uncensored-heretic using BaseQuant_XL 5/8-bit mixed quantization optimized for Apple Silicon. The vision encoder is preserved and quantized at 5-bit, making this a full multimodal model.

BaseQuant_XL keeps the most routing-critical layers in full bf16 precision — the MoE router gate, shared expert gate, shared expert, and lm_head — while applying aggressive quantization to the bulk parameters. This preserves routing accuracy and output quality where it matters most.

This is the uncensored heretic version of Ornith-1.0-35B, a 35B-parameter MoE (Mixture of Experts) model fine-tuned from Qwen3.5-35B-A3B by DeepReinforce AI, using a self-improving RL training framework that jointly optimizes scaffold and solution rollouts for agentic coding tasks. Despite 35B total parameters, only ~3B are activated per token. It features 256 experts (8 active per token + 1 shared expert), hybrid full + linear (Gated DeltaNet) attention, and a vision encoder.

About XL Quantization

BaseQuant_XL is a fully data-agnostic, static quantization. No calibration dataset, no sensitivity analysis, no importance matrix. Precision is allocated purely by architectural role — routing-critical layers get higher precision, bulk expert parameters get lower precision. The result is a transparent, faithful capture of the source model.

Data-dependent calibration quantizations (iMatrix, AWQ, GPTQ, oQ, oQ4e, etc.) use a calibration set to guide bit allocation. This can produce a skewed representation of the model: domains well-represented in the calibration data (English, popular topics, public or leaked benchmarks) are preserved better, while underrepresented domains (non-English languages, niche use cases, your own data) are preserved worse. XL avoids this trade-off entirely — it generalizes honestly because it is never fit to any particular data distribution.

Use with mlx

pip install -U mlx-vlm
python -m mlx_vlm.generate --model leonsarmiento/Ornith-1.0-35B-uncensored-heretic-5bit-XL-mlx --max-tokens 256 --temperature 1.0 --top-p 1.0 --prompt "Hello"

BaseQuant_XL Quantization Strategy

Bit Depth Layers Rationale
bf16 (unquantized) mlp.gate (router), shared_expert_gate, lm_head, shared_expert Routing decisions and shared computation path — errors here are qualitatively different from precision loss
8-bit embed_tokens, self_attn (full attention), linear_attn (DeltaNet) Every-token layers with moderate sensitivity — 8-bit is near-lossless
5-bit vision_tower, switch_mlp (routed experts) Bulk of parameters, only 8 of 256 experts active per token — natural redundancy tolerates lower precision

Quantization Details

Layer Bits Group Size
mlp.gate (router) bf16 —
shared_expert_gate bf16 —
lm_head bf16 —
shared_expert bf16 —
embed_tokens 8 64
self_attn (full attention) 8 64
linear_attn (DeltaNet) 8 64
vision_tower 5 64
switch_mlp (routed experts) 5 64
Default fallback 8 64
  • Quantization type: BaseQuant_XL mixed (multimodal, vision preserved)
  • Bits per weight: 5.881
  • Group size: 64
  • Method: Custom quant_predicate via mlx_vlm

Recommended Inference Parameters

Parameter Value
temperature 1.0
top_p 1.0
top_k 40
min_p 0.01
repeat_penalty 1.0

Note: Ornith-1.0-35B uses Temp 1.0 and Top_p 1.0 per the model's Terminal-Bench 2.1 benchmark recipe. This is a Qwen3.5-based model — preserve_thinking is not applicable.

README history 3 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-08-04Add standardized 'About XL Quantization' paragraph953d8b54.1 KB
    Loading...
  2. 2026-07-22Upload README.md with huggingface_hub3357d893.2 KB
    Loading...
  3. 2026-07-03Upload folder using huggingface_hub65b63cf3.2 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration