library_name: mlx
base_model: llmfan46/Ornith-1.0-35B-uncensored-heretic
tags:
- mlx
- mlx-vlm
- moe
- multimodal
- vision
- coding
- agentic
- uncensored
- heretic
- basequant-xl
pipeline_tag: image-text-to-text
leonsarmiento/Ornith-1.0-35B-uncensored-heretic-5bit-XL-mlx
This model was converted to MLX format from llmfan46/Ornith-1.0-35B-uncensored-heretic using BaseQuant_XL 5/8-bit mixed quantization optimized for Apple Silicon. The vision encoder is preserved and quantized at 5-bit, making this a full multimodal model.
BaseQuant_XL keeps the most routing-critical layers in full bf16 precision — the MoE router gate, shared expert gate, shared expert, and lm_head — while applying aggressive quantization to the bulk parameters. This preserves routing accuracy and output quality where it matters most.
This is the uncensored heretic version of Ornith-1.0-35B, a 35B-parameter MoE (Mixture of Experts) model fine-tuned from Qwen3.5-35B-A3B by DeepReinforce AI, using a self-improving RL training framework that jointly optimizes scaffold and solution rollouts for agentic coding tasks. Despite 35B total parameters, only ~3B are activated per token. It features 256 experts (8 active per token + 1 shared expert), hybrid full + linear (Gated DeltaNet) attention, and a vision encoder.
About XL Quantization
BaseQuant_XL is a fully data-agnostic, static quantization. No calibration dataset, no sensitivity analysis, no importance matrix. Precision is allocated purely by architectural role — routing-critical layers get higher precision, bulk expert parameters get lower precision. The result is a transparent, faithful capture of the source model.
Data-dependent calibration quantizations (iMatrix, AWQ, GPTQ, oQ, oQ4e, etc.) use a calibration set to guide bit allocation. This can produce a skewed representation of the model: domains well-represented in the calibration data (English, popular topics, public or leaked benchmarks) are preserved better, while underrepresented domains (non-English languages, niche use cases, your own data) are preserved worse. XL avoids this trade-off entirely — it generalizes honestly because it is never fit to any particular data distribution.
Use with mlx
pip install -U mlx-vlm
python -m mlx_vlm.generate --model leonsarmiento/Ornith-1.0-35B-uncensored-heretic-5bit-XL-mlx --max-tokens 256 --temperature 1.0 --top-p 1.0 --prompt "Hello"
BaseQuant_XL Quantization Strategy
| Bit Depth | Layers | Rationale |
|---|---|---|
| bf16 (unquantized) | mlp.gate (router), shared_expert_gate, lm_head, shared_expert |
Routing decisions and shared computation path — errors here are qualitatively different from precision loss |
| 8-bit | embed_tokens, self_attn (full attention), linear_attn (DeltaNet) |
Every-token layers with moderate sensitivity — 8-bit is near-lossless |
| 5-bit | vision_tower, switch_mlp (routed experts) |
Bulk of parameters, only 8 of 256 experts active per token — natural redundancy tolerates lower precision |
Quantization Details
| Layer | Bits | Group Size |
|---|---|---|
mlp.gate (router) |
bf16 | — |
shared_expert_gate |
bf16 | — |
lm_head |
bf16 | — |
shared_expert |
bf16 | — |
embed_tokens |
8 | 64 |
self_attn (full attention) |
8 | 64 |
linear_attn (DeltaNet) |
8 | 64 |
vision_tower |
5 | 64 |
switch_mlp (routed experts) |
5 | 64 |
| Default fallback | 8 | 64 |
- Quantization type: BaseQuant_XL mixed (multimodal, vision preserved)
- Bits per weight: 5.881
- Group size: 64
- Method: Custom
quant_predicateviamlx_vlm
Recommended Inference Parameters
| Parameter | Value |
|---|---|
temperature |
1.0 |
top_p |
1.0 |
top_k |
40 |
min_p |
0.01 |
repeat_penalty |
1.0 |
Note: Ornith-1.0-35B uses Temp 1.0 and Top_p 1.0 per the model's Terminal-Bench 2.1 benchmark recipe. This is a Qwen3.5-based model —
preserve_thinkingis not applicable.