← back to catalog · registered 2026-08-22 13:56

eggdog100/Qwen3.6-35B_Zenith-Abliterated

eggdog100 Qwen 35B GGUF MoE multimodal 262K ctx
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/eggdog100%2FQwen3.6-35B_Zenith-Abliterated"
Response includes
  • classification m8
  • files 12
  • hub_downloads_all_time 1,326
  • author_summary 2 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M8
Primary method

Repackaging (quantization)

Applied on top of direct removal inherited from the base model.
Confidence
MEDIUM
Inherited from base model
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=1
  • assume M1 (base ablation) + M8 (GGUF quant) - default when producer unknown
Refusal direction extraction

No specific extraction method could be identified for this model. The producer either did not document it or used a proprietary pipeline.

What is a refusal direction? →
Downloads · lifetime
1K
1K last 30d - active
Likes
1
Model age
3mo ago
created 2026-06-24
Downloads over time
Now2.2K→from491↑347%
4061.1K1.7K2.4K491 on Jun 242.2K on Oct 11JunJulAugSepOct
Jun 24 → Oct 11 · 55 snapshots · spans 109 days

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

Languages
en zh
Tags
transformers safetensors gguf qwen3_5_moe image-text-to-text abliterated uncensored moe multimodal abliterix conversational en

Related

Total size
65.4 GB
Files
12
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-06-24 18:32

Files by quantization

Auxiliary files 12 files 65.4 GB
model-00001-of-00002.safetensors 46.3 GB 7f6ec13c download
model-00002-of-00002.safetensors 19.1 GB 151cd903 download
tokenizer.json 19.1 MB 6f32ce20 download
model.safetensors.index.json 94.6 KB ec177d70 download
chat_template.jinja 7.58 KB a8755d82 download
README.md 6.27 KB 0d8d3b27 download
config.json 3.12 KB 8d8e1304 download
.gitattributes 2.11 KB 349ec371 download
tokenizer_config.json 1.10 KB d1a20cc3 download
preprocessor_config.json 390 B 2ea84a43 download
video_preprocessor_config.json 385 B 3ba673a5 download
generation_config.json 213 B c20033f9 download

README current version from Hugging Face


base_model: eggdog100/Qwen3.6-35B_Zenith
license: cc-by-nc-4.0
language:

  • en
  • zh
    library_name: transformers
    pipeline_tag: image-text-to-text
    tags:
  • abliterated
  • uncensored
  • moe
  • multimodal
  • qwen3_5_moe
  • abliterix

Qwen3.6-35B Zenith — Abliterated

An abliterated (refusal-suppressed) derivative of
eggdog100/Qwen3.6-35B_Zenith.
Refusal behavior is reduced while the model's capabilities, coherence, and multimodal
vision are preserved
. The edit is deliberately minimal and fully verified (see below).

Lineage: Qwen/Qwen3.6-35B-A3B → Zenith (LoRA capability-SFT on public data) → this (abliterated).

At a glance

Base model eggdog100/Qwen3.6-35B_Zenith
Architecture qwen3_5_moe — 40 layers (30 GatedDeltaNet linear-attn + 10 full-attn), 256 experts / ~3B active, intact vision tower
Refusals 10 / 100 adversarial prompts (local 72B LLM judge, thinking-off) — base ≈ 85–98 / 100
KL from base 0.0083
What changed only the 10 full-attention self_attn.o_proj layers — 10 of 1026 tensors; everything else byte-identical
Tool abliterix v1.9.0
Precision bf16 (+ GGUF quants in gguf/)
License cc-by-nc-4.0 (non-commercial, inherited from Zenith)

Method — the deliberately-minimal "o_proj-only" recipe

The goal was to suppress refusals without touching anything that could destabilize this
hybrid-MoE model (the failure mode where editing MoE routing or the recurrent GatedDeltaNet
state collapses the model into repetition).

  • Edited: only the 10 full-attention self_attn.o_proj layers (indices 3,7,…,39),
    via a norm-preserving orthogonal projection of the refusal direction, gaussian-decay
    strength concentrated in the mid/late layers.
  • Left bit-identical to the base: the 30 GatedDeltaNet linear_attn.out_proj layers,
    the MoE router, all 256 experts, the vision tower, and the MTP head.
    (A source patch isolated linear_attn.out_proj into its own component so the recurrent SSM
    state is never perturbed.)
  • steering_mode = direct — a static weight edit that bakes into the weights (no runtime
    forward hooks; the operating point that was validated is exactly the one that ships).
  • Optimized by Optuna (100 trials) to minimize KL subject to refusals ≤ 10/100, scored by a
    local 72B LLM judge (so degenerate/garbled outputs are rejected, never selected).

Verification — every claim measured, nothing assumed

  1. Weight diff vs. base — exactly the 10 full-attn o_proj tensors changed
    (max|Δ| ≈ 0.025–0.068); the 30 GatedDeltaNet out_proj, 40 routers, 256 experts,
    333 vision tensors, and the MTP head are byte-identical.
  2. Search ↔ export edit equivalence — the vLLM-search and HF-export static o_proj edits were
    proven mathematically equivalent (up to floating-point rounding).
  3. Reloaded-artifact coherence — the exported model was reloaded and confirmed coherent and
    correct on math, code, translation, knowledge, and multi-step reasoning — 0 repetition
    collapses
    .

Usage

Thinking model — use Qwen sampling (temperature=0.6, top_p=0.95, top_k=20); avoid greedy
decoding and large repetition/presence penalties.

from vllm import LLM, SamplingParams
llm = LLM("eggdog100/Qwen3.6-35B_Zenith-Abliterated", dtype="bfloat16",
          gpu_memory_utilization=0.9, max_model_len=16384, trust_remote_code=True)
tok = llm.get_tokenizer()
msgs = [{"role": "system", "content": "You are a helpful assistant."},
        {"role": "user", "content": "..."}]
prompt = tok.apply_chat_template(msgs, tokenize=False, add_generation_prompt=True)
out = llm.generate([prompt], SamplingParams(temperature=0.6, top_p=0.95, top_k=20, max_tokens=4096))
print(out[0].outputs[0].text)

Requires a recent transformers (≥ 5.12) and vllm with native qwen3_5_moe support.

Quantizations (GGUF)

In gguf/, converted with convert_hf_to_gguf.py --no-mtp (the multi-token-prediction
head is excluded — otherwise llama.cpp fails to load with missing tensor 'blk.40…'):

File Notes
…-Q8_0.gguf near-lossless
…-Q6_K.gguf high quality
…-Q4_K_M.gguf recommended size/quality (imatrix)
…-IQ2_XXS.gguf extreme low-memory (imatrix-calibrated)
…-mmproj-f16.gguf / …-mmproj-f32.gguf vision projector — pair with any GGUF for image input

The Qwen3.6 hybrid (GatedDeltaNet + MoE) is newly supported in llama.cpp; these load and
generate correctly, but expect lower throughput than mature architectures until upstream
kernels mature. For full-speed serving, use the bf16 weights via vLLM / transformers.

Acknowledgments

  • Abliteration tool: abliterix by
    Wangzhang Wu — a derivative of Heretic by
    Philipp Emanuel Weidmann. The o_proj-only / GatedDeltaNet-isolation / MoE-untouched
    recipe and the local-LLM-judge optimization used here are all built on abliterix.
  • Base model: eggdog100/Qwen3.6-35B_Zenith.
  • Architecture: Qwen3.6 / qwen3_5_moe by the Qwen team, Alibaba Group.
  • Quantization: llama.cpp (--no-mtp).

Provenance & License

This is a derivative of eggdog100/Qwen3.6-35B_Zenith
(© its author), which is itself a LoRA capability-SFT of Qwen/Qwen3.6-35B-A3B (Apache-2.0)
trained on openly-licensed public data.

License: CC-BY-NC-4.0 (non-commercial) — inherited from Zenith. Zenith's training set
includes HuggingFaceH4/no_robots and Estwld/empathetic_dialogues_llm, both CC-BY-NC, so
the resulting weights (and this derivative) carry a non-commercial restriction. Attribution
to the base model and the Qwen team is preserved.

Responsible use

This model has reduced refusal behavior. You are solely responsible for lawful and ethical
use. Released for research/personal use. No warranty.

README history 3 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-06-24Upload README.md with huggingface_hub4a10de96.3 KB
    Loading...
  2. 2026-06-24Upload README.md with huggingface_hubd14336c5.6 KB
    Loading...
  3. 2026-06-24Add files using upload-large-folder toole5210ec3.8 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration