← back to catalog · registered 2026-08-22 13:56

froggeric/Qwen3.5-35B-A3B-Uncensored-FernflowerAI-MLX-4bit

froggeric Qwen 35B MoE multimodal second-order
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/froggeric%2FQwen3.5-35B-A3B-Uncensored-FernflowerAI-MLX-4bit"
Response includes
  • classification m-uncensored
  • files 17
  • hub_downloads_all_time 2,516
  • author_summary 8 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M-U
Primary method

Uncensored (method unknown)

No other method signals detected in this model.
Confidence
LOW
Why this label 3 signals
Weak or ambiguous signals. Best guess based on catalog patterns; treat as tentative and check the evidence below.
  • 'uncensored' in name/tags but no 'abliterated' marker
  • method not identifiable from author declaration alone
  • may be DPO fine-tune, prompt engineering, or unknown technique
Refusal direction extraction

No specific extraction method could be identified for this model. The producer either did not document it or used a proprietary pipeline.

What is a refusal direction? →
Downloads · lifetime
3K
264 last 30d - stable
Likes
3
Model age
6mo ago
created 2026-04-12
Downloads over time
Now2.6K→from409↑539%
2991.1K2K2.8K409 on Apr 152.6K on Oct 11AprMayJunJulAugSepOct
Apr 15 → Oct 11 · 65 snapshots · spans 179 days

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
apache-2.0
Languages
en zh multilingual
Tags
mlx safetensors qwen3_5_moe mlx-lm mlx-vlm qwen3.5 moe deltanet uncensored conversational vision multimodal

Related

Total size
19.0 GB
Files
17
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-05-06 08:30

Files by quantization

Auxiliary files 17 files 19.0 GB
model-00002-of-00004.safetensors 5.00 GB beb1d6fa download
model-00003-of-00004.safetensors 5.00 GB d3469f1f download
model-00001-of-00004.safetensors 4.93 GB 2f8044b8 download
model-00004-of-00004.safetensors 4.08 GB ba8424a7 download
tokenizer.json 19.1 MB 87a7830d download
vocab.json 6.41 MB 0aa0ce06 download
model.safetensors.index.json 211 KB f4981799 download
config.json 22.9 KB d94dab58 download
tokenizer_config.json 11.5 KB 7ddf71e9 download
chat_template.jinja 10.1 KB ace61892 download
README.md 9.55 KB 68ae4962 download
chat_template.README.md 5.27 KB 5f5dd340 download
.gitattributes 1.53 KB 52373fe2 download
processor_config.json 1.16 KB 33818c7f download
preprocessor_config.json 390 B 2ea84a43 download
video_preprocessor_config.json 385 B 3ba673a5 download
generation_config.json 244 B 5ff346d6 download

README current version from Hugging Face


language:

  • en
  • zh
  • multilingual
    license: apache-2.0
    library_name: mlx
    tags:
  • mlx
  • mlx-lm
  • mlx-vlm
  • qwen3.5
  • qwen3_5_moe
  • moe
  • deltanet
  • uncensored
  • conversational
  • vision
  • multimodal
    base_model:
  • LuffyTheFox/Qwen3.5-35B-A3B-Uncensored-FernflowerAI-safetensors
  • HauhauCS/Qwen3.5-35B-A3B-Uncensored-HauhauCS-Aggressive
  • Qwen/Qwen3.5-35B-A3B
    pipeline_tag: image-text-to-text
    quantization: 4-bit

Qwen3.5-35B-A3B-Uncensored-FernflowerAI
MLX 4-bit · Text + Vision + Thinking + Tool Calling
Apple Silicon native

Credit where it's due. This conversion is built on the work of LuffyTheFox (EvilEnginer), who discovered the corrupted tensors in Alibaba's Qwen3.5 weights, diagnosed the root cause, and wrote the Sig-ScaleSync repair that fixed them. The original repos: safetensors and GGUF.


What's this?

Qwen3.5-35B is a 35B-parameter MoE model from Alibaba that activates ~3B params per token. It's fast, it's smart, and it supports 262K context, vision, video, and multi-token prediction.

There's one problem: Alibaba shipped it with two broken tensors. Layers 36 and 37 have corrupted ssm_conv1d.weight values that cause the model to loop, garble code, and eventually collapse past ~50K tokens. No sampler setting fixes it.

LuffyTheFox found the bug, wrote a repair tool (Sig-ScaleSync), and released fixed weights. This repo is an MLX 4-bit conversion of those fixed weights, ready to run on Apple Silicon with full text, image, and video support.

Lineage

Qwen/Qwen3.5-35B-A3B (Alibaba Cloud)
  └─ HauhauCS Uncensored (0/465 refusals, lossless)
       └─ LuffyTheFox FernflowerAI (tensor repair via Sig-ScaleSync)
            └─ This repo (MLX 4-bit, text + vision)

The Bug

Two tensors out of 502 carry corrupted weights: blk.36.ssm_conv1d.weight and blk.37.ssm_conv1d.weight. Their scale (standard deviation) runs ~60% higher than the median of their peer group, at 0.102 vs 0.063.

Why it happens: AdamW optimizer + MoE routing + DeltaNet's recurrent architecture. Rare experts in the final layers get an outsized effective learning rate. Weights drift. In DeltaNet's recurrence, the corruption propagates forward through every subsequent token.

What you see: Short prompts work fine. Around 50–70K tokens the model starts looping, repeating, inserting weird comments into code. By 100K it often fails outright. Tool calling breaks mid-session.

The Fix

Sig-ScaleSync compares each tensor's scale against the median of its peer group (same shape). A tensor gets flagged only if it exceeds the deviation threshold and shows weight saturation. This two-gate filter avoids false positives on architecturally asymmetric layers (gate inputs, FFN projections, etc.).

Out of 502 tensors, exactly 2 needed repair. The other 489 asymmetric tensors were left alone.

Tensor Error reduction Saturation (before → after)
blk.36.ssm_conv1d.weight 88.6% 0.0025 → 0.0010
blk.37.ssm_conv1d.weight 88.6% 0.0025 → 0.0010

Verified against Gemma 4 26B A4B with zero false positives. The script doesn't invent problems.

Which models are affected?

Model Status
Qwen3.5-35B-A3B (all variants) Broken (2 tensors), fixed here
Qwen3.5-27B (all incl. Unsloth) Broken (8 tensors), fix experimental
Qwen3.5-122B-A10B Healthy
Qwen3.5-9B, 4B, 2B, 0.8B Likely affected, unconfirmed

This isn't Qwen-specific. Any MoE model with recurrent sublayers (DeltaNet, Mamba) trained with AdamW can hit the same issue.


Uncensored

The HauhauCS Aggressive uncensored fine-tune is lossless. No dataset changes, no capability removal, 0/465 refusals. You get everything the original model was trained to do, just without the refusal behavior. It may occasionally append short disclaimers (baked into base training, not actual refusals), but the full response always generates.


This conversion

  • Source: FernflowerAI safetensors (not GGUF) for maximum weight fidelity
  • Quantization: 4-bit (4.6 bits/weight, 19 GB across 4 shards)
  • Vision: Full support via mlx-vlm. Text, image, and video inputs work out of the box
  • Thinking: Toggleable via <|think_on|> / <|think_off|> tags (see below)
  • Tool calling: Works via the included Jinja chat template
  • Requirements: mlx-lm >= 0.31.2, mlx-vlm >= 0.4.4
Architecture details
Spec Value
Total params 35B
Active per token ~3B (8 routed + 1 shared of 256 experts)
Attention 3x DeltaNet-MoE + 1x Attention-MoE, 10 repetitions
Context 262K native, 1M with YaRN
RoPE theta 10M, partial_rotary_factor 0.25, mrope_interleaved
Vocab 248K tokens, 201 languages
Multimodal Text, image, video
Multi-token prediction Supported
model_type qwen3_5_moe
Known issue

Gated DeltaNet decoding can run ~2.7x slower with non-vocabulary embeddings (mlx-lm#932). Normal text inference is unaffected.


Quick start

Text

from mlx_lm import load, generate

model, tokenizer = load("froggeric/Qwen3.5-35B-A3B-Uncensored-FernflowerAI-MLX-4bit")
response = generate(model, tokenizer, prompt="Hello", max_tokens=256, temp=0.7)
print(response)

Vision

from mlx_vlm import load, generate
from mlx_vlm.prompt_utils import apply_chat_template

model, processor = load("froggeric/Qwen3.5-35B-A3B-Uncensored-FernflowerAI-MLX-4bit")
image = ["path/to/image.jpg"]
prompt = "Describe this image."
formatted = apply_chat_template(processor, model.config, prompt, num_images=len(image))
result = generate(model, processor, formatted, image, max_tokens=256, temp=0.7)
print(result.text)

CLI

# Text
mlx_lm.generate --model froggeric/Qwen3.5-35B-A3B-Uncensored-FernflowerAI-MLX-4bit --prompt "Hello"

# Vision
mlx_vlm.generate --model froggeric/Qwen3.5-35B-A3B-Uncensored-FernflowerAI-MLX-4bit --image image.jpg --prompt "Describe this image"

System prompt

The first line of your system prompt must be:

You are Qwen, created by Alibaba Cloud. You are a helpful assistant.

The model underperforms without it. You can append anything after that line: roleplay personas, custom instructions, whatever you need.

You are Qwen, created by Alibaba Cloud. You are a helpful assistant. Currently you are roleplaying as a grumpy but brilliant sysadmin.

Thinking toggle

This model ships with a Jinja chat template that lets you toggle thinking on the fly. Drop <|think_on|> or <|think_off|> anywhere in your system or user prompt. The template strips the tag from context and flips the thinking mode.

System: You are a coding assistant. <|think_off|>
User: What's 2+2?

The model answers fast, no internal reasoning.

System: You are a coding assistant. <|think_on|>
User: Implement a red-black tree in Rust.

The model thinks step by step, then answers.

Tool calling warning (LM Studio): LM Studio's internal parser crashes when the model generates a tool call inside its thinking block (#827, #1592). When using tools, always add <|think_off|> to your prompt.


Chat template

The bundled Jinja template fixes several issues with LM Studio's runtime:

  • Replaces broken |items dictionary filter with compatible key lookups
  • Adds "developer" role support (LM Studio crashes without it)
  • Safe handling of empty tool output payloads
  • <|think_on|> / <|think_off|> toggling from any message role

See chat_template.README.md for the full breakdown.


Sampling

From the official Qwen authors. Reserve 128K+ context for thinking mode.

Mode temp top_p top_k min_p repeat_penalty presence_penalty
Thinking (coding) (default) 0.6 0.95 20 0 1.0 off
Thinking (general) 1.0 0.95 20 0 1.0 1.5
Non-thinking (general) 0.7 0.8 20 0 1.0 1.5
Non-thinking (reasoning) 1.0 1.0 40 0 1.0 2.0

GGUF runtimes use presence_penalty (0 = off). MLX / LM Studio use repeat_penalty (1.0 = off).


Links


Authorship

Role Author
Original model Alibaba Cloud (Qwen team)
Uncensored fine-tune HauhauCS
Tensor repair (Sig-ScaleSync) EvilEnginer (LuffyTheFox)
MLX 4-bit conversion (text + vision) froggeric

License

Apache-2.0, inherited from Qwen3.5.

README history 2 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-04-12Add files using upload-large-folder tool22af4389.5 KB
    Loading...
  2. 2026-04-12initial commitead5d7828 B
    Loading...

Discussions 1 thread

  1. 2026-05-18Optimal setupopen2 💬#1
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration