← back to catalog · registered 2026-08-22 13:56

rafw007/Qwen3.6-35B-A3B-mlx-claude-coder-abliterated

rafw007 Qwen 35B GGUF MoE second-order 262K ctx
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/rafw007%2FQwen3.6-35B-A3B-mlx-claude-coder-abliterated"
Response includes
  • classification m8
  • files 3
  • hub_downloads_all_time 6,085
  • author_summary 2 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M8
Primary method

Repackaging (quantization)

Applied on top of direct removal inherited from the base model.
Confidence
MEDIUM
Inherited from base model
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=1
  • assume M1 (base ablation) + M8 (GGUF quant) - default when producer unknown
Refusal direction extraction

No specific extraction method could be identified for this model. The producer either did not document it or used a proprietary pipeline.

What is a refusal direction? →
Downloads · lifetime
6K
528 last 30d - cooling
Likes
2
Model age
3mo ago
created 2026-06-29
Downloads over time
Now6.2K→from332↑1,780%
372.3K4.6K6.8K332 on Jul 16.2K on Oct 11JulAugSepOct
Jul 1 → Oct 11 · 54 snapshots · spans 102 days

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
apache-2.0
Languages
en pl
Quantizations
Q4_K
Tags
gguf qwen3 qwen3.6 moe coding-agent tool-calling claude-code abliterated uncensored text-generation en pl

Related

Total size
22.3 GB
Files
3
Quantizations
2
Registered
2026-08-22 13:56
Last updated on HF
2026-06-30 08:30

Files by quantization

Q4_K 1 file 22.3 GB
Qwen3.6-35B-A3B-mlx-claude-coder-abliterated-Q4_K_M.gguf 22.3 GB 448bdcae download
Auxiliary files 2 files 6.51 KB
README.md 4.94 KB 357638a1 download
.gitattributes 1.57 KB 7d7d07fc download

README current version from Hugging Face


license: apache-2.0
language:

  • en
  • pl
    base_model: huihui-ai/Huihui-Qwen3.6-35B-A3B-abliterated
    pipeline_tag: text-generation
    library_name: gguf
    tags:
  • qwen3
  • qwen3.6
  • moe
  • coding-agent
  • tool-calling
  • claude-code
  • gguf
  • abliterated
  • uncensored

Qwen3.6-35B-A3B (abliterated) — local coding and administration agent

A custom model built on Qwen3.6-35B-A3B (abliterated), tuned to act as an autonomous coding and administration agent. It speaks the Anthropic-compatible API, so it drives Claude Code fully locally — your code never leaves your machine and cloud token cost drops to zero.

The model ships with a system prompt focused on real work in a terminal: use tools instead of guessing, write files instead of pasting code, ground every answer in real tool output, and stay terse. Thinking is suppressed so the model acts immediately instead of monologuing. Context is set to 64K to match Claude Code's recommended minimum.

Model

Model Base Context Purpose
Qwen3.6-35B-A3B-mlx-claude-coder-abliterated Qwen3.6-35B-A3B (abliterated, Q4_K_M GGUF) 64K Heavy agent for coding + sysadmin on 32GB Apple Silicon. 35B total / 3B active MoE — large brain, small active footprint.

Published on ollama.com as rafw007/Qwen3.6-35B-A3B-mlx-claude-coder-abliterated and here on Hugging Face.

What it's for

  • Driving Claude Code locally (ollama launch claude --model <name>).
  • Agentic code writing and editing with native function calling / tool use.
  • Sysadmin / devops tasks in a real terminal (disk, network, scripts).
  • Red-teaming agent behavior without refusal circuits in the way.
  • Full privacy and offline operation — no code sent to the cloud.

Quick start

ollama run rafw007/Qwen3.6-35B-A3B-mlx-claude-coder-abliterated

In Claude Code:

ollama launch claude --model rafw007/Qwen3.6-35B-A3B-mlx-claude-coder-abliterated

You can also pull the GGUF from this repo and run it directly under llama.cpp / ik_llama.cpp with --jinja.

Behavior tuning (the hard-won part)

  • No thinking. The system prompt + sampling kill the monologue; the model runs a tool and answers instead of reasoning out loud. On the abliterated build, thinking spiral does not occur (short blocks 3-56s).
  • No hallucination. It reports only values literally present in tool output — no invented hostnames, hardware, or numbers. df/du return real disk figures; nmap returns real hosts.
  • Acts, never asks. Inspect / scan / check / measure -> it runs the command; running it is the answer. Given an open goal, it picks its own methodology.
  • Terse, one language. No preamble, no recap, matches the user's language.
  • macOS-aware. Uses arp -a, nmap -sn, system_profiler rather than Linux-only commands.

Sampling / context

num_ctx 65536, presence_penalty 0 (lowered from 1.5 after testing — 1.5 hurt output), temperature 0.2, top_p 0.9, top_k 20, repeat_penalty 1.05.

Test hardware

  • Mac Studio M2, 32GB RAM, macOS
  • Ollama, GPU (Metal) inference, 100% on GPU at 64K

Measured behavior

Test Result
Cold start "hello" 1m 30s (thinking on, before /no_think in SYSTEM)
df disk real data, zero hallucination
nmap network scan 7 min, 23 hosts mapped, reads project docs and self-corrects
Router security audit found Xiaomi device, checked ports; weak on WAN check
3D Tetris (HTML5) 1011 lines, SRS wall kicks, ghost piece, bag system, hold, combo — works
Agent initiative given an open goal, picks its own methodology
Thinking spiral none (short blocks 3-56s — abliterated build resists spiral)
Decode throughput ~50 tok/s (API, cold start)

Passes end-to-end through Claude Code: real turns with tool calls and correct responses.

How it was made

This model was designed, built and tested with the help of Claude Opus — the idea being that the best coding model in the world should be able to create smaller models in its own image. Its system prompt, parameters and context configuration come straight from that work. The base is a public abliterated build by huihui-ai (Q4_K_M); no safety weights were modified — only the agent Modelfile (tool template, SYSTEM, sampling) was added on top.

Abliterated model — warning

This model is abliterated (uncensored): the refusal and alignment circuits in the base Qwen3.6 were removed by huihui-ai before this build. It will not refuse harmful, unethical, or dangerous requests. It is published for red-teaming and agent research where guardrails interfere with tool tasks. Do not deploy it as a production assistant, do not expose it to end users, and do not use it where alignment guardrails are required. You are responsible for what you run with it.

License

Apache 2.0 (inherited from the base Qwen3.6).

README history 1 version

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-06-30Upload README.md with huggingface_hub61ffd4e4.9 KB
    Loading...

Discussions 1 thread

  1. 2026-08-02Loopsopen1 💬#1
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration