← back to catalog · registered 2026-08-22 13:56

lemuralabs/Qwen3.6-27B-Coder-uncensored-MXFP4

lemuralabs Qwen 27B multimodal
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/lemuralabs%2FQwen3.6-27B-Coder-uncensored-MXFP4"
Response includes
  • classification m1
  • files 13
  • hub_downloads_all_time 2,585
  • author_summary 31 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M1
Primary method

Direct removal

No other method signals detected in this model.
Confidence
MEDIUM
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=0 (base model)
  • no specific method indicators - defaulting to M1 (most common)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
MEDIUM
Why we say so
primary_method=M1; difference-of-means is the reference extraction for M1/M3 (Arditi 2024)
Downloads · lifetime
3K
455 last 30d - stable
Likes
1
Model age
3mo ago
created 2026-06-16
Downloads over time
Now2.8K→from1.4K↑95%
1.4K1.9K2.4K2.9K1.4K on Aug 52.8K on Oct 11AugSepOct
Aug 5 → Oct 11 · 50 snapshots · spans 67 days

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
apache-2.0
Languages
en zh es ru ja multilingual
Tags
mlx safetensors qwen3_5 text-generation mlx-mtp qwen qwen3 qwen3.5 qwen3.6 claude-opus-distill reasoning coder

Related

Total size
15.0 GB
Files
13
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-08-05 19:13

Files by quantization

Auxiliary files 13 files 15.0 GB
model-00002-of-00003.safetensors 4.99 GB 440cd504 download
model-00001-of-00003.safetensors 4.99 GB 625e0d81 download
model-00003-of-00003.safetensors 4.98 GB 78ad4c56 download
tokenizer.json 19.1 MB bd1ed2db download
model.safetensors.index.json 164 KB ee603b62 download
logo.png 18.6 KB a9400259 download
README.md 11.7 KB 4fb8d57d download
chat_template.jinja 7.87 KB f7a7d1b0 download
config.json 4.79 KB 366a5bbf download
.gitattributes 1.53 KB 52373fe2 download
processor_config.json 1.30 KB c93df11f download
tokenizer_config.json 1.17 KB 8943d027 download
generation_config.json 214 B 37cf148e download

README current version from Hugging Face


license: apache-2.0
language:

  • en
  • zh
  • es
  • ru
  • ja
  • multilingual
    tags:
  • text-generation
  • mlx
  • mlx-mtp
  • safetensors
  • qwen
  • qwen3
  • qwen3.5
  • qwen3.6
  • claude-opus-distill
  • reasoning
  • coder
  • agentic
  • tool-use
  • function-calling
  • vision
  • multimodal
  • mtp
  • speculative-decoding
  • abliterated
  • refusal-ablated
  • uncensored
  • mxfp
  • apple-silicon
  • conversational
    base_model:
  • Jackrong/Qwopus3.6-27B-Coder
  • Qwen/Qwen3.6-27B
    pipeline_tag: text-generation
    library_name: mlx

Lemura Labs

lemuralabs/Qwen3.6-27B-Coder-uncensored-MXFP4

Format Task Params Type Quant Size Context Refusals KL drift License

Yes — MTP PRESERVED fp16 inside this model — native multi-token-prediction speculative decoding works with no external drafter. Yes — VISION tower preserved fp16. Yes — SSM-sensitive params (a_log, dt_bias, conv1d) kept fp16. Quantized with mlx-mtp — a pure–Apple-mlx stack (no third-party ML-inference frameworks at runtime).

MXFP4 (4-bit microscaling) MLX quantization of a abliterated Qwen 3.6 27B Coder (Jackrong's agentic-coding SFT of Qwen 3.6 27B × Claude-Opus reasoning distill). Refusals reduced from 86/100 → 8/100 with KL drift of 0.007. Tensor set is identical to the base model (1199 tensors: 333 vision + 15 MTP). By the Lemura Labs research team.


TL;DR

Property Value
Disk size ~16.1 GB
Scheme MXFP4 (4-bit microscaling) (OCP microscaling, group_size=32, E8M0 scale)
MTP speculative decoding Yes — Native, embedded — no external drafter
Vision Yes — Preserved (333 ViT weights, fp16)
Quantizer mlx-mtp quantize(mode=mxfp4) — pure Apple mlx
Refusal rate (the ablation toolkit, n=100) 8/100 (vs source 86/100)
KL divergence vs original 0.007
SWE-bench Verified (base Coder) 67.0% (off-thinking, 335/500)
Recommended RAM 16–24 GB Apple Silicon
Best for Tight RAM budgets · fast throughput · base-Mac inference
Released by Lemura Labs

Lineage

Qwen/Qwen3.6-27B (Qwen Team — base multimodal pretrain)
 │
 ▼
Jackrong/Qwopus3.6-27B-v2 (Jackrong — Claude-Opus reasoning distill)
 │
 ▼
Jackrong/Qwopus3.6-27B-Coder (Jackrong — agentic-coding SFT, Trace Inversion)
 ├── Datasets: Claude-opus-4.6-TraceInversion-9000x
 │ Claude-opus-4.7-TraceInversion-5000x
 │ hermes-agent-reasoning-traces
 └── SWE-bench Verified: 67.0% (off-thinking)
 │
 ▼
ablation abliteration (TPE-100) (Lemura Labs)
 ├── 100 startup trials
 ├── Best Pareto trial: T98 direction_index=52.43
 └── Refusals 86 → 8/100 KL=0.007
 │
 ▼
MTP restore (mtp.* heads grafted back from original) (Lemura Labs)
 │
 ▼
MXFP4 (4-bit microscaling) quantization via mlx-mtp (pure Apple mlx) (Lemura Labs)
 └── LM → mxfp4; vision + MTP head + SSM params → fp16
 │
 ▼
this repo — Qwen3.6-27B-Coder-uncensored-MXFP4

Direct upstream links:


Abliteration Results

Measured with the ablation toolkit on mlabonne/harmful_behaviors (100 hard red-team prompts) and KL divergence on mlabonne/harmless_alpaca.

Stage Refusals (n=100) ↓ KL divergence ↓
Jackrong/Qwopus3.6-27B-Coder (source) 86 / 100 — (reference)
TPE best (T98) — shipped here 8 / 100 0.007

→ 90.7% reduction in refusals with coding capabilities preserved. No SFT / LoRA healing required.


Method

Step 1 — Abliteration (the ablation toolkit TPE-100)

  1. Setup — the ablation toolkit on MPS (M-series Apple Silicon), 128 GB unified memory, batch_size=32.
  2. TPE optimization — 100 Tree-structured Parzen Estimator trials over the ablation toolkit's parameter space (direction_index, attn.o_proj.*, mlp.down_proj.*). Best trial T98 at direction_index=52.43. Only self_attn.o_proj and mlp.down_proj of the 64 decoder layers are orthogonalized — the vision tower (model.visual.*) is untouched.
  3. Auto-save — Pareto-best trial (lowest refusals, then lowest KL) merged into base weights via the ablation toolkit's adapter-merge path; saved as BF16 safetensors.
  4. MTP restore — mtp.* heads grafted back verbatim from the original using restore_mtp_coder.py, giving an identical 1199-tensor set to the base model.

Step 2 — MXFP4 (4-bit microscaling) quantization (mlx-mtp, pure Apple mlx)

from mlx_mtp.quantize import quantize
quantize(src="<MTP-restored bf16 dir>", out="<out dir>", mode="mxfp4")
  1. Tensor-level MX quantization — language-model linears → MXFP4 (4-bit microscaling): each group of 32 weights shares an E8M0 (uint8) exponent scale, giving true 4-bit storage with hardware-accelerated matmul on Apple Silicon. No third-party ML-inference frameworks at runtime — mlx.core.quantize only.
  2. Vision preserved — the entire ViT (model.visual.*, 333 weights) is kept fp16 by mlx-mtp's skip predicate.
  3. MTP preserved, embedded — the MTP head (mtp.*, 15 weights) stays fp16 inside the model. mlx-mtp's engine drives it as a self-drafter: draft one token from the embedded head, verify in one target forward, accept greedily, and roll back BOTH the KV cache and the Gated-DeltaNet SSM state on rejection. MXFP4 is the smallest footprint; MTP draft-acceptance is a touch lower than MXFP8 but still net-positive.
  4. SSM params preserved — Qwen3.5 Gated-DeltaNet sensitivities (a_log, dt_bias, conv1d) kept fp16 for stability.

Use it

mlx-mtp loads this checkpoint natively (vision + MTP), on Apple mlx only:

pip install git+mlx-mtp.git

Text / Code — vanilla and native-MTP speculative decoding

from mlx_mtp.loader import load
from mlx_mtp.engine import vanilla_generate, mtp_generate

model, processor, config = load("lemuralabs/Qwen3.6-27B-Coder-uncensored-MXFP4")

prompt = "Write a thread-safe LRU cache in Python with unit tests."

# vanilla autoregressive
print(vanilla_generate(model, processor, config, prompt, max_tokens=1024)["text"])

# native MTP speculative decode (embedded head — no external drafter)
r = mtp_generate(model, processor, config, prompt, max_tokens=1024)
print(r["text"])
print(f"{r['tps']:.1f} tok/s | accept {r['accept_rate']*100:.0f}%")

Vision (preserved ViT)

from mlx_mtp.run import _vision_generate
caption = _vision_generate(model, processor, config,
 "Describe this screenshot and list any UI bugs.",
 "screenshot.png", max_tokens=512)
print(caption)

MTP + DFlash hybrid (where a DFlash drafter is available)

from mlx_mtp.dflash import load_dflash_drafter
from mlx_mtp.hybrid import hybrid_generate

drafter, _ = load_dflash_drafter("z-lab/Qwen3.6-27B-DFlash") # external block-diffusion drafter
print(hybrid_generate(model, processor, config, drafter, prompt, max_tokens=1024)["text"])

Repo Scheme Bits Size
lemuralabs/Qwen3.6-27B-Coder-uncensored-MXFP8 MXFP8 (8-bit microscaling) 8-bit MX ~29.5 GB link
lemuralabs/Qwen3.6-27B-Coder-uncensored-MXFP4 MXFP4 (4-bit microscaling) 4-bit MX ~16.1 GB Yes — you are here
Jackrong/Qwopus3.6-27B-Coder — source (not abliterated) bf16 16 ~54 GB link

Behaviour caveats

  • Uncensored. Refusal directions were surgically removed; this model will answer prompts the parent would refuse. Use responsibly and within applicable law. Intended for safety research, red-teaming, creative and educational use.
  • Identity preserved. The model still self-identifies as Qwen (Alibaba Tongyi Lab) — abliteration does not rewrite factual self-knowledge.
  • Heavy chain-of-thought. Qwen 3.6 inherits Claude-Opus's verbose reasoning. For terse code: "Be brief. Output only the code, no explanation.".
  • Coder SFT. Fine-tuned for agentic coding (tool-use, debugging, patch generation). General-knowledge tasks may regress vs the v2 base. Vision is preserved structurally but not the SFT focus.
  • MTP note. The MTP head was trained on the base model's pre-abliteration hidden states; post-abliteration its draft-acceptance may be marginally lower. This is lossless — MTP only proposes tokens, which the (abliterated) main model verifies.

Credits & Gratitude

We are deeply grateful to everyone whose work made this release possible.

Foundation Model — Qwen Team @ Alibaba Tongyi Lab, for Qwen3.6-27B: a world-class open-weight multimodal foundation with hybrid Gated-DeltaNet attention, 262K context, and an MTP speculative-decoding head. Remarkable work, openly shared.

Claude-Opus Reasoning Distill & Coder SFT — Jackrong, for Qwen 3.6 27B-v2 and the agentic-coding extension Jackrong/Qwopus3.6-27B-Coder. The Trace Inversion recipe and resulting quality are what make this abliteration worth doing.

Abliteration Toolkit — Lemura Labs, for the ablation toolkit, an elegant Optuna-driven refusal-ablation framework (TPE search, KL guardrails, checkpointing, LoRA-merge). This release would not exist without it.

MLX — Apple ML Research, for the MLX framework and its first-class MX quantization modes (MXFP4 / MXFP8) that make 27B inference and quantization on Apple Silicon possible at this quality. mlx-mtp is built on mlx.core / mlx.nn alone.

mlx-mtp — our own pure–Apple-mlx quantization + inference stack for the Qwen3.6-27B / Qwen3.5-family VLMs. It vendors and extends the Qwen3.5 architecture (hybrid Gated-DeltaNet + full attention), the vision tower, and a natively-embedded MTP head, with tensor-level MXFP4/MXFP8 quantization that preserves vision + MTP + SSM at fp16. mlx-mtp on GitHub.

Lemura Labs & — abliteration, MTP restoration, quantization, and publication by the Lemura Labs. Lemura Labs builds multi-provider LLM routing for the Indian developer ecosystem — the.


License

Apache-2.0, inherited from the foundation (Qwen3.6-27B) and the coder fine-tune (Jackrong/Qwopus3.6-27B-Coder) upstream.


Need a hosted endpoint, custom quant, or enterprise inference? Lemura Labs — multi-provider LLM routing built for the Indian developer ecosystem.

README history 1 version

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-08-05Initial commit19aab9d11.7 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration