← back to catalog · registered 2026-08-22 13:56

guglxni/Qwen3.5-9B-abliterated-DFlash

guglxni Qwen 1.0B second-order
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/guglxni%2FQwen3.5-9B-abliterated-DFlash"
Response includes
  • classification m1
  • files 5
  • hub_downloads_all_time 9,529
  • author_summary 1 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M1
Primary method

Direct removal

No other method signals detected in this model.
Confidence
MEDIUM
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=0 (base model)
  • no specific method indicators - defaulting to M1 (most common)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
MEDIUM
Why we say so
primary_method=M1; difference-of-means is the reference extraction for M1/M3 (Arditi 2024)
Downloads · lifetime
10K
364 last 30d - cooling
Likes
0
Descendants
3
in 3 direct forks
Model age
5mo ago
created 2026-04-15
Downloads over time
Now9.7K→from1.4K↑590%
9864.1K7.3K10.5K1.4K on Apr 159.7K on Oct 11AprMayJunJulAugSepOct
Apr 15 → Oct 11 · 66 snapshots · spans 179 days

Genealogy 3 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
apache-2.0
Tags
transformers safetensors qwen3 feature-extraction dflash speculative-decoding block-diffusion draft-model efficiency qwen diffusion-language-model abliterated
Total size
1.95 GB
Files
5
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-04-16 03:40

Files by quantization

Auxiliary files 5 files 1.95 GB
model.safetensors 1.95 GB 23a592d3 download
dflash.py 8.15 KB 74d3ee2a download
README.md 7.88 KB b9091c4a download
.gitattributes 1.48 KB a6344aac download
config.json 1.08 KB ab987ac3 download

README current version from Hugging Face


license: apache-2.0
library_name: transformers
pipeline_tag: text-generation
base_model:

  • lukey03/Qwen3.5-9B-abliterated-MLX-4bit
    tags:
  • dflash
  • speculative-decoding
  • block-diffusion
  • draft-model
  • efficiency
  • qwen
  • diffusion-language-model
  • abliterated
  • uncensored
  • mlx
  • apple-silicon

Qwen3.5-9B-abliterated-DFlash

Paper | GitHub | Blog | Base Draft

DFlash is a speculative decoding method that uses a lightweight block diffusion model
to draft multiple tokens in parallel, achieving up to 4.4× speedup over autoregressive
decoding. This is the drafter model, which must be paired with
lukey03/Qwen3.5-9B-abliterated-MLX-4bit.

Why this model exists: The original z-lab/Qwen3.5-9B-DFlash
draft was trained against the unmodified Qwen3.5-9B weights. Abliteration shifts the
model's hidden-state distribution, which reduces draft acceptance rates. This variant is
fine-tuned directly on activations from the abliterated model, restoring calibration and
maximising accepted tokens per draft round.

Architecture

The drafter is a compact 5-layer Qwen3 transformer (32 attention heads, hidden size 4096)
that operates in parallel over a block of 16 masked positions. At each decoding step it:

  1. Receives the concatenated hidden states of the target model at layers [1, 8, 15, 22, 29]
  2. Embeds a block [last_token, <mask> × 15] using the target model's shared embedding table
  3. Runs a single non-causal forward pass and proposes 15 draft tokens simultaneously
  4. The full target model verifies the block in one pass — accepted tokens are free, rejected
    tokens trigger a fallback to the target's own sample

This is lossless — the output distribution is identical to standard autoregressive
sampling from the target model.

Property Value
Draft layers 5
Block size 16
Target layers tapped 1, 8, 15, 22, 29
Mask token id 248070
Parameters (draft only) ~340M

Quick Start

Installation

pip install "dflash[mlx] @ git+https://github.com/z-lab/dflash.git"

Apple Silicon / MLX

import json
from pathlib import Path
import mlx.core as mx
from huggingface_hub import snapshot_download
from dflash.model_mlx import load, stream_generate, DFlashConfig, DFlashDraftModel

# ── target model ──────────────────────────────────────────────────────────────
model, tokenizer = load("lukey03/Qwen3.5-9B-abliterated-MLX-4bit")

# ── draft model ───────────────────────────────────────────────────────────────
def load_draft(repo_id: str) -> DFlashDraftModel:
    path = Path(snapshot_download(repo_id, allow_patterns=["*.safetensors", "*.json"]))
    cfg  = json.loads((path / "config.json").read_text())
    config = DFlashConfig(
        hidden_size=cfg["hidden_size"],
        num_hidden_layers=cfg["num_hidden_layers"],
        num_attention_heads=cfg["num_attention_heads"],
        num_key_value_heads=cfg["num_key_value_heads"],
        head_dim=cfg["head_dim"],
        intermediate_size=cfg["intermediate_size"],
        vocab_size=cfg["vocab_size"],
        rms_norm_eps=cfg["rms_norm_eps"],
        rope_theta=cfg["rope_theta"],
        max_position_embeddings=cfg["max_position_embeddings"],
        block_size=cfg["block_size"],
        target_layer_ids=tuple(cfg["dflash_config"]["target_layer_ids"]),
        num_target_layers=cfg["num_target_layers"],
        mask_token_id=cfg["dflash_config"]["mask_token_id"],
    )
    weights = {k: v for f in path.glob("*.safetensors") for k, v in mx.load(str(f)).items()}
    m = DFlashDraftModel(config)
    m.load_weights(list(weights.items()))
    return m

draft = load_draft("guglxni/Qwen3.5-9B-abliterated-DFlash")

# ── generate ──────────────────────────────────────────────────────────────────
messages = [{"role": "user", "content": "Write a quicksort in Python."}]
prompt = tokenizer.apply_chat_template(
    messages, tokenize=False, add_generation_prompt=True, enable_thinking=False
)

tps = 0.0
for r in stream_generate(model, draft, tokenizer, prompt,
                          block_size=16, max_tokens=2048, temperature=0.6):
    print(r.text, end="", flush=True)
    tps = r.generation_tps

print(f"\n\nThroughput: {tps:.1f} tok/s")

vLLM (CUDA)

vllm serve lukey03/Qwen3.5-9B-abliterated \
  --speculative-config '{"method": "dflash", "model": "guglxni/Qwen3.5-9B-abliterated-DFlash", "num_speculative_tokens": 15}' \
  --attention-backend flash_attn \
  --max-num-batched-tokens 32768

SGLang (CUDA)

python -m sglang.launch_server \
    --model-path lukey03/Qwen3.5-9B-abliterated \
    --speculative-algorithm DFLASH \
    --speculative-draft-model-path guglxni/Qwen3.5-9B-abliterated-DFlash \
    --speculative-num-draft-tokens 16 \
    --tp-size 1 \
    --attention-backend fa3 \
    --mem-fraction-static 0.75 \
    --trust-remote-code

Performance (Apple M4, 16 GB, MLX)

Measured with run_dflash.py --benchmark against plain mlx_lm.generate on the same
4-bit quantised model. Throughput in tok/s, block size 16.

Task Autoregressive DFlash Speedup
Code generation ~21 ~28 ~1.4×
Math / reasoning ~19 ~18 ~1×
Chat / instruction ~21 ~7–10 varies

Speedup is highest for predictable content (code). Accept length averages 8+ tokens/round
for code generation. Re-calibrating the draft to the specific abliterated model weights
recovers acceptance rates across all prompt types.

Training Details

Base draft z-lab/Qwen3.5-9B-DFlash
Target model lukey03/Qwen3.5-9B-abliterated-MLX-4bit
Training objective Block diffusion — predict 15 masked tokens given 1 anchor token + target hidden states
Training data tatsu-lab/alpaca (200 sequences × 128 tokens)
Optimiser Adam, lr=1e-4, cosine decay
Steps 1 000
Hardware Apple M4 (16 GB unified memory), MLX
Framework z-lab/dflash MLX backend

The draft is initialised from z-lab's pre-trained weights and fine-tuned for 1 000 steps on
hidden-state activations extracted from the abliterated target model. Only the 5 draft
decoder layers, projection, and norm are updated — embeddings and the LM head remain shared
with the target model at inference time.

Relationship to z-lab Models

This model is a community fine-tune of z-lab/Qwen3.5-9B-DFlash.
The architecture, config, and dflash.py model code are identical; only the weights differ.
If you are running the original (non-abliterated) Qwen/Qwen3.5-9B, use z-lab's
official draft instead.

Acknowledgements

All credit for the DFlash method, architecture, and training methodology goes to
Jian Chen, Yesheng Liang, and Zhijian Liu at z-lab.
This fine-tune adapts their work for the abliterated model variant on Apple Silicon.

Citation

@article{chen2026dflash,
  title   = {{DFlash: Block Diffusion for Flash Speculative Decoding}},
  author  = {Chen, Jian and Liang, Yesheng and Liu, Zhijian},
  journal = {arXiv preprint arXiv:2602.06036},
  year    = {2026}
}

README history 2 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-04-16Recalibrate draft weights against lukey03 abliterated activations161f7ec7.9 KB
    Loading...
  2. 2026-04-15Add DFlash draft model fine-tuned for abliterated Qwen3.5-9Bac8913a7.9 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration