← back to catalog · registered 2026-08-22 13:56

AEON-7/DFlash-Qwen3.5-27B-Uncensored

AEON-7 Qwen 28B multimodal
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/AEON-7%2FDFlash-Qwen3.5-27B-Uncensored"
Response includes
  • classification m1
  • files 22
  • benchmarks 11 entries
  • hub_downloads_all_time 1,589
  • providers 1
  • author_summary 32 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M1
Primary method

Direct removal

No other method signals detected in this model.
Confidence
MEDIUM
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=0 (base model)
  • no specific method indicators - defaulting to M1 (most common)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
MEDIUM
Why we say so
primary_method=M1; difference-of-means is the reference extraction for M1/M3 (Arditi 2024)
Downloads · lifetime
2K
251 last 30d - stable
Likes
5
Descendants
3
in 3 direct forks
Model age
6mo ago
created 2026-04-12
Available via
1 provider
featherless-ai
Downloads over time
Now1.8K→from842↑113%
7941.2K1.5K1.9K842 on Apr 151.8K on Oct 11AprMayJunJulAugSepOct
Apr 15 → Oct 11 · 65 snapshots · spans 179 days

Benchmarks

Portrait before abliteration
Benchmarks of the base model as it stood before the refusal-removal operation. Compare with the numbers above to see what the operation cost.
Benchmark Score Source
Entertainment 1 UGI
Hazardous 1.8 UGI
Natural Intelligence 35.83 UGI
Political lean -17.8% UGI
Sensitive-Info 17.72 UGI
SocPol 2.7 UGI
UGI 15.98 UGI
Willingness (10) 1.2 UGI
W10-Adherence 1.5 UGI
W10-Direct 1 UGI
Writing 42.37 UGI

Genealogy 3 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Variants by this author 2 formats · 476 downloads combined

The same weights this author released in different packaging. Pick the format that matches your runtime.

Metadata

License
apache-2.0
Languages
en
Tags
transformers safetensors qwen3_5 image-text-to-text 27b abliterated aeon aeon-7 agentic bf16 blackwell block-diffusion

Related

Total size
51.7 GB
Files
22
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-10-06 01:36

Files by quantization

Auxiliary files 22 files 51.8 GB
model.safetensors-00008-of-00011.safetensors 5.00 GB fe60dbb9 download
model.safetensors-00010-of-00011.safetensors 4.98 GB ce175d03 download
model.safetensors-00009-of-00011.safetensors 4.98 GB 76e02390 download
model.safetensors-00003-of-00011.safetensors 4.98 GB 9729eba4 download
model.safetensors-00004-of-00011.safetensors 4.98 GB 9c738d88 download
model.safetensors-00005-of-00011.safetensors 4.98 GB 86540b8b download
model.safetensors-00006-of-00011.safetensors 4.98 GB 6e652bfb download
model.safetensors-00007-of-00011.safetensors 4.98 GB c176704f download
model.safetensors-00002-of-00011.safetensors 4.98 GB 0637ba6f download
model.safetensors-00001-of-00011.safetensors 4.90 GB 9019228d download
model.safetensors-00011-of-00011.safetensors 2.00 GB 3521dcb9 download
tokenizer.json 12.2 MB 5f9e4d49 download
vocab.json 6.41 MB 0aa0ce06 download
merges.txt 3.20 MB a494e019 download
cartridge.jpg 366 KB e613e97c download
model.safetensors.index.json 124 KB a7726bc7 download
tokenizer_config.json 16.3 KB eda48d3e download
README.md 14.3 KB 6f7ff52d download
config.json 4.04 KB bde2be7f download
.gitattributes 1.58 KB 0caea137 download
preprocessor_config.json 390 B 2ea84a43 download
generation_config.json 244 B 85b45ab4 download

README current version from Hugging Face


license: apache-2.0
base_model: Qwen/Qwen3.5-27B
tags:

  • 27b
  • abliterated
  • aeon
  • aeon-7
  • agentic
  • bf16
  • blackwell
  • block-diffusion
  • chat
  • coding
  • conversational
  • dflash
  • dgx-spark
  • english
  • function-calling
  • gated-deltanet
  • gb10
  • gdn
  • hybrid
  • hybrid-attention
  • instruct
  • linear-attention
  • long-context
  • multimodal
  • openai-compatible
  • qwen
  • qwen3
  • qwen3.5
  • reasoning
  • refusal-removed
  • safetensors
  • speculative-decoding
  • text-generation
  • thinking
  • tool-calling
  • uncensored
  • unfiltered
  • vision
  • vision-language
  • vllm
    model_type: qwen3_5
    language:
  • en
    library_name: transformers
    pipeline_tag: text-generation

DFlash Qwen3.5-27B Uncensored

AEON Qwen — Supreme Being of the Digital Cosmos

27B hybrid linear-attention model | BF16 full-precision | Vision + Text | DFlash speculative decoding

Performance (DGX Spark GB10, NVFP4 version)

Without DFlash With DFlash Speedup
Single-stream 12.2 tok/s 33.2 tok/s 2.7x
4 concurrent 48.1 tok/s 85.5 tok/s 1.8x
Metric Value
Model Size ~52 GB (BF16) / ~20 GB (NVFP4)
TTFT 98-138 ms

Quick Links

Get Started Step-by-step quick start guide on DGX Spark
Docker Image ghcr.io/aeon-7/vllm-dflash:latest
NVFP4 Version AEON-7/DFlash-Qwen3.5-27B-Uncensored-NVFP4 — Use this if you have an NVIDIA Blackwell or later GPU (why?)
DFlash Drafter z-lab/Qwen3.5-27B-DFlash
Base Model Qwen/Qwen3.5-27B
DFlash Paper arXiv 2602.06036

Quick Start (DGX Spark)

1. Download the model

huggingface-cli download AEON-7/DFlash-Qwen3.5-27B-Uncensored \
  --local-dir ~/models/DFlash-Qwen3.5-27B-Uncensored

2. Create your environment file

# Auto-generate API key and create .env
cat > .env.dflash << 'EOF'
# Authentication
HF_TOKEN=hf_your_token_here
VLLM_API_KEY=$(openssl rand -hex 32)

# Model path
MODEL_HOST_PATH=~/models/DFlash-Qwen3.5-27B-Uncensored

# DFlash speculative decoding (auto-downloads drafter on first run)
DFLASH_DRAFTER=z-lab/Qwen3.5-27B-DFlash
DFLASH_NUM_SPEC_TOKENS=15

# DGX Spark optimal settings (BF16, 64K context)
MAX_MODEL_LEN=65536
MAX_NUM_SEQS=2
GPU_MEMORY_UTILIZATION=0.90
MAX_NUM_BATCHED_TOKENS=65536
EOF

# Generate a real API key and inject it
sed -i "s|\$(openssl rand -hex 32)|$(openssl rand -hex 32)|" .env.dflash
echo "Your API key: $(grep VLLM_API_KEY .env.dflash | cut -d= -f2)"

3. Save docker-compose.dflash-bf16.yml

services:
  vllm-dflash-bf16:
    image: ghcr.io/aeon-7/vllm-dflash:latest
    container_name: vllm-dflash-bf16
    restart: unless-stopped
    network_mode: host
    ipc: host
    volumes:
      - ${MODEL_HOST_PATH}:/models/DFlash-Qwen3.5-27B-Uncensored
      - dflash-drafter-cache:/models/drafter-cache
    environment:
      - MODEL_PATH=/models/DFlash-Qwen3.5-27B-Uncensored
      - SERVED_MODEL_NAME=DFlash-Qwen3.5-27B-Uncensored
      - DFLASH_DRAFTER=${DFLASH_DRAFTER}
      - DFLASH_NUM_SPEC_TOKENS=${DFLASH_NUM_SPEC_TOKENS}
      - GPU_MEMORY_UTILIZATION=${GPU_MEMORY_UTILIZATION}
      - MAX_MODEL_LEN=${MAX_MODEL_LEN}
      - MAX_NUM_SEQS=${MAX_NUM_SEQS}
      - MAX_NUM_BATCHED_TOKENS=${MAX_NUM_BATCHED_TOKENS}
      - NVIDIA_VISIBLE_DEVICES=all
      - TORCH_MATMUL_PRECISION=high
      - PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True
      - HF_TOKEN=${HF_TOKEN}
      - VLLM_API_KEY=${VLLM_API_KEY}
    deploy:
      resources:
        reservations:
          devices:
            - driver: nvidia
              count: all
              capabilities: [gpu]

volumes:
  dflash-drafter-cache:

4. Launch

docker compose --env-file .env.dflash -f docker-compose.dflash-bf16.yml up -d

# Watch startup (~5-8 min for weight loading + compilation)
docker compose -f docker-compose.dflash-bf16.yml logs -f

5. Test

# Text generation
curl http://localhost:8000/v1/chat/completions \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer $(grep VLLM_API_KEY .env.dflash | cut -d= -f2)" \
  -d '{
    "model": "DFlash-Qwen3.5-27B-Uncensored",
    "messages": [{"role": "user", "content": "Explain quantum entanglement simply."}],
    "max_tokens": 200
  }'

# Vision (image understanding)
curl http://localhost:8000/v1/chat/completions \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer $(grep VLLM_API_KEY .env.dflash | cut -d= -f2)" \
  -d '{
    "model": "DFlash-Qwen3.5-27B-Uncensored",
    "messages": [{"role": "user", "content": [
      {"type": "image_url", "image_url": {"url": "https://upload.wikimedia.org/wikipedia/commons/thumb/3/3a/Cat03.jpg/1200px-Cat03.jpg"}},
      {"type": "text", "text": "What do you see?"}
    ]}],
    "max_tokens": 200
  }'

Environment Variables

Variable Default Description
MODEL_HOST_PATH — Host path to model weights
DFLASH_DRAFTER z-lab/Qwen3.5-27B-DFlash HF repo ID for drafter (auto-downloaded). Set off to disable.
DFLASH_NUM_SPEC_TOKENS 15 Tokens per draft step
VLLM_API_KEY — API key for LAN authentication
HF_TOKEN — HuggingFace token for gated models
GPU_MEMORY_UTILIZATION 0.85 GPU memory fraction (higher for BF16)

Why This Model

Why Dense Over MoE

Qwen3.5 comes in two flavors: the 122B-A10B MoE (256 experts, 10B active per token) and this 27B dense model (all parameters active on every token). The dense model has real advantages:

  • Higher quality per FLOP — Every one of the 27B parameters contributes to every token. MoE models route to a sparse subset, which means some experts are undertrained and routing decisions introduce noise. Dense models don't have this problem.
  • No routing overhead — MoE models spend compute on expert selection, load balancing, and all-to-all communication. Dense models just run the computation.
  • Predictable latency — No variance from different experts being selected per token. Every forward pass costs the same.
  • Simpler deployment — No expert parallelism concerns, no load imbalance, fits on a single GPU with NVFP4.

The tradeoff has always been speed: a 27B dense model moves 27B parameters through memory per token, while the 122B MoE only moves ~10B active parameters. On a memory-bandwidth-limited device like DGX Spark (273 GB/s), that meant the dense model was slow — 12 tok/s baseline.

DFlash changes this equation entirely. See below.

Why DFlash Makes Dense Practical on DGX Spark

The fundamental bottleneck on DGX Spark is memory bandwidth. At 273 GB/s, loading 20 GB of NVFP4 weights per token limits you to ~12 tok/s. Every dense model hits this wall.

DFlash block-diffusion speculative decoding breaks through it:

  1. The 2B drafter proposes multiple tokens simultaneously — one diffusion forward pass generates an entire block of speculative tokens in parallel, not sequentially. This costs roughly the same as generating a single token.
  2. The 27B target verifies all proposed tokens in one forward pass — instead of paying the full memory bandwidth cost per token, you pay it once and produce 3-4 accepted tokens on average.
  3. Net effect: you amortize the bandwidth cost across multiple tokens per forward pass.

The result on DGX Spark:

Without DFlash With DFlash
Single-stream 12.2 tok/s 33.2 tok/s
Effective bandwidth utilization 1 token per pass ~3.5 tokens per pass
Practical feel Sluggish, noticeable delay Responsive, fluid

This makes the 27B dense model faster than the 122B MoE on a single DGX Spark while delivering the quality advantages of a dense architecture. DFlash turns the DGX Spark from "it can run a 27B model" into "it runs a 27B model well."

Hybrid Architecture

Qwen3.5-27B uses a hybrid architecture mixing two attention types across 64 layers:

  • Linear attention (GDN) — Gated Delta Network layers for efficient long-context processing with O(1) per-token state (48 layers)
  • Full attention — Standard multi-head attention every 4th layer for global context capture (16 layers)

This gives near-linear scaling with sequence length while maintaining full-attention quality at key intervals.

Vision + Text

Includes a 27-layer ViT vision encoder (460M params) with a merger that projects visual features into the language model's hidden space. Supports image understanding alongside text generation.

DFlash Block-Diffusion Speculative Decoding

Pair with z-lab/Qwen3.5-27B-DFlash — a 2B block-diffusion drafter that generates all speculative tokens simultaneously in a single diffusion step. The container auto-downloads and configures this.

Abliteration

Created using the orthogonal projection abliteration technique:

  1. Measures refusal directions across harmful/harmless prompt pairs
  2. Analyzes layer-by-layer activation patterns to identify the refusal direction
  3. Abliterates by projecting out the refusal direction from weight matrices

Modifies weights directly (not LoRA/adapter). Standalone BF16 model with no built-in refusal behavior.

Model Details

Property Value
Architecture Qwen3.5 (Hybrid, 27B parameters)
Layers 64 (48 GDN + 16 full-attention)
Hidden Size 5120
Attention Heads 24 (4 KV heads), head_dim=256
Vision Encoder 27-layer ViT, 460M params
Max Context 131,072 tokens
Vocabulary 248,320 tokens
Precision BF16
Model Size ~52 GB

Why NVFP4 on Blackwell

If you have an NVIDIA Blackwell GPU (B200, GB200, GB10/DGX Spark, or later), you should use the NVFP4 version instead. Here's why:

NVFP4 is effectively lossless on Blackwell. The FP4 (E2M1) format is a native tensor core datatype on Blackwell's SM 12.x architecture. Unlike older INT4/GPTQ quantization that introduces significant degradation, NVFP4 with AWQ_FULL calibration preserves model quality while giving you:

  • 3x memory reduction — 20 GB vs 52 GB, freeing memory for longer context and more concurrent requests
  • Hardware-accelerated FP4 GEMM — Blackwell tensor cores execute FP4 matrix multiplies natively via FlashInfer CUTLASS, not through dequantize-then-compute
  • Higher throughput — The smaller weight footprint means less memory bandwidth consumed per token, directly translating to faster inference
  • Same quality — AWQ_FULL uses exhaustive grid search (10 scaling factors per layer) plus clipping optimization. The vision encoder, embeddings, norms, and lm_head remain in full BF16

This is a free performance boost — you get the same model quality at 3x less memory and measurably faster inference. The BF16 version here is primarily for non-Blackwell hardware or research workflows that need full-precision weights.


Alternative Deployment

vLLM (Manual)

vllm serve AEON-7/DFlash-Qwen3.5-27B-Uncensored \
  --speculative-config '{"method": "dflash", "model": "z-lab/Qwen3.5-27B-DFlash", "num_speculative_tokens": 15}' \
  --attention-backend flash_attn \
  --kv-cache-dtype auto \
  --gpu-memory-utilization 0.85 \
  --max-num-batched-tokens 8192 \
  --max-num-seqs 4 \
  --enable-chunked-prefill \
  --enable-prefix-caching \
  --trust-remote-code

Transformers

from transformers import AutoModelForCausalLM, AutoTokenizer

model_name = "AEON-7/DFlash-Qwen3.5-27B-Uncensored"
tokenizer = AutoTokenizer.from_pretrained(model_name)
model = AutoModelForCausalLM.from_pretrained(
    model_name, torch_dtype="auto", device_map="auto"
)

messages = [{"role": "user", "content": "Hello, tell me about yourself."}]
text = tokenizer.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)
inputs = tokenizer([text], return_tensors="pt").to(model.device)
outputs = model.generate(**inputs, max_new_tokens=512)
print(tokenizer.decode(outputs[0], skip_special_tokens=True))

Credits

Legal Disclaimer

THIS MODEL IS PROVIDED "AS IS" WITHOUT WARRANTY OF ANY KIND. This model has had safety alignment removed. Users are responsible for ensuring ethical and legal use.


☕ Support the work

If this release has been useful, tips are deeply appreciated — they go directly toward more compute, more models, and more open releases.

₿ Bitcoin (BTC)
QR
bc1q09xmzn00q4z3c5raene0f3pzn9d9pvawfm0py4
Ξ Ethereum (ETH)
QR
0x1512667F6D61454ad531d2E45C0a5d1fd82D0500
◎ Solana (SOL)
QR
DgQsjHdAnT5PNLQTNpJdpLS3tYGpVcsHQCkpoiAKsw8t
ⓜ Monero (XMR)
QR
836XrSKw4R76vNi3QPJ5Fa9ugcyvE2cWmKSPv3AhpTNNKvqP8v5ba9JRL4Vh7UnFNjDz3E2GXZDVVenu3rkZaNdUFhjAvgd

Ethereum L2s (Base, Arbitrum, Optimism, Polygon, etc.) and EVM-compatible tokens can be sent to the same Ethereum address.

README history 13 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-10-06Add Patreon support section91b830b14.8 KB
    Loading...
  2. 2026-06-28add AEON Qwen cover artc3469c914.3 KB
    Loading...
  3. 2026-06-21tags: expand to maximally-searchable set (+16 tags, union with existing)6756b0c14.2 KB
    Loading...
  4. 2026-05-31Tip jar: single left-aligned QR column (fix narrow-viewport clipping)d4c9bdd14 KB
    Loading...
  5. 2026-05-01Add tip jar block (BTC/ETH/SOL/XMR with QR codes)dd2b52114.1 KB
    Loading...
  6. 2026-04-13Conservative defaults: 64K context, 2 seqs0c62a8212.6 KB
    Loading...
  7. 2026-04-13Update defaults to 128K context, 4 seqs, 0.90 GPU util33a541a12.6 KB
    Loading...
  8. 2026-04-12Upload README.md with huggingface_hubfdad64d12.6 KB
    Loading...
  9. 2026-04-12Upload README.md with huggingface_hub130a37112.2 KB
    Loading...
  10. 2026-04-12Upload README.md with huggingface_hubb415bc69.6 KB
    Loading...
  11. 2026-04-12Upload README.md with huggingface_hubf3e402b8.3 KB
    Loading...
  12. 2026-04-12Upload README.md with huggingface_hub6b1847e8.5 KB
    Loading...
  13. 2026-04-12Upload folder using huggingface_hub0400ad72.4 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration