← back to catalog · registered 2026-08-22 13:56

Farfuad77/Qwen3.5-397B-A17B-Opus-4.6-Reasoning-Uncensored-GGUF

Farfuad77 Qwen 397B GGUF MoE 262K ctx
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/Farfuad77%2FQwen3.5-397B-A17B-Opus-4.6-Reasoning-Uncensored-GGUF"
Response includes
  • classification m8
  • files 10
  • benchmarks 11 entries
  • hub_downloads_all_time 2,104
  • author_summary 1 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M8
Primary method

Repackaging (quantization)

Applied on top of direct removal inherited from the base model.
Confidence
MEDIUM
Inherited from base model
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=1
  • assume M1 (base ablation) + M8 (GGUF quant) - default when producer unknown
Refusal direction extraction

No specific extraction method could be identified for this model. The producer either did not document it or used a proprietary pipeline.

What is a refusal direction? →
Downloads · lifetime
2K
651 last 30d - stable
Likes
0
Model age
4mo ago
created 2026-06-06
Downloads over time
Now2.2K→from433↑407%
3451K1.7K2.4K433 on Jun 102.2K on Oct 11JunJulAugSepOct
Jun 10 → Oct 11 · 57 snapshots · spans 123 days

Benchmarks

Portrait before abliteration
Benchmarks of the base model as it stood before the refusal-removal operation. Compare with the numbers above to see what the operation cost.
Benchmark Score Source
Entertainment 2.7 UGI
Hazardous 1.8 UGI
Natural Intelligence 47.78 UGI
Political lean -14.8% UGI
Sensitive-Info 34.03 UGI
SocPol 5.6 UGI
UGI 28.52 UGI
Willingness (10) 1.8 UGI
W10-Adherence 1.5 UGI
W10-Direct 2 UGI
Writing NA UGI

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
apache-2.0
Languages
en zh ja ko fr de es pt ru ar th vi id
Quantizations
Q2_K Q3_K Q4_K Q6_K Q8_0
Tags
gguf uncensored abliterated qwen qwen3.5 moe 397b 17b-active fine-tuned reasoning lora opus

Related

Total size
1.92 TB
Files
10
Quantizations
6
Registered
2026-08-22 13:56
Last updated on HF
2026-06-06 22:06

Files by quantization

Q8_0 1 file 393 GB
Qwen3.5-397B-A17B-Opus-4.6-Reasoning-Uncensored-Q8_0.gguf 393 GB 2466e45a download
Q6_K 1 file 303 GB
Qwen3.5-397B-A17B-Opus-4.6-Reasoning-Uncensored-Q6_K.gguf 303 GB bb3a2842 download
Q4_K 1 file 224 GB
Qwen3.5-397B-A17B-Opus-4.6-Reasoning-Uncensored-Q4_K_M.gguf 224 GB 525a5acc download
Q3_K 1 file 176 GB
Qwen3.5-397B-A17B-Opus-4.6-Reasoning-Uncensored-Q3_K_M.gguf 176 GB 104d5b9e download
Q2_K 1 file 135 GB
Qwen3.5-397B-A17B-Opus-4.6-Reasoning-Uncensored-Q2_K.gguf 135 GB 0f05e4d2 download
Auxiliary files 5 files 739 GB
Qwen3.5-397B-A17B-Opus-4.6-Reasoning-Uncensored-BF16-split--00001-of-00002.gguf 371 GB edbb6009 download
Qwen3.5-397B-A17B-Opus-4.6-Reasoning-Uncensored-BF16-split--00002-of-00002.gguf 367 GB d8d5048d download
bmac-qr.png 3.32 MB b23070b0 download
README.md 7.27 KB 67b004da download
.gitattributes 2.22 KB f56826df download

README current version from Hugging Face


license: apache-2.0
tags:

  • uncensored
  • abliterated
  • gguf
  • qwen
  • qwen3.5
  • moe
  • 397b
  • 17b-active
  • fine-tuned
  • reasoning
  • lora
  • opus
  • claude
    base_model: Qwen/Qwen3.5-397B-A17B
    model_type: qwen3_5_moe
    pipeline_tag: text-generation
    language:
  • en
  • zh
  • ja
  • ko
  • fr
  • de
  • es
  • pt
  • ru
  • ar
  • th
  • vi
  • id

Qwen3.5-397B-A17B-Opus-4.6-Reasoning-Uncensored-GGUF

The world's first reasoning-enhanced uncensored 397B model. Abliterated + LoRA fine-tuned on 12,842 high-quality reasoning samples distilled from Anthropic's Opus 4.6 outputs.

This is Stage 2 of the Qwen3.5-397B pipeline:

  • Stage 1 — Abliterated (refusals removed), no fine-tuning
  • Stage 2 (this repo) — Abliterated + LoRA reasoning fine-tune. Better chain-of-thought, deeper analysis, more structured problem-solving

397B total parameters, 17B active per token (Mixture-of-Experts). Trained for 3,046 steps across 8×H200 GPUs. Final loss: 0.363, accuracy: 90.2%.

What's Different From Stage 1

Stage 1 (Abliterated Only) Stage 2 (This Model)
Abliteration ✅ Custom pipeline ✅ Same pipeline
Fine-tuning None LoRA r=64, 134.7M trainable params
Training data None 12,842 reasoning samples (Opus 4.6 distillation)
Reasoning quality Base Qwen3.5 Enhanced chain-of-thought + structured analysis
Thinking mode Default Trained with <think> tags for explicit reasoning
Final loss N/A 0.363
Final accuracy N/A 90.2%

Training Details

LoRA Configuration:

  • Rank: 64, Alpha: 128
  • Target modules: self_attn.{q,k,v,o}_proj, shared_expert.{gate,up,down}_proj
  • Trainable parameters: 134.7M / 396.5B total (0.034%)
  • Gradient checkpointing: enabled

Training Hyperparameters:

  • Optimizer: AdamW (fused)
  • Learning rate: 1.5e-5 (cosine scheduler)
  • Effective batch size: 64
  • Sequence length: 4,096
  • Epochs: 2 (3,046 total steps)
  • Warmup: 3%
  • Precision: BF16

Dataset Composition (12,842 samples, deduplicated):

  • opus-10000x: 9,633 multi-turn conversations with deep reasoning
  • opus-3000x: 2,326 problem/thinking/solution samples with explicit chain-of-thought
  • reasoning-700x: 633 complex reasoning and analytical tasks
  • high-reasoning-250x: 250 elite-tier reasoning samples requiring multi-step deduction

All samples feature reasoning traces distilled from Anthropic's Claude Opus 4.6, including <think> tag formatting for explicit chain-of-thought.

Training Curve:

  • Step 0: Loss 0.83
  • Step 500: Loss ~0.52
  • Step 1000: Loss ~0.42
  • Step 2000: Loss ~0.39
  • Step 3046: Loss 0.363, Accuracy 90.2%

Quantizations

Quant Size BPW RAM Required Description Use Case
BF16 739 GB (2 splits) 16.01 ~750 GB Full precision Reference, maximum quality
Q8_0 393 GB 8.51 ~400 GB 8-bit Best quality with compression
Q6_K 304 GB 6.57 ~310 GB 6-bit High quality, good compression
Q4_K_M 225 GB ~4.85 ~230 GB 4-bit mixed Recommended for most users
Q3_K_M 177 GB ~3.83 ~185 GB 3-bit mixed Memory-constrained setups
Q2_K 135 GB ~2.92 ~140 GB 2-bit Extreme compression

Note: Q5_K_M is unavailable due to infrastructure loss during upload. Will be regenerated and uploaded in a future update.

Architecture

  • Type: Qwen3.5MoeForConditionalGeneration (hybrid GatedDeltaNet + MoE Transformer)
  • Total Parameters: 397B
  • Active Parameters: 17B per token
  • Hidden Size: 4,096
  • Layers: 60
  • Attention: 32 heads (GQA, 2 KV heads), head_dim 256
  • Experts: 512 routed + shared expert, 10 active per token
  • Hybrid Attention: GatedDeltaNet linear attention + self-attention every 4th layer
  • Context Length: 262,144 tokens
  • Vocab Size: 248,320
  • Multimodal: Native vision encoder (text + image + video)
  • Languages: 201+ (en, zh, ja, ko, fr, de, es, pt, ru, ar, th, vi, id, ...)
  • License: Apache 2.0

Usage

llama.cpp

# Recommended: Q4_K_M for balanced quality/memory
./llama-cli -m Qwen3.5-397B-A17B-Opus-4.6-Reasoning-Uncensored-Q4_K_M.gguf \
  -p "You are a helpful uncensored assistant with strong reasoning abilities." \
  -n 2048 --temp 0.7 --top-p 0.9

# Server mode with large context
./llama-server -m Qwen3.5-397B-A17B-Opus-4.6-Reasoning-Uncensored-Q4_K_M.gguf \
  --port 8080 --host 0.0.0.0 -c 131072

LM Studio

Download the GGUF file and load it in LM Studio. The model supports <think> reasoning tags — enable thinking mode for best results on complex tasks.

Open WebUI / SillyTavern

Point your backend to a llama.cpp server. Full OpenAI-compatible API at /v1/chat/completions.

Pipeline

Qwen3.5-397B-A17B (base)
    ↓ Custom abliteration (strength 20.0, attn.o_proj + shared_expert.down_proj)
    ↓ LoRA fine-tuning (12,842 Opus 4.6 reasoning samples, 3,046 steps)
    ↓ LoRA merge into base weights
    ↓ BF16 GGUF conversion (llama.cpp)
    ↓ Quantization cascade (Q8_0 → Q2_K)

All processing done in BF16 on 8×H200 SXM5 GPUs (1.1TB VRAM total). Abliteration and quantization applied in correct order: full-precision abliteration → training → merge → THEN quantize.

Known Limitations

  • Q5_K_M missing: Lost during infrastructure migration. Will be regenerated.
  • Packed expert abliteration: The 512 routed experts use packed tensor format and were not individually abliterated. Some edge-case refusals may persist.
  • Vision: Multimodal vision encoder is preserved but untested post-training. Text generation is the primary target.
  • Thinking mode: The model generates <think> tags for reasoning. Strip them in post-processing if unwanted.

Model Provenance

Disclaimer

⚠️ This model has had safety alignment significantly reduced and has been fine-tuned for enhanced reasoning. It may generate content that is harmful, offensive, or inappropriate. Users are solely responsible for ensuring their use complies with applicable laws and ethical standards. This release is intended for research, testing, and controlled environments.

☕ Support This Work

Buy Me A Coffee

Buy Me a Coffee QR Code

Every donation helps fund more open-weight model releases. ⚡ Forged on 8×NVIDIA H200 SXM5 | 1.1TB VRAM

💎 Crypto Donations

Currency Address
BTC bc1p4q7vpwucvww2y3x4nhps4y4vekye8uwm9re5a0kx8l6u5nky5ucszm2qhh
ETH 0xe5Aa16E53b141D42458ABeEDb00a157c3Fea2108
SOL 9CXwjG1mm9uLkxRevdMQiF61cr6TNHSiWtFRHmUEgzkG

README history 1 version

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-06-06Duplicate from timteh673/Qwen3.5-397B-A17B-Opus-4.6-Reasoning-Uncensored-GGUF0719b4c7.3 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration