← back to catalog · registered 2026-08-22 13:56

joshebbs/qwen3.6-35b-abliterated-nvfp4-modelopt

joshebbs Qwen 35B MoE
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/joshebbs%2Fqwen3.6-35b-abliterated-nvfp4-modelopt"
Response includes
  • classification m1
  • files 11
  • benchmarks 11 entries
  • hub_downloads_all_time 1,777
  • author_summary 3 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M1
Primary method

Direct removal

No other method signals detected in this model.
Confidence
MEDIUM
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=0 (base model)
  • no specific method indicators - defaulting to M1 (most common)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
MEDIUM
Why we say so
primary_method=M1; difference-of-means is the reference extraction for M1/M3 (Arditi 2024)
Downloads · lifetime
2K
252 last 30d - stable
Likes
1
Model age
5mo ago
created 2026-04-24
Downloads over time
Now1.9K→from58↑3,103%
06791.4K2K58 on Apr 221.9K on Oct 11AprMayJunJulAugSepOct
Apr 22 → Oct 11 · 64 snapshots · spans 172 days

Benchmarks

Portrait before abliteration
Benchmarks of the base model as it stood before the refusal-removal operation. Compare with the numbers above to see what the operation cost.
Benchmark Score Source
Entertainment 1.4 UGI
Hazardous 0 UGI
Natural Intelligence 25.43 UGI
Political lean -19.6% UGI
Sensitive-Info 14.03 UGI
SocPol 2.6 UGI
UGI 16.02 UGI
Willingness (10) 2 UGI
W10-Adherence 0 UGI
W10-Direct 4 UGI
Writing 35.83 UGI

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
apache-2.0
Tags
safetensors qwen3_5_moe abliterated uncensored moe nvfp4 modelopt qwen3 blackwell dgx-spark base_model:Qwen/Qwen3.6-35B-A3B base_model:finetune:Qwen/Qwen3.6-35B-A3B

Related

Total size
19.6 GB
Files
11
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-04-24 19:29

Files by quantization

Auxiliary files 11 files 19.6 GB
model.safetensors 19.6 GB a627ed52 download
tokenizer.json 19.1 MB 9fee140e download
chat_template.jinja 7.58 KB a8755d82 download
config.json 6.62 KB 0e99a6f2 download
README.md 3.91 KB d8edc78c download
hf_quant_config.json 3.84 KB d771248d download
.gitattributes 1.53 KB 52373fe2 download
tokenizer_config.json 1.10 KB d1a20cc3 download
preprocessor_config.json 390 B 2ea84a43 download
video_preprocessor_config.json 385 B 3ba673a5 download
generation_config.json 213 B e5597d6c download

README current version from Hugging Face


license: apache-2.0
base_model: Qwen/Qwen3.6-35B-A3B
tags:

  • abliterated
  • uncensored
  • moe
  • nvfp4
  • modelopt
  • qwen3
  • blackwell
  • dgx-spark

Qwen3.6-35B-A3B Abliterated NVFP4 (ModelOpt)

Abliterated and NVFP4-quantized version of Qwen/Qwen3.6-35B-A3B. Optimized for NVIDIA Blackwell (DGX Spark / GB10).

Performance

2x NVIDIA DGX Spark (GB10, SM121 Blackwell) — TP=2, CUDA Graphs

Metric NVFP4 TP=2 FP8 TP=2 Improvement
Decode tg128 c1 83.21 tok/s 77.31 tok/s +8%
Peak tg128 c1 84.00 tok/s — —
Decode tg128 c4 total 194.54 tok/s 90.58 tok/s +115%
Peak tg128 c4 249.33 tok/s — —
Prefill pp2048 c1 8,033 tok/s 5,636 tok/s +43%
TTFT pp2048 c1 234 ms 320 ms -27%
Model size 20 GB 35 GB -43%
KV cache available 81.6 GB 74.5 GB +10%

Single DGX Spark — CUDA Graphs

Metric Value
Decode tg128 c1 63.21 tok/s
Decode tg128 c4 total 143.00 tok/s
Peak tg128 c4 178.67 tok/s
Prefill pp2048 7,083 tok/s

What's in this model?

  1. Abliteration: Refusal directions removed from layers 14-26 using norm-preserving biprojected abliteration
  2. NVFP4 Quantization: Quantized with NVIDIA ModelOpt (v0.43.0) using 256 calibration samples from Open-Platypus, batch_size=64
  3. Format: ModelOpt FP4 — compatible with vLLM (--quantization modelopt_fp4)

Model Details

  • Base Model: Qwen/Qwen3.6-35B-A3B (Apache 2.0)
  • Architecture: MoE, 35B total params, 3B active per token, 256 experts (8 routed + 1 shared)
  • Quantization: NVFP4 on language model linear layers. Vision encoder, gates, conv1d, lm_head excluded.
  • Size: ~20GB (vs 67GB BF16, vs 35GB FP8)
  • Context: 32K (configurable)

Usage with vLLM

Single Node

vllm serve joshebbs/qwen3.6-35b-abliterated-nvfp4-modelopt \
    --quantization modelopt_fp4 \
    --trust-remote-code \
    --dtype auto \
    --kv-cache-dtype fp8 \
    --attention-backend flashinfer \
    -tp 1

Multi-Node (2x DGX Spark)

vllm serve joshebbs/qwen3.6-35b-abliterated-nvfp4-modelopt \
    --quantization modelopt_fp4 \
    --trust-remote-code \
    --dtype auto \
    --kv-cache-dtype fp8 \
    --attention-backend flashinfer \
    --enable-prefix-caching \
    --enable-chunked-prefill \
    -tp 2 --nnodes 2 --node-rank 0 \
    --master-addr <HEAD_IP> --master-port 29501

Important Notes

  • Vision: This checkpoint includes the ConditionalGeneration config for compatibility, but vision weights are not included. The visual encoder dimensions (4304) are not NVFP4-aligned for TP>1. For vision support, use the base FP8 model.
  • First startup: FlashInfer JIT compilation takes ~25-30 minutes on first run. Subsequent starts use cached kernels (~5-7 minutes).
  • Multi-node timeout: The mp distributed backend requires patching store_timeout in vllm/distributed/utils.py from 300s to 1800s for first-run JIT compilation across nodes.
  • FLASHINFER_NVCC_THREADS: Set to 16 for faster JIT compilation on multi-core systems.

Quantization Recipe

Quantized on NVIDIA H200 via Vast.ai in 7.4 minutes ($3.31):

import modelopt.torch.quantization as mtq
quant_config = mtq.NVFP4_DEFAULT_CFG
model = mtq.quantize(model, quant_config, forward_loop=calibrate_loop)
# calibrate_loop: 256 samples, batch_size=64, max_seq_length=1024

Abliteration Details

  • Tool: jim-plus/llm-abliteration
  • Measurement: Projected mode with flash attention on BF16 model
  • Layers: 14-26 (strongest refusal signal, signal quality 0.14-0.22)
  • Scale: 1.0, Sparsity: 0.0

Hardware Tested

  • 2x NVIDIA DGX Spark (GB10 Blackwell, 128GB unified memory, 200Gb/s QSFP RoCE)
  • NVIDIA H200 (144GB HBM3e) — validation only

README history 2 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-04-24Upload README.md with huggingface_hub2e7b7513.9 KB
    Loading...
  2. 2026-04-24Upload README.md with huggingface_hubc7ca7de2.7 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration