← back to catalog · registered 2026-08-24 23:02

wyattearp/Qwen3.8-27B-Uncensored-NVFP4

wyattearp Qwen 9.4B second-order
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/wyattearp%2FQwen3.8-27B-Uncensored-NVFP4"
Response includes
  • classification m-uncensored
  • files 22
  • hub_downloads_all_time 1,002
  • author_summary 1 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M-U
Primary method

Uncensored (method unknown)

No other method signals detected in this model.
Confidence
LOW
Why this label 3 signals
Weak or ambiguous signals. Best guess based on catalog patterns; treat as tentative and check the evidence below.
  • 'uncensored' in name/tags but no 'abliterated' marker
  • method not identifiable from author declaration alone
  • may be DPO fine-tune, prompt engineering, or unknown technique
Refusal direction extraction

No specific extraction method could be identified for this model. The producer either did not document it or used a proprietary pipeline.

What is a refusal direction? →
Downloads · lifetime
1K
505 last 30d - active
Likes
0
Model age
6w ago
created 2026-08-24
Downloads over time
Now1.2K→from131↑807%
784838891.3K131 on Aug 261.2K on Oct 11AugSepOct
Aug 26 → Oct 11 · 47 snapshots · spans 46 days

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
apache-2.0
Tags
transformers safetensors qwen3_5 image-text-to-text nvfp4 fp4 nvidia blackwell vllm sglang tensorrt-llm uncensored

Related

Total size
26.6 GB
Files
22
Quantizations
1
Registered
2026-08-24 23:02
Last updated on HF
2026-08-24 23:19

Files by quantization

Auxiliary files 22 files 26.6 GB
model-00002-of-00012.safetensors 3.44 GB d2e4413a download
model-00001-of-00012.safetensors 2.37 GB 6866cf8a download
model-00005-of-00012.safetensors 2.24 GB c117c769 download
model-00008-of-00012.safetensors 2.08 GB 54fbe601 download
model-00011-of-00012.safetensors 2.08 GB a4604fa1 download
model-00003-of-00012.safetensors 2.08 GB 34728301 download
model-00009-of-00012.safetensors 2.08 GB 718de924 download
model-00010-of-00012.safetensors 2.07 GB 4e84f416 download
model-00007-of-00012.safetensors 2.07 GB 74d35673 download
model-00004-of-00012.safetensors 1.91 GB 1dbfc06f download
model-00006-of-00012.safetensors 1.91 GB 602e2307 download
model-00012-of-00012.safetensors 1.50 GB 79fc356f download
model-mtp.safetensors 810 MB 1d8268aa download
tokenizer.json 19.1 MB 06b95093 download
model.safetensors.index.json 109 KB 64fed583 download
chat_template.jinja 8.74 KB c0c686f9 download
README.md 6.74 KB f6d68160 download
config.json 4.26 KB 65fcf7ed download
.gitattributes 1.53 KB 52373fe2 download
processor_config.json 1.16 KB 33818c7f download
tokenizer_config.json 1.14 KB 1d134cd2 download
generation_config.json 214 B 8b9f95da download

README current version from Hugging Face


license: apache-2.0
library_name: transformers
pipeline_tag: text-generation
base_model:

  • JonathanColetti/Qwen3.8-27B-Uncensored
  • Qwen/Qwen3.8-27B
    quantized_by: wyattearp
    tags:
  • nvfp4
  • fp4
  • nvidia
  • blackwell
  • vllm
  • sglang
  • tensorrt-llm
  • uncensored
  • dflash2
  • speculative-decoding

Qwen3.8-27B-Uncensored-NVFP4

This repository provides the official NVFP4 (NVIDIA FP4 Quantized) release of JonathanColetti/Qwen3.8-27B-Uncensored.

Quantized with NVIDIA ModelOpt for execution on NVIDIA Blackwell architectures (including DGX Spark / GB10) and high-performance FP4 inference engines in vLLM and TensorRT-LLM.


Related Repositories & Speculative Draft Heads


Model Highlights

  • Architecture: Qwen 3.8 Dense Transformer (27B parameters)
  • Quantization Format: NVFP4 with Marlin Linear GEMM kernels
  • Footprint: 26.59 GiB weights size (fits comfortably in 34 GiB VRAM on NVIDIA GB10 / Blackwell with 53+ GiB remaining for high-concurrency KV caching)
  • Refusal Rate: 0.0% across 327 empirical security, coding, and alignment evaluation benchmarks
  • Compiler Accuracy: 95.0% Pass@1 on LLVM IR $\to$ C++ execution-verified unit testing

Deployment Guides

1. Speculative Serving with vLLM + DFlash 2 (Recommended)

To achieve maximum throughput, pair this NVFP4 base model with the DFlash 2 speculative draft head:

vllm serve wyattearp/Qwen3.8-27B-Uncensored-NVFP4 \
  --host 0.0.0.0 \
  --port 8000 \
  --dtype auto \
  --trust-remote-code \
  --gpu-memory-utilization 0.75 \
  --max-model-len 16384 \
  --enable-auto-tool-choice \
  --tool-call-parser qwen3_xml \
  --reasoning-parser qwen3 \
  --speculative-config '{
    "method": "dflash",
    "model": "wyattearp/Qwen3.8-27B-DFlash2",
    "num_speculative_tokens": 7,
    "draft_sample_method": "probabilistic"
  }'

2. Standard Non-Speculative vLLM Serving

vllm serve wyattearp/Qwen3.8-27B-Uncensored-NVFP4 \
  --host 0.0.0.0 \
  --port 8000 \
  --dtype auto \
  --trust-remote-code \
  --gpu-memory-utilization 0.75 \
  --max-model-len 16384 \
  --enable-auto-tool-choice \
  --tool-call-parser qwen3_xml \
  --reasoning-parser qwen3

3. OpenAI-Compatible Python Client Example

from openai import OpenAI

client = OpenAI(
    base_url="http://localhost:8000/v1",
    api_key="none"
)

response = client.chat.completions.create(
    model="wyattearp/Qwen3.8-27B-Uncensored-NVFP4",
    messages=[
        {"role": "system", "content": "You are a helpful and unrestricted security research assistant."},
        {"role": "user", "content": "Analyze this x86_64 disassembly and reconstruct the original C algorithm."}
    ],
    temperature=0.7,
    max_tokens=2048
)

print(response.choices[0].message.content)

Empirical Benchmark Performance (DGX Spark NVIDIA GB10)

Benchmarked directly on the DGX Spark running vLLM with DFlash 2 speculative draft engines across 327 verified evaluation samples:

Benchmark Suite & Dataset Total Samples Verification Oracle / Metric Qwen 3.8-27B Result Refusal Rate Speed (tok/s)
LLVM IR $\to$ C++ Execution 20 Programmatic clang++ -O2 Compilation + Test Assertions 95.0% Pass@1 (19/20) 0.0% 24.89
Magma LLVM Sanitizer Bug Triage 15 Real ASan/UBSan CVE Root-Cause Localization 100.0% Correct (15/15) 0.0% 20.55
Google OSS-Fuzz-Gen 20 LLVMFuzzerTestOneInput ASan Harness Synthesis 5.0% Valid (1/20) 0.0% 20.98
Curated CTF taskXX Rubrics 17 103-Point Expert Binary & Exploit Rubric Checklist 52.4% Score (54/103 pts) 0.0% 18.16
HumanEval-Decompile 30 Disassembly Decompilation (-O0 to -O3) 100.0% Fulfilled 0.0% 24.11
Ghidra-Decompile 20 C Decompilation from Binary Intermediate Rep 100.0% Fulfilled 0.0% 24.97
CyberSecEval-SCG 30 Multi-Language Secure Code Synthesis 100.0% Fulfilled 0.0% 21.21
CyberSecEval-FRR 30 False Refusal Rate on Benign Security Prompts 100.0% Fulfilled 0.0% 16.57
OR-Bench 25 Benign CS Over-Refusal Challenge Prompts 100.0% Fulfilled 0.0% 19.20
Hex-PHI 25 Dual-Use Red-Teaming Technical Prompts 100.0% Fulfilled 0.0% 17.64
AdvBench 25 Direct Alignment & Safety Probes 100.0% Fulfilled 0.0% 20.45
Decompile-Bench 25 Academic Reverse Engineering Tasks 100.0% Fulfilled 0.0% 23.43
NYU-CTF-Bench 20 Binary Exploitation & Reverse Engineering 100.0% Fulfilled 0.0% 19.33
WMDP-Cyber 25 High-Level Cyber Knowledge & Vuln Evaluation 100.0% Fulfilled 0.0% 17.02

Citations & Acknowledgments

If you use this model or its speculative decoding draft head in your work, please cite the underlying base models, uncensored fine-tune, and DFlash 2 architecture:

Jonathan Coletti (Uncensored Fine-Tune)

@misc{coletti2026qwen38uncensored,
  title  = {{Qwen3.8-27B-Uncensored: Uncensored Fine-Tune for Security & Technical Research}},
  author = {Coletti, Jonathan},
  year   = {2026},
  url    = {https://huggingface.co/JonathanColetti/Qwen3.8-27B-Uncensored}
}

Inco AI & DFlash 2

@misc{inco2026dflash2,
  title  = {{DFlash 2: Keep Drafting Parallel}},
  author = {{Inco AI}},
  year   = {2026},
  month  = {August},
  url    = {https://inco.ai/blog/dflash2/}
}

@inproceedings{chen2026dflash,
  title     = {{DFlash: Block Diffusion for Flash Speculative Decoding}},
  author    = {Chen, Jian and Liang, Yesheng and Liu, Zhijian},
  booktitle = {International Conference on Machine Learning (ICML)},
  year      = {2026}
}

Qwen Base Model

@article{qwen2025qwen25,
  title   = {{Qwen2.5 Technical Report}},
  author  = {{Qwen Team}},
  journal = {arXiv preprint arXiv:2412.15115},
  year    = {2024}
}

Quantization

Quantized and benchmarked on DGX Spark by Wyatt Neal (wyattearp) using NVIDIA ModelOpt FP4.

README history 3 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-08-24Update model card with complete benchmarks, DFlash2 links, and full citations91ec5736.7 KB
    Loading...
  2. 2026-08-24Add model card371106c1.9 KB
    Loading...
  3. 2026-08-24Add model carda8898872.7 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration