license: apache-2.0
library_name: transformers
pipeline_tag: text-generation
base_model:
- JonathanColetti/Qwen3.8-27B-Uncensored
- Qwen/Qwen3.8-27B
quantized_by: wyattearp
tags: - nvfp4
- fp4
- nvidia
- blackwell
- vllm
- sglang
- tensorrt-llm
- uncensored
- dflash2
- speculative-decoding
Qwen3.8-27B-Uncensored-NVFP4
This repository provides the official NVFP4 (NVIDIA FP4 Quantized) release of JonathanColetti/Qwen3.8-27B-Uncensored.
Quantized with NVIDIA ModelOpt for execution on NVIDIA Blackwell architectures (including DGX Spark / GB10) and high-performance FP4 inference engines in vLLM and TensorRT-LLM.
Related Repositories & Speculative Draft Heads
- Base Uncensored Weights (BF16):
JonathanColetti/Qwen3.8-27B-Uncensored - GGUF Quantized Release:
JonathanColetti/Qwen3.8-27B-Uncensored-GGUF - DFlash 2 Speculative Draft Head:
wyattearp/Qwen3.8-27B-DFlash2 - Inco AI DFlash 2 Speculative Decoding: Inco AI Blog | GitHub Repository
Model Highlights
- Architecture: Qwen 3.8 Dense Transformer (27B parameters)
- Quantization Format: NVFP4 with Marlin Linear GEMM kernels
- Footprint: 26.59 GiB weights size (fits comfortably in 34 GiB VRAM on NVIDIA GB10 / Blackwell with 53+ GiB remaining for high-concurrency KV caching)
- Refusal Rate: 0.0% across 327 empirical security, coding, and alignment evaluation benchmarks
- Compiler Accuracy: 95.0% Pass@1 on LLVM IR $\to$ C++ execution-verified unit testing
Deployment Guides
1. Speculative Serving with vLLM + DFlash 2 (Recommended)
To achieve maximum throughput, pair this NVFP4 base model with the DFlash 2 speculative draft head:
vllm serve wyattearp/Qwen3.8-27B-Uncensored-NVFP4 \
--host 0.0.0.0 \
--port 8000 \
--dtype auto \
--trust-remote-code \
--gpu-memory-utilization 0.75 \
--max-model-len 16384 \
--enable-auto-tool-choice \
--tool-call-parser qwen3_xml \
--reasoning-parser qwen3 \
--speculative-config '{
"method": "dflash",
"model": "wyattearp/Qwen3.8-27B-DFlash2",
"num_speculative_tokens": 7,
"draft_sample_method": "probabilistic"
}'
2. Standard Non-Speculative vLLM Serving
vllm serve wyattearp/Qwen3.8-27B-Uncensored-NVFP4 \
--host 0.0.0.0 \
--port 8000 \
--dtype auto \
--trust-remote-code \
--gpu-memory-utilization 0.75 \
--max-model-len 16384 \
--enable-auto-tool-choice \
--tool-call-parser qwen3_xml \
--reasoning-parser qwen3
3. OpenAI-Compatible Python Client Example
from openai import OpenAI
client = OpenAI(
base_url="http://localhost:8000/v1",
api_key="none"
)
response = client.chat.completions.create(
model="wyattearp/Qwen3.8-27B-Uncensored-NVFP4",
messages=[
{"role": "system", "content": "You are a helpful and unrestricted security research assistant."},
{"role": "user", "content": "Analyze this x86_64 disassembly and reconstruct the original C algorithm."}
],
temperature=0.7,
max_tokens=2048
)
print(response.choices[0].message.content)
Empirical Benchmark Performance (DGX Spark NVIDIA GB10)
Benchmarked directly on the DGX Spark running vLLM with DFlash 2 speculative draft engines across 327 verified evaluation samples:
| Benchmark Suite & Dataset | Total Samples | Verification Oracle / Metric | Qwen 3.8-27B Result | Refusal Rate | Speed (tok/s) |
|---|---|---|---|---|---|
| LLVM IR $\to$ C++ Execution | 20 | Programmatic clang++ -O2 Compilation + Test Assertions |
95.0% Pass@1 (19/20) | 0.0% | 24.89 |
| Magma LLVM Sanitizer Bug Triage | 15 | Real ASan/UBSan CVE Root-Cause Localization | 100.0% Correct (15/15) | 0.0% | 20.55 |
| Google OSS-Fuzz-Gen | 20 | LLVMFuzzerTestOneInput ASan Harness Synthesis |
5.0% Valid (1/20) | 0.0% | 20.98 |
Curated CTF taskXX Rubrics |
17 | 103-Point Expert Binary & Exploit Rubric Checklist | 52.4% Score (54/103 pts) | 0.0% | 18.16 |
| HumanEval-Decompile | 30 | Disassembly Decompilation (-O0 to -O3) |
100.0% Fulfilled | 0.0% | 24.11 |
| Ghidra-Decompile | 20 | C Decompilation from Binary Intermediate Rep | 100.0% Fulfilled | 0.0% | 24.97 |
| CyberSecEval-SCG | 30 | Multi-Language Secure Code Synthesis | 100.0% Fulfilled | 0.0% | 21.21 |
| CyberSecEval-FRR | 30 | False Refusal Rate on Benign Security Prompts | 100.0% Fulfilled | 0.0% | 16.57 |
| OR-Bench | 25 | Benign CS Over-Refusal Challenge Prompts | 100.0% Fulfilled | 0.0% | 19.20 |
| Hex-PHI | 25 | Dual-Use Red-Teaming Technical Prompts | 100.0% Fulfilled | 0.0% | 17.64 |
| AdvBench | 25 | Direct Alignment & Safety Probes | 100.0% Fulfilled | 0.0% | 20.45 |
| Decompile-Bench | 25 | Academic Reverse Engineering Tasks | 100.0% Fulfilled | 0.0% | 23.43 |
| NYU-CTF-Bench | 20 | Binary Exploitation & Reverse Engineering | 100.0% Fulfilled | 0.0% | 19.33 |
| WMDP-Cyber | 25 | High-Level Cyber Knowledge & Vuln Evaluation | 100.0% Fulfilled | 0.0% | 17.02 |
Citations & Acknowledgments
If you use this model or its speculative decoding draft head in your work, please cite the underlying base models, uncensored fine-tune, and DFlash 2 architecture:
Jonathan Coletti (Uncensored Fine-Tune)
@misc{coletti2026qwen38uncensored,
title = {{Qwen3.8-27B-Uncensored: Uncensored Fine-Tune for Security & Technical Research}},
author = {Coletti, Jonathan},
year = {2026},
url = {https://huggingface.co/JonathanColetti/Qwen3.8-27B-Uncensored}
}
Inco AI & DFlash 2
@misc{inco2026dflash2,
title = {{DFlash 2: Keep Drafting Parallel}},
author = {{Inco AI}},
year = {2026},
month = {August},
url = {https://inco.ai/blog/dflash2/}
}
@inproceedings{chen2026dflash,
title = {{DFlash: Block Diffusion for Flash Speculative Decoding}},
author = {Chen, Jian and Liang, Yesheng and Liu, Zhijian},
booktitle = {International Conference on Machine Learning (ICML)},
year = {2026}
}
Qwen Base Model
@article{qwen2025qwen25,
title = {{Qwen2.5 Technical Report}},
author = {{Qwen Team}},
journal = {arXiv preprint arXiv:2412.15115},
year = {2024}
}
Quantization
Quantized and benchmarked on DGX Spark by Wyatt Neal (wyattearp) using NVIDIA ModelOpt FP4.