← back to catalog · registered 2026-09-28 10:57

coolthor/Qwen3.8-27B-Uncensored-NVFP4-FastLLM

coolthor 27B multimodal second-order
curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/coolthor%2FQwen3.8-27B-Uncensored-NVFP4-FastLLM"
Response includes
  • classification m-uncensored
  • files 17
  • author_summary 7 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M-U
Primary method

Uncensored (method unknown)

No other method signals detected in this model.
Confidence
LOW
Why this label 3 signals
Weak or ambiguous signals. Best guess based on catalog patterns; treat as tentative and check the evidence below.
  • 'uncensored' in name/tags but no 'abliterated' marker
  • method not identifiable from author declaration alone
  • may be DPO fine-tune, prompt engineering, or unknown technique
Refusal direction extraction

No specific extraction method could be identified for this model. The producer either did not document it or used a proprietary pipeline.

What is a refusal direction? →
Downloads · 30-day
0
Likes
0
Model age
today
created 2026-09-28

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
apache-2.0
Languages
en zh
Tags
transformers safetensors qwen3_5 image-text-to-text qwen3.8 nvfp4 compressed-tensors fastllm rtx-2080-ti turing abliterated uncensored

Related

Total size
19.1 GB
Files
17
Quantizations
1
Registered
2026-09-28 10:57
Last updated on HF
2026-09-28 10:29

Files by quantization

Auxiliary files 17 files 19.2 GB
model-00001-of-00002.safetensors 9.27 GB d9d57d98 download
model-00002-of-00002.safetensors 9.08 GB 95d07688 download
model-mtp-extra.safetensors 810 MB 7a9fd4ee download
tokenizer.json 12.2 MB 0997f410 download
vocab.json 6.41 MB 0aa0ce06 download
merges.txt 3.20 MB a494e019 download
model.safetensors.index.json 272 KB 92e4d5a6 download
tokenizer_config.json 17.5 KB 5de744b3 download
LICENSE 11.3 KB f938136e download
chat_template.jinja 8.74 KB c0c686f9 download
README.md 5.60 KB 10080ea7 download
config.json 4.95 KB 6955b5a0 download
hf_quant_config.json 3.28 KB a306db26 download
.gitattributes 1.53 KB 52373fe2 download
preprocessor_config.json 390 B 2ea84a43 download
video_preprocessor_config.json 385 B 3ba673a5 download
generation_config.json 202 B 023756cf download

README current version from Hugging Face


license: apache-2.0
base_model: orcarouter/Qwen3.8-27B-Uncensored
base_model_relation: quantized
pipeline_tag: image-text-to-text
library_name: transformers
language:

  • en
  • zh
    tags:
  • qwen3.8
  • nvfp4
  • compressed-tensors
  • fastllm
  • rtx-2080-ti
  • turing
  • abliterated
  • uncensored
  • vision-language
  • mtp

Qwen3.8-27B-Uncensored-NVFP4-FastLLM

An NVFP4 build of orcarouter/Qwen3.8-27B-Uncensored with every linear layer in NVFP4. It is 20.6 GB, down from 55.6 GB in BF16 and 29 GB in FP8.

I made it to serve the model on two RTX 2080 Ti 22 GB cards with FastLLM. Turing has no FP4 hardware. FastLLM stores the NVFP4 weights as-is and unpacks them to FP16 inside the kernel. Decoding is limited by weight reads, so the smaller file is faster on these cards even with the extra unpacking.

Write-up with the full numbers, the launch command and the debugging notes: NVFP4 Without FP4 Hardware: Qwen3.8-27B on Two RTX 2080 Tis at 151.7 tok/s (中文: 2080 Ti 沒有 FP4 也能跑 NVFP4).

⚠️ Inherited disclaimer

The base model has had its safety alignment largely removed by abliteration. Quantizing it does not change that. Everything in the base model's disclaimer applies here. The model will comply with harmful requests. It is meant for research and controlled use, you are responsible for how you use it, and you should add your own moderation before exposing it to anyone else.

What's in it

Weights Format
All linear layers: attention q/k/v/o, GDN projections (in_proj_qkv/z/a/b, out_proj), MLP NVFP4 (E2M1 values, one FP8-E4M3 scale per 16 values, one FP32 global scale per tensor)
lm_head, embed_tokens, vision tower, conv1d, MTP head BF16, unchanged
  • The layout matches lyf/Qwen3.8-27B-Heretic-ARA-NVFP4-MTP-VL exactly: tensor names, dtypes, shapes, shard map, quantization_config and hf_quant_config.json (2,687 tensors, total_size 20,558,935,392). Only the base model differs.
  • As in ModelOpt's recipe, q/k/v share one global scale per layer, the four GDN input projections share one, and MLP gate/up share one.
  • No calibration: input_global_scale is set to 1.0. FastLLM on sm_75 computes with FP16 activations and does not read it. If you run this on an engine that does W4A4 with real activation scales, expect worse results than a calibrated checkpoint.

How it was made

I converted it with a numpy-only script that runs on the CPU, no GPU needed: qwen38-nvfp4-convert.py.

hf download orcarouter/Qwen3.8-27B-Uncensored --local-dir ./bf16
hf download lyf/Qwen3.8-27B-Heretic-ARA-NVFP4-MTP-VL --include "*.json" "recipe.yaml" --local-dir ./nvfp4-ref
python3 qwen38-nvfp4-convert.py ./bf16 ./out --reference ./nvfp4-ref

To check the converter, I dequantized one attention, one GDN and one MLP layer of lyf's release to BF16 and quantized them again. The packed weights, block scales and global scales came out byte-identical to lyf's.

Results on two RTX 2080 Ti 22 GB (TP=2, FastLLM + DFlash2)

Same machine, same FastLLM build, temperature 0, 512 tokens, median of 3 runs, tok/s:

code math prose no speculation (code)
orcarouter FP8 116.1 135.5 74.2 32.9
this NVFP4 151.7 178.8 101.9 45.1

With fp8 KV cache and --max_batch 4: 2 concurrent streams run at 107.8 tok/s each, 4 streams at 74.3 each (about 297 combined).

140-question eval (GSM8K 50, MATH-500 30, MuSR 20, HumanEval+ 40):

no thinking thinking (effort low)
orcarouter FP8 132 127
this NVFP4 131 128

Trade-off: prefilling a 54K-token prompt takes 50.6 s instead of 40.5 s. Prefill is compute-bound, so the unpacking is pure overhead there.

I have only tested this with FastLLM on sm_75. I have not tested vLLM, SGLang or Blackwell GPUs.

Serving with FastLLM

Tested on FastLLM at upstream a2bf07fd with d4b04876 reverted and PRs #749 and #756 merged (build steps in the previous post). The draft model is z-lab/Qwen3.8-27B-DFlash2.

export CUDA_VISIBLE_DEVICES=0,1
export FASTLLM_CUDA_GRAPH=0
export FASTLLM_DRAFT_QUANT=nvfp4
export FASTLLM_COOPERATIVE_LONG_PREFILL=1   # from PR #756; leave out on stock FastLLM

ftllm server -p ./Qwen3.8-27B-Uncensored-NVFP4-FastLLM \
  --tp 2 --max_batch 4 --chunked_prefill_size 4096 \
  --gpu_mem_ratio 0.95 --kv_cache_dtype fp8_e4m3 \
  --speculative_algorithm dflash \
  --speculative_draft_model_path ./Qwen3.8-27B-DFlash2 \
  --speculative_num_draft_tokens 8 \
  --prefix_cache true --port 8080

Don't pass --tokens: with it set, FastLLM clamps --max_batch back to 1 for this model.

Credits

Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.