← back to catalog · registered 2026-08-22 13:56

Vtuber-plan/Qwen3.8-27B-Uncensored-NVFP4

Vtuber-plan Qwen 12B
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/Vtuber-plan%2FQwen3.8-27B-Uncensored-NVFP4"
Response includes
  • classification m-uncensored
  • files 20
  • hub_downloads_all_time 2,012
  • author_summary 2 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M-U
Primary method

Uncensored (method unknown)

No other method signals detected in this model.
Confidence
LOW
Why this label 3 signals
Weak or ambiguous signals. Best guess based on catalog patterns; treat as tentative and check the evidence below.
  • 'uncensored' in name/tags but no 'abliterated' marker
  • method not identifiable from author declaration alone
  • may be DPO fine-tune, prompt engineering, or unknown technique
Refusal direction extraction

No specific extraction method could be identified for this model. The producer either did not document it or used a proprietary pipeline.

What is a refusal direction? →
Downloads · lifetime
2K
290 last 30d - stable
Likes
1
Model age
7w ago
created 2026-08-18
Downloads over time
Now2.1K→from1K↑107%
9641.4K1.8K2.2K1K on Aug 192.1K on Oct 11AugSepOct
Aug 19 → Oct 11 · 48 snapshots · spans 53 days

Metadata

License
other
Languages
en hi ja ko pt
Tags
transformers safetensors qwen3_5 image-text-to-text Qwen3.5 NVFP4 quantized modelopt text-generation conversational en hi

Related

Total size
19.2 GB
Files
20
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-08-20 03:22

Files by quantization

Auxiliary files 20 files 19.2 GB
model-00002-of-00005.safetensors 4.66 GB 9cf911de download
model-00003-of-00005.safetensors 4.64 GB 8f810481 download
model-00001-of-00005.safetensors 4.63 GB 5fce46f1 download
model-00004-of-00005.safetensors 4.55 GB 9e4e31c3 download
model-00005-of-00005.safetensors 710 MB 6e2f47c8 download
tokenizer.json 12.2 MB 0997f410 download
vocab.json 6.41 MB 0aa0ce06 download
merges.txt 3.20 MB a494e019 download
model.safetensors.index.json 232 KB 2cb79163 download
quality_comparison.png 84.3 KB 82c10011 download
chat_template.jinja 21.6 KB d996db59 download
tokenizer_config.json 17.5 KB 5de744b3 download
config.json 14.7 KB bd403404 download
LICENSE 11.3 KB f938136e download
hf_quant_config.json 9.78 KB 9c88a1e7 download
README.md 3.11 KB d9018fe0 download
.gitattributes 1.57 KB 0d353d45 download
preprocessor_config.json 390 B 2ea84a43 download
video_preprocessor_config.json 385 B 3ba673a5 download
generation_config.json 214 B 0bc3addd download

README current version from Hugging Face


language:

  • en
  • hi
  • ja
  • ko
  • pt
    license: other
    library_name: transformers
    pipeline_tag: text-generation
    tags:
  • Qwen3.5
  • NVFP4
  • quantized
  • modelopt

Qwen3.8-27B-Uncensored-NVFP4

NVFP4 (4-bit per-block) quantized version of the Qwen3.8-27B Uncensored model, produced with
NVIDIA TensorRT Model Optimizer.

Quality Comparison (Original vs BF16 vs NVFP4)

Preliminary evaluation with lm_eval + sglang (greedy, temperature 0). MMLU / CMMLU / C-Eval use a
20-question-per-subtask sample; GSM8K uses the full test set. We compare:
the original Qwen3.8-27B, our Uncensored BF16 (before quantization), and the
Uncensored NVFP4 (4-bit) checkpoint.

Quality comparison

Benchmark Original Uncensored BF16 Uncensored NVFP4 NVFP4 − BF16 Std Err
MMLU (sample) 0.8368 0.8307 0.8316 +0.0009 ±0.011
CMMLU (sample) 0.7075 0.7716 0.7761 +0.0045 ±0.011
C-Eval (sample) 0.7609 0.7837 0.7956 +0.0119 ±0.013
GSM8K (strict) 0.7036 0.7627 0.7786 +0.0159 ±0.012
GSM8K (flexible) 0.7263 0.7870 0.8014 +0.0144 ±0.011

Key takeaways:

  • Quantization preserves quality. Unlike the Original vs Uncensored gap, NVFP4 tracks BF16
    almost exactly — all NVFP4−BF16 deltas are within ±1 standard error, i.e. no measurable
    degradation
    from 4-bit quantization (~4× weight compression).
  • Uncensored is generally stronger on these benchmarks than the stock original. The uncensored
    checkpoint scores higher on CMMLU / C-Eval / GSM8K and comparable on MMLU. This is not caused by
    quantization — the same difference already exists between the uncensored BF16 model and the
    stock original, so it reflects the uncensoring/finetuning itself.
  • The small positive NVFP4−BF16 deltas are within noise and should not be read as
    "NVFP4 is better than BF16"; the practical takeaway is that 4-bit quantization is effectively
    lossless on these tasks.

Quantization

  • Format: NVFP4 weights, FP8 KV cache, group size 16
  • Tool: NVIDIA ModelOpt 0.45.0
  • Calibration: ultrachat_200k + nvidia/Nemotron-SFT-Multilingual-v2 (code/math/stem across Japanese, Korean, Portuguese, Hindi)
  • Excluded modules: lm_head, embeddings, linear-attention conv1d/in_proj_a/in_proj_b, and MTP layers

Shards

All safetensors shards are ≤ 5 GB (5 shards), so the repository can be cloned and uploaded without
Hugging Face large-file (>5 GB) restrictions.

Usage

Load with Hugging Face transformers:

from transformers import AutoModelForCausalLM, AutoTokenizer

model_id = "Vtuber-plan/Qwen3.8-27B-Uncensored-NVFP4"
tokenizer = AutoTokenizer.from_pretrained(model_id, trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained(model_id, trust_remote_code=True)

Note: this NVFP4 checkpoint is intended for deployment with frameworks that support the
ModelOpt NVFP4 format (e.g. TensorRT-LLM). Plain transformers/BF16 inference will not dequantize
it natively and requires the corresponding quantization backend.

README history 4 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-08-20Expand quality comparison to include Original Qwen3.8-27B columnb0f5a993.1 KB
    Loading...
  2. 2026-08-19Add preliminary BF16 vs NVFP4 quality comparison chart + tablee70a4bc2.5 KB
    Loading...
  3. 2026-08-18Add Qwen3.8-27B-Uncensored-NVFP4 (ModelOpt NVFP4, ultrachat_200k + Nemotron-S...aa7438a1.4 KB
    Loading...
  4. 2026-08-18initial commitd10fd0b28 B
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration