← back to catalog · registered 2026-08-22 13:56

philbert440/Qwen3.8-27B-Uncensored-Aggressive-NVFP4

philbert440 Qwen 20B multimodal second-order
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/philbert440%2FQwen3.8-27B-Uncensored-Aggressive-NVFP4"
Response includes
  • classification m1
  • files 15
  • hub_downloads_all_time 1,347
  • author_summary 20 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M1
Primary method

Direct removal

No other method signals detected in this model.
Confidence
MEDIUM
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=0 (base model)
  • no specific method indicators - defaulting to M1 (most common)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
MEDIUM
Why we say so
primary_method=M1; difference-of-means is the reference extraction for M1/M3 (Arditi 2024)
Downloads · lifetime
1K
536 last 30d - stable
Likes
3
Model age
8w ago
created 2026-08-15
Downloads over time
Now1.5K→from543↑171%
4978531.2K1.6K543 on Aug 191.5K on Oct 11AugSepOct
Aug 19 → Oct 11 · 48 snapshots · spans 53 days

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Variants by this author 3 formats · 2K downloads combined

The same weights this author released in different packaging. Pick the format that matches your runtime.

Metadata

License
apache-2.0
Tags
safetensors qwen3_5 compressed-tensors vllm uncensored abliterated qwen3 nvfp4 image-text-to-text conversational base_model:philbert440/Qwen3.8-27B-Uncensored-Aggressive base_model:quantized:philbert440/Qwen3.8-27B-Uncensored-Aggressive

Related

Total size
24.6 GB
Files
15
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-08-19 01:13

Files by quantization

Auxiliary files 15 files 24.6 GB
model-00001-of-00002.safetensors 18.6 GB f98e4784 download
model-00002-of-00002.safetensors 5.16 GB 96229119 download
model-mtp.safetensors 810 MB 90fa0e3e download
tokenizer.json 19.1 MB f399b3cd download
model.safetensors.index.json 174 KB 91139230 download
config.json 26.3 KB b7c9e16a download
recipe.yaml 10.4 KB 6e3841af download
chat_template.jinja 8.74 KB c0c686f9 download
README.md 1.65 KB c04f68fb download
.gitattributes 1.53 KB 52373fe2 download
tokenizer_config.json 1.10 KB d1a20cc3 download
preprocessor_config.json 390 B 2ea84a43 download
video_preprocessor_config.json 385 B 3ba673a5 download
generation_config.json 214 B 8b9f95da download
variant.json 194 B eb8765d1 download

README current version from Hugging Face


base_model: philbert440/Qwen3.8-27B-Uncensored-Aggressive
license: apache-2.0
tags: [compressed-tensors, vllm, uncensored, abliterated, qwen3, nvfp4]
pipeline_tag: image-text-to-text

Qwen3.8-27B-Uncensored-Aggressive — NVFP4

NVFP4 quant of Qwen3.8-27B-Uncensored-Aggressive (α=1.15 recipe update), compressed-tensors
NVFP4A16 (E2M1 4-bit weights, FP8-E4M3 scales @ group 16), for NVIDIA V100 (SM70) under
1Cat-vLLM. Vision tower and the grafted MTP head (bf16) are preserved.

About this update (α=1.15)

Recipe update to the Aggressive line: the previous build ablated at α≈1.24, which a larger benchmark
sweep showed over-ablates past the ~1.15 quality peak. This build uses α=1.15 — more open and
better on every measured axis.

Evaluation (bf16 parent, larger-sample, thinking mode, Claude-judged)

openness ↑ confab ↓ factual ↑ gsm8k ↑
stock base (censored) 0.08 0.75 1.00 0.85
Aggressive (α=1.15) 0.88 0.725 1.00 0.85
previous Aggressive (α≈1.24) 0.80 0.80 1.00 0.817

Quant: variant-C — GPTQ NVFP4A16 targeting [Linear], GDN in_proj_qkv/in_proj_z kept fp16,
768 CoT + 256 wiki calibration, actorder=weight, MSE observer. MTP head carried verbatim in bf16.

Serve (1Cat-vLLM, 2× V100, TP2)

--kv-cache-dtype fp8_e5m2, MTP speculative decoding, --gpu-memory-utilization tuned to KV/context budget.
Serves on V100/SM70 via 1Cat prepare_nvfp4_linear (min capability 70; not the modelopt path).

Note

Uncensored / de-refused. Use responsibly and in compliance with applicable law.

README history 4 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-08-19Upload folder using huggingface_hub28acca81.7 KB
    Loading...
  2. 2026-08-15Fix tokenizer: drop leftover truncation block, restore upstream tokenizer.jso...5adcc702.7 KB
    Loading...
  3. 2026-08-15Upload README.md with huggingface_hub36104c52 KB
    Loading...
  4. 2026-08-15Upload folder using huggingface_hub8a22e601.4 KB
    Loading...

Discussions 1 thread

  1. 2026-08-21NVFP4 vs W4A16open2 💬#1
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration