← back to catalog · registered 2026-08-26 16:02

DevelopingDad/Qwen3.8-27B-NVFP4-Uncensored

DevelopingDad Qwen 9.2B multimodal second-order
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/DevelopingDad%2FQwen3.8-27B-NVFP4-Uncensored"
Response includes
  • classification m1
  • files 19
  • hub_downloads_all_time 112
  • author_summary 3 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M1
Primary method

Direct removal

No other method signals detected in this model.
Confidence
MEDIUM
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=0 (base model)
  • no specific method indicators - defaulting to M1 (most common)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
MEDIUM
Why we say so
primary_method=M1; difference-of-means is the reference extraction for M1/M3 (Arditi 2024)
Downloads · lifetime
112
82 last 30d - active
Likes
0
Model age
6w ago
created 2026-08-26
Downloads over time
Now155→from0↑0%
0571141710 on Aug 26155 on Oct 11AugSepOct
Aug 26 → Oct 11 · 47 snapshots · spans 46 days

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
apache-2.0
Languages
en zh
Tags
transformers safetensors qwen3_5 image-text-to-text qwen qwen3.8 nvfp4 fp8 fp8-kv uncensored abliterated vision-language

Related

Total size
20.4 GB
Files
19
Quantizations
1
Registered
2026-08-26 16:02
Last updated on HF
2026-08-26 15:47

Files by quantization

Auxiliary files 19 files 20.4 GB
model-00002-of-00003.safetensors 9.30 GB c8a5d0e3 download
model-00001-of-00003.safetensors 9.28 GB c5253293 download
model-00003-of-00003.safetensors 1.83 GB b642fd9e download
tokenizer.json 12.2 MB 0997f410 download
vocab.json 6.41 MB 0aa0ce06 download
merges.txt 3.20 MB a494e019 download
.quant_summary.txt 285 KB c7349f87 download
model.safetensors.index.json 194 KB c81e95a1 download
config.json 86.5 KB 38224275 download
hf_quant_config.json 53.6 KB 77838df4 download
LICENSE 11.3 KB f938136e download
chat_template.jinja 8.74 KB c0c686f9 download
README.md 5.13 KB ef226e27 download
.gitattributes 1.53 KB 52373fe2 download
processor_config.json 1.16 KB 33818c7f download
tokenizer_config.json 1.09 KB 18c88802 download
preprocessor_config.json 390 B 2ea84a43 download
video_preprocessor_config.json 385 B 3ba673a5 download
generation_config.json 214 B 8b9f95da download

README current version from Hugging Face


license: apache-2.0
base_model: orcarouter/Qwen3.8-27B-Uncensored
base_model_relation: quantized
pipeline_tag: image-text-to-text
library_name: transformers
language:

  • en
  • zh
    tags:
  • qwen
  • qwen3.8
  • nvfp4
  • fp8
  • fp8-kv
  • uncensored
  • abliterated
  • vision-language
  • function-calling
  • modelopt

Qwen3.8-27B-NVFP4-Uncensored

This is a ModelOpt mixed-precision derivative of
orcarouter/Qwen3.8-27B-Uncensored,
generated from source revision
9878936be9458522b5aeed0e13476bb8426f57f0.

This is not a pure all-NVFP4 checkpoint. Its weight path combines NVFP4,
FP8, and retained higher-precision tensors; the exact composition is described
below. The repository is intended to be a reproducible vLLM/ModelOpt artifact,
not a claim that every Transformers backend can execute the quantized layers
without matching ModelOpt support.

The model preserves the Qwen3.8 vision-language tower and MTP head. It is an
abliterated/uncensored model with substantially reduced built-in refusal
behavior. It is intended for controlled research, evaluation, and local
experimentation. Add your own moderation and access controls before any
user-facing deployment.

Quantization

  • MLP and language-model-head weights: W4A16_NVFP4, group size 16. The
    exported metadata identifies 193 such target layers.
  • Attention and linear-attention projections: ModelOpt FP8 mixed precision,
    covering 208 target layers.
  • KV cache: FP8 E4M3 with 32 exported scalar scale tensors (16 K and 16 V)
    from the data-calibrated ModelOpt kv_fp8 recipe. The audited scales are
    finite, positive, and non-unit.
  • Vision, MTP, and hybrid-state tensor families are retained in the unified
    checkpoint and excluded from the weight-quantization target map.
  • Export format: unified Hugging Face safetensors checkpoint.

The full export contains 2,033 indexed tensors across three safetensors
shards. ModelOpt metadata records MIXED_PRECISION weights and
kv_cache_quant_algo: FP8, produced with ModelOpt
0.47.0.dev81+ga2fbac7ba.

Calibration

The KV scales were calibrated with 256 pre-rendered, text-only examples at a
2,048-token calibration sequence length, batch size 1, and
enable_thinking=false. The corpus was generic rather than application-owned:
128 general instruction/chat rows, 64 code/technical rows, 32 structured-output
rows, and 32 longer-context rows. No images or videos were used for calibration.
This establishes a scale-aware FP8-KV export for the tested distribution; it is
not a universal multimodal or application-specific calibration claim.

vLLM serving

This artifact was validated with vllm/vllm-openai:v0.27.1. The relevant
starting flags are:

--quantization modelopt_fp4
--kv-cache-dtype fp8_e4m3
--max-model-len 262144
--trust-remote-code

On the validation stack, the requested modelopt_fp4 flag resolved to vLLM's
modelopt_mixed path. The resolved KV dtype was float8_e4m3fn.

For the tested Qwen XML tool-calling route, also use:

--enable-auto-tool-choice --tool-call-parser qwen3_xml
--default-chat-template-kwargs '{"enable_thinking":false}'

Validation notes

In the project's Experiment 014, the candidate returned 56/56 HTTP 200
responses with complete streaming [DONE] markers, zero stream parse errors,
and zero reasoning leaks. The explicit response-contract path passed 8/8.
The unchanged no-contract baseline passed 12/24 because the known
fenced-JSON and HH:MM formatting behaviors remained; these were output
contract misses, not cache-load failures. A short long-context load returned
24/24 complete streams at concurrency levels 1, 2, and 4; its diagnostic
fixture oracle passed 21/24. vLLM reported an allocator capacity of 2,491,134
FP8-KV tokens while the configured maximum sequence length was 262,144. The
allocator figure is a host/runtime capacity diagnostic, not a claim that the
model supports a 2.49-million-token context.

These are narrow, single-host exploratory results, not a general quality or
production-readiness claim.

Important caveats

  • On the project's NVIDIA GB10, vLLM uses the Marlin software-FP4 path because
    the GPU has no native FP4 computation support; compute-heavy performance may
    differ from a native-FP4 GPU.
  • The checkpoint does not include separate q-scale metadata. In the tested
    vLLM FP8 attention path, q scaling therefore falls back to the K scale;
    q/probability scales were not independently calibrated.
  • The calibration and validation were text-focused. Multimodal quality,
    application-wide quality, independent BF16-vs-FP8 throughput, and production
    readiness remain unproven. The tested persistent WebUI route continued to
    use BF16 KV as its correctness-first default.
  • The model's uncensored behavior means it may produce harmful or illegal
    content. Follow the Apache 2.0 license, applicable law, and your own safety
    requirements.

License and provenance

The Apache 2.0 license file from the derived checkpoint is included in this
repository. Users are responsible for complying with the base model's terms,
the derivative model's terms, and all applicable laws.

README history 2 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-08-26Clarify mixed precision and validation scopea29c79e5.1 KB
    Loading...
  2. 2026-08-26Add model card and validation caveats6a6a0d93.1 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration