← back to catalog · registered 2026-08-22 13:56

pyros-vault/Qwen3.8-27B-Uncensored-NInfer

pyros-vault Qwen 27B multimodal second-order
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/pyros-vault%2FQwen3.8-27B-Uncensored-NInfer"
Response includes
  • classification m1
  • files 6
  • hub_downloads_all_time 51,601
  • author_summary 5 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M1
Primary method

Direct removal

No other method signals detected in this model.
Confidence
MEDIUM
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=0 (base model)
  • no specific method indicators - defaulting to M1 (most common)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
MEDIUM
Why we say so
primary_method=M1; difference-of-means is the reference extraction for M1/M3 (Arditi 2024)
Downloads · lifetime
52K
Likes
8
Model age
7w ago
created 2026-08-18
Downloads over time
Now54.9K→from86↑63,721%
020.1K40.2K60.4K86 on Aug 1954.9K on Oct 11AugSepOct
Aug 19 → Oct 11 · 48 snapshots · spans 53 days

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
apache-2.0
Tags
ninfer ninfer-4090 qwen3.8 multimodal speculative-decoding mtp quantized rtx-4090 sm89 abliterated uncensored red-teaming

Related

Total size
0 B
Files
6
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-08-21 14:58

Files by quantization

Auxiliary files 6 files 17.0 GB
Qwen3.8-27B-Uncensored.ninfer 17.0 GB 314e2812 download
LICENSE 11.3 KB 29f81d81 download
README.md 8.80 KB 60f9ff43 download
Qwen3.8-27B-Uncensored.ninfer.conversion.json 3.42 KB 4cdc98b4 download
.gitattributes 1.55 KB a8f306d1 download
SHA256SUMS 96.0 B 497d1f23 download

README current version from Hugging Face


library_name: ninfer
pipeline_tag: image-text-to-text
inference: false
license: apache-2.0
base_model:

  • orcarouter/Qwen3.8-27B-Uncensored
    base_model_relation: quantized
    tags:
  • ninfer
  • ninfer-4090
  • qwen3.8
  • multimodal
  • speculative-decoding
  • mtp
  • quantized
  • rtx-4090
  • sm89
  • abliterated
  • uncensored
  • red-teaming

Qwen3.8-27B-Uncensored for NInfer-4090

This repository contains a single-file, mixed-precision .ninfer conversion of
orcarouter/Qwen3.8-27B-Uncensored, built for
UDPSendToFailed/ninfer-4090 on a 24 GB RTX 4090 (sm_89).

This is a deployment-format conversion only. It does not add training, change the upstream behavior, or independently establish the upstream model's "uncensored" or "abliterated" characteristics. The artifact is not a Safetensors or GGUF checkpoint and is not intended for Transformers, llama.cpp, or unrelated NInfer forks.

Quick facts

Item Value
Artifact Qwen3.8-27B-Uncensored.ninfer
File size 18,210,531,328 bytes / 16.96 GiB
SHA-256 314e2812942b2b078f341530f6f5f5e46d79b08d668b77ed4d8f58e12fda41f4
NInfer identity qwen3.8-27b / groupwise-int
Conversion recipe qwen3_8_27b-v1
Direct source snapshot orcarouter/Qwen3.8-27B-Uncensored@adb2e501
Base model Qwen/Qwen3.8-27B
Target runtime UDPSendToFailed/ninfer-4090, feat/rtx-4090-sm89-native
Converted on Windows 11, RTX 4090 24 GB, Python 3.11.13, PyTorch 2.13.0+cu130

The conversion report records NInfer revision
6d3fd16.
The verification below used the later source revision
e4233d6, whose cached-token reporting change is merged upstream in
UDPSendToFailed/ninfer-4090#5.

What is inside the .ninfer bundle

The file is a self-contained runtime bundle: mixed-precision model weights plus six embedded frontend resources (tokenizer.json, tokenizer and generation configs, chat_template.jinja, and image/video processor configs). Q4G64_F16S, for example, means 4-bit weights quantized in groups of 64 with FP16 scales; the Q5/Q6 variants use the same group size, while W8G32_F16S uses 8-bit weights in groups of 32.

BF16 FP32 I32 Q4G64 Q5G64 Q6G64 W8G32
582 96 1 183 246 1 9

NInfer materializes only the components selected at startup. In the verified runs, resident weights were 15.92 GiB for baseline text, 16.19 GiB with Vision, and 16.67 GiB with MTP. The 16.96 GiB artifact size therefore is not the same thing as per-mode VRAM usage; KV state, workspaces, CUDA Graphs, and concurrency add further runtime memory.

Download

hf download pyros-vault/Qwen3.8-27B-Uncensored-NInfer `
  --local-dir .\Qwen3.8-27B-Uncensored-NInfer

Verify the artifact:

Get-FileHash '.\Qwen3.8-27B-Uncensored-NInfer\Qwen3.8-27B-Uncensored.ninfer' -Algorithm SHA256

Serving with the intended NInfer runtime

Build a current native sm_89 version of
UDPSendToFailed/ninfer-4090. .ninfer is a registered, model-bound format; release binaries or other forks without the matching Qwen3.8 target may not load this file.

The following serving examples use the same conservative memory profile as the CLI verification below: a 4,096-token KV capacity with CUDA Graphs disabled. Sampling and thinking behavior remain request-configurable.

Text with MTP4

.\ninfer-serve.exe '.\Qwen3.8-27B-Uncensored.ninfer' `
  --kv-dtype rk4v4-e8 `
  --spec mtp --draft-tokens 4 --lm-head-draft `
  --max-context 4096 --kv-capacity 4096 --prefill-chunk 512 `
  --preserve-thinking --no-cuda-graph

For baseline decoding, omit --spec mtp --draft-tokens 4 --lm-head-draft.

Vision with MTP4

.\ninfer-serve.exe '.\Qwen3.8-27B-Uncensored.ninfer' `
  --vision --vision-max-tokens 1024 `
  --kv-dtype rk4v4-e8 `
  --spec mtp --draft-tokens 4 --lm-head-draft `
  --max-context 4096 --kv-capacity 4096 --prefill-chunk 512 `
  --preserve-thinking --no-cuda-graph

The server exposes OpenAI-compatible /v1/chat/completions and /v1/responses endpoints plus Anthropic-compatible /v1/messages at http://127.0.0.1:8080. Prefix reuse is enabled by default; current upstream builds expose it as usage.prompt_tokens_details.cached_tokens.

This Qwen3.8 artifact supports MTP with one to five draft tokens. It does not contain a DFlash model; NInfer's DFlash path is specific to the 35B-A3B target.

Local verification results

These are bounded single-run integration measurements, not a benchmark suite. Tests used greedy decoding, thinking disabled, rk4v4-e8 KV, --max-context 4096, --kv-capacity 4096, --prefill-chunk 512, and --no-cuda-graph on one RTX 4090.

The baseline and MTP rows used the same prompt and produced the same 106-token answer, including the same Fibonacci implementation and assertions. This makes their decode-rate ratio meaningful for this one workload, but not a general performance guarantee.

Path Observed result Prompt / generated tokens Decode MTP statistics Planned device total
Baseline text Correct Fibonacci function and three assertions 37 / 106 46.68 tok/s off 16.36 GiB
MTP4 text Same output as baseline 37 / 106 155.42 tok/s (3.33x) 91.30% accepted; 4.65 tok/round; 0 fallback 17.13 GiB
Vision Exact NIFER VISION 731;3;左侧 428 / 14 46.32 tok/s off 16.65 GiB

The pinned NInfer fixture (image, message JSON) contains the title NIFER VISION 731, three red circles, and a blue square to the left of a green triangle. Its short decode rate is included for completeness and should not be compared directly with the coding run.

Conversion details

The dedicated converter preflighted 1,199 BF16 source tensors across 18 shards and emitted 1,118 runtime tensors plus six embedded tokenizer/template/processor resources. The 1,124 stored objects use a mixture of BF16, FP32, groupwise Q4/Q5/Q6, and W8 formats. This is a quantized deployment artifact, not a lossless copy of the BF16 checkpoint.

Conversion completed in 88.98 seconds on the RTX 4090. The complete manifest and environment are retained in Qwen3.8-27B-Uncensored.ninfer.conversion.json.

The converter report records the local source directory but not its Hugging Face revision. The adjacent Hugging Face download metadata pins the converted source files to commit adb2e5014317d59bb2093d90eaf3f8ef2b1975fe; the six direct frontend resources also matched NInfer's registered Qwen3.8 hashes. The direct source currently uses Hugging Face auto-gating, so users should review its access terms themselves.

Important limitations

  • Only the documented 4k profile is locally verified here. The upstream config declares 262,144 positions, but this card does not claim a long-context validation.
  • Vision is opt-in. NInfer defaults to an 8,192-token Vision scratch capacity; start with --vision-max-tokens 1024 on a 24 GB card and increase only when needed.
  • Quantization can change outputs. Quality and performance depend on the prompt, sampler, KV dtype, context length, and runtime revision.
  • Upstream behavior is not independently certified. "Uncensored" and "abliterated" are upstream labels, not a safety or capability evaluation by pyros-vault. Use appropriate isolation and review for red-team workloads.
  • No hosted inference. The Hugging Face Inference API cannot execute .ninfer files.

License and credits

The direct source and Qwen base declare Apache-2.0. A copy of the license is included in LICENSE. Users remain responsible for reviewing the direct source's terms and complying with applicable laws.

NInfer runtime code is not redistributed here; obtain it from its separately licensed
upstream repository.

Converted, tested, and packaged by pyros-vault.

README history 5 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-08-21docs: document NInfer bundle and source snapshotc2ce5e58.8 KB
    Loading...
  2. 2026-08-21Expand NInfer-4090 model card and verification63bb54f7.3 KB
    Loading...
  3. 2026-08-19Update README.mdfdc54e31.4 KB
    Loading...
  4. 2026-08-19Update README.md702ebe81.4 KB
    Loading...
  5. 2026-08-18Upload NInfer model, README and conversion metadata48e724d1.1 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration