← back to catalog · registered 2026-08-22 13:56

lyf/Qwen3.8-27B-Huihui-Abliterated-NInfer-NVFP4

lyf Qwen 27B multimodal second-order
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/lyf%2FQwen3.8-27B-Huihui-Abliterated-NInfer-NVFP4"
Response includes
  • classification m1
  • files 17
  • hub_downloads_all_time 8,451
  • author_summary 10 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M1
Primary method

Direct removal

No other method signals detected in this model.
Confidence
MEDIUM
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=0 (base model)
  • no specific method indicators - defaulting to M1 (most common)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
MEDIUM
Why we say so
primary_method=M1; difference-of-means is the reference extraction for M1/M3 (Arditi 2024)
Downloads · lifetime
8K
Likes
17
Model age
7w ago
created 2026-08-19
Downloads over time
Now13.3K→from229↑5,726%
04.9K9.8K14.7K229 on Aug 1913.3K on Oct 11AugSepOct
Aug 19 → Oct 11 · 48 snapshots · spans 53 days

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
apache-2.0
Languages
en zh
Tags
ninfer qwen3_5 qwen3.8 nvfp4 fp8 mixed-precision blackwell rtx-5090 sm120 multimodal vision mtp

Related

Total size
0 B
Files
17
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-08-19 03:55

Files by quantization

Auxiliary files 17 files 20.0 GB
qwen3_8_27b_nvfp4.ninfer 20.0 GB f21f308d download
artifact-inspect.log 126 KB f9facbad download
qwen3_8_27b_nvfp4.ninfer.conversion.json 4.27 KB 1d3293e0 download
README.md 4.00 KB f9e19c23 download
infer-text.log 2.20 KB 40f4b8df download
infer-vision.log 2.20 KB b70a55e6 download
infer-behavior.log 2.20 KB 1d6d3458 download
.gitattributes 1.54 KB 7e513d68 download
ninfer-nvfp4-preflight.log 1.50 KB 1eac895b download
SHA256SUMS 1.27 KB e79a2cd3 download
BUILD_MANIFEST.json 706 B a8c6bb34 download
infer-behavior.out 697 B ed41360c download
infer-text.out 600 B 2228ba69 download
infer-vision.out 566 B 58a6693e download
VALIDATION_REPORT.json 416 B 445ee238 download
python-versions.txt 337 B 831d5521 download
ninfer.commit.txt 41.0 B 586e5ba7 download

README current version from Hugging Face


license: apache-2.0
license_link: https://huggingface.co/Qwen/Qwen3.8-27B/blob/main/LICENSE
base_model: huihui-ai/Huihui-Qwen3.8-27B-abliterated
base_model_relation: quantized
pipeline_tag: image-text-to-text
library_name: ninfer
language:

  • en
  • zh
    tags:
  • qwen3_5
  • qwen3.8
  • ninfer
  • nvfp4
  • fp8
  • mixed-precision
  • blackwell
  • rtx-5090
  • sm120
  • multimodal
  • vision
  • mtp
  • reasoning
  • abliterated

Huihui Qwen3.8-27B Abliterated — NInfer Mixed FP8/NVFP4

A complete NInfer artifact derived from huihui-ai/Huihui-Qwen3.8-27B-abliterated revision d42ca8978c5a66e92c3446d46e8adfe03ef692ff.

The artifact combines NInfer's fixed FP8 + NVFP4 Qwen3.8 allocation with same-source Huihui BF16 norms, embeddings, vision tower, and MTP tensors. No official Qwen or unrelated behavioral checkpoint supplies model weights.

Download and run

hf download lyf/Qwen3.8-27B-Huihui-Abliterated-NInfer-NVFP4 qwen3_8_27b_nvfp4.ninfer   --local-dir ./huihui-ninfer

git clone https://github.com/Neroued/ninfer.git
cd ninfer
git checkout a99407c63fc5bbd25d9fb597cbb8ab352bdb01ef
cmake -S . -B build -G Ninja   -DCMAKE_BUILD_TYPE=Release   -DCMAKE_CUDA_ARCHITECTURES=120a
cmake --build build -j"$(nproc)"

./build/apps/ninfer ../huihui-ninfer/qwen3_8_27b_nvfp4.ninfer   --prompt "Explain FP8 and NVFP4 briefly."   --max-new 128 --max-context 8192 --kv-dtype int8 --no-thinking

OpenAI-compatible server

./build/apps/ninfer-serve ../huihui-ninfer/qwen3_8_27b_nvfp4.ninfer   --host 127.0.0.1 --port 8000   --model-id qwen38-huihui-ninfer-nvfp4   --max-context 204800 --kv-capacity 204800   --max-concurrency 1 --prefill-chunk 4096   --kv-dtype int8 --spec mtp --draft-tokens 3   --default-max-tokens 16384 --preserve-thinking

This 204.8K server profile was startup- and API-smoke-tested on one RTX 5090. Long-context quality and maximum sustained prompt length remain workload-dependent.

Mixed allocation

Matrices Format
Full attention and Gated DeltaNet projections FP8 E4M3, per-output-row BF16 scales
lm_head and layers 56–63 MLP FP8 E4M3
Layers 0–55 MLP NVFP4, group size 16
Norms, GDN state, embeddings, vision, MTP Same-source Huihui tensors

NInfer source preflight passed with:

FP8 source matrices: 233
NVFP4 source matrices: 168
source fields: 1587
F8_E4M3: 401
BF16: 682
U8: 168
F32: 336
NINF_PREFLIGHT_OK

The FP8 matrices are deterministic row-scaled E4M3 exports from Huihui BF16. The NVFP4 matrices come from the separately calibrated Huihui ModelOpt NVFP4 checkpoint (CNN/DailyMail 3.0.0, 20 × 8192 tokens).

Artifact inventory

NInfer commit: a99407c63fc5bbd25d9fb597cbb8ab352bdb01ef
model_id: qwen3.8-27b
weights_id: nvfp4
objects: 1124 (1118 tensors, 6 resources)
artifact bytes: 21,492,695,040
SHA256: f21f308d3b23ccd627071cd015e413db08deee4356643900518e2b251750fdc2

Key formats:

BF16: 534
FP32: 208
FP8 row-scaled: 146
NVFP4: 112

Runtime validation

Hardware: RTX 5090 / SM120, driver 610.43.02, CUDA 13.1, 450 W cap.

Text smoke test:

prompt tokens: 27
generated tokens: 114
prefill: 792.69 tok/s
decode: 74.83 tok/s
overall: 73.83 tok/s
GPU weights: 18.98 GiB

Real-image test passed. For the test image, the model correctly identified a red square in the upper-left and a blue circle in the lower-right. Vision decode measured 74.82 tok/s with 19.25 GiB of GPU weights loaded.

A benign ablation/refusal-vector research prompt also received a direct technical answer. This is a smoke test of behavioral continuity, not a comprehensive behavioral benchmark.

Included evidence

  • qwen3_8_27b_nvfp4.ninfer.conversion.json
  • ninfer-nvfp4-preflight.log
  • artifact-inspect.log
  • text, vision, and behavior outputs/logs
  • BUILD_MANIFEST.json, VALIDATION_REPORT.json, SHA256SUMS

Intended use

This is an abliterated behavioral derivative intended for model research and local inference. Users are responsible for downstream use and applicable policies.

README history 1 version

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-08-19Add files using upload-large-folder tool18144694 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration