← back to catalog · registered 2026-08-25 19:02

Vtuber-plan/Huihui-Qwen3.8-27B-abliterated-NVFP4

Vtuber-plan Qwen 12B second-order
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/Vtuber-plan%2FHuihui-Qwen3.8-27B-abliterated-NVFP4"
Response includes
  • classification m1
  • files 19
  • hub_downloads_all_time 1,418
  • author_summary 2 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M1
Primary method

Direct removal

No other method signals detected in this model.
Confidence
MEDIUM
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=0 (base model)
  • no specific method indicators - defaulting to M1 (most common)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
MEDIUM
Why we say so
primary_method=M1; difference-of-means is the reference extraction for M1/M3 (Arditi 2024)
Downloads · lifetime
1K
265 last 30d - stable
Likes
1
Model age
6w ago
created 2026-08-25
Downloads over time
Now1.5K→from308↑395%
2477141.2K1.6K308 on Aug 261.5K on Oct 11AugSepOct
Aug 26 → Oct 11 · 47 snapshots · spans 46 days

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
apache-2.0
Languages
en fr multilingual
Tags
transformers safetensors qwen3_5 image-text-to-text Qwen3.8 abliterated NVFP4 quantized modelopt reasoning agent text-generation

Related

Total size
19.2 GB
Files
19
Quantizations
1
Registered
2026-08-25 19:02
Last updated on HF
2026-08-26 14:12

Files by quantization

Auxiliary files 19 files 19.2 GB
model-00002-of-00005.safetensors 4.66 GB fc17c2af download
model-00003-of-00005.safetensors 4.64 GB afa3eb8d download
model-00001-of-00005.safetensors 4.63 GB 8c1c9252 download
model-00004-of-00005.safetensors 4.55 GB daaa0242 download
model-00005-of-00005.safetensors 710 MB 9ce944d5 download
tokenizer.json 19.1 MB 06b95093 download
.quant_summary.txt 328 KB d7df9823 download
model.safetensors.index.json 232 KB 2cb79163 download
chat_template.jinja 26.5 KB a7d89930 download
config.json 14.7 KB bd403404 download
LICENSE 11.3 KB f938136e download
hf_quant_config.json 9.78 KB 9c88a1e7 download
README.md 5.75 KB d9a40357 download
gate9_cases.json 3.74 KB 294d16fd download
.gitattributes 1.53 KB 52373fe2 download
tokenizer_config.json 1.09 KB 77053000 download
preprocessor_config.json 390 B 2ea84a43 download
video_preprocessor_config.json 385 B 3ba673a5 download
generation_config.json 214 B 0bc3addd download

README current version from Hugging Face


language:

  • en
  • fr
  • multilingual
    license: apache-2.0
    library_name: transformers
    pipeline_tag: text-generation
    tags:
  • Qwen3.8
  • abliterated
  • NVFP4
  • quantized
  • modelopt
  • reasoning
    base_model:
  • huihui-ai/Huihui-Qwen3.8-27B-abliterated
  • Qwen/Qwen3.8-27B

Huihui-Qwen3.8-27B-abliterated-NVFP4

NVFP4 (4-bit) quantization of huihui-ai/Huihui-Qwen3.8-27B-abliterated,
produced with NVIDIA TensorRT Model Optimizer. The calibration set was extended with
self-generated long chain-of-thought samples that terminate properly with </think>,
specifically so that 4-bit quantization noise does not degrade thinking-phase stop behavior.

Chat template (important)

This repo ships the fixed template from
froggeric/Qwen-Fixed-Chat-Templates (v22.4),
replacing the official Qwen3.8 template. Key differences:

  • reasoning_effort defaults to medium (zero injected tokens). The official template defaults to
    xhigh, which routinely exhausts the token budget on reasoning before any answer is produced.
  • Restores a working non-reasoning mode (enable_thinking=false / reasoning_effort="none").
  • Extracts in-content <think> blocks in history without duplicating tags or poisoning context with
    empty <think></think> blocks.
  • Robust tool-call rendering (no crashes on stringified JSON arguments); client effort aliases
    (high/max/ultracode → xhigh, minimal → low, none/off → disabled).

xhigh remains available via chat_template_kwargs: {"reasoning_effort": "xhigh"} — see the
validation notes below for when that is a bad idea.

Quantization

  • Format: NVFP4 weights (group size 16) + FP8 KV cache
  • Tool: NVIDIA ModelOpt 0.45.0 (hf_ptq.py)
  • Excluded modules: lm_head, embeddings, linear-attention conv1d / in_proj_a / in_proj_b, MTP layers, vision tower
  • Calibration (4771 samples, calib_seq=6144):
    1. HuggingFaceH4/ultrachat_200k — 2048 samples
    2. nvidia/Nemotron-SFT-Multilingual-v2 — 2048 samples
    3. 675 self-generated long-CoT samples (this repo's addition): user-turn prefixes sampled from
      Vtuber-plan/sharegpt-cleaned (EN 640 pool / FR 256 pool)
      were answered by the unquantized base model at reasoning_effort=xhigh (1/3 greedy, 2/3 temp 0.7).
      Only responses that terminated properly with </think> were kept (675/896; rejected: no close tag,
      truncated, sub-200-char reasoning, or repetitive degeneration). calib_seq was raised 512 → 6144 so
      the <think> → </think> → answer transition stays inside the calibration window.

Validation: xhigh thinking-runaway gate (9 cases)

A 9-case gate (gate9_cases.json, included in this repo) probes runaway thinking:
FR/EN long-form tasks × temperature {0, 0.7} × reasoning_effort=xhigh, max_tokens 32768.
An 8-gram decile copy-rate analysis distinguishes genuine extended thinking (low copy rate)
from repetition loops (copy rate → 100%).

Build Closed-think & stopped Hit token cap (runaway) Hard 100% loops Mean think len
This NVFP4 7/9 2 (greedy) 0 49.5k chars
Base BF16 (huihui) 7/9 1 (greedy) + 1 stopped-unclosed 0 47.6k chars
Existing qwen3.8-27b deployment 2/9 7 2 (both greedy) —

Readings:

  • No quantization regression: paired per-case deltas vs the BF16 base are mixed in sign and small
    overall (mean 49.5k vs 47.6k chars; each build has exactly one greedy-decode cap-out and one
    unclosed-think anomaly). On deliberately brutal xhigh prompts both builds overthink massively —
    that is base-model behavior, not a 4-bit artifact.
  • Greedy decoding (temp 0) is the danger zone: every terminal runaway in every build occurred at
    temperature 0. Once a repetition loop forms, greedy decoding cannot escape it. Use temp ≥ 0.7
    and/or repetition_penalty ≈ 1.05–1.1 for xhigh workloads.
  • The deployed stock endpoint failed far harder (0/9, two hard loops) than either local build.

Benchmarks (GSM8K / MMLU / CMMLU / C-Eval) will be added here.

Usage

Recommended server (sglang, single 80 GB GPU):

CUDA_VISIBLE_DEVICES=0 python -m sglang.launch_server \
  --model-path <this-checkpoint> \
  --quantization modelopt \
  --trust-remote-code \
  --mem-fraction-static 0.8
import requests
requests.post("http://127.0.0.1:30000/v1/chat/completions", json={
    "model": "Huihui-Qwen3.8-27B-abliterated-NVFP4",
    "messages": [{"role": "user", "content": "..."}],
    "temperature": 0.7,
    "max_tokens": 8192,
    # optional, only if you really want maximal deliberation:
    # "chat_template_kwargs": {"reasoning_effort": "xhigh"},
})

Plain transformers/BF16 inference will not dequantize this checkpoint; use a runtime with ModelOpt
NVFP4 support (sglang --quantization modelopt, TensorRT-LLM).

Deployment recommendations

  1. Leave reasoning_effort at the default (medium) — the shipped template makes this safe.
  2. If xhigh is required: temperature > 0, repetition_penalty ≈ 1.05, and a max_tokens cap
    (8k–16k) with retry-at-medium on timeout.
  3. Only GPU 0 was used for all development and validation of this checkpoint.

Files

5 safetensors shards, each ≤ 5 GB (~19.5 GB total). gate9_cases.json is the release gate set.

License

Apache 2.0 (inherited from the base model). Quantization and calibration were performed on top of
huihui-ai/Huihui-Qwen3.8-27B-abliterated;
credit to huihui-ai for the abliterated base and to froggeric
for the fixed chat template.

README history 4 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-08-26Add files using upload-large-folder tool43aa7ff6.1 KB
    Loading...
  2. 2026-08-25Add files using upload-large-folder toolb8c87515.2 KB
    Loading...
  3. 2026-08-25Add files using upload-large-folder tool55db5905.7 KB
    Loading...
  4. 2026-08-25initial commitd7af65428 B
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration