language:
- en
- fr
- multilingual
license: apache-2.0
library_name: transformers
pipeline_tag: text-generation
tags: - Qwen3.8
- abliterated
- NVFP4
- quantized
- modelopt
- reasoning
base_model: - huihui-ai/Huihui-Qwen3.8-27B-abliterated
- Qwen/Qwen3.8-27B
Huihui-Qwen3.8-27B-abliterated-NVFP4
NVFP4 (4-bit) quantization of huihui-ai/Huihui-Qwen3.8-27B-abliterated,
produced with NVIDIA TensorRT Model Optimizer. The calibration set was extended with
self-generated long chain-of-thought samples that terminate properly with </think>,
specifically so that 4-bit quantization noise does not degrade thinking-phase stop behavior.
Chat template (important)
This repo ships the fixed template from
froggeric/Qwen-Fixed-Chat-Templates (v22.4),
replacing the official Qwen3.8 template. Key differences:
reasoning_effortdefaults tomedium(zero injected tokens). The official template defaults toxhigh, which routinely exhausts the token budget on reasoning before any answer is produced.- Restores a working non-reasoning mode (
enable_thinking=false/reasoning_effort="none"). - Extracts in-content
<think>blocks in history without duplicating tags or poisoning context with
empty<think></think>blocks. - Robust tool-call rendering (no crashes on stringified JSON arguments); client effort aliases
(high/max/ultracode→xhigh,minimal→low,none/off→ disabled).
xhigh remains available via chat_template_kwargs: {"reasoning_effort": "xhigh"} — see the
validation notes below for when that is a bad idea.
Quantization
- Format: NVFP4 weights (group size 16) + FP8 KV cache
- Tool: NVIDIA ModelOpt 0.45.0 (
hf_ptq.py) - Excluded modules:
lm_head, embeddings, linear-attentionconv1d/in_proj_a/in_proj_b, MTP layers, vision tower - Calibration (4771 samples,
calib_seq=6144):HuggingFaceH4/ultrachat_200k— 2048 samplesnvidia/Nemotron-SFT-Multilingual-v2— 2048 samples- 675 self-generated long-CoT samples (this repo's addition): user-turn prefixes sampled from
Vtuber-plan/sharegpt-cleaned (EN 640 pool / FR 256 pool)
were answered by the unquantized base model atreasoning_effort=xhigh(1/3 greedy, 2/3 temp 0.7).
Only responses that terminated properly with</think>were kept (675/896; rejected: no close tag,
truncated, sub-200-char reasoning, or repetitive degeneration).calib_seqwas raised 512 → 6144 so
the<think> → </think> → answertransition stays inside the calibration window.
Validation: xhigh thinking-runaway gate (9 cases)
A 9-case gate (gate9_cases.json, included in this repo) probes runaway thinking:
FR/EN long-form tasks × temperature {0, 0.7} × reasoning_effort=xhigh, max_tokens 32768.
An 8-gram decile copy-rate analysis distinguishes genuine extended thinking (low copy rate)
from repetition loops (copy rate → 100%).
| Build | Closed-think & stopped | Hit token cap (runaway) | Hard 100% loops | Mean think len |
|---|---|---|---|---|
| This NVFP4 | 7/9 | 2 (greedy) | 0 | 49.5k chars |
| Base BF16 (huihui) | 7/9 | 1 (greedy) + 1 stopped-unclosed | 0 | 47.6k chars |
Existing qwen3.8-27b deployment |
2/9 | 7 | 2 (both greedy) | — |
Readings:
- No quantization regression: paired per-case deltas vs the BF16 base are mixed in sign and small
overall (mean 49.5k vs 47.6k chars; each build has exactly one greedy-decode cap-out and one
unclosed-think anomaly). On deliberately brutal xhigh prompts both builds overthink massively —
that is base-model behavior, not a 4-bit artifact. - Greedy decoding (temp 0) is the danger zone: every terminal runaway in every build occurred at
temperature 0. Once a repetition loop forms, greedy decoding cannot escape it. Use temp ≥ 0.7
and/orrepetition_penalty ≈ 1.05–1.1for xhigh workloads. - The deployed stock endpoint failed far harder (0/9, two hard loops) than either local build.
Benchmarks (GSM8K / MMLU / CMMLU / C-Eval) will be added here.
Usage
Recommended server (sglang, single 80 GB GPU):
CUDA_VISIBLE_DEVICES=0 python -m sglang.launch_server \
--model-path <this-checkpoint> \
--quantization modelopt \
--trust-remote-code \
--mem-fraction-static 0.8
import requests
requests.post("http://127.0.0.1:30000/v1/chat/completions", json={
"model": "Huihui-Qwen3.8-27B-abliterated-NVFP4",
"messages": [{"role": "user", "content": "..."}],
"temperature": 0.7,
"max_tokens": 8192,
# optional, only if you really want maximal deliberation:
# "chat_template_kwargs": {"reasoning_effort": "xhigh"},
})
Plain transformers/BF16 inference will not dequantize this checkpoint; use a runtime with ModelOpt
NVFP4 support (sglang--quantization modelopt, TensorRT-LLM).
Deployment recommendations
- Leave
reasoning_effortat the default (medium) — the shipped template makes this safe. - If
xhighis required:temperature > 0,repetition_penalty ≈ 1.05, and amax_tokenscap
(8k–16k) with retry-at-medium on timeout. - Only GPU 0 was used for all development and validation of this checkpoint.
Files
5 safetensors shards, each ≤ 5 GB (~19.5 GB total). gate9_cases.json is the release gate set.
License
Apache 2.0 (inherited from the base model). Quantization and calibration were performed on top of
huihui-ai/Huihui-Qwen3.8-27B-abliterated;
credit to huihui-ai for the abliterated base and to froggeric
for the fixed chat template.