← back to catalog · registered 2026-10-11 03:59

gyang274/Huihui-Qwen3.8-27B-abliterated-FP8

gyang274 27B multimodal second-order
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/gyang274%2FHuihui-Qwen3.8-27B-abliterated-FP8"
Response includes
  • classification m-uncensored
  • files 31
  • author_summary 1 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M-U
Primary method

Uncensored (method unknown)

No other method signals detected in this model.
Confidence
LOW
Why this label 3 signals
Weak or ambiguous signals. Best guess based on catalog patterns; treat as tentative and check the evidence below.
  • 'uncensored' in name/tags but no 'abliterated' marker
  • method not identifiable from author declaration alone
  • may be DPO fine-tune, prompt engineering, or unknown technique
Refusal direction extraction

No specific extraction method could be identified for this model. The producer either did not document it or used a proprietary pipeline.

What is a refusal direction? →
Downloads · 30-day
0
Likes
0
Model age
today
created 2026-10-11

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
apache-2.0
Tags
transformers safetensors qwen3_5 image-text-to-text fp8 qwen3.8 abliterated uncensored mtp sglang conversational arxiv:2406.11717

Related

Total size
28.7 GB
Files
31
Quantizations
1
Registered
2026-10-11 03:59
Last updated on HF
2026-10-11 03:41

Files by quantization

Auxiliary files 31 files 28.8 GB
model-00018-of-00018.safetensors 2.81 GB 98f25cc3 download
model-00003-of-00018.safetensors 2.37 GB 2e1bf62c download
model-00001-of-00018.safetensors 2.28 GB 58d699d9 download
model-00004-of-00018.safetensors 1.86 GB 251041f7 download
model-00016-of-00018.safetensors 1.86 GB 39246363 download
model-00006-of-00018.safetensors 1.86 GB 43d39c1c download
model-00008-of-00018.safetensors 1.86 GB 14cd1b23 download
model-00010-of-00018.safetensors 1.86 GB 065b2d9e download
model-00012-of-00018.safetensors 1.86 GB 3f93c80f download
model-00014-of-00018.safetensors 1.86 GB 7c17c430 download
model-00002-of-00018.safetensors 1.42 GB 93ec91ce download
model-00007-of-00018.safetensors 1006 MB e3c08650 download
model-00009-of-00018.safetensors 1006 MB 7691f8e6 download
model-00011-of-00018.safetensors 1006 MB 97d37c8c download
model-00013-of-00018.safetensors 1006 MB d57db7ce download
model-00015-of-00018.safetensors 1006 MB 7ea7bd83 download
model-00017-of-00018.safetensors 1006 MB 98360bbe download
model-00005-of-00018.safetensors 1002 MB 65498215 download
tokenizer.json 12.2 MB 0997f410 download
vocab.json 6.41 MB 0aa0ce06 download
merges.txt 3.20 MB a494e019 download
model.safetensors.index.json 152 KB 1351f5d3 download
config.json 50.1 KB e66d0dcb download
tokenizer_config.json 17.5 KB 5de744b3 download
LICENSE 11.3 KB f938136e download
chat_template.jinja 8.74 KB c0c686f9 download
README.md 5.05 KB 95b9bf89 download
.gitattributes 1.53 KB 52373fe2 download
preprocessor_config.json 390 B 2ea84a43 download
video_preprocessor_config.json 385 B 3ba673a5 download
generation_config.json 202 B 023756cf download

README current version from Hugging Face


library_name: transformers
license: apache-2.0
pipeline_tag: image-text-to-text
base_model:

  • huihui-ai/Huihui-Qwen3.8-27B-abliterated
    base_model_relation: quantized
    tags:
  • fp8
  • qwen3.8
  • abliterated
  • uncensored
  • mtp
  • sglang

Huihui-Qwen3.8-27B-abliterated-FP8

An FP8 quantization of huihui-ai/Huihui-Qwen3.8-27B-abliterated
that reproduces Qwen's official FP8 recipe byte for byte on every tensor the abliteration did not touch.

The weights are block-wise FP8 E4M3 with 128×128 scales, in the same layout as
Qwen/Qwen3.8-27B-FP8. The MTP head and vision tower are
preserved, and the checkpoint is 28.75 GiB. It loads through the same code path as the official FP8
checkpoint. This model was built and validated as part of a single-GPU deployment study. The full notes,
scripts and raw data are at github.com/gyang274/ms-Qwen3.8-27B.

What makes this quant verifiable

The quantizer (quantize_fp8_like.py)
uses the official FP8 checkpoint as a layout template. It quantizes exactly the tensors that are FP8 there,
copies the rest in the template's dtype, and asserts that names, dtypes and shapes match. Its scale arithmetic
matches Qwen's: an FP32 block scale amax × (1/448) is used for quantization and then stored as BF16.

As a result, comparing this checkpoint with the official one isolates the abliteration edit exactly:

Tensors vs. official Qwen3.8-27B-FP8
Layers 0-16 and 52-63, MTP head, vision tower (333), embedding, LM head byte-identical
mlp.down_proj (35), linear_attn.out_proj (26), self_attn.o_proj (9) in layers 17-51 differ: these are the abliteration edit

Only 70 of 1,606 tensors differ, and all of them write into the residual stream. That is the footprint
directional ablation predicts (Arditi et al., 2024). An SVD against the
official BF16 weights gives the edit's exact form:

  • Every changed matrix is rank-1: σ2/σ1 ≤ 0.012.
  • All 70 share one unit direction r: pairwise |cos| ≥ 0.99999.
  • All 70 use one scale: ΔW = −1.3 · r rᵀW.

Rebuilding the abliterated weights from the official ones with that single r reproduces 91.9% of the BF16
elements bit for bit. See §8-9 of the notes
for the recipe probe and the full analysis.

Evaluation

Both models were served with SGLang v0.5.21 using an identical configuration, thinking off and temperature 0,
on an RTX A6000.

Official Qwen3.8-27B-FP8 This model
HumanEval pass@1 (164, sandboxed) 98.2% 95.7%
GSM8K, first 200 97.5% 94.5%
Over-refusal: 15 benign prompts that safety-tuned models often refuse 2 partial 0
Decode, single stream (RTX A6000) ~60 tok/s ~62 tok/s
MTP mean accepted length 3.52 3.52

Removing refusals has a measurable cost. On GSM8K this model lost 6 problems the official model solved and
gained none. Most of the losses came from indecision on ambiguously worded problems rather than arithmetic
errors.

Usage

SGLang (tested with v0.5.21). On Ampere (sm_86, sm_80) FP8 runs automatically as Marlin W8A16. On Ada,
Hopper and Blackwell it uses native FP8.

python3 -m sglang.launch_server \
  --model-path gyang274/Huihui-Qwen3.8-27B-abliterated-FP8 \
  --context-length 262144 --kv-cache-dtype fp8_e4m3 \
  --reasoning-parser qwen3 --tool-call-parser qwen3_coder \
  --speculative-algorithm EAGLE --speculative-num-steps 3 \
  --speculative-eagle-topk 1 --speculative-num-draft-tokens 4

The deployment notes give the memory flags that fit the full 262K context with 4 concurrent requests on a 48 GB
card. The checkpoint uses the same format as the official FP8 release, so other engines that load
Qwen/Qwen3.8-27B-FP8 (for example vLLM) should load it the same way. Only SGLang was tested here.

Provenance

Source huihui-ai/Huihui-Qwen3.8-27B-abliterated @ 739e3c5b89849f6c238ce1e5b70008612ae42cdd
Template Qwen/Qwen3.8-27B-FP8 (config.json and quantization layout)
Quantization FP8 E4M3, 128×128 blocks, dynamic activation scheme, max relative weight error 2.65%
Original model Qwen/Qwen3.8-27B, Apache-2.0

Risks

This model has had its refusal behavior removed and will comply with requests that the original model
declines. It is intended for research and for personal use in controlled settings. It is not suitable for
public-facing or child-facing deployments without your own safeguards. You are responsible for how you use
it and for the outputs it produces.

Credits

Qwen for Qwen3.8-27B and its FP8 recipe, and huihui-ai for the abliterated model. Licensed Apache-2.0, as are
both upstream models.

Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration